The Science Machine
Why biomedical research is stagnating and how to fix it.
TL;DR: At ScienceMachine, our goal is to radically accelerate the development of new medicines. We do so by developing AI to automate the complex, slow processes behind drug development. Helping regulated biopharma companies scale their internal processes in a compliant, auditable, and reliable way.
Research productivity is declining across the board. In semiconductors, agriculture, and the economy at large, sustaining a given rate of improvement requires ever more researchers and capital. 1 Biology is even worse. Drug discovery has absorbed enormous gains in technology and capital, just to maintain a constant rate of innovation. New-drug approvals per inflation-adjusted R&D dollar roughly halved every nine years for the past 70. 2 3
This stagnation is a time bomb. As our populations age and our already-embattled healthcare systems start to wane, we will need much more of the only deflationary technology we have in healthcare: therapeutics. Therapeutics have the ability to massively reduce the cost of delivering care to patients at scale. No other technology does this. So it is important to realise, stagnation is not an option here.
Why stagnation is happening
There are many specific explanations, but the general one is complexity. It will come as a surprise to no one that biology is complex. But how and why that complexity matters is more nuanced. Diabetes is no more complex a disease today than it was 70 years ago. But bringing treatments for it to market has become much more expensive.
For instance, for a disease where there already exists an effective treatment, any new entrant is likely to bring a small marginal benefit. This means that any trial to prove superiority will require a much larger sample size. This is usually prohibitive. So drug companies instead focus on subsets of patients where they can achieve larger effect sizes and run smaller trials. But this brings with it a lot of extra complexity: which patients? why? how do you define them? how do you find them? how do you recruit them? etc.
Moreover, regulators have, over time, layered rules and auditing requirements to avoid adverse events and limit bad behaviour from market participants. While some of these could be streamlined (as China has done quite effectively 4), most are irreducible in the sense that they are the logical consequence of wanting safe and efficacious drugs. If we were to wind back the clocks and start over with the same goal, we would likely end up in a similar spot.
So you can break down complexity into two components: biological and regulatory. And both of these have grown at an exponential rate — so much so that the massive increases in investment were necessary just to maintain the rate of innovation.
What to do about it
When we first founded ScienceMachine, we were obsessed by the problem of The Long Tail and how coding agents were so uniquely positioned to help solve it. What we did not realise at the time was that it was just another symptom of the same complexity problem. Briefly, The Long Tail tries to explain why software for life sciences has been a challenging category to build in. It poses that biotech is hyper-diverse, with practically infinite edge cases, so building traditional software for it was really hard. But it did not explain why it was so diverse. Fundamentally it is because biology is really complex and the rising bar for drug quality pushes us further out into the complexity frontier.
We also did not consider regulatory complexity as the other axis of the issue. To ensure safe and efficacious drugs, and to avoid bad behaviour from malicious actors, regulators require additional experiments and significant documentation. This interacts with research in interesting ways. It also often means scientists spend their time doing things other than science. Method validation reports, for example, document in detail a particular experimental assay and whether it is fit for clinical use. It is hundreds, if not thousands, of pages long. So big in fact that Microsoft Word frequently freezes. It is a critical piece of documentation, as its information could significantly impact e.g. toxicity readouts in trials.
Any solution that wants to make a dent in our productivity conundrum, then, will need to take both of these components into account: the intelligence to interpret biological complexity and the infrastructure to be compliant by default. A machine that will be smart enough to untangle the gnarliest science, and built in a way society can actually rely on and use its findings. A Science Machine, if you will.
Building the Science Machine
This is what we are dedicating ourselves to: building reliable, compliant and accurate AI that can be relied upon to execute complex scientific work autonomously, ask for human review at critical checkpoints, and deliver results in a transparent and traceable way. All of this done in a secure and private environment, with appropriate access controls and tools for collaboration.
The models provide the fluid intelligence, and it is our job to use the best models for the task at hand. As we are hoping to publish, different models are better for different tasks. We view the model being used as an implementation detail that the scientist should not need to worry about. They just need accurate and reliable results that they can use in their work. Asking them to learn when to use which model, with which reasoning effort, and how to prompt it would just load them with even more complexity.
Science is fundamentally a collaborative enterprise, and as such we need better infrastructure for working together on scientific projects. Systems and concepts that have become commonplace in software development have yet to arrive in the life sciences: version control, forking, audit logging, CI/CD, RBAC, and “PR” review are all basics in the SDLC, but have either been too difficult to make work in the domain, or just simply not been relevant. The intelligence now enables many of these things as we can build a unified workspace where all of the work can get done.
This collaborative infrastructure also make compliance a non-issue, as everything is tracked, logged and versioned by default. Every report, figure or table can be traced back to the work that produced it, down to the exact lines of code used. This makes cheating impossible to get away with, without burdening scientists with more complex audit logging and QC work. We get to have our cake, and eat it too!
Therapeutic abundance
The Science Machine will not only stave off the collapse of the healthcare system, it will usher in an era of therapeutic abundance. A new golden age of drug development akin to the 1950s, but with all of today’s safety and efficacy guarantees we’ve come to expect. Data rolls off the lab instrument, is automatically analysed and contextualised, interpreted, submitted to the FDA, and reviewed by the regulator on the same day. The machine will be the digital exoskeleton of the mind to move scientific mountains. To assist us in our fight against complexity.
The Science Machine will thaw the biotech winter and create exceptional returns for the industry. But more importantly, it is in our view the path to longevity for both individuals and society. Imagine a world where cancer is no worse than a UTI, or where we can prevent Alzheimer’s with a vaccine. And where the cost of delivering healthcare to an aging population is actually decreasing.