Article
Compute Is Not Evidence
Three pharmaceutical companies have now claimed the largest AI machine in the industry. In the same fortnight, the first honest scoreboard for AI-discovered drugs showed the advantage thinning out at exactly the point where independent checking begins.

The argument for scale in AI drug discovery has been simple and, until recently, hard to dispute. Bigger models, trained on more data, running on more compute, will find better molecules. Better molecules will become better drugs. The bottleneck is capacity, so buy capacity.
The last two weeks bought a great deal of capacity. They also produced the first serious look at what happens to AI-derived assets once they leave the model and enter a clinic, and the answer complicates the scale argument in a specific way.
Three companies in nine months have claimed the same superlative.
On July 20, Bristol Myers Squibb announced it would deploy an Nvidia DGX SuperPOD built on Vera Rubin NVL72 systems, calling the result the most powerful single-owned Nvidia infrastructure in life sciences. The hardware is real and the efficiency gain is substantial, up to ten times the performance per megawatt of the system it replaces.
It is also the third such claim in nine months, as Brittany Trang noted at STAT. Eli Lilly committed up to a billion dollars with Nvidia over five years in January. Roche launched a hybrid-cloud version in March.
Robert Plenge, chief research officer at BMS, framed the purpose precisely: better decisions come from better evidence, faster.
He is right, and the sentence does two jobs the industry keeps collapsing into one. Compute makes evidence faster. It does not, by itself, make evidence better. Those are separate investments, and only one is being funded at scale.
The market is pricing discovery and paying almost nothing for confirmation.
Capital said the same thing twice. On July 14, Chai Discovery raised 400 million dollars at a 3.8 billion dollar valuation, close to a tripling in eight months, on the strength of platform deals with Eli Lilly and Pfizer. A week later, Dimension closed an 800 million dollar third fund, 60 percent larger than its second, reaching 1.65 billion under management.
Neither valuation rests on a clinical readout, because there is not yet one to rest on. Both rest on model capability and signed deployment contracts. That is a defensible way to price a discovery platform. It is also a precise statement of what the market rewards today: the ability to produce candidates, not the ability to show the reasoning behind them was sound.
Phase 1 is the wrong place to look for evidence that AI discovery works.
PharmaVoice published the most useful thing in the window on July 15. Reviewing how AI-derived candidates are actually performing, Kelly Bilodeau found a consistent pattern: they clear Phase 1 at higher rates than conventional ones, per work in Nature Biotechnology, and that advantage appears to narrow in Phase 2, per work in Nature Medicine.
Read quickly, this looks like early promise fading. Read carefully, it is a story about what each phase actually tests.
Phase 1 asks whether a molecule is tolerable in humans. That is largely a chemistry question, and the properties involved, solubility, pharmacokinetics, off-target liability, are computable. Generative chemistry is genuinely good here, because the model optimizes directly against measurable objectives. A Phase 1 advantage is real, and it is evidence that AI optimization works on the things AI can already measure.
Phase 2 asks a different question. It asks whether the target hypothesis is correct in human biology. That hypothesis is the claim the AI actually made. And Phase 2 is the first moment at which anyone tests that claim independently of the system that produced it.
So the narrowing is not a disappointment. It is what an unverified claim looks like meeting its first outside check, years and many millions of dollars after it was made.
Verge Genomics makes the point in the negative. Its AI-identified ALS candidate, VRG50635, never moved past Phase 1. The molecule behaved. The hypothesis did not.
Insilico is the counterexample, and the reason is not the platform.
It would be easy to read the above as a case against AI-driven discovery. It is not. One program this cycle shows what the other posture looks like.
Insilico Medicine initiated a Phase 3 trial of rentosertib on July 7, a 320-patient, 52-week study in idiopathic pulmonary fibrosis. The target, TNIK, was prioritized by an AI biology engine; the molecule was generated by a generative chemistry platform. It is the furthest any AI-discovered drug has travelled.
What makes it credible is not the platform. Many companies have platforms. It is that the chain is externally checkable at every link: the discovery-to-clinic path published in Nature Biotechnology, the medicinal chemistry in the Journal of Medicinal Chemistry, randomized Phase 2a results in Nature Medicine, and registered trial identifiers a stranger can look up. An outsider can interrogate the reasoning without asking Insilico to vouch for it.
Recursion is doing a version of the same thing, reporting a 43 percent median reduction in polyp burden for REC-4881 in an ongoing Phase 1b/2 that completes in 2027, with the data in the public record rather than a deck.
That is the distinction that matters. Not whether a company used AI, but whether the claim it produced can be inspected by someone who was not in the room.
The one investment aimed at checking is a line item on a list of 85.
Something encouraging happened this cycle, and it was almost invisible. On July 16, Joanne Eglovitch reported at Regulatory Focus that CDER had added four guidances to its agenda since February. One is a supplemental guidance on computer software assurance for AI-based systems in drug manufacturing and clinical investigations.
Computer software assurance is a better idea than its name suggests. FDA finalized the underlying framework in February, superseding a version issued the previous September. Its core move is to scale the depth of your evidence to the consequence of the software being wrong, rather than validating everything to the same depth and calling the paperwork proof. It replaces ritual with proportionality.
Extending that logic to AI is the closest the agency has come to treating model validation as a discipline in its own right, distinct from both software validation and clinical validation. It is the right instinct.
It is also one entry on a list that now runs to 85 documents, at a center that states plainly it is not obliged to follow the list.
The regulator cannot yet document its own AI.
The timing is unfortunate. PharmaVoice reported on July 24 that after Marty Makary's departure, chief AI officer Jeremy Walsh and acting chief information officer Sridhar Mantha have both left. Tala Fakhouri, now at Parexel and previously an FDA AI policy official, described the resulting governance picture as unclear and expects a drift back toward fragmented, center-by-center practice.
Meanwhile the agency's own use of AI keeps growing. Internal AI use cases rose 148 percent between 2024 and 2025, per the Bipartisan Policy Center. Elsa, the agency-wide tool, runs on retrieval-augmented generation confined to curated databases, a sound choice for reducing fabrication. But little is published about how it informs review work, so sponsors cannot see the standard they are held to.
Sponsor-facing AI policy has not changed, which is the good news. The harder fact is that the institution that will adjudicate your evidence cannot presently document its own, and Fakhouri's estimate is that even the fastest guidance takes about a year.
The throughline.
Last cycle the argument was that the moat had moved to data and the gap had moved to proof. This cycle the gap did not close. It got funded on one side only.
Look at what was bought in fourteen days. More compute at BMS. More capital at Chai and Dimension. More programs moving into later phases. More topics on a guidance agenda. Every one of those purchases raises the rate at which this industry generates claims. Not one raises the rate at which anyone can confirm them.
The ladder has not changed. Traceability, a record that a decision was made, is the first rung. Reproducibility, the same inputs yielding the same outputs, is the second. Neither is safety, because a logged and repeatable process can be reliably wrong. Verifiability, an independent party confirming the decision was correct, is the third. The Phase 2 data is what that third rung looks like when it arrives late: a delayed, expensive, first-time check on a claim made years earlier by a system nobody could interrogate at the time.
Insilico's publication record is the alternative in miniature. Not a better model. A claim built so that outsiders can check it without permission.
The moat moved to data. The gap moved to proof. The next advantage will not belong to whoever owns the largest machine or the largest fund. It will belong to whoever can show that every datapoint, and every decision built on it, is correct and accountable from the raw signal to the final claim, and can show it before Phase 2 does it for them.
That is what evidence infrastructure is for.
