Article
A Clearance Is Not Proof
In one fortnight the AI tools got sharper and the deals got bigger. Not one of them made a single AI decision easier to verify. That gap is the whole game.

For most of the last two years, the industry has treated an FDA clearance as the finish line. Get the letter, and the AI is proven. The reasoning is understandable. A clearance is hard to win, it involves a regulator, and it arrives with a paper trail. Surely that is proof.
It is not. A clearance tells you a review happened. It does not, by itself, let anyone outside the room confirm what the model actually does, on what data, checked by whom. Those are different things, and this cycle showed the distance between them inside a single filing.
The clearance is a receipt, not a proof.
On June 25, a clinical AI company called UpDoc announced what it described as the first FDA-cleared platform built on patient-facing large language models. It is the kind of claim that travels: the first medical LLM a patient can talk to, blessed by the agency.
Then you check the record. The only matching entry in the FDA's public 510(k) database is a bounded insulin-dosing tool for adults with type 2 diabetes, cleared back in December. That record names no foundation model, no prompt configuration, no version, and no validation harness. If a broader, LLM-driven device has in fact been cleared, it is not yet in the public file.
So the marketing describes a general-purpose medical LLM. The public evidence describes a narrow, specific insulin calculator. The clearance is real. It is also a receipt that a review took place, not a document anyone can use to verify the model's behavior.
The distinction is worth stating plainly. Traceability is a record that something happened. Verifiability is the ability of an independent party to confirm it is correct. A clearance, an audit log, a signed report: these are traceability. They tell you the process ran. They do not tell you the process was right.
The 2020 SolarWinds compromise is the cleanest reminder. The updates were signed with legitimate certificates. The paperwork was immaculate. The software was also trojanized. A record that looks correct is not the same as a system that is correct, and self-attestation is not verification. In a regulated AI workflow, the same logic holds: the letter on the wall is not proof of the decision underneath it.
The workbenches are real. That is not the same as proof.
The capability news this cycle was genuine, and worth conceding before turning on it.
On June 30, Anthropic launched Claude Science, a research environment that wires more than 60 databases and tools into a single workbench, and announced that it will run its own drug-discovery program aimed at neglected diseases. It is the last of the three frontier labs to plant a flag; OpenAI shipped its GPT-Rosalind model in April, and Google works through Isomorphic. Anthropic's newest model, Fable 5, is being pitched as making parts of drug design ten times faster.
Take the claim at face value. Assume the workbench does everything it says. You have compressed the front of the pipeline, the ideation, the literature search, the hypothesis. That is valuable. It is also the cheap part.
The expensive part is downstream, and it did not move. Zero AI-designed drugs have FDA approval. A faster workbench generates claims faster; it does not generate verified claims, and the distance between a claim and a claim a regulator will accept is exactly where the last two years of AI in biotech have gotten stuck. The tool makers now want to be drug makers, which is a reasonable ambition and changes nothing about that arithmetic. Better tooling at the top of the funnel does not clear the bottleneck at the bottom.
You cannot acquire your way past the evidence.
The capital told the same story. On July 2, Insilico Medicine signed a collaboration with Takeda, about $60 million up front and up to $600 million in total, to run its discovery platform across Takeda's core therapeutic areas. It is a rational deal. It is also the third or fourth version of the same deal this year.
Set it against the scoreboard. Isomorphic's $2.1 billion raise in May has still not produced a single dosed patient. Roughly 173 AI-discovered programs sit in clinical development, and 15 to 20 of them may reach pivotal Phase III trials this year, the cohort that will finally test whether the thesis was real.
Capital buys discovery access. It does not buy the chain of evidence that survives an inspection. The investable asset has moved from the molecule to the substrate that produces molecules, and that is a sound bet. But the substrate still has to emit output an outsider can verify, and no amount of upfront payment shortcuts that step. Insilico, to its credit, is one of the few players with an asset actually in Phase II, which is precisely the point: the scarce thing is not the discovery, it is the proof that the discovery holds.
One clearance this cycle did it the other way.
It would be easy to read all of this as a case against AI in regulated medicine. It is not. It is a case for building the evidence layer into the artifact, and one clearance this cycle showed what that looks like.
On July 13, Qure.ai's chest-X-ray triage software cleared the FDA with two features that matter more than the clearance itself. First, a Predetermined Change Control Plan, which means the model is allowed to update after market, but only inside a boundary the agency agreed to in advance. Second, built-in explainability: the system does not merely flag a finding, it shows where and why, with visual localization and region-of-interest labels a radiologist can check.
Put UpDoc and Qure.ai side by side. Same regulator, same fortnight, opposite postures on verifiability. One asks you to trust the letter. The other builds the means of checking into the product, and accepts the harder obligation that comes with a change-control plan: you now have to prove the model still works after it changes, continuously, not once on the day it shipped. That obligation is not a burden to be minimized. It is the actual product.
The throughline.
Last cycle I wrote that the moat had moved to data, and the gap had moved to proof. This cycle the gap did not close. It got more expensive to ignore.
Every headline in the window was a capability or a deal: a workbench, a model, a partnership, a clearance. The scarce input underneath all of them is the same, a claim you can substantiate, from raw signal to final decision. The frontier labs solved for capability. The pharma balance sheets solved for access. Even the FDA solved for its own throughput, folding more than 40 of its internal data systems into a single platform this spring so its reviewers could query across them. Everyone optimized the parts of the stack that were already strong.
Nobody, this cycle, shipped the part that is missing. There is a ladder here, and the field keeps climbing the first two rungs and calling it the summit. Traceability, a record that a decision was made, is rung one. Reproducibility, the same inputs producing the same outputs, is rung two. Neither is safety, because a signed, repeatable process can be reliably wrong. Verifiability, an independent party confirming the decision was correct, is the third rung, and it is the one almost no one is building.
The moat moved to data. The gap moved to proof. The next advantage will not belong to whoever ships the fastest workbench or signs the largest deal. It will belong to whoever can prove that every datapoint, and every decision built on it, is correct and accountable from the raw signal to the final claim.
That is what evidence infrastructure is for.
