Article

Validation Is Not Verification

In one week of August, two algorithmic outputs reached patients. Both cleared the evidence bar the system asks for. Neither can be checked in the case that matters.

On August 19, Merck and Moderna reported that intismeran autogene plus Keytruda met its primary endpoint of recurrence-free survival and a key secondary endpoint of distant metastasis-free survival in INTerpath-001, a 1,137-patient trial in resected stage IIB-IV melanoma. It is the first positive Phase 3 readout for an individualized neoantigen therapy. The next day, the FDA cleared C2N Diagnostics' PrecivityAD2 for symptomatic adults 40 and older, the first Alzheimer's blood test authorized for patients that young.

Both are genuine advances. Both also share a property that the coverage did not examine: the thing that was validated is not the thing the clinician has to act on.


The trial validated the pipeline, not the dose

Every dose of intismeran autogene is manufactured for one patient. It encodes up to 34 neoantigens selected from the mutations in that individual's tumor. No two patients receive the same product.

Randomization answers a specific question well: across a population, does this process produce benefit relative to the comparator. INTerpath-001 answered it. What randomization cannot answer is whether the 34 peptides chosen for a particular patient were the right 34 for that patient.

Conventional pharmaceutical quality control assumes a fungible product. There is a reference lot, a retained sample, a specification, and an assay that any competent third-party laboratory can run to confirm that this vial matches that standard. Individualized therapy dissolves that assumption. There is no reference lot because there is no lot. There is no comparator dose because every dose is unique. The selection step that determines what is in the vial is a computational judgment, and the evidence that it was a good judgment is statistical, distributed across 1,137 people, and unavailable in any single case.

This is not an argument that the therapy should not have been approved on this evidence. It is an observation that the evidentiary structure the industry inherited from small molecules has quietly stopped applying, and nothing has replaced it.


A cleared diagnostic can still be an unopenable function

PrecivityAD2 measures plasma amyloid-beta 42/40 and phosphorylated tau 217 by high-resolution mass spectrometry, then folds both analytes into a proprietary Amyloid Probability Score 2. In the clearance-supporting cohort of 1,142 symptomatic adults, positive predictive value was 97.6 percent and negative predictive value 93.1 percent against amyloid PET or cerebrospinal fluid findings.

The interesting number is the third one. An intermediate band, reported as "likely positive," captured 17.3 percent of patients and carried a positive predictive value of 77.3 percent. Roughly one patient in six lands in a category where the test is right about three times in four.

The measurement side of this is exemplary. Mass spectrometry is traceable, and the analyte values are reproducible by any laboratory with the instrument and the method. The mapping from those values to the category is not disclosed. A clinician can reproduce the inputs and cannot derive the output.

That gap is the whole subject. FDA clearance is a statement about how the device performed on a cohort. It is not a warrant that any particular result is correct, and it does not equip the person reading the report to interrogate why a patient landed in the intermediate band rather than the negative one. The score is traceable in the sense that it can be attributed to a system and a version. It is not verifiable in the sense that a skeptic outside C2N could confirm it without taking the company's word for the function.


The only serious verification instrument built this cycle came from a nonprofit

On August 20, Arc Institute opened the 2026 Virtual Cell Challenge. The design deserves attention from people who will never enter it.

There is no challenge training set this year. Models must predict CRISPR-interference knockdown responses in six cell lines they have never seen perturbed, given only the unperturbed state of those cells and a list of target genes. Three lines drive the live leaderboard. Three are held back for final scoring, with Arc's own measurements withheld as ground truth. Final rankings aggregate six metrics, and Arc stated the reason plainly: year one showed that a narrow scoring surface invites optimizing the metric instead of the biology.

Read that as an architecture rather than a competition and it is a complete verification apparatus. An independent party holds the answer. The claimant cannot see it. The scoring function is deliberately hardened against gaming. Nothing in the process requires trusting the modeler's account of their own performance.

It is instructive to set this against GenBio AI's preview of AIDO Cell two days earlier, positioned as a stateful simulator spanning DNA, RNA, protein and whole-cell behavior rather than transcriptomics alone. The claim may well be correct. Co-founder David Baker was careful to say the field has not solved cellular biology. But the claim is currently graded by its author, which is the default condition of nearly every model release in this field.

The difference between the two is not honesty. It is that one of them cost money. Arc ran Perturb-seq across six cell lines using CRISPR-interference and high-throughput sequencing to generate ground truth it then refused to publish. That is the expense. A nonprofit paid it, and the prize pool is one hundred thousand dollars.


The industry has already named the constraint and funded something else

Writing on August 11, Moe Alsumidaie assembled the figures now circulating among pharmaceutical digital leaders. AI-designed candidates are clearing Phase I at 80 to 90 percent against a historical industry rate of 40 to 65 percent. Of 117 tracked AI-enabled assets that have entered interventional trials, 8 have cleared Phase II. Vikram Singh of Gilead compressed it into a sentence: speed is bought, trust is earned in Phase III. Greg Meyers of Bristol Myers Squibb argued that AI is moving rather than removing bottlenecks in discovery.

One statistic in that collection is doing more work than the rest. A ZS survey of 115 pharmaceutical and biotechnology digital leaders found that 68 percent attribute stalled AI programs to data quality and governance rather than model capability.

That is a verifiability finding wearing different clothes. Data quality and governance is what an organization calls the problem when it cannot establish that a record is what it claims to be. The constraint the industry keeps reporting is not that the models are weak. It is that nobody can confirm the provenance of what went in or the integrity of what came out.


The case against treating any of this as a problem

Three objections are worth taking seriously.

The first is that medicine has always relied on instruments clinicians cannot derive from first principles. Nobody asks a physician to reconstruct a mass spectrometer's calibration curve. This is true, and it is the strongest objection. It is also incomplete. A regulated assay has a reference standard, defined analytical performance, and external proficiency testing programs that let independent laboratories confirm they get the same answer on the same sample. A proprietary composite score has none of those. The comparison flatters the score by borrowing the assay's infrastructure.

The second is that trade secrecy funds the research. Also true. Nothing here requires publishing a model. Verification does not mean disclosure, which is precisely the point the previous cycle made. A held-out benchmark verifies without revealing anything about the model's internals.

The third is that randomized trials are the strongest evidence instrument medicine has, and demanding case-level confirmation misunderstands what statistical evidence is for. This is correct as a matter of statistics and beside the point as a matter of practice. The claim here is narrower: where an output is individualized and consequential, aggregate validation should not be read as case-level assurance. It routinely is.


Disclosure was the cheap part

The previous cycle ended on the observation that disclosure is not verification. This cycle supplies the reason the gap persists, and it is not ignorance.

Traceability is cheap. A system logs what it did, and the log costs almost nothing to produce. Verification requires a second party with resources, custody of ground truth, and enough independence to publish an unflattering result. Arc built exactly that and had to run wet-lab experiments to do it. No regulator built it. No consortium built it. No vendor built it, because a vendor cannot: the entire value of held-out ground truth is that the claimant does not hold it.

Two algorithmic outputs entered clinical practice in the same week, and the only serious checking apparatus produced in that window was a machine learning competition run by a nonprofit in Palo Alto. The asymmetry is not an accident. It is a pricing failure. Until somebody is structurally obliged to pay for independent verification, the field will keep producing validated systems that nobody can check.

That is what evidence infrastructure is for.