Article
Disclosure Is Not Verification
Europe's flagship artificial intelligence deadline arrived on 2 August. It requires machines to be labelled. It does not require anyone to show that the machines are right. Four developments in the two weeks around it show what that distinction costs in life sciences.

Europe regulated the label and deferred the claim
On 2 August 2026 the European Commission began enforcing Article 50 of the AI Act. A system that talks to a person now has to say it is not a person. Content generated or manipulated by a model has to carry a visible label and a machine-readable mark. Deepfakes have to be disclosed. Fines run to 15 million euros or 3 percent of global turnover.
That is the whole of what took effect. The obligations scheduled for the same date, the ones governing high-risk systems, did not arrive. The Digital Omnibus on AI, endorsed by the European Parliament on 16 June and approved by the Council on 29 June, entered into force on 27 July. It moved standalone high-risk systems under Annex III to 2 December 2027, and high-risk AI embedded in regulated products under Annex I, which is where medical devices and in vitro diagnostics sit, to 2 August 2028.
The substance did not shrink. Providers will still owe data governance, technical documentation, robustness evidence, human oversight and postmarket monitoring. What changed is when, and the reason is the more interesting fact: the harmonised standards meant to make compliance demonstrable are unfinished. Regulators deferred the evidence requirements because the machinery for producing and assessing that evidence does not exist yet.
So the most ambitious AI law in the world reached its first hard enforcement date and shipped a disclosure regime. A label tells a reader that a machine was involved. It says nothing about whether the machine was correct.
Grail's submission shows what happens when the strongest test fails
On 7 August, Grail announced that the FDA's Molecular and Clinical Genetics Panel will meet on 23 September to review the premarket approval application for Galleri, the first PMA for a multi-cancer early detection test. Galleri reads methylation patterns from blood and uses machine learning to predict the origin of a detected cancer signal. This is an AI-derived clinical claim heading for the highest-evidence pathway the device framework has.
The February result matters. NHS-Galleri, a randomised trial of roughly 140,000 participants and the only randomised trial of a multi-cancer test in its intended-use population, did not show a statistically significant reduction in cancers diagnosed at stage III or IV.
Look at what the submission is made of. Per Grail's own release, the package rests on results from 25,490 consented participants in PATHFINDER 2 with one year of follow-up, plus more than 70,000 participants drawn from the intervention arm of the first screening round of NHS-Galleri. It is also supported by an analysis comparing the version of Galleri used in those studies to the updated version actually submitted for approval.
Three things are happening at once. The randomised comparison that could have falsified the clinical benefit claim did not support it, and the evidence now before the panel is single-arm performance extracted from inside that same trial. Single-arm performance can establish that the test detects signals accurately. It cannot establish that detecting them earlier helps anyone, which is what the randomisation was designed to answer. And the model that generated that evidence is not the model under review; a bridging analysis carries the weight of connecting them.
Each move is ordinary regulatory practice and none is improper. Together they describe a structural problem. When the strongest available check comes back negative, the evidence base migrates toward artefacts that cannot deliver a negative answer, and toward comparability arguments resting on the sponsor's own account of what changed between model versions. In September a panel will weigh a claim about a model it cannot independently run against a trial whose answer it already has.
The generation rate is about to move again
On 5 August, Jeff Dean left Google after twenty-seven years, with Sanjay Ghemawat, Oriol Vinyals and Quoc Le, to found Discovery Loop. The company is a public benefit corporation with a single stated purpose: automating the experimental loop of science, running thousands of cycles in parallel with minimal human involvement. It starts with machine learning research, then generalises to domains including drug discovery. Alphabet is a founding investor and cloud partner.
These four built much of the computational infrastructure this era runs on, so the ambition deserves to be taken at face value. If it works even partially, the rate at which hypotheses are proposed, tested in silico and reported moves by an order of magnitude.
Nothing announced in this window moves the rate at which those results can be independently confirmed. In a regulated industry, a finding that cannot be checked by anyone other than its producer does not become more useful when it arrives a thousand times faster. It becomes more of the same liability, arriving at a review process that already cannot keep up.
The same week, Novo Nordisk named Amazon Web Services its preferred cloud and AI partner, opening a London co-innovation hub aimed at compressing the path from drug target to first human dose. The metric offered was that more than 25,000 employees have used tools from the partnership: an adoption number standing in for an outcome number. Compressing target to first dose is testable. Building the counterfactual is genuinely difficult, which is precisely why nobody has published one.
The industry cannot agree on a metric for the thing it can count
BioPharma Dive asked a small question on 7 August with a large answer. Bristol Myers Squibb, Eli Lilly and Roche have each claimed the most powerful AI machine in the industry. How can all three be true? Nvidia's senior director of business development for life sciences answered plainly: they use different yardsticks. One measures operations per second, another counts GPUs, a third compares a single owned system against a hybrid-cloud factory. Four superlative claims, no shared measure, no way for an outsider to rank them.
Compute is the most countable thing in this field. If the industry cannot agree on a unit for hardware, agreeing on units for whether a model's output is trustworthy enough to act on is remote unless someone builds those units deliberately.
One pilot treats checking as a design requirement
The counter-example arrived just before this window. On 22 July the FDA named Dexcom the first participant in its TEMPO pilot, which offers enforcement discretion on premarket authorisation for certain digital health devices in exchange for collecting and reporting real-world performance data, tied to the CMS Innovation Center's ACCESS Model.
This is the most structurally interesting regulatory move of the period, because it treats evidence generation as a continuing obligation attached to market access rather than a gate passed once. Access and checking are coupled by design.
The limit is equally clear. The manufacturer collects and reports the data on its own device: a well-instrumented account of what a system did, provided by its owner. It does not let an outside party confirm that account is accurate without trusting the party that produced it. That is the difference between a good audit trail and an independent check, and it is the gap the pilot will have to close.
Compute is not evidence, and disclosure is not verification
Last cycle the argument was that compute is not evidence: buying the capacity to generate claims is not the same as buying the ability to substantiate them. This cycle extends it. Disclosure is not verification either.
Both are the same error at different ends of the pipeline. Compute is an input mistaken for a result. A label is a provenance marker mistaken for a check. Neither tells you whether a claim is true, and both feel like progress because both are cheap and countable. Verification is neither.
Three distinct things get conflated. Traceability records what a system did, and the record comes from the system itself. Reproducibility means the same inputs produce the same outputs, which is necessary and badly insufficient, because a system can be deterministically and repeatedly wrong. Verifiability means an independent party can confirm a specific result is correct without rerunning the producer's pipeline and without taking the producer's word for it.
Read across this window and the pattern is uniform. Europe mandates a provenance marker, which is traceability. Grail's bridging analysis asks a panel to accept the sponsor's account of a model change, which is traceability. TEMPO collects manufacturer-reported data, which is traceability with better instrumentation. Discovery Loop will produce results at a rate that makes even reproduction economically impossible at scale.
Nobody is buying the third thing. Not because the industry is careless, but because it does not yet exist as purchasable infrastructure the way compute, cloud contracts and compliance labelling do. The regulators deferring the AI Act's high-risk regime said as much: the standards and assessment tools are not ready.
That is the actual gap. Not in capability, capital or regulatory intent, but in the artefacts. A claim produced by a model needs to travel with something an outsider can check on its own terms: a portable, signed, content-addressed record bound to that specific claim, verifiable without the producer's cooperation and without rerunning anything. That is what turns a label into a proof and an audit trail into an independent check.
That is what evidence infrastructure is for.
