On AI and the regulatory interface
Model Validation, Process Validation, and the Decision That Is Not Yours to Make
Published 2026-08-24 · FDA guidance status is stated explicitly throughout — draft guidance describes current thinking and is not binding. Not legal or regulatory advice. Practice varies by Center, office, and division.
The question that changed
Andreas Bender, Jack Scannell, David Shaywitz and thirteen co-authors have a Perspective out this month in Nature Reviews Drug Discovery on AI in drug discovery. The useful part isn't any single finding. It's that they change the question the field has been asking.
They separate two things that routinely get run together. Model validation asks whether a model performs well. Process validation asks whether it improves the decisions someone actually makes. Their conclusion is blunt: "models need to be linked to processes they are used in."
They pick AlphaFold to make the point, which takes some nerve, because AlphaFold is the field's flagship. "AlphaFold2 achieved a score of ~90% with regards to folding accuracy." Then you take those predicted structures and use them for virtual screening — the next thing anyone in drug discovery would actually do — and the Perspective reports "mixed results (often comparable to traditional homology models)." Model validation, triumphant. Process validation, unproven.
They give the value of a model as a product rather than a score:
Model value = Improvement in clinical success rate × Project applicability domain
Look at what drops out. A model that changes no decision is worth zero, whatever its benchmark says. The leaderboard number doesn't appear anywhere in the expression.
You can take that reframing and leave the rest of the paper. The question stops being is our model good? and becomes what is it good for — in this program, at this decision?
Where process validation runs out
Here's where I'd push their framework a step further than they take it.
Their test is whether the model improves a consequential decision. That holds up as long as the decision is yours. Which experiment to run, which compound to advance, which program to stop — you set the evidentiary bar, you apply it, you live with the result.
At the regulatory interface it stops holding. Some of those decisions aren't the sponsor's to make. The agency decides whether the evidence supports the claim, and your internal conviction, however well calibrated, is an input to that decision rather than a substitute for it.
So two sentences that feel interchangeable inside a company come apart:
- "It improved our decision."
- "It produced evidence that survives adjudication."
The first is process validation in Bender's sense. The second is a different test, run by a different party, against criteria you don't control. A model can pass one and fail the other, and nothing in an internal review will catch it — because internally, everything worked.
The endpoint is where this bites hardest. The Perspective reaches clinical endpoints and treats them as a labelling problem: "Clinical end point can be chosen arbitrarily (such as survival, tumour size or surrogate end points) and compared in different ways (such as versus standard of care or placebo), hence clinical success depends on end point definition."
True as far as it goes. But in a registration setting the endpoint isn't chosen arbitrarily. It's negotiated, justified and accepted — or it isn't. The word carrying the argument there is adjudicated, not chosen.
What the agency is actually doing about this
FDA got somewhere structurally similar from the opposite direction, and earlier than most of the field noticed.
Its January 2025 draft guidance on using artificial intelligence to support regulatory decision-making for drugs and biologics — and I want to be exact that this is a draft, not operative policy — sets out a risk-based credibility framework. How much evidence a model has to carry depends on two things: the context of use, meaning the specific question the model is being asked in the specific setting, and how much of the answer rests on that model output as against the other evidence bearing on the same question.
Same insight as process validation, in regulatory grammar. Credibility isn't a property of a model. It's a property of a model applied to a question in a context, with a stated consequence if it's wrong.
Two features of the draft matter if you're planning around it.
First, it excludes drug discovery outright — along with operational uses that don't touch patient safety, drug quality or the reliability of a study's results. If your models never touch a regulatory decision, this framework isn't aimed at you. Which is, I think, the strongest available argument for carrying the process-validation discipline downstream anyway: the moment your model informs something the agency will weigh, credibility expectations engage in a way they didn't upstream. Shaywitz's own cited prior work in that Perspective — the 2025 piece with Subha Madhavan on managing complexity in clinical development — points the same way.
Second, it's a draft. Draft guidance describes current thinking. It isn't binding, and planning as though it were is its own error. What it does tell you reliably is how the agency reasons, and that is worth a great deal when you're deciding what evidence to generate two years before you need it.
Four evidentiary burdens hiding inside one model
The practical consequence is easy to state and easy to miss. The same model output can occupy several very different regulatory positions, and the evidence each one demands differs by more than an order of magnitude.
Take one model — say it scores patients on likelihood of response from a multi-omic profile. Four uses.
- An operational tool. It helps find, screen or stratify patients. It makes no claim about efficacy. The burden is real but modest: the tool mustn't distort the population in ways that undermine what the trial can conclude, and its role belongs in the protocol rather than in a footnote. Watch for the selection tool quietly deciding what question the trial is capable of answering.
- An exploratory biomarker. Reported, hypothesis-generating, not driving decisions inside the trial. The burden here is honesty about status — prespecification, and discipline about not letting an exploratory analysis migrate into the efficacy narrative somewhere between the data lock and the slide deck.
- A measure used to make decisions within the program. Dose escalation, expansion go/no-go, stopping a cohort. Now the model steers the development path. The burden is analytical validity plus enough clinical grounding that a decision taken on it can be defended later. This is exactly where Bender's process validation lives, and where careful teams already do good work.
- A measure intended to support a regulatory decision. A surrogate endpoint, a marker defining the intended population, an output linked to a companion diagnostic. Here the burden escalates sharply, and it escalates in kind rather than in degree — new types of evidence, not more of the same: analytical validation, clinical validation, a defined context of use, and, for anything standing in for clinical benefit, evidence that it is reasonably likely to predict, or has been shown to predict, that benefit. The agency has to accept it. Conviction won't get you there.
These are positions an output occupies, not fixed properties of the output. Which is why one model can slide from one position to the next without anyone deciding it should. FDA's BEST resource — Biomarkers, EndpointS and other Tools — is the shared vocabulary for exactly these distinctions, and it treats the escalation as graded context-of-use risk rather than a single cliff.
Three and four sit next to each other on an org chart. They are nowhere near each other in evidentiary terms. A program can build a genuinely excellent instrument for the third use, work out late that its story needs the fourth, and find that the evidence for the fourth was never generated — because nothing in the internal process ever asked for it.
That gap is avoidable, and far cheaper to close early than late. It isn't a modelling problem at all. It comes down to having decided, at the outset, which of the four positions the output is meant to occupy.
Why this is sharper in oncology
The Perspective is explicit that clinical leverage concentrates at the first phase of efficacy evaluation, "usually phase II," while AI effort concentrates preclinically — an asymmetry they describe as looking for the keys where the light is.
Their staging is deliberately non-oncology. Outside oncology, phase I usually enrols healthy volunteers, so variability is low. Oncology doesn't work that way. Patients are enrolled from the start, efficacy signal often begins to appear in dose expansion, and the first adequately powered look at efficacy can sit inside what is still nominally a phase I. The leverage argument holds. The stage boundary blurs. Which means the decisions that carry regulatory weight arrive earlier than the non-oncology framing suggests.
Two developments push the same way.
Dose selection is now a regulatory question rather than an internal one. FDA's August 2024 final guidance on dosage optimization in oncology — the regulatory expression of the agency's Project Optimus initiative — recommends a dosage justified by comparative data rather than by the maximum tolerated dose. So a model informing dose selection is now informing something the agency will examine directly. Category three drifting toward category four, with nobody deciding to move it.
And intermediate endpoints in oncology are contested territory already. Anything a model produces that functions as a stand-in for clinical benefit lands in an area where the evidentiary expectations are actively debated, well documented and unforgiving.
The problem nobody solves alone
Underneath all of this sits an incentive asymmetry that no single company can fix, and it deserves naming.
A company that invents a molecule captures most of the value it creates. A company that validates a biomarker or a translational model creates value its competitors also get to enjoy. Scannell and Shaywitz have both written about this: the work most needed to improve predictive validity is the work least rewarded by any one sponsor's returns. Hence their argument for disease foundations, public funders and precompetitive consortia.
Regulatory qualification is one of the few existing mechanisms built to offset that asymmetry rather than lament it. Once a tool is qualified for a defined context of use, other sponsors can rely on it in that context — which reduces, though doesn't eliminate, the case each still has to build for fit to their own program. A private cost becomes something closer to shared infrastructure, and a formal evidentiary standard attaches to a proxy instead of every sponsor inventing one.
It's slow. It's demanding. It's the wrong route for most programs most of the time. But it means the answer to "who pays to validate the proxy?" isn't only philanthropy. Precompetitive collaboration and a qualification pathway may be the same answer approached from two directions, one economic and one regulatory.
What I would not conclude from any of this
The authors of the Perspective are careful, and that care is the part I'd copy. They write that the field is at "an 'absence of evidence' (and not necessarily an 'evidence of absence')" when it comes to AI translating into clinical impact — then immediately add that progress in technical capability "is much more encouraging."
They apply the same discipline to numbers that flatter their own thesis. Reported phase I success rates of up to 90% for AI-associated programs rest on few data points, and most of those programs are grounded in established biology and chemistry, which lowers the chance of running into safety problems the second time around. Rising approval rates have several explanations with nothing to do with AI: a shift toward biologic modalities with higher average success rates, better intellectual property protection, more rare-disease approvals.
None of this says AI doesn't work in drug development. It says the field has been optimising against benchmarks rather than decisions, and measuring itself where the light is good. That's a critique of aim, not of capability — and aim is fixable.
The same discipline should apply to the argument I've made here, so let me say what would refute it. If programs routinely took an instrument built for the third category, used it to support a registrational claim without generating new clinical-validation evidence, and the agency accepted it, then the distinction I've drawn would be wrong. I haven't seen that happen. But that is the observation to weigh it against, and it's a fair test.
One more thing readers should know. By the Perspective's own competing-interests declaration, its lead author is a shareholder, options holder or consultant across a range of companies in the AI drug discovery sector, and several co-authors hold comparable positions. The critique runs against those commercial interests. That makes it more credible, not less.
A note for anyone reading this about their own program
The argument above is general. This part is practical, and it's commentary rather than analysis — take it as a set of questions rather than a prescription.
If you're running or funding an AI-enabled oncology program, four questions will tell you fairly quickly whether you have something worth your attention.
Which decision does this model change? Not what it predicts. Which decision, made by whom, that would go differently without it. If nobody can name the decision, what you have is a capability rather than a tool — which is fine, as long as nobody is telling a board or an investor otherwise.
Which of the four positions does the output occupy, and which will it need to occupy in eighteen months? Two separate questions, and the second is the one that gets skipped. A program can be honest about today's answer and still be building toward a claim it has generated no evidence for.
If the answer is category four, what is the context of use, and what would the agency have to accept? Write it down as a sentence: this output, used for this purpose, in this population, with this consequence if it is wrong. If the sentence is hard to write, that difficulty is the finding.
What evidence has to exist by the time you ask, and when does generating it have to start? Evidentiary work has lead times measured in program-years. The cost of finding this out late isn't that the work is hard. It's that the window where it was cheap has closed.
None of these questions are glamorous and all of them are answerable early. Most of what makes them expensive is asking them at the wrong time — after the model is built, after the protocol is written, after the story has been told to people who will remember it.
That's the work I do with companies: not evaluating the model, but establishing what it is being asked to carry, and whether the evidence to carry it exists or is scheduled to. If that's a live question in your program, it's worth a conversation.
Sources
- Bender A, Thomas MC, Scannell JW, Shaywitz DA, et al. — Artificial intelligence in drug discovery — what it is, where we stand and the path forward. Nature Reviews Drug Discovery, published online 7 August 2026. PMID 42567977.
- FDA — Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products (January 2025, draft guidance — describes current thinking, not binding).
- FDA — Drug Development Tool (DDT) Qualification Programs and Biomarker Qualification Program — Context of Use. Statutory basis: 21st Century Cures Act §3011 (FD&C Act §507).
- FDA — BEST (Biomarkers, EndpointS, and other Tools) Resource (FDA-NIH Biomarker Working Group).
- FDA — Optimizing the Dosage of Human Prescription Drugs and Biological Products for the Treatment of Oncologic Diseases (final, August 2024).
- Madhavan S, Shaywitz DA — AI: an essential tool for managing the burgeoning complexity of clinical development in pharmaceutical R&D. Drug Discovery Today 30, 104271 (2025). Cited as reference 14 of the Perspective above.
Jesús Gómez-Navarro, M.D., is a medical oncologist and drug development executive, and founder of OncAdios LLC. He advises AI-oncology and biotech companies as a fractional CMO, board director, or co-founder.
← Back to Writing · The pre-IND meeting → · The Calibration Sprint →