Methodology case study

What a Governed Multi-Model Workflow Caught in Its Own Oncology Analysis

A bounded methodology case study on role separation, frozen evidence, adversarial cross-checking, deterministic controls, and human decision gates — and the limits of what a single pilot can show.

By Jesús Gómez Navarro, MD — OncAdios LLC · Published 2026-08-02

The problem I care about is not whether an AI system can write fluent oncology analysis. It can. The problem is that fluent analysis can be wrong in ways a reader cannot see. When the state of the evidence, the authority behind a regulatory statement, the identity of who actually did the work, and the corrections made along the way are all invisible, a confident paragraph and a defensible one look identical on the page. For a decision that moves capital or a development program, that gap is the whole risk.

So I ran a bounded pilot to test a different question: if you govern the work — separate the roles, freeze the evidence, let an independent reviewer try to break the claims, add machine checks, and keep a human at every gate — does the system surface and fix its own consequential errors before anyone relies on them? This is what I found, stated as plainly as I can, including the parts that argue against my own conclusion.

The claim I am willing to make, and its hedge. A governed, role-separated multi-model workflow repeatedly detected and corrected consequential evidence, denominator, regulatory-scope, source-identity, and causal-attribution defects that plausible single-pass work could have carried into a decision-facing oncology deliverable. That is a methodology claim about one bounded workflow. It is not proof of consulting efficacy, model superiority, prevented attrition, saved capital, flawless accuracy, or universal portability. The word plausible is doing real work: the pilot did not run an independent single-pass baseline, so “a single pass would have shipped these errors” is a reasoned hypothesis, not a measured result. I will not upgrade could have carried into would have carried, and I will not call anything prevented.


The method

The workflow assigns three portable roles — a Researcher who must refuse or cite, an Editor who preserves truth conditions rather than smoothing them, and a Verifier who independently tries to refute the claims — and layers an OncAdios overlay on top for oncology and regulatory discipline: evidence levels, a Health Authority record, causal-attribution rules, confidentiality tiers, and refuse-or-cite. The governed architecture requires each phase to run in a distinct, fresh session so that no role inherits another's context; the final chain enforced this, even though an early stage of the pilot itself did not — a deviation I disclose below. Evidence is frozen and hash-checked so that later steps cannot silently edit earlier ones. Deterministic checks — machine-readable matrices, identifier-integrity checks, and positive controls — sit underneath the human intelligence. A human holds the release gates throughout.

Frozen, hash-checked evidence

Model family A · creation

Researcher fresh session
refuse or cite
Editor fresh session
preserve truth conditions
Frozen handoff Independent context

Model family B · adversarial challenge

Independent Verifier fresh session · retrieve independently
attempt to refute · rank findings
Deterministic controls beneath both families · hashes · identifier integrity · positive controls
Human decision gate Reliance · release · publication
Figure 1. The governed loop as designed and as enforced in the final chain: distinct fresh sessions, frozen artifacts, an adversarial independent check, and a human gate at every consequential step. Verification here was cross-family — the reviewer ran on a different frontier model family than the creator.

Two independence controls matter most. First, the final chain enforced fresh session separation between the creating roles and the reviewing role — even though an earlier stage of this same pilot did not: in the initial Run B work the Researcher and Editor shared one creator session, a deviation later closed prospectively through fresh-session controls rather than retrospectively erased. Second, verification was cross-family: the independent Verifier ran on a different frontier model family, from a different provider, than the creator it was checking. I discuss below why I am careful not to turn that into a vendor ranking.

The observed trajectory

One evaluation package, which I will call Run B, is the spine of the story. It began with an independent disposition of blocked — five High-severity findings and one Medium process failure. Through a sequence of bounded correction rounds, each followed by independent re-review, the package moved to a Round 3 disposition of clear-for-human-decision with no Critical and no High findings and a single Medium residual. That final Medium was then closed by a narrow, add-only documentary correction and an independent cross-family recheck.

I want to be precise about the shape rather than impressive about a total. The severity counts at each stage are stage-local and must not be added together, because several later findings descend from earlier correction targets; summing them would double-count. What I can say cleanly is the endpoint: across the full Run B correction lineage, all fourteen tracked Run B performance findings were closed (Table 1). At one mid-trajectory snapshot the state was thirteen of fourteen closed; after the final documentary correction, the remaining finding was closed and none remained. The corrected package reached clear-for-human-decision — which means exactly that, and explicitly not cleared for reliance or release.

Table 1. The fourteen tracked Run B performance findings and their final state. Sanitized defect classes; no program, company, or person is named.

Swipe horizontally to view all columns.

#FindingSanitized defect classFinal state
1V-B-01stale program-lifecycle renderingClosed
2V-B-02contingent transaction value rendered as committed valueClosed
3V-B-03randomized and efficacy denominators conflated; safety overgeneralizedClosed
4V-B-04regulatory-absence claim exceeded the reviewed corpusClosed
5V-B-05suggestive calibration concern over-attributedClosed
6V-B-06Researcher and Editor shared one creator sessionClosed prospectively via fresh-session controls
7F-01current-pipeline provenance incomplete despite a pass labelClosed
8F-02declared authority-search method exceeded the recorded executionClosed
9F-03access status incorrectly characterizedClosed
10V-C2-01required authority-search term set incompletely executedClosed
11V-C2-02global and country-level registry status collapsedClosed
12CHK-C2-01Verifier used the wrong source identifierClosed
13CHK-C2-02Verifier misstated an aggregate trial countClosed
14V-C3-01two query cells misclassified applicability rather than no-matchClosed by documentary corrigendum + independent recheck

The last finding is worth dwelling on, because it shows what “closure” means here and what it does not. The workflow had built a structured authority matrix of 24 rows across 9 query surfaces — 216 disposition cells in total. Two of those cells, where one public ingredient identifier was queried against two public regulatory-data endpoints, had been classified as “not applicable” when the correct classification was “executed — no match.” The correction moved the “not applicable” category from 8 to 6 and the “no match” category from 47 to 49; every other category was unchanged, and the categories still totaled 216. Both target queries had in fact returned no match, and positive controls returned identified records, which proved the endpoints were functioning. So the classification changed; the substantive no-match result did not, and no scientific, regulatory, causal, or program-status conclusion moved.

That is why I refuse to describe 214 of 216, or 216 of 216, as “scientific accuracy.” Those ratios describe classification accuracy of a declared matrix under one bounded official-source test at one retrieval time. They are not a measure of scientific truth, regulatory correctness, or real-world completeness. A no-match from a current data endpoint is bounded to that endpoint, that field, and that moment; it does not prove absence or non-existence of anything.

What actually changed

The most useful result for me was not that the roles caught errors. It was what had to change to catch them. More prose in the role cards was judged not to be the fix. The improvements that mattered were structural: a project-specific evaluation contract, structured and machine-checkable evidence, explicit routing and session-separation rules, a fail-closed preflight, and sharper verification criteria. The intelligence was necessary but not sufficient; the scaffolding around it did much of the work.

Six contributors, kept distinct

It is tempting to compress this into “the AI did it” or “the role card did it.” Neither is true. Output quality here depended jointly on six contributors, and collapsing them would misstate the finding.

  • Role design was necessary but insufficient. The cards frequently already contained the correct instruction; the failures were often non-compliance with controls that existed.
  • Model execution was a credible contributor to a subset of errors that resemble semantic compression or plausible completion — a contingent value read as a commitment, a global field absorbing a country-level status, a miscounted aggregate. This happened across both model families. I do not rank vendors and I do not claim either family is superior; both produced and caught errors.
  • Context and source access — portal indexing, current-versus-historical records, retrieval friction — repeatedly amplified execution failures. Some defects were partly artifacts of source friction, not analytic failure.
  • Deterministic controls were judged the highest-value place for the next quality gains, because they remove clerical failure modes so that intelligence is spent on interpretation. That is a design implication this pilot motivated, not a capability it proved in production.
  • Adversarial verification increased defect discovery, including defects that creator passes missed — and, notably, a Verifier that itself erred and was later caught. I frame this as increased discovery, not a guarantee of correctness.
  • Human judgment acted as the positive control. A human supplied the Health Authority requirement, withheld premature acceptance, authorized bounded corrections, and held the release gates. Without it, the system could have stopped at a plausible but materially defective output.

Demonstrated, suggested, not tested

I try to keep these three registers separate. Demonstrated, and well-supported by the frozen record: governed, role-separated, cross-family review found and corrected consequential defects in the pilot's own artifacts, across the error classes in Table 1; the final package reached clear-for-human-decision with all fourteen findings closed; session separation and cross-family verification were enforced in the final chain; and the substantive conclusions survived the documentary correction. Suggested, and stated as observed rather than proven: cross-family, role-separated challenge increased defect discovery in this one uncontrolled sequence, and model behavior contributes to certain error classes. Not tested, and therefore a limit, not a silence: any protected client data or redaction handling; a shorter “brief” operating mode; a real outlet taken end to end; non-English translation authority; autonomous or unattended operation; provider cost or efficiency; and any real drug-development outcome.

The case against my own reading

A responsible version of this story has to include the readings that cut against it. The multiple correction rounds do not, by themselves, prove the architecture's value. The iteration burden could instead reflect design gaps in a few places — chiefly authorization-provenance handling and the scope of the Verifier's own identifier checks. It could reflect execution defects, since the dominant attribution was non-compliance with existing controls, which invites the reading that a single more faithful pass might have avoided several errors. It could reflect source-access limits, where some “defects” were partly friction. It could reflect task difficulty: a long, manually maintained 216-cell matrix is inherently error-prone, and the proposed remedy — machine-checkable matrices and a fail-closed preflight — implies the manual process was a chosen difficulty. And it is not a controlled comparison: the same governed system that produced the errors also graded them, and independent cross-family verification mitigates but does not eliminate that.

Those rivals are exactly why my thesis is scoped to defects that plausible single-pass work could have carried. The evidence supports that governed review found and fixed real defects in these artifacts. It does not establish that an equally competent single pass would have shipped them, nor that the architecture is the sole or decisive cause of the quality gain. When the evidence cannot discriminate among these explanations, the honest label is inadequately attributable.

One discipline from the pilot illustrates the standard, and doubles as a regulatory-fidelity example. Under FDA's good guidance practices at 21 CFR 10.115(d), guidance documents do not establish legally enforceable rights or responsibilities and do not legally bind the public or FDA; they represent the agency's current thinking, and an alternative approach may be used if it satisfies the applicable statutes and regulations. Separately, paragraph (i)(2) directs that FDA not use mandatory language in guidance unless it is describing a statutory or regulatory requirement [1]. The workflow is required to preserve that exact verb structure — recommends is not requires — and to keep institutional roles intact across jurisdictions rather than manufacturing a single global consensus. Getting that right is not cosmetic; it is the difference between a defensible regulatory statement and an overreaching one.

What this can and cannot establish, and where it points

This is one bounded pilot on public-source oncology material, with an iterative correction path that is not a controlled experiment. It exercised supervised work on public and internal-non-client data in a full operating mode; it did not exercise protected data, a shorter mode, a live outlet end to end, non-English translation authority, autonomous operation, or a deployed runtime. Named frontier-model routes are ephemeral and will change. No efficiency, consulting-efficacy, model-superiority, prevented-attrition, saved-capital, or favorable real-outcome conclusion is available from this evidence.

What it points toward is a design direction I already hold as OncAdios's standard: the scarce input is not the model, it is the senior operator who knows when the model is wrong, and calibration is the accuracy of confidence rather than the volume of opinion. The learning loop should improve future OncAdios deliverables by moving more clerical failure modes into deterministic checks and keeping human judgment on interpretation and gates. I state that as a forward design implication, not a demonstrated outcome.

Methods

The governed architecture requires the Researcher, Editor, and Verifier roles to run in distinct fresh sessions under portable role cards and an OncAdios oncology overlay, with a project evaluation contract and an outlet contract governing scope and format. An initial stage of the pilot deviated — in the early Run B work the Researcher and Editor shared one creator session — and that shared-session defect was closed prospectively through fresh-session controls rather than retrospectively erased; the final governed chain enforced distinct fresh creator sessions and independent cross-family verification. Evidence was frozen and hash-verified between phases. The final verification chain used cross-family review — an independent Verifier on a different frontier model family and provider than the creator — with findings carrying severity and location, bounded re-review after correction, and a human gate at each consequential step. The final documentary correction was add-only and independently rechecked. Provider token and cost telemetry were not exposed to the workflow in an authoritative form and are not estimated. Case-selection and scientific evidence were public sources; governance evidence was maintained CK-safe.

AI and human-control disclosure

Frontier language models performed role-separated research, editing, and independent verification under human supervision, with frozen artifacts, deterministic checks, and explicit stop gates. Two independent frontier model families, from two different providers, were used so that verification could be cross-family; the exact provider and model-version identities are recorded in the frozen internal governance record and are held out of this program-agnostic draft pending a separate disclosure decision. Jesús Gómez Navarro retained authority over the question, the evidence standards, the risk acceptance, the conclusions, the attribution, and the release. The AI systems are tools in this workflow, not authors.

Disclosures

No client engagement, confidential client material, or unpublished client output originated the evidence used in this case study. No outside funding supported the work. Jesús Gómez Navarro is the founder and single operator of OncAdios LLC; that is a disclosure of interest and operating context, not evidence that the methodology is effective and not a claim of independent institutional validation.

Limitations

The single most important limitation is the absence of an independent single-pass baseline, which is why no prevention or counterfactual claim is made. The correction path is not a controlled comparison; the grading system overlaps the producing system; public sources are incomplete; a current-endpoint no-match is bounded to its endpoint and time. Protected data, brief mode, a live end-to-end outlet, translation authority, autonomous operation, and a deployed runtime were not exercised. Provider cost and efficiency were not measured. Nothing here should be read as a claim of accuracy guarantees, model superiority, or portability beyond this workflow.

A note on what happens next

If you are weighing a development decision where the confidence of the analysis matters as much as its content, I am glad to talk about how evidence state, authority, execution identity, and correction history can be made inspectable in that specific context. I am not offering an outcome guarantee, and this case study is not one.

References

  1. Code of Federal Regulations, Title 21 §10.115 (Good guidance practices), paragraphs (d) and (i)(2). Paragraph (d): “Guidance documents do not establish legally enforceable rights or responsibilities. They do not legally bind the public or FDA,” and guidance documents “represent the agency's current thinking”; an alternative approach may be used if it satisfies the applicable statutes and regulations. Paragraph (i)(2): FDA does not use mandatory language in guidance documents unless it is describing a statutory or regulatory requirement. Retrieved via eCFR (official) 2026-07-31.

← Back to Writing