Publication
Essay · Version 4.0 · Designed edition

"The AI Said So" Is Not a Defense

The record, not the draft, is the product. AI-assisted work will be bought, trusted, defended, and litigated on the account of how it was made, and most organizations are optimizing the wrong artifact. Here is the full argument (the economics, the trust repricing, the law, the evidence, the architecture) and the ten-minute Record Test that tells you where you stand.

Along the way: what an ontology actually is, and the four different things vendors mean when they say the word; the genome principle, and why a system should be modeled on the work rather than on the data; context of use, the declaration regulators now expect; and why a system that is right once is not yet reliable enough to sign.

By Hicham Naim & Blaise Jacholkowski · 3 September 2026

Article cover: The AI Said So Is Not a Defense

Picture a Tuesday next March. The vice president of market access at a mid-size biotech sits down with the final draft of a payer evidence dossier, the cost-effectiveness narrative that will go to health technology assessment bodies in three countries. Most of the draft was produced with AI assistance. Nobody hid that; the work went noticeably faster than it did a year ago. Now comes the moment no technology has changed in thirty years: her name goes on it.

She pauses over a comparator effectiveness figure and asks the question every signatory in a regulated industry asks: Where did this come from?

A year ago the answer lived in a chain she could reconstruct, slowly: the analyst who pulled the number, the email in which a colleague challenged it, the tracked change that revised it, the meeting that accepted it. Scattered and ugly, but hers, and reconstructable. This year the number came out of a model. The draft is better than last year's. The answer to where did this come from is worse. She signs anyway, four days late, after her team rebuilds by hand the provenance the tool never kept.

Multiply her four days across every signed deliverable in a launch program, and you have found the place where the economics of AI in life sciences are actually decided. It is not the drafting. It was never the drafting.

She has a second question, and it is the one her toolchain answers least well: who did this work? Not which sources; which workers. Part of the draft came from a colleague, part from an agent inside one vendor’s assistant, part from a different agent inside another, prompted by someone who has since moved teams. Nobody chose this arrangement and nobody is accountable for it. Her organization is already a hybrid of people and machines, at every level, doing work that carries her name, and it has no way to see itself as one. That is not a tooling problem. It is a workforce with no payroll and no record, and it is the condition the rest of this essay is about.

This essay argues that the binding constraint on AI in regulated work is neither model capability nor user adoption but a missing artifact: a record of how the work was made that leaves the person who signs it safer than she would have been working alone. That record is now being specified, with dates, by legislators and regulators on both sides of the Atlantic. A growing body of peer-reviewed research shows what makes such a record trustworthy. And producing it requires a category of infrastructure most vendors do not have and most buyers do not yet know to demand: a machine-readable model of the work itself, rooted in outcomes rather than in data or documents.

We built such a system. We are about to test it in the real world, against work a professional actually signs, because that is the only test that counts. Here is why.


The idea in brief: the problem, the bigger problem, and the shift

CHAPTER 01

The Great Repricing of Trust

Before the law and the evidence, the economics, because the wall organizations are hitting was not built by regulators, and treating it as a compliance story is how executives end up solving the smallest version of the problem.

For a century, professional work carried an implicit warranty. A dossier from a serious firm, a memo from a senior analyst, an assessment with a partner's name on it: the artifact itself was evidence of the process behind it, because no other process could have produced it. Fluency implied labor; labor implied judgment; judgment implied someone you could hold accountable. That chain just snapped. When anything can produce fluent, confident, well-structured prose in seconds, the artifact stops vouching for the process, and everyone whose job is to accept work starts asking, out loud, what they used to assume.

Six external askers converge on the person who signs the work

And there is a seventh asker, the one this essay is really about: you, at the moment your name goes on the work. The vice president's pause over the comparator figure was not about EMA. It was about her signature, the discovery that the modern toolchain had made her faster and her name less safe in the same motion. Hers is the only version of the question with two halves: where did this come from, and who did this work; and the second half is the one no current toolchain can answer at all, because the workers it would have to name were never counted. Trust used to be borrowed from the institution and the letterhead. Increasingly it must be shown, deliverable by deliverable, and the professionals and firms that can show it will be believed at a premium exactly as large as everyone else's discount.

CHAPTER 02

The Bottleneck Nobody Budgeted For

Consider what happens when a medical affairs organization deploys an AI drafting assistant, a scene common enough that most readers in the industry will recognize it. Output volume rises. Then the quality function does the only rational thing available to it: it re-verifies everything, manually, because nothing in the output distinguishes solid retrieval from fluent invention. Syndicated industry surveys (commissioned, not peer-reviewed, and cited here as directional rather than decisive) put the share of medical affairs organizations manually re-checking every AI output at around half; and clinical-document timelines were dominated by review cycles rather than writing time long before AI arrived. The argument does not rest on that figure. It rests on the mechanism, which holds at any share above nil, and an essay that asks you to demand sourcing should say plainly which of its own numbers are soft. The drafting accelerated; the bottleneck never moved. The saving cancels arithmetically, and the pilot quietly ends.

Three schematic tracks: drafting effort falls but review grows to absorb it, unless the run leaves a record

The instinctive diagnosis (the model isn't good enough yet) is wrong, and acting on it wastes budget cycles. Even a substantially better model changes nothing for the reviewer, because the reviewer's problem is not output quality. It is epistemic: she cannot see what the system consulted, what grounds each assertion, or what a previous human already challenged and resolved. Given only conclusions, a competent reviewer in a regulated function must re-derive everything. Given an evidence trail, she can spot-check. The difference between re-deriving and spot-checking is the entire economic case for AI in regulated work, and it is a property of the record, not of the model.

The incumbents' own products now concede this diagnosis. The review agents shipping across the life-sciences software landscape implement a telling pattern: AI finds problems; a human declares there are none. That is review-as-a-workflow: the right instinct, aimed at the right person. What it leaves unbuilt is the artifact the signatory keeps: a structured, portable record of what was checked, what was found, who decided, and what a court, an auditor, or next year's team can inspect.

CHAPTER 03

The Law Is Writing the Spec

On 2 April 2026 the FDA's drug center issued a warning letter to a manufacturer that had used AI agents to write its drug product specifications, its procedures, and its master production and control records. The firm's explanation, as the agency records it, is the sentence this essay is named after: it had not known of the legal requirement, because "the AI agent you used … never told you it was required."

The agency's answer runs to four lines and is worth reading as a specification rather than a rebuke. If AI helps create the documents, "you must review the AI generated documents to ensure they were accurate and actually compliant with CGMP." And, prospectively: "any output or recommendations from an AI agent must be reviewed and cleared by an authorized human representative of your firm's QU."

Read that twice. A drug regulator has now held, in an enforcement document against a named firm, that an AI system's assertion is not a defense, and that what closes the gap is a named human in a specified function clearing the output. Not a better model. Not a policy document. A person, a role, and a record of the clearing. Everything else in this section was, until that letter, a forecast about where the law was heading. It is no longer only a forecast.

For a decade, "responsible AI" was a genre of principles documents. In the last twenty-four months it became text with dates. The regulatory face of the wall is not its most important; a client's procurement letter moves markets faster than any directive, because liability, not AI regulation, is the binding pressure: procurement, insurance, and outside counsel respond to expensive faster than they respond to obligatory. But the regulatory face is the most dated, and dates turn a trust problem into a planning problem.

Seven dated regulatory texts on a single rule, each annotated with what it presumes already exists, with the April 2026 FDA warning letter set below as the requirement applied to a named firm

Read the column of presumptions, not the column of demands; that's where the story is. EMA and FDA assume you can account for what the system did, with the documentation made before the output shipped. Article 86 assumes an explanation exists to be given, and an explanation duty is, operationally, a record duty: you cannot explain what you did not capture. Its sibling Article 12 drops the pretense and says it outright; the article mandating automatic logs for high-risk systems is titled, simply, "Record-keeping." Be precise about the calendar there, because it cuts against the easy version of this argument: a 2026 simplification regulation moved the high-risk obligations, Article 12 among them, out to December 2027, leaving Article 86's right to an explanation applying while the provider duties behind it wait. Anyone selling urgency off the AI Act is selling a date that has already moved once. NICE assumes you know, deliverable by deliverable, where AI touched the work; Canada's drug agency has taken the same position, and the professional societies are ahead of both; ISPOR's good-practice reports have asked for method transparency in machine-assisted evidence work since before it was fashionable. And the revised Product Liability Directive carries the asymmetry worth sitting with: a record that doesn't exist cannot be disclosed, but its absence can be held against you.

Three observations matter for executives. First, the liability point above: watch what procurement and insurers do, not what compliance memos say. Second, the strictest text is the most portable. Annex 22 is scoped to manufacturing, and non-deterministic generative systems largely sit outside its intended use. But its logic is not a manufacturing rule. It is a substitution rule: when a machine takes over part of a process a human performed, measure it against the human baseline and keep evidence per run. That logic will not stay inside the GMP fence, because nothing about it is specific to manufacturing. We should be explicit about our own position here, since we are borrowing the argument: what we build is for non-GxP work and makes no GxP claim. We invoke Annex 22 for its reasoning, not its jurisdiction, and we hold ourselves to that reasoning anyway. Third, none of this is new law, and the oldest text is the sharpest. Part 11 of Title 21, in force since 1997, already requires that an electronic signature record carry "the meaning (such as review, approval, responsibility, or authorship) associated with the signature," and already requires "secure, computer-generated, time-stamped audit trails." The MHRA's data-integrity guidance states the attribution principle in five words: data must be "attributable to the person generating the data." That clause is the whole problem with machine-generated work, written decades before the machines arrived. Software has the same idea running in production: in-toto, a graduated CNCF project, exists to make "transparent to the user what steps were performed, by whom and in what order," and the W3C's provenance model has defined provenance as information about "entities, activities, and people involved in producing a piece of data" since 2013. The account is not a novel artifact we are proposing. It is an old artifact that never had a place to put a machine. Fourth, the regulators did not copy one another. FDA, EMA, and the EU drafters arrived at record-shaped requirements independently, because all of them are downstream of the same industrial fact: in regulated work, accountability attaches to named persons, and named persons need records. When independent institutions converge, executives should treat the convergence as a forecast.

CHAPTER 04

Beyond the Law: the Rooms Where the Record Pays

It would be a mistake, the most natural mistake available, to file everything above under compliance. The law makes the record mandatory; the market is quietly making it profitable, and the profitable cases will move budgets faster than the mandatory ones. Walk through the rooms where the question "how was this made?" is already being asked with no regulator present. (A fuller register, twenty-six cases across six families, is in the Appendix; these are the ones that change P&Ls.)

The client's room: trust, priced. Supplier-qualification language asking for production documentation on request is beginning to circulate; we have met it in draft supplier standards rather than in any published survey, so read it as a leading indicator and not a measured trend. The firms that can answer them stop competing on rate cards: a deliverable that arrives with its record is a different product from the same deliverable without one, and buyers know it before their lawyers do. In requests for proposals, "what record does your work leave?" is becoming a selection question, and for the supplier who can answer, disqualifying to everyone else in the room.

The board's room: the basis, on file. A board that relies on AI-assisted analysis for a strategic decision owns a hindsight problem: when the decision is challenged (by shareholders, by a buyer, by events), "on what basis did the board decide?" must have a better answer than a deck and a memory. The same record that satisfies a quality function is what lets a director say: here is what we knew, here is who reviewed it, here is who disagreed, and here is why we proceeded. Duty of care is about to acquire a machine-readable form.

The investor's room: the evidence estate. In diligence (venture, M&A, licensing) the claims are only as valuable as the account of how they were produced. A company whose evidence carries its trail gets read at face value; one whose evidence is "probably right" gets discounted, and the discount lands on the valuation. For firms whose product is their method, the stakes double: an acquirer testing a method estate is testing its provenance and reproducibility, whether the expertise survives the departure of the people who currently carry it. A recorded method is an asset; an unrecorded one is a payroll dependency.

The professional's own room: the name. Somewhere below every organizational case sits the person on whom it all lands: the employee asked, after something goes wrong, why did you trust it? An unrecorded toolchain leaves her with nothing but her memory and her sincerity. A record leaves her with what she checked, what she was shown, and what she decided. The professionals who understand this first will choose their tools accordingly; and, over a career, the record travels with them: their precedents, their method, their account of how they work, portable between employers in a way institutional trust never was.

The operations room: the margin on memory. The least dramatic case may be the most valuable. Work that leaves a record is work that can be re-run: the annual update becomes deltas instead of a rebuild; the next program traverses what the last one learned instead of paying to rediscover it; the organization's knowledge stops walking out the door at every handover. Defense is the floor. The end of paying twice (for verification, for rediscovery, for rebuilds) is the business case. And there is a prior version of the same case that most organizations have not costed at all: they cannot say how much of last quarter's knowledge work was done by people and how much by machines, in which tools, under whose direction. A workforce you cannot count is a workforce you cannot plan, price, or defend in front of anyone.

One question, many rooms, and only one of them has a judge in it. Compliance is the smallest version of this story. Trust is the currency, and the record is how it's minted, run by run.
CHAPTER 05

What the Evidence Shows, and Where It Stops

The claim that structure improves AI output has acquired numbers; and so has the cost of its absence.

Start with the absence. A study in JMIR Mental Health, published in late 2025 and testing GPT-4o (a current, capable, everywhere-deployed model) asked it for research citations across mental-health topics and then checked every one: 19.9 percent were outright fabricated, and 45.4 percent of the genuine ones contained bibliographic errors. Roughly two out of three references could not be trusted as given. The contamination has reached the literature itself. A 2026 audit published in The Lancet checked biomedical references for citations that resolve to nothing. Note what it measured: fabricated references, not AI authorship; the cause is inferred from the timing. Reporting on the audit describes a steep multi-year rise tracking LLM adoption, with the correction mechanism visibly not keeping pace. The error is not merely entering the record. It is staying in it.

The legal profession is running the same experiment in public and with names attached. A tracked database of decisions in which a court found reliance on hallucinated content listed 2,039 cases when we checked it, a figure that moves weekly, which is itself the point. Two of them are worth reading rather than counting. England's Divisional Court, in Ayinde and Al-Haroun, held that the duty to verify "rests on lawyers who use artificial intelligence … or rely on the work of others who have done so"; delegation does not move the duty, in either direction. And a California appellate court, in Noland, put the same point in four words: "Attorneys cannot delegate that role to AI." Professionals with licenses are discovering, one judge at a time, what an unrecorded toolchain does to a career.

Three measurements of one failure: a current model, and the courts

Now the structure, and here we are going to be less dramatic than this genre usually is. Microsoft's OG-RAG work, published at EMNLP 2025, structured retrieval around a formal domain model rather than flat document chunks and reported that it "increases the recall of accurate facts by 55%", with response correctness up 40% across four different models; and, notably for the economics of Section II, attribution of answers to their context roughly 30% faster. In clinical deployment the effects are real and smaller: a controlled radiology study in npj Digital Medicine found that grounding a local model in the relevant guidance "eliminated hallucinations (0% vs 8%)". Different systems, different scopes, one mechanism: a model answering over a structure fails at a lower rate than a model answering from its own distribution, because the structure constrains what can be asserted and makes every assertion attributable. What the literature does not support is an order-of-magnitude claim, and we are not going to make one; an essay about fabricated citations does not get to round its own numbers up.

A second line of research shows why grounding alone still does not clear the signature bar. Benchmarks that measure whether an agent succeeds repeatedly rather than once (the pass^k discipline introduced with τ-bench, a Sierra AI preprint rather than a peer-reviewed paper) define success as "the chance that all k i.i.d. task trials are successful." On that definition a leading model scoring 61% on a single attempt in the retail domain falls below 25% across eight. And the degradation is not confined to conversational agents: work presented at ICLR 2026 found that the best model tested "sees its accuracy fall below 50% within 15 turns" despite near-perfect accuracy on the first step, and that models grow measurably worse after being shown their own earlier mistakes. The bottleneck is execution over length, not reasoning, which is precisely the argument for a record kept per unit of work rather than a check performed on the finished output.

Then there is the control most organizations are actually relying on, which is a human reading the output. The evidence on that is the most uncomfortable in this section. A study in Radiology put 27 radiologists in front of mammograms with AI-suggested assessments. When the suggestion was correct, they rated correctly 80% of the time. When the suggestion was wrong, correct ratings fell to 19.8% among inexperienced readers, 24.8% among the moderately experienced, and 45.5% among the very experienced. Seniority helped, and still left the most experienced readers wrong more often than right. This is the automation-bias literature, and it is forty years old: the canonical review records omission error rates of 41% and commission rates near 65% when automation recommended wrongly, and an experiment in which 75% of pilots followed a wrong engine-shutdown recommendation. The systematic review of the clinical literature identifies accountability among the few reliable mitigators; but the finding is sharper than that: externally imposed accountability helped in some of the review's studies and not others, while what worked consistently was a person's own felt sense of being answerable. Which is the whole thesis of this essay, arrived at from the opposite direction: a signature is not a control. A signature with a record behind it, and a name attached to the consequence, is closer to one.

Reliability, in the sense a signatory needs, is a distributional property. No single output reveals it, however polished, and no single reviewer reliably catches it. Only the record of many runs can, which is one more reason the record cannot be an afterthought bolted onto a pilot.

The honest synthesis: grounding is worth an order of magnitude on truthfulness; repetition is where even grounded systems fail; and neither benchmark tells you whether this deliverable, produced this Tuesday for this dossier, deserves this signatory's name. That final question requires machinery the benchmarks do not measure.

CHAPTER 06

Why "We Have an Ontology" Isn't an Answer Yet

Ask a vendor how their system will meet this bar and the answer increasingly contains the word ontology, a formal, machine-readable model of a domain. The word is doing heavy lifting, and buyers should learn to sort its meanings, because the prestigious ones do not solve the problem above.

The most mature lineage is the knowledge ontology: the Basic Formal Ontology and OBO Foundry ecosystem in biomedicine, SNOMED and its relatives in clinical terminology, decades of principled curation answering the question what exists, and what is it called? A second lineage models the enterprise (TOVE and DEMO in the academic line, APQC's process classification and BIZBOK's business architecture in the practitioner line), answering what does this organization do? A third is an engineering discipline: OWL for expressing ontologies, SHACL for validating data against them, answering is this data well-formed? The most commercially significant recent lineage is the decision-centric operational ontology of the Palantir school, which binds an organization's data, logic, and actions into objects software can operate on. It deserves to be taken seriously rather than caricatured: it genuinely operationalizes the enterprise. But its root is the data estate: it organizes what the organization's systems already record. Nothing in it asks what must become true.

Knowledge ontologies describe what exists; operational ontologies execute what is decided; a model rooted in outcomes derives execution from what the work must achieve, and emits the record.

Call it what it functionally is, living institutional knowledge: a structure rooted at the top in what the work must accomplish (the access decision defended, the dossier accepted, the evidence gap closed) and derived downward: what must be true, which work objects produce those truths, which capabilities and knowledge the work requires, which gates govern it, and which humans hold authority at each gate. Outcome-rooted at the top, execution-grade at the leaf, record-emitting throughout.

Four ontology lineages and the question each one answers, with a fifth, rooted in outcomes, set below the rule

The difference is practical, not philosophical. A knowledge ontology grounds a system's statements, a measurable service, as Section V showed. But the signatory's question is not only is this statement grounded? It is was this work done right: by what method, reviewed by whom, with what dissent, against what acceptance criteria? Only a structure that models the work itself can answer that, because only such a structure has entities for the things the answer is made of: jobs, units of work, review gates, deciders, and the records themselves as first-class citizens rather than log exhaust.

CHAPTER 07

The Genome, Not the Map

Our design principle for such a structure, which we call the genome principle, borrows deliberately from biology. A genome is not a blueprint in the architectural sense, a static drawing that specifies where every component goes. It is a dynamic instruction set: it encodes what can be produced, under what conditions, and how the products of one gene regulate the expression of others. The genome does not describe the organism. It constitutes the organism's capacity to develop, respond, and adapt. A structure built on this principle is not documentation layered on top of the organization; it is the substrate the organization's intelligence runs on: traversed by agents, enriched by every execution cycle, governed at every promotion.

Concretely, that means layers derived top-down and fed back bottom-up. Strategic direction resolves into desired outcomes, defined independently of any current process, because an outcome defined by reference to the existing process has lost its power to challenge it. Outcomes pass through an intelligence core, the translation layer that reads which conditions apply, which evidence standards, which payer archetypes, which constraints. This layer is our candidate answer to a question executives have asked for decades: why does strategy fail to become execution? Not because the strategy is wrong or the execution incompetent, but because the translation between them has never been made explicit; it lives in the heads of experienced practitioners, and when they leave, it leaves with them. The translation resolved, an execution engine composes the work from independently governed ingredients and runs it. And an intelligence layer at the bottom measures what happened, retains what matters, and feeds learning back up: patterns continuously, outcome recalibrations tactically, and, only past a deliberately high evidence threshold, challenges to the strategy itself.

Two properties of this architecture matter most for the argument of this essay. First, it is what makes the organization's improvement governed: a successful deliberation is promoted into a reusable pattern only by a named steward's documented decision, every pattern firing generates an audit record, and every pattern carries a validity indicator with a review cadence, because automatic learning in a complex organization is indistinguishable from automatic error propagation. Compounding is earned, not assumed. Second, and this is where the genome metaphor pays off, a genome is not valuable because it exists. It is valuable because of what it expresses. What this genome expresses, on every governed run, is the record.

One more thing follows from building it this way, and it is the premise underneath everything above: AI of this kind is not a tool. It is non-human capital: a workforce an organization hires, organizes, governs, develops, and pairs with its people. Tools get deployed; capital gets managed. That reframe is why the architecture resolves into two operating systems rather than one product. An Enterprise OS keeps institutional knowledge alive, with every piece of work chained to the outcome it serves and the strategy above it, learning from every run, without waiting for anyone's model to retrain. A Workforce OS makes blended human–AI teams real: escalation, delegation, and a named human at every gate, with the mix you intended and the mix that actually ran both inspectable. Either alone is a fragment: institutional knowledge without a workforce is a library that does not execute; a workforce without institutional knowledge is agents without context, which never compound. The record is what the pair expresses, run by run.

Two operating systems side by side, what each keeps and what each alone fails to be, resting on the record

Two consequences of taking the workforce framing literally. The first is that agents in such a system are custodial: they hold work on behalf of a principal, which means every action an agent takes carries a delegated authority (from whom, scoped to what, revocable when). An agent is not a feature that ran; it is a worker that acted under a mandate, and the mandate is part of the record. The second is what this buys an organization. Capacity can flex because part of the workforce is non-human; that is the elastic organization, and it is the reason the record is not a compliance artifact but an operating requirement. You cannot elastically scale a workforce you cannot account for. Governance is not the tax on elasticity; it is the mechanism that makes elasticity survivable.

CHAPTER 08

Anatomy of a Signable Record

Work through the signatory's position seriously and the design falls out almost mechanically. Every governed run must leave three artifacts.

The draft: the work product itself, at whatever grain it was commissioned.

The evidence trail: what was consulted, what was asserted on what grounds, which sources support which claims. This is the artifact that converts review from re-derivation to spot-checking, which is where the canceled saving comes back.

The review record, the artifact nothing in the current tool landscape structurally produces: who decided, with what authority, on what evidence, with what dissent. The record, in genome terms, is the phenotype, what the structure expresses. That is also why it cannot be retrofitted onto a system that never modeled the work: every field of a real record is a traversal endpoint in the structure, which outcome the run served, under which constraints it was permitted, which capability and actor produced it, what the gate decided and who dissented. Bolt a "record" onto an architecture without those entities and the fields have nothing to bind to; you get a log, not a statement.

Three artifacts of a signable deliverable in sequence: the draft, the evidence trail, and the review record

And here the usual vocabulary is too small for the artifact. A review record names one chapter, the human decision at the gate, and then gets used for the whole thing. Underneath the review sits the harder object: the chain of every actor, every action, every input and every version that produced the work, in order, with the authority each one acted under. The review is what a signatory shows a reviewer. The chain is what answers a court, an auditor, or next year's team. Together they are what this essay has been calling the account of how the work was made; and the account, not the review alone, is the artifact.

In our architecture these are entity law rather than logging convention. The ontology specifies a run record carrying per-step confidence and an explainability reference: the Annex 22 expectations, adopted for knowledge work. Once a research or reasoning turn settles, it is scored by the protocol we call VERIFY, on six dimensions: validate, evidence, review, identify, fact-check and yield. The reviewer sees those scores with claim-level evidence, which is the whole point; a reviewer should not have to work out for herself which claims are thin. Two honesties about it: VERIFY is an overlay on a settled answer rather than a gate on every path, and the quick path does not run it at all. A gate produces a gate record binding a named, qualified human to the decision, with dissent as a structured, first-class field, and the gate, not the score, is what the signature rests on.

Which raises the question the next few years will actually turn on: who appears in that chain. Today the honest answer in most organizations is “some humans, and some tools nobody logged.” In a governed hybrid workforce the actors are named on the same footing, and naming a machine takes more fields than naming a person. A human reviewer is a name and a qualification. An agent is a name, a version, the configuration it ran under and the knowledge it held at that moment, because “reviewed by the evidence agent” means nothing unless that exact agent can be re-run, and “v1.8” means nothing unless v1.8 is pinned. Each agent action also carries its delegated authority, which is what a quality function and a workers' council will both reach for first, and for opposite reasons.

Two commitments hold that structure honest, and they are commitments rather than features. The first is a non-delegable set: decisions no agent may close, declared in advance rather than discovered after an incident. Publishing that list is not a limitation to be buried; it is the only credible answer to “does the AI decide?” The second is a distinction we would keep even if the technology stopped requiring it: agents review, humans sign. Agents can draft, check, challenge and recommend, and a system where they do none of that is wasting them. But the signature is a legal act performed by a person who can be asked why, and it does not delegate, which is no longer only our position. It is what a California appellate court said to a profession and what the FDA said to a manufacturer in the same twelve months: the output may be the machine's, the clearing must be a named human's. Blur that line and two things break at once: the assurance a quality function needs, and the protection a workforce is owed.

A vertical ledger of one governed run: two human and two agent actors, their actions in the order they happened, an escalation outside an agent mandate, and a gate that only a human signs

Each governed workflow carries a context-of-use declaration (what question this configuration answers, what decisions it may influence), which is the FDA's framework adopted as internal law rather than awaited as mandate. Gates are engineered against rubber-stamping: the gate presents the trail, not merely the conclusion, because a reviewer shown only conclusions will eventually approve anything. And because gates are discrete moments while failure is continuous, runtime circuit breakers monitor behavior between gates: the pass^k problem, addressed where it bites.

Note what this architecture refuses to claim. It does not certify outputs; externally the honest verb is evaluate, because certification language writes checks only regulators and courts can cash. Nor is the account an audit trail, and the difference is not cosmetic. An audit trail is consulted when something has gone wrong; the account is what makes ordinary work defensible on the day it ships, and what lets the next run start from the last one. Forensic completeness is the standard the chain has to meet when someone finally tests it; it is not the reason to build one. It does not remove the human; it makes the human decision the recorded, load-bearing center of the process. And it does not pretend that AI reviewing AI closes the loop. The reviewer is the customer of the record; the question is whether what she keeps is a dashboard screenshot or a structured object she owns.

CHAPTER 09

The Costs Vendors Won't Mention

An argument for record-emitting AI owes executives the cost side, because the research is equally clear about it.

Curation is expensive, permanently. Production knowledge-graph and structured-model maintenance runs at a large multiple of the initial build. A model that tries to describe everything decays into an expensive fiction. The design consequence is scope discipline: model the work (outcomes, work objects, gates, records) at full rigor; bind domain knowledge at the point of use rather than duplicating decades of biomedical curation; and let the record itself, not a modeling team's ambition, drive what gets added.

Structure compounds; content must not. The tempting model (aggregate customer content, learn from it, compound) is forbidden twice in this industry: by contract, as data providers increasingly prohibit retained derived models, and by trust, since no quality function should accept its dossiers training a vendor. What can compound honestly is structure: the shape of the work, gate configurations, failure modes, the growing library of codified methods. Content, including the records, stays with the customer, under the customer's tenancy. The record belongs to the signatory. That is not a concession; it is the point.

And the central claim still has to be measured, ours included. Here is the sentence this genre of essay usually omits. Nobody will let an agent produce work a named human must sign until the record of how it was produced is better than the record that human would have left alone, and that is a comparison this category has yet to run in the open. We have specified that record and built parts of it; we have not measured it against a real signature. The measurement has a protocol, and a name, PROOF: a protocol written down before anything is run, a bank of instances held back rather than shown in advance, and a dated record attached to each subject. It scores five axes, which are the five questions a buyer actually asks: performance, reasoning, operations, oversight and frugality, or whether the work is better, smarter, faster, safer and cheaper. The commitment is to publish the protocol and the score, never an approval, because approval was never ours to give. No public result against a real signature has been published. A company whose product is the honest record does not get to exempt its own claims from it. The two protocols are one posture at two altitudes: VERIFY scores a settled turn before a reviewer signs it; PROOF measures the platform in the open. Neither asks anyone to take our word for anything.

The PROOF protocol: a published protocol, a held-out instance bank and a dated per-subject record, scored on five axes
CHAPTER 10

The Record Test

Which brings the argument to its instrument. Take one AI-assisted deliverable your organization produced this quarter, one that someone signed, submitted, or sent to a client, and ask six questions of the system that produced it. Not of your people; people heroically fill gaps by hand, which is precisely how the gaps stay invisible. Ask what the system itself can show.

The Record Test: six questions put to the system that produced a signed deliverable

Scoring is binary and unforgiving: any "no" means that deliverable is a draft, not a defense. Six yeses mean you already hold the record the next five years of regulation describe, and you should be telling that story to your clients loudly.

And when a vendor is in the room, add two: What is the declared context of use: what question does this configuration answer, and what happens when a task falls outside it? And: show me your measurement against the process you replaced, or the pre-registered protocol, with a date, on which you will produce one. The first is FDA's framework turned into a purchasing question; the second is Annex 22's substitution rule turned into a negotiating position. A vendor who can't answer either is selling you drafts.

The test, in the wild. Three failures from rooms you will recognize; composites, all of them, and all of them happened in some form to someone this year. On what evidence? A digital-health CEO is eight minutes into diligence when an associate asks for the source behind the adherence claim on slide 14; the analyst who ran the AI-assisted scan has left; the claim is probably right, and "probably right" costs more in a diligence room than "wrong but sourced," because wrong-but-sourced gets corrected while probably-right gets discounted, and the discount lands on the valuation. Who disagreed? A biostatistician flags a limitation in the indirect comparison, in an email, politely, once; a year later an HTA assessor raises precisely that limitation, and the firm's record shows something worse than the flaw: unanimity. The one person who was right is invisible, and she has learned that dissent evaporates, so next time she won't send the email. Could you run it again? A client returns with the happiest possible request (same landscape as last year, updated for the new entrants), and nobody can reconstruct how last year's was made; "update" silently becomes "rebuild," at rebuild cost. A reproducible deliverable would have made this week's work the deltas, and the margin on deltas is the best margin in professional services.

Run the test on your current stack. The same six questions apply to us; it is vendor-agnostic by construction, and a vendor who flinches at it is telling you something worth knowing.

CHAPTER 11

What Leaders Should Do Now

If you lead a function in pharma or biotech (medical affairs, HEOR, market access, regulatory): reframe your AI evaluations around the reviewer, not the drafter. Before the next pilot, quantify where your document time actually goes; if review dominates, an assistant that accelerates drafting is optimizing your smallest term. Put the test, and the two vendor questions, into procurement language now, before the December liability regime makes it your counsel's language. And treat your own review practice as the baseline it legally is: under the substitution logic regulators are converging on, your current record is the thing any AI system must beat. Knowing what your best reviewers actually leave behind is suddenly a strategic asset.

If you run a consulting or advisory firm, the boutique and mid-tier firms that sell expertise by the hour: the same shift is an opening rather than a threat. Your method (how your firm actually structures a landscape assessment, a value dossier, an evidence-gap analysis) is codifiable, and codified methods running under governance produce exactly the records the demand side will be required to prefer. The firms that win will not be the ones with a chatbot; they will be the ones whose method runs as a service, at whatever capacity demand requires, leaving a record their client's signatory can defend. Headcount stops being the ceiling on the franchise.

If you build AI products for regulated industries: the roadmap implication is uncomfortable and clear. Evidence trails, named-decider gates, structured dissent, context-of-use declarations, and per-run records are not enterprise features to add after product-market fit. They are the product. Retrofitting a record onto an architecture that never modeled the work is the expensive path; the cheap path is to make the record the thing the system emits by construction.

Where the record changes what you do next, from three chairs

And whoever you are: assign the record an owner. In most organizations nobody owns the account of how AI-assisted work is made; quality owns process, IT owns tools, the business owns output, and the record falls into the gap between them. It is about to become the most-requested artifact in your building. Someone should be responsible for the fact that it exists. And give that person a second brief while you are at it: count the workforce you already have. Not the headcount; the whole of it, people and machines, across every assistant and agent already in use at every level of the organization. Most leadership teams cannot produce that number today. It is the first thing an operating system for a hybrid organization makes visible, and it is difficult to govern, price or defend anything downstream of it until someone can.


The industry has spent three years asking what AI can write. That was the wrong artifact all along. In regulated work the draft was never the scarce thing; defensibility was: the signature, and everything a signature has to rest on. The law now says the record is required. The research says structure is what makes it trustworthy. The economics say the reviewer is where the value is won or lost. The organizations that internalize this will not merely deploy AI more safely than their competitors; they will leave, run by run, a better account of their own expertise than any of us has ever kept, and expertise that leaves a record is expertise that compounds.

And for the seventh asker, the one who looks back from the mirror before the signature goes on, the choice is now concrete: a decade of getting faster while your name gets more exposed, or a decade of getting faster while it gets safer. The next three years will be decided by what AI can prove. The next ten minutes can tell you where you stand: six questions, one deliverable you've already shipped.


Hicham Naim and Blaise Jacholkowski are co-founders of VITAL.expert, a governed AI-workforce platform for non-GxP life-sciences work, operated by Crossroads Catalyst GmbH. The platform is an LLM-agnostic agentic Operating System with a living Ontology, designed for professionals working in the pharma, medtech and healthtech industries.

Appendix: The Defensibility Case Register

The six faces of Exhibit 01, unpacked: twenty-six documented moments in which someone with standing asks "how was this made?", and the answer is either a record or a shrug. Maintained as a living register; the version below is current as of publication.

Twenty-six rooms where someone with standing asks how this was made
The VERIFY protocol as built: six dimensions named Validate, Evidence, Review, Identify, Fact-check and Yield

References

In-text hyperlinks serve as point-of-claim citations; the consolidated list follows Harvard style.

Canada's Drug Agency (no date) Position statement on the use of artificial intelligence in the generation and reporting of evidence. Ottawa: CDA-AMC. Available at: https://www.cda-amc.ca/sites/default/files/MG%20Methods/Position_Statement_AI_Renumbered.pdf (Accessed: 2 September 2026).

European Commission (2025) Draft EudraLex Volume 4, Annex 22: Artificial Intelligence, stakeholder consultation, July 2025. Brussels: European Commission. Available at: https://health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en (Accessed: 2 September 2026).

European Medicines Agency (2024) Reflection paper on the use of artificial intelligence (AI) in the medicinal product lifecycle. Amsterdam: EMA. Available at: https://www.ema.europa.eu/en/use-artificial-intelligence-ai-medicinal-product-lifecycle-scientific-guideline (Accessed: 2 September 2026).

European Union (2024a) Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 12 and 86. Available at: https://artificialintelligenceact.eu/article/86/ (Accessed: 2 September 2026).

European Union (2024b) Directive (EU) 2024/2853 of the European Parliament and of the Council on liability for defective products. Available at: https://eur-lex.europa.eu/eli/dir/2024/2853/oj/eng (Accessed: 2 September 2026).

Faegre Drinker (2026) Transposing the EU's new Product Liability Directive: a member state progress report. Available at: https://www.faegredrinker.com/en/insights/publications/2026/6/transposing-the-eus-new-product-liability-directive-a-member-state-progress-report (Accessed: 2 September 2026).

'Fabricated citations: an audit across 2·5 million biomedical papers' (2026) The Lancet. Available at: https://www.thelancet.com/journals/lancet/article/PIIS0140-6736%2826%2900603-3/fulltext (Accessed: 2 September 2026). News coverage: Nature (2026) ‘Surge in fake citations uncovered by audit of 2.5 million biomedical-science papers’, 13 May 2026; Orrall, A. (2026) ‘One in 277 PubMed-indexed papers in 2026 shows fabricated references, says analysis’, Retraction Watch, 7 May 2026.

'Influence of topic familiarity and prompt specificity on citation fabrication in large language models' (2025) JMIR Mental Health, 12, e80371. Available at: https://mental.jmir.org/2025/1/e80371 (Accessed: 2 September 2026).

National Institute for Health and Care Excellence (2025) Use of AI in evidence generation: NICE position statement (ECD11). London: NICE. Available at: https://www.nice.org.uk/corporate/ecd11/chapter/our-position-on-the-use-of-ai-in-evidence-generation-and-reporting (Accessed: 2 September 2026).

Norton Rose Fulbright (2026) AI in litigation: update on Gen AI sanctions in 2026. Available at: https://www.nortonrosefulbright.com/en/knowledge/publications/792d8bf3/ai-in-litigation-update-on-gen-ai-sanctions-in-2026 (Accessed: 2 September 2026).

'OG-RAG: ontology-grounded retrieval-augmented generation for large language models' (2025) Proceedings of EMNLP 2025. Available at: https://aclanthology.org/2025.emnlp-main.1674/ (Accessed: 2 September 2026). Preprint: arXiv:2412.15235.

Padula, W.V. et al. (2022) 'Machine learning methods in health economics and outcomes research — the PALISADE checklist: a good practices report of an ISPOR task force', Value in Health, 25(7). Available at: https://www.ispor.org/heor-resources/good-practices/article/machine-learning-methods-in-health-economics-and-outcomes-research-the-palisade-checklist (Accessed: 2 September 2026).

U.S. Food and Drug Administration (2025) Considerations for the use of artificial intelligence to support regulatory decision-making for drug and biological products, draft guidance, January 2025. Silver Spring, MD: FDA. Available at: https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug (Accessed: 2 September 2026).

Yao, S. et al. (2024) 'τ-bench: a benchmark for tool-agent-user interaction in real-world domains', arXiv preprint, arXiv:2406.12045. Available at: https://arxiv.org/abs/2406.12045 (Accessed: 2 September 2026). Not peer-reviewed; cited as a preprint.

Ayinde v London Borough of Haringey; Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin). Divisional Court, 6 June 2025. Available at: https://www.judiciary.uk/wp-content/uploads/2025/06/Ayinde-v-London-Borough-of-Haringey-and-Al-Haroun-v-Qatar-National-Bank.pdf (Accessed: 12 September 2026).

Charlotin, D. (2026) AI Hallucination Cases [database]. Available at: https://www.damiencharlotin.com/hallucinations/ (Accessed: 12 September 2026; 2,039 cases listed; the count changes weekly).

Dratsch, T. et al. (2023) 'Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance', Radiology, 307(4). DOI: 10.1148/radiol.222176.

European Medicines Agency and US Food and Drug Administration (2026) Guiding principles of good AI practice in drug development, 14 January 2026. Available at: https://www.ema.europa.eu/en/documents/other/guiding-principles-good-ai-practice-drug-development_en.pdf (Accessed: 12 September 2026).

European Parliament and Council (2026) Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 … (Digital Omnibus on AI), in force 27 July 2026. Available at: https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng (Accessed: 12 September 2026).

Goddard, K., Roudsari, A. and Wyatt, J.C. (2012) 'Automation bias: a systematic review of frequency, effect mediators, and mitigators', Journal of the American Medical Informatics Association, 19(1), pp. 121–127. DOI: 10.1136/amiajnl-2011-000089.

in-toto Authors (no date) in-toto [software supply-chain integrity framework, CNCF graduated project]. Available at: https://in-toto.io/ (Accessed: 12 September 2026).

Medicines and Healthcare products Regulatory Agency (2018) 'GXP' data integrity guidance and definitions, Revision 1, March 2018. London: MHRA.

Noland v Land of the Free, L.P., No. B331918 (Cal. Ct. App., 2d Dist., Div. 3, 12 September 2025). Available at: https://law.justia.com/cases/california/court-of-appeal/2025/b331918.html (Accessed: 12 September 2026).

Parasuraman, R. and Manzey, D.H. (2010) 'Complacency and bias in human use of automation: an attentional integration', Human Factors, 52(3), pp. 381–410. DOI: 10.1177/0018720810376055.

Sinha, A. et al. (2026) 'The illusion of diminishing returns: measuring long horizon execution in LLMs', Proceedings of ICLR 2026. Preprint: arXiv:2509.09677.

US Food and Drug Administration (1997) 21 CFR Part 11 — Electronic records; electronic signatures. Available at: https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11 (Accessed: 12 September 2026).

US Food and Drug Administration, Center for Drug Evaluation and Research (2026) Warning letter to Purolea Cosmetics Lab, ref. 320-26-58, 2 April 2026. Available at: https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/purolea-cosmetics-lab-722591-04022026 (Accessed: 12 September 2026).

Wada, A. et al. (2025) 'Retrieval-augmented generation elevates local LLM quality in radiology contrast media consultation', npj Digital Medicine, 8. DOI: 10.1038/s41746-025-01802-z.

W3C (2013) PROV-DM: the PROV data model. W3C Recommendation, 30 April 2013. Available at: https://www.w3.org/TR/prov-dm/ (Accessed: 12 September 2026).

Lineages referenced in Section VI: BFO / OBO Foundry; SNOMED; TOVE, DEMO, APQC PCF, BIZBOK; W3C OWL 2 and SHACL; Palantir Foundry ontology documentation. Industry review-burden figures are from syndicated industry surveys on file with the strategy's evidence appendix.