BioCentury
ARTICLE | Guest Commentary

The War of the Claudes is coming for the rooms where medicines get made

To safely scale drug development, AI must strengthen teams, make decisions auditable, and build human expertise rather than bypass it

September 2, 2026 12:43 AM UTC

In 2014, a small Melbourne not-for-profit, Medicines Development for Global Health, licensed a molecule from the World Health Organization that industry had walked away from. I was one of a handful of consultants they brought in to help.

Moxidectin had real promise against river blindness but no commercial future, so its development remained unfinished despite a global disease burden of around 20 million people. MDGH built a first-of-its-kind financing agreement to secure expert time from around the world and pair it with a very small in-house team. Then the team got creative. It revived an abandoned manufacturing program, ran clinical pharmacology studies that did double and triple duty, and bridged the remaining evidence gaps with modeling.

It is the most creative development package I have ever seen, and it worked. FDA approval. A priority review voucher. Then the World Health Organization Essential Medicines List.

We did not have large language models then. What moxidectin shows is what real multidisciplinary judgment and integration look like when a team gets everything right at once, and how rarely that alignment happens. The question now is whether AI can help teams achieve that alignment more consistently without tempting them to substitute generated answers for judgment.

Moving a molecule into the body of a waiting patient is among the most complex and consequential things human beings do. I say that as someone who has watched it succeed, and watched it fail rather more often.

So, when we talk about scaling drug development with AI, we should be precise about what we are scaling. Not just the process. Process is the more tractable part. The harder challenge is scaling human creativity, human judgment, and human integration.

AI is already widely used in early drug development: target discovery, molecule design, imaging. Those applications lend themselves more readily to AI, and thank goodness for it. Decision support is a different beast entirely.

In discovery, you are asking the machine whether it can generate or identify something useful. If it is wrong, the answer can often be discarded before much has been invested. Move along the development path, and there is a patient on the other end of the decision, then a regulator, and every call has to be defensible. The investment climbs fast. The consequence of being wrong climbs faster.

Somewhere along that path, the question quietly changes from “can AI generate something useful?” to “can I trust this decision?”

And trust, at that end, is not a feeling. It has to be demonstrable to a regulator and defensible to a board, two of the least sentimental audiences on earth.

Earning it takes four things. Hear them as engineering choices, not aspirations, whether applied to organizational processes, internal systems, or external tools. Left as hopes, every one of them quietly fails.

Hold the AI to the standard you hold the medicine to

A line often attributed to W. Edwards Deming puts it better than I can: “In God we trust; all others must bring data.”

Science slop is real. Nature estimates that tens of thousands of papers published in 2025 may cite references that an AI simply invented. An audit of nearly 7,000 journal submissions found writing quality declined as AI use rose, and AI-assisted papers were rejected more often, not less.

So evaluate any AI-supported process against current practice and ask honestly whether it saves time after its outputs have been checked or quietly creates more work. Measure the effects on total time, cost, and probability of success. Then publish the results where possible, so every claim survives contact with someone who would rather like to prove you wrong.

The stakes are not abstract. In April, FDA issued what appears to be its first warning letter citing over-reliance on AI. The company’s defense was, in essence, that the AI never told them validation was required. The regulator was unmoved.

AI will not take the blame for you.

Point it at the team, not at the individual

I call this one the War of the Claudes. Two capable people prepare for a meeting, each in a private conversation with their own AI. They arrive, and within 10 minutes it is “my Claude said this,” and, funny, “my Claude said that.” Two confident, fluent, completely irreconcilable answers, neither with an evidentiary trail anyone else can inspect.

The fix is to change where you point the thing. Rather than leaving each person to question a private AI in isolation, bring AI-supported claims into a common evidentiary record that everyone can interrogate together and document how the team reaches its decision. Disagreement, which you want, is then grounded in a single object everyone can stress test. The win was never a better chatbot. The win is creative collision in the room.

That does not mean asking a single system to generate the team’s answer. A single system can produce correlated errors, where everyone trusts the same blind spot at once, and a false sense of confidence emerges simply because an answer is shared rather than private. A common record is not a substitute for human challenge; it has to be built to invite disagreement, with clear accountability for whoever signs off on the final call.

You cannot review your way to accountability

AI has made authoring deceptively easy. It will draft you a 90-page investigator’s brochure over a coffee. The problem is the other side. No human can truly review 90 pages of fluent, confident, plausible text at the speed it is produced. We think we can. We cannot.

The evidence is humbling.

One study handed AI-fabricated abstracts to expert reviewers who were told in advance that fakes were present and instructed to look for them. They still missed roughly one in three. There is a name for the mechanism: automation bias. When something is fluent, quick, and mostly right, we defer to it. In some studies people reversed their own correct judgments because the machine disagreed. If final review cannot reliably catch the slop, you have to architect against it at the start.

That means breaking AI-assisted work into bounded pieces, small enough that a human can hold real critical attention on each one. It means a full chain of custody: who wrote what, how much of it was AI, and the source behind every claim. Made concrete, it looks as dull as this: “Section 4.2, drafted by Mary, AI assisted, three sources cited.” Dull is defensible.

In this vein, Anthropic announced in August that Claude would be watermarking AI authoring in compliance with the EU AI Act, though the company itself cautions that a detected mark shows only that Claude may have been involved, not who wrote which section. Watermarking is not a substitute for section-level provenance and human accountability.

Build capability without dependency

This one keeps me up at night because the damage stays invisible for years. AI can accelerate the development of expertise or quietly erode it.

Picture our profession as a pyramid: a broad base of juniors learning their craft, holding up a narrow tip of gray-haired experts. AI makes that tip fly. It fools experts too, but experience makes them better equipped to recognize when something is wrong. Novices have fewer such defenses.

AI also hollows out the base. Why grind through the training when the machine hands you the answer? The pyramid becomes a diamond. It takes about 15 years in this business to learn which questions to ask. If nobody is made to earn that experience now, who will have the judgment to make the hard calls five or 10 years from now?

Nowhere is that base thinner than in global health, where there was never a big commercial machine to train the next generation. Who leads the next moxidectin team, if we let the base hollow out?

Microsoft and Carnegie Mellon researchers found last year that the more people trusted AI, the less critical thinking they did, with the effect strongest in exactly the people least equipped to catch the errors.

Give a Stradivarius to a world-class musician and you get sublime music. Give the same instrument to a novice and you get noise.

So compress the climb, but do not let people skip it. Expose the reasoning and the judgment calls behind every output, not just the tidy answer. Grow the next generation on purpose.

Because the goal was never a faster novice. The goal is the next expert, the one who knows what to ask.

Which brings us back to what made moxidectin possible. Those four principles are how we close what I call the governance gap: hold processes, systems, and tools to the standard we hold medicine to; point it at the team rather than the individual; build the chain of custody that makes real review possible; and grow the next generation of experts on purpose. 

Get that right, and the lesson from the moxidectin example is not that the outcome can be repeated to order but that the judgment behind it can be built deliberately by giving people the time, support, and accountability to earn it. The patients who are waiting stand to benefit most.

That is the ethical imperative for AI in drug development: not a way around expertise, but a way of building more of it, deliberately, faster.

 _______________

Craig Rayner is co-founder and CEO of Oktopi, which develops AI-supported tools for drug-development decision-making. He co-founded and was CEO of drug-development consultancy d3 Medicine and held roles at Roche, CSL, and Moderna. 

Signed commentaries do not necessarily reflect the views of BioCentury.