The reasoning problem

R&D
Biomedical researchers discuss results

Spend enough time inside biopharmaceutical R&D and you notice that the decisions that matter most rarely happen during the documented steps; they happen in the pauses between them.

A senior scientist scans a body of literature and quietly sets aside ten supporting studies to focus on one contradictory finding nobody else flagged. A translational team halts a promising programme because a signal buried deep in the data introduces a level of uncertainty they're not willing to carry forward. An experienced researcher looks at a target that checks every box on paper and decides the underlying evidence isn't mature enough to justify moving forward. It might take another eighteen months before that changes.

None of those decisions appear in any SOP. They rarely get written down at all. They are the product of years of exposure to failed experiments, ambiguous signals, and patterns that only become visible after you have enough experience to recognise them. For most of biopharma's history, organisations treated this kind of reasoning as something that would transfer on its own through mentorship, proximity, or the slow absorption of instincts that experienced researchers carried with them.

That assumption is getting harder to sustain.

When institutional knowledge walks out the door

The volume problem in biomedical research has moved past inconvenient and into something more structurally serious. Even in narrow therapeutic areas, comprehensive manual synthesis is no longer a practical baseline. Information accumulates faster than teams can process it. Important signals get buried. Gaps that should be visible early stay hidden until they become expensive.

The recent layoff waves across biotech and pharma have compounded this in ways that don't get discussed much. When experienced researchers leave, the scientific reasoning they carried often goes with them. In some organisations, entire therapeutic areas have been effectively reset after restructuring. Teams inherit programmes without inheriting the reasoning that shaped them. Organisations document workflows extensively; what they almost never document is how a senior scientist actually moved through a workflow: which signals they weighted, what level of uncertainty they considered acceptable, why they trusted one study over three others that appeared more statistically robust.

The SOPs describe the steps. They don't describe the thinking.

Research teams were once stable enough that this gap was manageable. Institutional knowledge transferred slowly, and the habits of scientific judgement passed informally between generations, that system was imperfect, but it functioned. It is now under pressure from several directions at once, and the scientific judgement that walked out with departing researchers reduced headcount and something harder to replace: the accumulated reasoning of the people who knew how to navigate uncertainty.

Decision quality is the real measure

Much of the conversation around AI in pharmaceutical R&D has centred on efficiency. Faster literature searches, automated summaries, accelerated reporting cycles. Those gains are real. But scientific organisations are ultimately judged by the quality of their decisions, and speed doesn't change the terms of that judgement.

Consider what happens when a research team misses contradictory evidence early in the process; the target advances and resources accumulate behind it. By the time the contradiction surfaces, sometimes years later or only in clinical trials, the costs are staggering: failed programmes, patient safety risks, years of scientific effort redirected toward something that was never going to work. In most of these cases, the contradictory information existed in the literature the entire time. Nobody surfaced it, and nobody weighed it against what the team already believed.

This shows up in target review boards as decisions that looked defensible at the time, but weren't. When the primary measure for AI tools is time saved, organisations end up optimising for a secondary concern. Reducing uncertainty earlier by capturing a contradictory signal before a programme advances to trials, or surfacing a safety issue while there's still room to act, is where the real value sits. Indication exploration done thoroughly can compress months of manual research. That compression only produces better decisions, though, if the reasoning behind it is sound, traceable, and built to scientific standards, not the general-purpose kind.

Scientists understand this distinction viscerally. Most are trained to interrogate evidence reflexively: checking where a conclusion came from, how conflicting findings were handled, and whether important context was set aside. Systems that summarise fluently, but obscure their reasoning, lose credibility quickly in research settings. The question scientists ask about any output is whether it can be defended. Sounding plausible counts for little.

Workflow design is harder to fix than it looks, and most organisations underestimate that.

Building backward from decisions

Scientific workflows in R&D are typically built forward. It starts with a task, the steps are followed, and an output is produced. What most workflow design doesn't capture is the decision architecture underneath: the gates where experienced scientists evaluate whether evidence is sufficient to proceed, what acceptable uncertainty looks like at each stage, and what specific signals would constitute a reason to stop.

Building workflows backward from decisions changes the design problem entirely. Instead of documenting steps, judgement is documented: what the evidence needs to show before a programme moves forward, what level of confirmation is acceptable at a target assessment review, what gaps are serious enough to pause a programme. Getting experienced scientists to articulate this is harder than it sounds. Many researchers who have been making these calls for decades have never had to describe them explicitly. The reasoning is fast, built over years of pattern recognition and hard-won familiarity with how these calls tend to play out. Turning that into something explicit enough to govern a workflow is a different kind of exercise and, for most researchers, an unfamiliar one.

Codifying scientific reasoning also raises a question that tends to get underweighted: how do you know the codification is trustworthy? It isn't useful if it can't be inspected, traced back to its sources, or audited when a decision is challenged. The goal is to scale reasoning in ways that are scientifically defensible. Evidence provenance, consistent output structure across teams and programmes, human checkpoints at key decision gates, where scientists evaluate the work instead of performing every step of it.

From execution to orchestration

The scientists who have worked furthest through this tend to describe a shift in how they experience their own role. Literature searches, evidence compilation, monitoring for updates, generating summaries: all increasingly handled by automated systems. The work that remains is the work that was always the point: defining the right questions, setting evidence standards, evaluating hypotheses, challenging assumptions, making the calls. It becomes less execution and more interpretation.

The easier it becomes to retrieve and organise evidence, the more important interpretation becomes. Whether a body of literature supports a defensible conclusion still requires a scientist, one who understands the experimental context, recognises methodological limitations, and can weigh uncertainty with the calibrated instinct that takes years to develop.

The question worth answering

Biopharmaceutical organisations have long struggled to preserve how experienced scientists reason through uncertainty. The steps they follow get documented. The judgement underneath them rarely does. AI-enabled scientific workflows are beginning to make that possible. They do it by encoding that reasoning in forms that can be applied consistently, inspected when needed, and handed to the next generation, rather than lost when senior people leave.

That leaves a question for every research organisation: does the scientific judgement depended on exist anywhere outside the heads of the people who currently hold it? In most cases, honestly, it doesn't. That's the problem worth solving.

About the author

Stavroula Ntoufa, PhD, is director of scientific affairs at Causaly. A molecular biologist by training with 14 years in research, she's published on bias in early scientific research and works directly with biopharma R&D teams on how scientific judgement is captured (and lost) in real workflows.

Image
Stavroula Ntoufa

Stavroula Ntoufa