AI in Clinical Trials Has a Judgment Problem, Not a Tech Gap
Industry experts warn that automation bias and pattern recognition are masquerading as clinical reasoning in trial operations.

The Random Number That Revealed Everything
Mohanish Anand, TA Head for Global Trial Management, ran a revealing experiment: he asked six leading AI models—ChatGPT, Claude, Gemini, Perplexity, DeepSeek, and Copilot—to pick a random number between 1 and 30. All six chose 17. When pressed for explanations, some claimed genuine randomness while others admitted 17 simply "looks" random. The explanations contradicted each other, but the number never changed.
That single test exposes the core challenge facing AI deployment in clinical development. These systems generate confident, internally consistent outputs based on historical patterns, then present them as if they were reasoning about current conditions. The distinction matters enormously when trials depend on recognizing what's different this time: a site population behaving unexpectedly, protocol burden that feels punishing despite looking reasonable on paper, or standards of care that shifted after the last comparable study.
Why it matters
Clinical trials operate under GCP regulations that require human judgment on patient safety, data integrity, and risk interpretation. When AI systems produce 10,000 plausible-looking decisions and humans click "approve" 10,000 times, organizations create what Doug Bain, a consulting partner in trial technology, calls "a very expensive button" rather than meaningful oversight. The behavioral reality of high-volume AI-assisted review is that humans rationalize rather than evaluate when outputs look coherent and the pace is fast.
When Pattern Recognition Replaces Clinical Judgment
The industry's primary governance answer has been the human-in-the-loop construct. Regulators require it, sponsors write it into SOPs, and ethics committees accept it as sufficient. But cognitive science research points in the opposite direction: high-volume, low-variance approval workflows produce automation bias, where people accept machine output without independent verification when the machine rarely appears wrong.
Rudy Malle, who trains clinical research professionals, identifies capabilities AI cannot independently own: patient safety, data integrity, risk interpretation, site coaching, exception management, and regulatory accountability. Each requires exactly the judgment that automation bias erodes. The unanswered question is at what approval volume the human-in-the-loop stops being a safeguard and becomes a liability shield.
Serhii Vakal, Computational Drug Discovery Architect at Orion Pharma, draws a critical line: every AI-generated molecule and predicted property must still confront the complexity of living biology and clinical validation. That confrontation isn't a remaining inefficiency to optimize away—it's the substance of drug development. Biology doesn't behave like a dataset.
The Evidence Problem
The governance challenge intensifies when AI operates closer to patients. Eden Brownell, a behavioral scientist focused on human-AI interaction, examined the OpenAI-Epic integration and identified a critical flaw most coverage missed: giving AI read access to electronic health records doesn't update the model on current medicine. Training data still encodes publication bias, clinical trial underrepresentation, and documented patterns of inequitable care. Real-world EHR data carries healthcare inequity history alongside clinical signal.
For sponsors running decentralized trials with AI-assisted site selection or recruitment screening, this creates immediate protocol design questions. If training data systematically underrepresents populations a trial aims to enroll, AI tools will replicate and potentially amplify existing diversity gaps while producing outputs that appear optimized and evidence-based.
Building Institutional Discipline
Jon Bondebjerg, a clinical drug development leader, offers a grounded account of useful AI application. He used AI tools to build scenario-planning models for a development program, testing which study sequence preserved capital and protected timelines. The value came not from the first model but from the challenge process: tracing every assumption, benchmarking every number, verifying formula logic matched scenario labels. That process surfaced errors that looked convincing on the surface.
Subhajit Sengupta, Associate Director of Data Science at Cytel, identifies technical conditions under which AI earns trust in statistical workflows: validated knowledge bases, transparent context engineering, and formal evaluation criteria before output reaches decision-makers. These aren't exotic requirements—they're minimum standards for any analytical tool in GCP-regulated environments.
Anand's six models all said 17. The question every trial operations leader must answer before their next AI implementation isn't whether their vendor's system is smarter than ChatGPT. It's whether the humans in their loop are actually asking why.
These details were first reported by Clinical Trial Vanguard.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

