AI Models Identify Cell Reactions But Struggle With Predictions
University of Virginia researchers found popular AI tools correctly predicted cellular drug responses only 6% of the time in biomedical testing.

AI Shows Mixed Results in Cellular Biology Research
Researchers at the University of Virginia School of Medicine have completed a systematic evaluation of leading artificial intelligence models—including ChatGPT, Gemini, and Claude—revealing significant limitations when applied to biomedical research questions.
The study, led by Jeff Saucerman, PhD, from UVA's Department of Biomedical Engineering, tasked these AI systems with explaining cellular communication processes, specifically focusing on heart cells. While the models demonstrated competence in identifying known cellular reactions—achieving accuracy rates up to 65%—their performance deteriorated sharply when asked to predict how cells would respond to disease states or experimental drug compounds.
In predictive scenarios, the AI tools achieved accuracy rates as low as 6%.
Why it matters
These accuracy gaps carry substantial consequences for drug development and patient care. Incorrect predictions in laboratory settings can redirect research efforts down unproductive paths, wasting millions in development costs and delaying potential treatments. For patients awaiting new therapies, AI-generated false positives could create misleading expectations about treatment timelines or efficacy.
The Knowledge Versus Reasoning Gap
Saucerman characterized the current state of AI capabilities in biomedical contexts as fundamentally uneven. The models excel at cataloging individual cellular components and established biological facts—essentially retrieving and organizing existing knowledge. However, they struggle with the integrative reasoning required to predict novel interactions.
"It's pretty good already at knowing the individual pieces of cells, but it's not very good at piecing them together, and that's what's needed to predict more substantially into the future what new drugs would do," Saucerman explained.
This distinction highlights a critical limitation: understanding components doesn't automatically translate to understanding systems. Cellular biology involves complex, dynamic interactions where multiple factors influence outcomes simultaneously—a type of reasoning that current AI architectures handle poorly.
Validation Requirements Before Clinical Application
The research team emphasized that rigorous, multi-level testing must precede any reliance on AI predictions in biomedical contexts. Saucerman noted that validation should employ different methodological approaches rather than accepting model outputs at face value.
"Just because a model makes a prediction doesn't mean we should trust it," he said.
For the foreseeable future, the UVA team recommends maintaining human oversight as the final authority in research decisions. While acknowledging that AI capabilities are advancing rapidly, the current generation of models requires human scientists to verify predictions before they inform experimental design or clinical decisions.
The findings underscore an emerging principle in AI deployment: these tools may serve best as assistants that accelerate information gathering rather than autonomous decision-makers in high-stakes scientific contexts.
These details were first reported by WVIR.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
