NSF Funds Open-Source Framework to Trace Medical AI Failures
USC-led team receives $900K to build automated system that pinpoints data errors causing model breakdowns in clinical settings.

Automated diagnosis for broken medical AI models
A three-year, $900,000 National Science Foundation grant will fund development of an open-source framework designed to automatically identify why medical AI systems fail — and how to fix them.
The project addresses a critical gap in clinical AI deployment: when models produce poor results, researchers currently have no systematic way to trace performance problems back to specific errors in training data. Instead, they must conduct manual, time-consuming investigations that are difficult to reproduce and can delay patient care.
Ruishan Liu, a WiSE Gabilan assistant professor at USC Viterbi School of Engineering with joint appointments in quantitative biology and radiation oncology, is leading the effort alongside Virginia Tech researchers Ruoxi Jia and Wenjie Xiong. The team will test the framework primarily in radiation oncology, which relies on complex multimodal data including CT scans, MRI images, treatment plans, and longitudinal health records.
Why it matters
Medical data is expensive to collect and requires patient consent and clinical expertise, making it impractical to simply discard problematic datasets. An automated system that identifies and repairs data issues could accelerate AI adoption in healthcare while reducing the risk of model failures that delay critical treatments. The open-source approach means the framework will be accessible to medical institutions beyond major research centers.
Three-part architecture for AI troubleshooting
The framework consists of three integrated components, according to details first reported by USC Viterbi School of Engineering.
The algorithmic foundation uses reasoning-based large language models to identify which data caused a failure, diagnose why it occurred, and recommend cost-effective fixes. This replaces manual troubleshooting with an automated pipeline that handles the context-dependent errors common in medical workflows.
A provenance infrastructure layer logs the complete history and transformation of every data sample, using efficient structures like Bloom filters to enable rapid querying. This creates a traceable record of each data point's origins, transforming opaque AI systems into transparent ones where failures can be traced to specific scanner protocols, processing steps, or human errors.
The Integrated Curation Agent orchestrates these components into a closed-loop workflow. Critically, it includes a human-in-the-loop feature allowing medical experts to review and approve high-stakes remediation decisions before implementation.
Beyond radiation oncology
While the initial testing will focus on radiation oncology's complex data requirements, Liu emphasized the framework could extend across medical specialties. The standardized methodology for maintaining data integrity is designed to work in any high-stakes domain where AI reliability is paramount.
The tool's open-source release will make it available to medical researchers and clinicians, potentially reducing the manual labor required to identify and resolve data issues while promoting greater trust, accountability, and transparency in clinical AI deployment.
Details of the project were announced by USC Viterbi School of Engineering.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call