AI Models Show 75% Agreement With Heart Teams on Complex CAD
Large language models aligned with multidisciplinary teams three-quarters of the time when given detailed clinical narratives, but far less with structured data alone.
Large language models can match multidisciplinary heart team recommendations for complex coronary artery disease management in three out of four cases—but only when provided with extensive clinical detail, according to new research from the University of Melbourne.
The single-center study examined 546 patients with complex CAD whose treatment plans were determined through multidisciplinary heart team discussions between 2019 and 2024. Researchers then fed each case through two commercial AI systems—ChatGPT-4o and Gemini 2.0—using prompts with varying levels of clinical information.
When the AI tools received detailed clinical narrative inputs mirroring the information available to human teams, they agreed with heart team recommendations 75% of the time. But when given only structured case summaries containing 30 clinical and angiographic variables, agreement plummeted to just 35%.
The findings, published in JSCAI, reveal how heavily AI performance depends on input quality—a critical consideration as these tools move toward clinical deployment.
Context determines AI reliability
"The output that we got from AI was entirely dependent on the context that we gave it," senior author Anoop Koshy, MBBS, PhD, told TCTMD, which first reported the research. The results were "quite surprising," he added, given the sophisticated capabilities of modern language models.
The study population included predominantly patients with triple-vessel disease (75%) and some with severe left main disease (12%). Heart teams ultimately recommended CABG for 67% of patients, medical therapy for 19%, and PCI for 13%. Nearly all patients (91%) followed the recommended treatment.
Both AI systems performed similarly across prompts. Adding information from European and US revascularization guidelines to the prompts did not improve model performance.
Disagreement may flag higher risk
Cases where ChatGPT-4o disagreed with the heart team showed higher rates of major adverse cardiac events—myocardial infarction, stroke, or death—over a median follow-up of 989 days. The association weakened after adjusting for clinical complexity and treatment type, but Koshy suggested discordance itself might identify patients warranting closer scrutiny.
"When there is disagreement, that kind of highlights an intrinsically high-risk patient population that we probably need to put a bit more thought into," he said.
The researchers noted such cases might benefit from heightened review, second opinions, or additional risk-mitigation strategies.
Why it matters
Patients increasingly turn to AI tools for medical guidance, often without physician oversight. This study demonstrates that confident-sounding AI recommendations can vary dramatically based on how questions are framed—even when the underlying clinical situation remains identical. For healthcare organizations exploring AI decision support, the findings underscore that deployment requires careful validation, format-aware integration, and recognition that these tools cannot yet replace the nuanced judgment of multidisciplinary teams. The 25% disagreement rate even with optimal prompting suggests meaningful gaps remain before AI can reliably guide complex treatment decisions.
The research team emphasized that while AI shows promise, additional study is needed before deploying these tools as clinical decision-making aids for complex patient populations. Details of the study were first reported by TCTMD.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call