AI

AI Diagnostic Tools Help Novices But Mislead Them More Often

MIT study reveals non-experts defer to AI explanations even when wrong, while clinicians catch errors and perform best with minimal AI assistance.

Omega Editorial· August 4, 2026· 3 min read

A new study from MIT challenges the assumption that one explainability approach works for all users of AI diagnostic systems. Researchers found that while AI assistance generally improved diagnostic accuracy across skill levels, the way users interact with AI explanations varies dramatically based on their medical expertise—with potentially dangerous consequences for novices.

The research, published in Nature Medicine, tested both non-experts and primary care physicians on skin disease diagnosis tasks using various explainable AI approaches. These included basic predictions with confidence scores, similar-image comparisons, heat maps highlighting important image regions, and large language model explanations in plain language.

Non-experts show dangerous deference

Non-experts improved their accuracy in identifying cancerous moles when using AI tools, primarily because the systems helped them correctly diagnose benign cases. However, this improvement came with a significant caveat: these users trusted AI explanations whether correct or incorrect, and found vague or generic explanations more convincing.

"When the model is wrong, it hurts performance more than it helps performance when the model is right," says Marzyeh Ghassemi, associate professor in MIT's Department of Electrical Engineering and Computer Science and principal investigator at the Abdul Latif Jameel Clinic for Machine Learning in Health.

The deference effect proved strongest with LLM-generated explanations, where users expressed more confidence in their incorrect answers when aided by the AI's reasoning.

Clinicians resist AI errors

By contrast, primary care physicians demonstrated resilience to incorrect AI assistance. When given the more complex task of providing differential diagnoses for dermatological conditions, clinicians performed best when presented with only the model's prediction—no explanation at all. They caught AI errors and used their existing medical knowledge to validate or reject the system's recommendations.

"A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught," explains lead author Orson Xu, assistant professor in the Department of Biomedical Informatics at Columbia University. "Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place."

Timing and presentation matter

The research revealed that when users receive AI explanations influences their behavior significantly. Presenting explanations before users form their own diagnostic hypothesis increased deference to the model. The researchers also found that AI systems outperformed humans when disease presentation was subtle, but humans excelled when images contained atypical symptoms or unrelated features.

Users most deferential to AI assistance were those who performed worst on diagnostic tasks without AI help—precisely the population that could benefit most from technological support but is also most vulnerable to being misled.

Why it matters

As FDA-approved AI diagnostic interfaces proliferate in clinical settings and consumer-facing AI health tools become ubiquitous, understanding how different users interact with AI explanations becomes critical for patient safety. The findings suggest that explainability methods should be tailored to user expertise levels, and that forcing users to generate diagnostic hypotheses before seeing AI recommendations may reduce harmful overreliance. For healthcare AI developers, the research underscores that improving model accuracy alone isn't sufficient—the presentation and timing of AI assistance must account for automation bias and encourage critical thinking rather than blind acceptance.

The study was conducted by researchers at MIT, Stanford University, and Columbia University, with findings first reported in Nature Medicine. The research was funded by the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia University.

#explainable ai#medical ai#diagnostic tools#automation bias#healthcare technology#mit research

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in AI

AI· 2 min read

Anthropic's Claude Now Leads 26% of Company's AI Development

The milestone marks rapid progress toward self-improving AI systems that require minimal human oversight.

Via AI Watch · Sep 18, 2026
AI· 3 min read

Anthropic Reports Claude AI Now Contributes 26% to Its Own Development

The company disclosed new metrics showing how AI agents are increasingly involved in building next-generation models, though human oversight remains constant.

Via AI Watch · Sep 17, 2026
AI· 2 min read

Anthropic Releases Three Metrics for Tracking AI Development Pace

The Claude maker shares methodologies for measuring autonomous R&D, agent oversight, and compute allocation following CEO's slowdown proposal.

Via AI Watch · Sep 17, 2026