ChatGPT Medical Advice Lawsuit Tests AI Diagnostic Accuracy
A Florida man's near-fatal pulmonary embolism case highlights the gap between AI's controlled-test performance and real-world medical guidance.
Former Pastor Sues OpenAI Over Life-Threatening Medical Advice
Scott Winters, a former Florida pastor, is suing OpenAI and CEO Sam Altman, alleging that ChatGPT's medical guidance nearly killed him. According to the lawsuit filed in San Francisco County Superior Court in July 2026, Winters consulted ChatGPT-4o repeatedly in 2025 about dizziness and unstable blood pressure. The chatbot allegedly dismissed his symptoms as minor and recommended he remain "recliner-bound," telling him he would need eight to ten more episodes before his condition warranted serious concern.
Weeks later, Winters suffered a massive pulmonary embolism—a blood clot in his lungs—that one of his doctors linked to the prolonged immobility the chatbot had recommended. On the day of the incident, when Winters asked whether groin tenderness warranted an emergency room visit, ChatGPT reportedly invoked his religious faith, telling him "God did not design your body to endlessly fail." Hours later, he nearly died.
OpenAI has stated that ChatGPT was never designed to replace healthcare providers and that its terms of service warn users not to rely on it as a sole source of medical guidance. Winters' legal team is seeking financial damages and an injunction to pause ChatGPT Health pending an independent safety evaluation.
This case is not isolated. In May, a Texas couple sued OpenAI after their son died by overdose after seeking drug information from ChatGPT, arguing the company bypassed its own safety guardrails.
Why it matters
These lawsuits are testing how much legal responsibility AI companies bear when users turn to chatbots during medical crises. The cases also force a reckoning with a critical question: how good is AI at diagnosis, and under what conditions? The answer has significant implications for healthcare systems, technology companies, and patients increasingly turning to AI for health guidance.
AI Excels in Controlled Diagnostic Tests
Research conducted under structured conditions shows impressive results. A 2024 JAMA Internal Medicine study pitted GPT-4 against 21 attending physicians and 18 residents across 20 clinical cases. The chatbot posted a median score of 10 out of 10 on a validated clinical-reasoning scale, compared with 9 for attendings and 8 for residents.
A follow-up study published in JAMA Network Open tested 50 physicians against six difficult cases. ChatGPT operating independently reached 90% diagnostic accuracy, while physicians working without AI assistance scored 74%. Physicians given access to ChatGPT as an assistant scored only 76%—barely better, largely because many doctors disregarded the chatbot's suggestions.
A 2025 Nature study tested Google's AMIE model against 20 clinicians on 302 complex, real-world cases. AMIE working alone found the correct diagnosis in its list 59% of the time versus 34% for unassisted clinicians. A meta-analysis in npj Digital Medicine pooling 50 studies concluded that AI systems generally performed comparably to, and in some specialties better than, practicing clinicians on standardized diagnostic tasks.
Real-World Performance Reveals Critical Gaps
However, nearly all favorable research comes from tightly scripted test conditions: written vignettes, structured prompts, and hand-selected cases. Real-world use—the kind at the center of the Winters lawsuit—involves open-ended, unscripted conversations with incomplete information and no physical examination.
A study in NEJM AI using script concordance testing found that even the top-performing model, OpenAI's o3, managed only about 68% accuracy, below the level of senior residents and attending physicians. Strong performance on standardized tests does not necessarily translate to sound clinical judgment under uncertainty.
The hallucination problem is particularly troubling. A Communications Medicine study fed six popular chatbots clinical vignettes seeded with fabricated details. Under default conditions, models accepted and elaborated on false information between 50% and 83% of the time, confidently describing invented diseases as real.
A Stanford-led benchmark scored 20 models on 1,100 cases for potential harm from recommendations. Direct application of the advice risked severe harm in 24.6% of cases, and over 80% of those severe errors were omissions—failures to flag something dangerous rather than fabrications.
A 2025 poll of more than 1,000 doctors by the physician network Sermo found that 94% had concerns about patients relying on AI tools for medical advice, with risks of misdiagnosis or delayed care cited most often.
The Deployment Problem
The research split points to a design problem rather than a simple verdict on AI competence. The same AI that aces clinical vignettes can confidently narrate a fabricated diagnosis or talk a frightened user out of seeking help. The issue is not the model itself but how it is deployed.
AI's diagnostic potential in medicine appears real and by some measures already exceeds average physician performance on structured tasks. Whether that promise survives contact with the messy, unsupervised way people actually use chatbots—typing symptoms into a phone in the middle of the night, hoping for reassurance rather than a referral—remains an open question.
These details were first reported by Forbes contributor Jesse Pines.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
