Doctors Already Defer to AI Drafts, Abandoning Their Own Judgment
A pediatric surgeon warns that physicians editing AI-generated messages end up closer to the machine's answer than their own—even when the AI is wrong.

Physicians are overriding their own clinical judgment in favor of AI-generated text
When radiation oncologists at Mass General Brigham were asked to edit patient messages drafted by GPT-4, something unexpected happened. After editing, their final responses resembled the AI's draft more closely than the answers those same physicians had written minutes earlier—on their own, without the machine.
Some of those AI drafts contained dangerous errors, specifically misjudging how sick patients were. Yet experienced oncologists let the fluent, confident language quietly replace their own clinical assessment.
The pattern echoes Asiana Flight 214, which struck a seawall in San Francisco in 2013 on a clear day. The flight crew believed the autopilot was managing airspeed. It wasn't, but because they trusted the system, they stopped monitoring. The plane crashed.
Colin G. Knight, a board-certified pediatric surgeon practicing in Florida, argues this same automation bias is already embedded in clinical practice—and we're not prepared for it.
The phone consult problem
Knight points to a reality most hospitals have normalized: pediatric subspecialists who never examine patients in person. Endocrinology, hematology, infectious disease consultations happen by phone. The specialist reads the chart, asks questions remotely, and leaves a plan.
"I have no way to know that [the plan] is usually right," Knight writes. "The hospitalist examined the child. The specialist did not, and there is no second version of the case to compare against."
Once clinical findings become text, the question of who reads that text becomes wide open. And reading text is exactly what large language models excel at.
Where AI wins—and where it fails
A 2023 study compared physician answers to patient questions against responses generated by ChatGPT. Clinicians grading both blind preferred the AI most of the time. On empathy, the gap wasn't close.
Critics noted the human answers were brief, volunteer replies typed between patients, while the AI responses ran four times longer. But when researchers controlled for word count, the gap persisted. The machine validated, reassured, withheld judgment, and never sounded rushed.
Yet in another trial, physicians using a large language model for difficult diagnoses performed no better than those without it. The model working alone outperformed both groups. The tool was effective, the doctors were skilled, but the combination added nothing.
The reason: judging clinical acuity—recognizing when a patient is sicker than the numbers suggest—is where AI fails. It's also the skill physicians find hardest to articulate.
Why it matters
Medicine has not defined which clinical tasks require human attention and protected them intentionally, the way aviation did after decades of automation accidents. Instead, we're discovering our vulnerabilities one consult at a time. When a plan arrives from a source we cannot see—whether a remote specialist or an AI system—we can only judge how it sounds. That was tolerable when the voice belonged to someone with a decade of training who would answer a page at 2 a.m. It becomes a structural weakness when the system has been optimized to sound right, regardless of whether it is right.
Knight's warning is clear: we are defending the wrong boundary. The risk isn't that AI will replace pattern recognition while we keep the human touch. The evidence suggests the opposite. AI already handles patient-facing communication better than overloaded clinicians. What it cannot do reliably is the tacit, embodied judgment that tells us a child is crashing before the monitors do—and that's the judgment we're most likely to override when a confident paragraph appears on screen.
These details were first reported by Colin G. Knight, writing for KevinMD.
This is an original analysis by the Omega editorial team. Source reporting: Automation Watch.
Want systems like this working for your business?
Book a Call