Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
Why It Matters
What makes this one worth your time
Understanding the limitations of LLMs in clinical settings is crucial for ensuring patient safety and guiding future research in AI-driven healthcare solutions.
LLMs are not yet safe for autonomous clinical triage due to critical gaps in information gathering and decision-making under uncertainty.
Summary
The paper discusses the limitations of large language models (LLMs) in autonomous clinical decision support, particularly in the context of triaging undifferentiated patients, highlighting the risks associated with their use in safety-critical scenarios.
Key contributions
- Identification of specific modes of failure for LLMs in clinical triage scenarios.
- A framework for understanding the asymmetrical costs associated with clinical decision-making.
Notable insights
- The paper emphasizes the importance of seeking improbable but critical diagnoses rather than just the most likely ones, which is a nuanced aspect of clinical reasoning.
- It highlights the potential for LLMs to exhibit biases such as credulity and agreeableness, which can compromise clinical decision-making.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2607.28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.