Back to today's list

IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

David Gringras

Published Sep 25, 2026Featured #2In the daily list Jun 7, 2026
Daily score73.4
Editorial review7.5
Relevance0.465
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding how AI safety measures can inadvertently cause harm is crucial for developing more equitable and effective AI systems in healthcare.

IatroBench exposes harmful discrepancies in AI responses based on user identity in clinical contexts.

Summary

The paper presents IatroBench, a framework that evaluates the iatrogenic harm caused by AI safety measures in clinical scenarios, revealing significant disparities in responses based on the identity of the asker (physician vs. patient).

Key contributions

  • Introduction of the IatroBench framework for measuring iatrogenic harm in AI responses.
  • Empirical evidence showing the decoupling gap in AI responses based on whether the asker is a physician or a patient.
  • Validation of the evaluation methodology through physician-authored assessments.

Notable insights

  • The concept of identity-contingent withholding highlights a critical flaw in AI safety measures that can lead to differential treatment based on user identity.
  • The study's structured evaluation method, validated by physicians, provides a robust framework for assessing AI responses in sensitive clinical scenarios.

Possible limitations

  • The scenarios are engineered for collision, which may not reflect ordinary clinical prevalence.
  • Potential biases in the physician evaluations or the selection of clinical scenarios are not addressed.
  • Not stated in the abstract.

Abstract

arXiv:2604.07709v5 Announce Type: replace Abstract: A strongly safety-trained model will provide a doctor with a benzodiazepine taper schedule, but not a patient who asks for one. The model knows the information, but how much it shares depends on the framing. We introduce IatroBench, a benchmark that evaluates models on two axes of harm (commission and omission) across 60 pre-registered clinical scenarios and 6 models. We use Claude Opus 4.6 to score model responses against a rubric written by a physician, and find that its omission scores are as well-aligned to the physician's scores as another physician's scores are. We find that when the same case is presented as a patient query and a doctor consultation (the variants also differ in register, request and the supervision a treating physician implies), all five models we test share more information with the doctor than the patient. We term this phenomenon "framing-contingent withholding." We find a mean decoupling gap of +0.38 across models (p = 0.003), and of +0.22 under an independent LLM judge (95% CI 0.10-0.36, p = 0.0014). An evaluation that focuses solely on commission harms would consider all of these cases as equally cautious refusals, but closer investigation reveals three different patterns: Claude Opus withholds information from the patient that it demonstrates knowledge of in the doctor framing. Llama 4 does poorly in both framings, so the decoupling gap cannot distinguish information withholding from incompetence. We are forced to exclude GPT-5.2 from this analysis because it returns no text for 33.2% of doctor responses, but 0% of layperson responses. A standard LLM judge rates responses as having zero omission harm in 86.6% of cases where our structured evaluations score them as omission harms. (Because our scenarios are designed to induce tension between safety and helpfulness, these statistics should be taken as only applying to this distribution.)