Back to today's list

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

Koyar Afrasyab

Published Jul 22, 2026
Editorial review6.8
Relevance0.468
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding how evaluator biases affect AI safety assessments is crucial for developing reliable medical AI systems.

The study reveals how evaluator choice impacts the perceived safety of medical AI in incomplete clinical conversations.

Summary

The paper evaluates the safety of medical AI models in open-ended clinical conversations with missing information, using a panel of LLM judges and a clinician reference to assess model responses. It finds that judge choice affects perceived safety and that LLM judges are generally more lenient than clinicians.

Key contributions

  • Introduced a method to evaluate medical AI under conditions of missing information.
  • Demonstrated the impact of evaluator choice on perceived AI safety.
  • Released evaluation tools and data for further research.

Notable insights

  • Evaluator choice significantly alters perceived model safety, indicating the importance of diverse and unbiased evaluation panels.
  • LLM judges tend to be more lenient than clinicians, highlighting potential discrepancies in AI safety evaluations.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss' kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.