Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation
Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
Why It Matters
What makes this one worth your time
This work addresses the critical need for reliable evaluation methods in medical AI, which can enhance patient safety and improve clinical decision-making.
A novel framework automates the evaluation of medical dialogue systems using evidence-based rubrics.
Summary
The paper presents a framework for automating the generation of evaluation rubrics for medical dialogue systems, leveraging retrieval-augmented methods to create instance-specific criteria based on authoritative medical evidence.
Key contributions
- Development of a retrieval-augmented multi-agent framework for automated rubric generation.
- Demonstration of improved Clinical Intent Alignment (CIA) scores compared to existing baselines.
- Provision of a scalable solution for both evaluating and refining medical LLM responses.
Notable insights
- The framework's use of retrieval-augmented techniques to ground evaluations in authoritative medical evidence is a clever approach to enhance reliability.
- The ability to generate fine-grained rubrics that adapt to user interactions represents a significant advancement in evaluation methodologies.
Possible limitations
- Potential limitations in the generalizability of the framework across diverse medical domains or languages are not addressed.
- Not stated in the abstract.
Abstract
arXiv:2601.15161v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed by generic metrics and LLM judges using general criteria, while expert-authored fine-grained rubrics are expensive and difficult to scale. In this paper, we propose a retrieval-augmented multi-agent framework for automatically generating instance-specific evaluation rubrics. Our approach grounds evaluation in authoritative medical evidence by decomposing retrieved content into atomic facts and synthesizing them with user interaction constraints to form fine-grained evaluation criteria. Evaluated on HealthBench and LLMEval-Med, our framework achieves Clinical Intent Alignment (CIA) scores of 50.20% and 31.90%, significantly outperforming the GPT-4o baseline and showing consistent improvements across English and Chinese medical benchmarks. In discriminative tests on HealthBench, our rubrics achieve a 7.8% point higher win rate than GPT-4o and increase the mean score difference from 4.972 to 8.658. Ablation studies further show that individual components contribute differently across datasets, with interaction-intent modeling providing the most consistent contribution to clinical-criterion coverage. Beyond evaluation, our rubrics guide response refinement, improving response quality by 9.2%. These results suggest that automated, knowledge-grounded rubric generation provides a scalable foundation for evaluating and improving medical LLMs. The code is available at https://anonymous.4open.science/r/Automated-Rubric-Generation-E716/.