Back to today's list

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs

Mahmood Bayeshi, Veysel Kocaman, Muhammed Ali Naqvi, Yigit Gul, David Talby

Published Jul 28, 2026Featured #1In the daily list Jul 29, 2026
Daily score75.9
Editorial review8.2
Relevance0.452
Freshness0.722

Why It Matters

What makes this one worth your time

This research demonstrates that structured workflows can enhance diagnostic reasoning in medical AI, potentially improving clinical outcomes and reducing costs.

A novel agentic workflow boosts a small model's diagnostic accuracy beyond larger LLMs.

Summary

The paper introduces the DeepLens Diagnosis Agent, a structured five-stage workflow that enhances the diagnostic accuracy of a small medical reasoning model by leveraging retrieval-augmented generation, achieving significant performance improvements over standard models.

Key contributions

  • Development of the DeepLens Diagnosis Agent with a five-stage workflow for medical diagnosis.
  • Demonstration of a +36-point accuracy improvement through structured workflow compared to the model used in isolation.
  • Cost-effective performance metrics showing the agent's efficiency relative to larger models.

Notable insights

  • The agent's structured pipeline allows for explicit evidence triangulation and auditability, which are critical in high-stakes medical settings.
  • The significant performance gain from workflow design alone suggests that model size is not the sole determinant of diagnostic effectiveness.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.22555v1 Announce Type: new Abstract: Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost. Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance.