Back to today's list

Response drift across frontier large language models

Mohammed Aledhari, Ali Aledhari, Fatimah Aledhari, Gowtham Venkat Eathamokkala, Mohamed Rahouti

Published Jul 24, 2026Featured #7In the daily list Jul 25, 2026
Daily score59.7
Editorial review6.8
Relevance0.466
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding response drift in LLMs is crucial for improving their reliability and performance across different applications and domains.

The study systematically evaluates response drift in LLMs, highlighting its universal and domain-dependent nature.

Summary

The paper investigates response drift in large language models (LLMs) by conducting a comprehensive human evaluation across ten models and multiple domains, revealing that all models exhibit drift with varying magnitudes and domain-specific profiles.

Key contributions

  • Comprehensive human evaluation of response drift across ten LLMs.
  • Identification of domain- and question-dependent drift profiles.
  • Demonstration of the inadequacy of automated metrics in explaining human judgments on drift.

Notable insights

  • Human evaluation is essential for accurately characterizing response drift in LLMs.
  • Automated metrics poorly correlate with human judgments on response drift.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.20454v1 Announce Type: cross Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation. Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments. Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%). Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85. Automated similarity metrics explain less than 2% of variance in human judgements. These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centred evaluation.