Back to today's list

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

Jinglan Gong, Jiefan Lu, Hewei Guo, Kehan Li, Zhiyuan Han, Jihang Jiang, Wenwen Tong, Lewei Lu

Published Jul 27, 2026Featured #6In the daily list Jul 28, 2026
Daily score61.0
Editorial review7.0
Relevance0.458
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and improving multi-turn dialogue capabilities in LLMs is crucial for developing more effective conversational agents, and this benchmark provides a structured way to evaluate these capabilities.

EYT-Bench offers a novel benchmark for evaluating multi-turn dialogue in LLMs with a unique three-party design.

Summary

The paper introduces EYT-Bench, a human-centered benchmark for evaluating multi-turn dialogue capabilities of large language models, focusing on persona consistency, intent tracking, emotional dynamics, and goal completion. It employs a decoupled three-party design involving a user simulator, a target model, and LLM judges, revealing insights into model performance that previous benchmarks missed.

Key contributions

  • Introduction of EYT-Bench, a new benchmark for multi-turn dialogue evaluation.
  • Decoupled evaluation protocol involving user simulation, target modeling, and judging.
  • Identification of performance differences in LLMs that previous benchmarks missed.

Notable insights

  • Decoupled three-party design allows for independent evaluation of user simulation, target modeling, and judging.
  • Reveals that state-of-the-art models perform similarly on subjective dimensions but differ significantly on objective intent-tracking.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.10428v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal completion across many turns. We introduce EYT-Bench, a human-centered benchmark whose evaluation protocol is built around a decoupled three-party design: a persona-grounded user simulator, a target model evaluated on both intent perception and response generation, and an independent, configurable ensemble of LLM judges. Across 3,400 dialogues with 17 target models, EYT-Bench reveals four findings that previous benchmarks miss: (i) state-of-the-art closed and open-source models are statistically indistinguishable on subjective dimensions, but separate by up to 9x on objective intent-tracking; (ii) reasoning is a phase transition for objective tracking on long-context personas but is essentially flat on subjective scores; (iii) persona format strongly affects trajectory spread, FICR (final-intent completion rate) saturates above 0.95 on Nemotron-USA but ranges from 0.53 to 0.88 on PersonaMem-v2; and (iv) the warm-up effect is observed in 16 of 17 models.