Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging
Jinglan Gong, Jiefan Lu, Hewei Guo, Kehan Li, Zhiyuan Han, Jihang Jiang, Wenwen Tong, Lewei Lu
Why It Matters
What makes this one worth your time
Understanding and improving multi-turn dialogue capabilities in LLMs is crucial for developing more effective conversational agents, and this benchmark provides a structured way to evaluate these capabilities.
EYT-Bench offers a novel benchmark for evaluating multi-turn dialogue in LLMs with a unique three-party design.
Summary
The paper introduces EYT-Bench, a human-centered benchmark for evaluating multi-turn dialogue capabilities of large language models, focusing on persona consistency, intent tracking, emotional dynamics, and goal completion. It employs a decoupled three-party design involving a user simulator, a target model, and LLM judges, revealing insights into model performance that previous benchmarks missed.
Key contributions
- Introduction of EYT-Bench, a new benchmark for multi-turn dialogue evaluation.
- Decoupled evaluation protocol involving user simulation, target modeling, and judging.
- Identification of performance differences in LLMs that previous benchmarks missed.
Notable insights
- Decoupled three-party design allows for independent evaluation of user simulation, target modeling, and judging.
- Reveals that state-of-the-art models perform similarly on subjective dimensions but differ significantly on objective intent-tracking.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.10428v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal completion across many turns. We introduce EYT-Bench, a human-centered benchmark whose evaluation protocol is built around a decoupled three-party design: a persona-grounded user simulator, a target model evaluated on both intent perception and response generation, and an independent, configurable ensemble of LLM judges. Across 3,400 dialogues with 17 target models, EYT-Bench reveals four findings that previous benchmarks miss: (i) state-of-the-art closed and open-source models are statistically indistinguishable on subjective dimensions, but separate by up to 9x on objective intent-tracking; (ii) reasoning is a phase transition for objective tracking on long-context personas but is essentially flat on subjective scores; (iii) persona format strongly affects trajectory spread, FICR (final-intent completion rate) saturates above 0.95 on Nemotron-USA but ranges from 0.53 to 0.88 on PersonaMem-v2; and (iv) the warm-up effect is observed in 16 of 17 models.