Back to today's list

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo

Published Sep 3, 2026Featured #3In the daily list Jul 31, 2026
Daily score69.1
Editorial review7.2
Relevance0.464
Freshness0.722

Why It Matters

What makes this one worth your time

This approach offers a scalable and potentially more objective method for evaluating LLMs, addressing issues with static benchmarks and subjective human preferences.

LivingArena evaluates LLMs by having them probe each other's knowledge gaps.

Summary

The paper introduces LivingArena, a framework for evaluating large language models (LLMs) by having them propose questions to each other, aiming to identify and exploit knowledge gaps. This method is designed to be contamination-resistant and scalable, producing a stable leaderboard of model performance based on their ability to challenge peers and answer questions accurately.

Key contributions

  • Introduction of a dynamic, peer-probing evaluation framework for LLMs.
  • Development of a contamination-resistant method for model evaluation.
  • Creation of a stable Elo leaderboard for model performance comparison.

Notable insights

  • Models can actively identify and exploit the cognitive boundaries of their peers.
  • The framework measures not just knowledge recall but also the ability to probe weaknesses.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.24780v2 Announce Type: replace Abstract: Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We ask whether LLMs can instead discover one another's weaknesses and turn those observations into an evaluation process. To study this question, we introduce \textbf{LivingArena}, an automated peer-probing framework in which models take turns testing one another. Using the interaction history, each questioner identifies potential weaknesses of its opponent and constructs targeted, verifiable questions to probe them. A 3,600-round tournament of ten models reveals a clear role asymmetry: strong answerers are not always reliable questioners, because they may generate internally inconsistent tests or fail to verify their own reference answers. After a questioner exposes an answerer's failure, it is more likely to pursue the same capability domain, while the answerer's weakness recurs on independently generated questions, including questions written by different models. These findings show that peer probing can reveal persistent model-specific weaknesses while separately evaluating answering and reliable test construction. By automating this process and allowing test difficulty to evolve with model capabilities, LivingArena provides a "living" benchmark for model development, red-teaming, and capability-aware multi-agent coordination. We publicly release our code: https://github.com/galaxyChen/LivingArena