Back to today's list

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Han-yu Wang

Published Sep 15, 2026Featured #2In the daily list Jun 28, 2026
Daily score72.0
Editorial review7.5
Relevance0.458
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the differences in reasoning patterns can inform the development of more effective AI systems and improve human-AI collaboration.

This study reveals a fundamental divergence in reasoning strategies between humans and large reasoning models.

Summary

The paper investigates the differences in reasoning patterns between large reasoning models (LRMs) and humans, specifically focusing on how response time correlates with problem difficulty and the allocation of deliberation resources during problem-solving.

Key contributions

  • Empirical analysis of human and LRM reasoning patterns using a matched corpus.
  • Identification of distinct deliberation policies between humans and LRMs.
  • Demonstration of how response time and token usage can be dissociated in evaluating reasoning performance.

Notable insights

  • LRMs exhibit a wrong-vs-right effect in token usage that contrasts with human behavior, suggesting different underlying cognitive strategies.
  • The study introduces a framework to separate difficulty registration from deliberation allocation, providing a nuanced view of reasoning processes.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2606.26502v5 Announce Type: replace Abstract: Large reasoning models (LRMs) tend to produce longer reasoning traces on problems that also take humans longer. This correspondence leaves open how the systems distribute further work on those problems. We distinguish *difficulty registration*, sensitivity to differences in problem difficulty, from *deliberation allocation*, the distribution of further work once difficulty is encountered. We examine both in item-matched data from three reasoning tasks. In visual abstraction (H-ARC), model trace length follows the human ordering of problems by duration. After item identity is controlled, successful human attempts last longer than failed attempts, while failed LRM attempts have longer traces than successful ones in the pooled model analysis. The estimated slopes follow the same pattern in intuitive reasoning (INTUIT). In relational reasoning (Cortes), successful attempts are longer in separate human and model analyses, while a joint fit on shared items finds a human-LRM difference. Longer human attempts include more grid actions. At comparable lengths, failed LRM traces contain more hedging on H-ARC and more repetition on Cortes. A resource-rational account relates these patterns to what further work is expected to achieve and how that progress is valued. The results identify a difference in the allocation of continued work that cross-problem duration alignment alone leaves undetected.