CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
Dengzhe Hou, Lingyu Jiang, Fangzhou Lin, Kazunori D Yamada
Why It Matters
What makes this one worth your time
Understanding cognitive ability structures in LLMs can enhance model evaluation and inform future developments in AI.
CogArena offers a new benchmark for assessing cognitive abilities in LLMs.
Summary
The paper introduces CogArena, a benchmark designed to evaluate cognitive ability structures in large language models (LLMs) through a multimethod framework, analyzing cognitive-task scores across various model families and interventions.
Key contributions
- Introduction of the CogArena benchmark with 13 paradigms for cognitive evaluation.
- Analysis of cognitive-task scores across 55 models, revealing insights into dimensionality and variance.
- Investigation of the effects of targeted scaffolds on cognitive-task performance across different model families.
Notable insights
- The study reveals that while cognitive-task scores correlate positively across paradigms, the common axis explains only half the variance, indicating complexity in cognitive profiling.
- The findings suggest that theory-aligned prompting yields only a minor tendency towards stable cognitive profiles, challenging assumptions about LLM capabilities.
Possible limitations
- The abstract indicates that the frozen confirmation criterion fails and that no scaffold-specific contrasts survive multiplicity correction, suggesting potential issues with the robustness of findings.
- Not stated in the abstract.
Abstract
arXiv:2607.24999v1 Announce Type: cross Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.