Back to today's list

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni

Published Aug 6, 2026
Editorial review6.8
Relevance0.496
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding LLMs' limitations in processing figurative language is crucial for improving their semantic reasoning capabilities, which is essential for applications requiring nuanced language comprehension.

ProverbIT exposes LLMs' limitations in understanding culturally embedded expressions.

Summary

The paper introduces ProverbIT, an Italian benchmark for evaluating large language models' ability to complete proverbs, revealing that while models can complete proverbs, they struggle with multiple-choice tasks without correct answers, indicating a reliance on memorization over semantic understanding.

Key contributions

  • Introduction of ProverbIT, a novel Italian benchmark for proverb completion.
  • Evaluation of 13 models on proverb completion and multiple-choice tasks.
  • Chain-of-Thought analysis revealing reliance on memorized patterns.

Notable insights

  • Models show a bias towards selecting literal synonyms in absence of correct answers.
  • LLMs often mention correct endings during reasoning but fail to recognize their absence in options.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.04670v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions. This paper introduces ProverbIT, a novel Italian benchmark comprising 100 multiple-choice questions designed to evaluate LLMs' ability to complete Italian proverbs. We assess 13 frontier models, including Large Reasoning Models (LRMs) and traditional LLMs, across three tasks: proverb completion, multiple-choice selection with correct answers, and multiple-choice selection without correct answers. Our evaluation reveals surprising results: while nearly all models demonstrate knowledge of the proverbs through successful completion tasks, performance drops dramatically when transitioning to multiple-choice formats without correct answers, with even state-of-the-art reasoning models showing substantial degradation. Through detailed Chain-of-Thought analysis of two LRMs, we uncover that models exhibit a strong bias toward selecting literal synonyms and frequently mention correct proverb endings during reasoning without successfully identifying their absence from the given options. These findings suggest that current LLMs rely heavily on memorized patterns rather than deeper semantic understanding of culturally grounded expressions, highlighting important limitations in their reasoning capabilities for figurative language comprehension.