Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
Hai Hu, Siyuan Song, Chongtian Shao, Kejia Zhang, Tianjian Zhu, Xiaojing Zhao
Why It Matters
What makes this one worth your time
Understanding the limitations of LLMs in reasoning and creativity, especially in non-English contexts, is crucial for developing more robust and versatile AI systems.
The paper challenges LLMs' reasoning abilities using Chinese xiehouyu riddles, highlighting potential overreliance on memorization.
Summary
The paper evaluates the reasoning and memorization capabilities of large language models (LLMs) using Chinese xiehouyu riddles, focusing on their ability to handle novel riddles created by linguists to avoid data contamination. It uses multiple-choice questions, free-form explanation generation, and new riddle creation to assess LLMs' understanding and creativity. The study finds that Chinese models may rely more on memorization due to larger training data, while English-centric models show less memorization. The paper suggests that LLMs' reasoning abilities may be overestimated and their creativity in language tasks is still inferior to humans.
Key contributions
- Evaluation of LLMs' reasoning and memorization using Chinese xiehouyu riddles.
- Introduction of novel xiehouyu riddles to test LLMs' reasoning without data contamination.
- Comparison of performance between Chinese and English-centric LLMs on language tasks.
Notable insights
- The use of novel xiehouyu riddles created by linguists helps avoid data contamination and tests true reasoning capabilities.
- The delta of accuracy between known and novel riddles serves as an index for assessing memorization versus reasoning.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.23440v1 Announce Type: cross Abstract: In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($\Delta_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $\Delta_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $\Delta_{acc}$ of 23.6\%, while English-centric models tested have a mean $\Delta_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.