Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand
Rafael da Silva, Jeff Eicher
Why It Matters
What makes this one worth your time
Understanding the cross-lingual comprehension gap is crucial for developing more equitable and effective language models that serve a global audience, especially for low-resource languages.
The paper quantifies how language models' understanding diminishes when content is presented in languages other than English.
Summary
The paper introduces the concept of the Cross-Lingual Comprehension Gap (CLCG) to measure how language models' comprehension varies when content is presented in languages other than English. Using a parallel corpus, the study evaluates five models across 18 languages, finding a notable reduction in response quality in non-English languages, particularly in low-resource languages.
Key contributions
- Introduction of the Cross-Lingual Comprehension Gap (CLCG) metric.
- Evaluation of language models across 18 languages using a professionally translated parallel corpus.
- Quantitative analysis of the impact of language resources on model comprehension.
Notable insights
- The study uses a within-item design to isolate language effects by varying only the passage language.
- The paper finds a negative association between language-level CLCG and resource class, highlighting the challenges faced by low-resource languages.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.06506v1 Announce Type: new Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.