Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+
Sherzod Hakimov, Karl Osswald, Jelle Psurek, Eszter Bukovszky, A. Altar L\"user, David Schlangen
Why It Matters
What makes this one worth your time
Understanding the performance disparities between commercial and open-weight LLMs across multiple languages can guide future development and deployment of language models in multilingual contexts.
Commercial LLMs outperform open-weight models in multi-language dialogue games, revealing the cost and performance gap in non-English languages.
Summary
The paper evaluates the performance of large language models (LLMs) in multi-turn dialogue games across 30 languages, including the 24 official EU languages. It compares nine LLMs, both open-weight and commercial, finding that commercial models outperform open-weight ones across all languages. The study highlights the challenges of achieving linguistic parity using public web text alone and notes the higher cost and lower performance of non-English languages.
Key contributions
- Systematic evaluation of LLMs across 30 languages using a multi-turn, reference-free dialogue game paradigm.
- Comparison of commercial and open-weight LLMs, highlighting performance disparities.
Notable insights
- Commercial LLMs outperform open-weight models even in languages with significantly less public web text.
- Non-English languages incur higher operational costs and lower performance compared to English.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this paradigm is multi-turn, reference-free and programmatically scored, and because the game mechanics are language-agnostic it extends to a new language by localising a fixed set of prompt and word-list files. Evaluating nine open-weight and commercial LLMs, we find that no open-weight model covers the EU-24 well: in every official language both commercial systems outscore every open-weight model, and the two weakest average below 40 points across the EU-24. The commercial systems stay ahead even in languages with four orders of magnitude less public web text, showing that linguistic parity is achievable, but not from public crawls alone. A model's home region lifts it without closing the gap: Chinese is the strongest of all 30 languages for two Chinese-developed models, yet the best Chinese score of any model belongs to a US commercial system. Coverage is also not parity of service. Pooled over models and languages, the median non-English language costs 31% more to run than English, and scores 10% lower.