Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
Why It Matters
What makes this one worth your time
Understanding the limitations and providing guidelines for LLM evaluation in diverse language settings is crucial for fair and accurate NLP advancements.
The paper critiques and offers recommendations for using LLMs as evaluators in multilingual and low-resource language contexts.
Summary
The paper investigates the use of LLM-as-a-Judge for evaluating natural language generation tasks in multilingual and low-resource language settings, identifying inconsistencies and overreliance on single models, and provides recommendations for improvement.
Key contributions
- Analysis of 650 papers to identify the use of LLM-as-a-Judge in multilingual and low-resource settings.
- Identification of inconsistencies in evaluation outcomes across these settings.
- Recommendations for improving LLM evaluation practices in multilingual contexts.
Notable insights
- There is a tendency to overtrust LLM judgments in multilingual settings.
- Most studies rely on a single judge model, which may not be sufficient for diverse language evaluations.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings.