Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
Why It Matters
What makes this one worth your time
Understanding the reliability of evaluation metrics is crucial for improving translation quality in digital humanities, particularly for historically significant languages.
This study reveals the limitations of current evaluation metrics in Classical Chinese translation.
Summary
The paper investigates the reliability of existing automatic evaluation metrics for translating Classical Chinese to English, introducing a diagnostic framework to assess error sensitivity and tolerance to variation, ultimately finding that all metrics have blind spots.
Key contributions
- A diagnostic framework for evaluating translation errors in Classical Chinese to English.
- An analysis of error sensitivity and tolerance in reference-based and reference-free metrics.
- Identification of the best-performing metric, MetricX24, in this context.
Notable insights
- The introduction of a diagnostic framework based on minimal pairs is a novel approach to assessing translation errors.
- The study highlights specific blind spots in existing metrics, which could guide future metric development.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.08283v2 Announce Type: replace-cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.