Back to today's list

Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models

Tomoya Mizumoto, Yusuke Fujita, Hao Shi, Lianbo Liu, Atsushi Kojima, Yui Sudo

Published Jun 25, 2026Featured #10In the daily list Jun 26, 2026
Daily score63.1
Editorial review7.2
Relevance0.456
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding dialectal variations is crucial for developing effective dialogue systems, especially in linguistically diverse contexts like Japan.

This study reveals how dialectal training enhances the robustness of speech language models.

Summary

The paper investigates the robustness of large language models (LLMs) and speech language models (SLMs) in understanding Japanese dialects, defining robustness as the performance ratio on dialectal versus standard inputs and demonstrating that training with dialectal data improves SLM performance.

Key contributions

  • Empirical evaluation of dialectal robustness in both LLMs and SLMs.
  • Establishment of a performance ratio metric for comparing dialectal and standard input processing.
  • Demonstration that dialectal training and fine-tuning enhance SLM performance.

Notable insights

  • The correlation between SLM and LLM robustness suggests that improvements in text-based models can directly benefit speech processing tasks.
  • Defining robustness as a performance ratio provides a clear metric for evaluating dialect comprehension across different model types.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2606.25436v1 Announce Type: cross Abstract: Dialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs.