Language Specific Knowledge: Do Models Know Better in X than in English?
Ishika Agarwal, Nimet Beyza Bozdag, Nisval Patel, Dilek Hakkani-T\"ur
Why It Matters
What makes this one worth your time
This research highlights the importance of language selection in multilingual models, which can enhance their effectiveness in diverse cultural contexts and improve accessibility for non-English speakers.
Language models can outperform English queries by leveraging language-specific knowledge.
Summary
The paper introduces the concept of Language Specific Knowledge (LSK) and demonstrates that language models can improve their question-answering performance when queried in languages other than English, including low-resource languages, by selecting the optimal language for specific queries.
Key contributions
- Introduction of the term Language Specific Knowledge (LSK) for optimizing query language selection.
- Empirical evaluation of language selection strategies across multiple datasets.
- Development of LSKExtractor, a method to aid in language selection for improved question answering.
Notable insights
- Language models may possess varying levels of expertise in different languages, which can be leveraged for better performance.
- The findings suggest that the relationship between language and model knowledge is not straightforward and can defy expectations.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2505.14990v3 Announce Type: replace Abstract: Often, multilingual language models are trained with the objective to map semantically similar content (in different languages) in the same latent space. In this paper, we show a nuance in this training objective, and find that by changing the language of the input query, we can improve the question answering ability of language models. We make two main contributions. First, we introduce the term Language Specific Knowledge (LSK) to denote queries that are best answered in an ``expert language'' for a given LLM, thereby enhancing its question-answering ability. We introduce the problem of language selection -- for some queries, language models can perform better when queried in languages other than English, sometimes even better in low-resource languages -- and the goal is to select the optimal language for the query. Second, we introduce a variety of simple to strong baselines to empirically motivate the language selection problem (including one of our own methods called LSKExtractor). During our evaluation, we employ three datasets that contain knowledge about both cultural and social behavioral norms. Overall, the results show that principled language selection can improve the performance of a language model, and that the expected question-to-language map is not always intuitive: Gemma models know most about China and Middle East in Spanish; Qwen models know most about authority and responsibility in Arabic and Chinese. Broadly, our research contributes to the open-source development of language models that are inclusive and more aligned with the cultural and linguistic contexts in which they are deployed.