MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
Why It Matters
What makes this one worth your time
This work provides a significant step forward in NLP for Yiddish, a historically rich but digitally underrepresented language, offering a template for similar efforts in other low-resource languages.
MameLoshnLM advances Yiddish NLP with a specialized language model and benchmark.
Summary
The paper introduces MameLoshnLM, an 8B-parameter language model specifically designed for Yiddish, along with a high-quality pretraining corpus called Oytser and a multi-task benchmark named Kashes. The model outperforms existing multilingual models in capturing Yiddish-specific linguistic features, highlighting the limitations of noisy multilingual data for low-resource languages.
Key contributions
- Development of MameLoshnLM, an 8B-parameter Yiddish language model.
- Creation of Oytser, a high-quality Yiddish pretraining corpus.
- Introduction of Kashes, a multi-task benchmark for evaluating Yiddish language models.
Notable insights
- The use of a high-quality, language-specific corpus (Oytser) improves model performance over general multilingual datasets.
- MameLoshnLM demonstrates that specialized models can better capture linguistic nuances of low-resource languages compared to general-purpose models.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.05850v1 Announce Type: new Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.