Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
Manuel Pita
Why It Matters
What makes this one worth your time
Understanding the limitations of national language models like AMALIA in accurately annotating complex constructs is crucial for their effective deployment in real-world applications.
The study questions the validity of AMALIA as a standalone annotator for moral constructs, highlighting its reliance on surface features.
Summary
The paper evaluates the validity of the AMALIA language model as a data annotator for the moral foundation of authority in European Portuguese, comparing its performance to larger open models and examining its reliance on surface correlates rather than theoretical constructs.
Key contributions
- Evaluation of AMALIA's performance against larger open models for annotating moral constructs.
- Introduction of the recovery gap method to assess the validity of model annotations.
Notable insights
- The recovery gap method is used to test whether a model's agreement with human coders is theoretically valid or based on surface shortcuts.
- Calibration's role in closing the performance gap between holistic and decomposed prompts is explored.
Possible limitations
- The study is limited to one construct and one corpus.
- Not stated in the abstract
Abstract
arXiv:2607.08731v2 Announce Type: replace-cross Abstract: National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weights create a presumption of trustworthiness. Such an instrument, built by and for a language community, looks like the natural choice for measuring what that community says and values. Whether such a model validly measures anything is untested at release. The evaluation of LLMs as measurement instruments is typically task-specific and stops at agreement with human coders. Agreement cannot distinguish an LLM instrument that measures a construct from one that reaches matching codes through surface correlates. We audit the presumption on a favourable case: AMALIA, Portugal's publicly funded 9B model, coding the moral foundation of authority in European Portuguese. The \textit{recovery gap} operationalizes the audit: decompose the codebook into its theory-defined clauses, recombine them through the theory's explicit rule, and measure how much of the original prompt's performance the stated theory reproduces. In a pre-registered, out-of-sample study on a transcreated (English to European Portuguese) corpus, AMALIA agrees with trained coders within six points of open models eight to thirteen times its size. Yet, the recovery gap shows that only about half of coding performance on authority can be attributed to the theory. A larger multilingual LLM closes the recovery gap on the same corpus, suggesting the shortfall lies in the annotator model, not the corpus or its translation. Sovereignty earns operational and performance trust; epistemic trust requires calibration -- and the audit method is inexpensive, and portable across models, languages and tasks.