Back to today's list

Validity of LLMs as data annotators: AMALIA on authority

Manuel Pita

Published Jul 11, 2026Featured #9In the daily list Jul 12, 2026
Daily score54.2
Editorial review6.5
Relevance0.490
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the limitations of national language models like AMALIA in accurately annotating complex constructs is crucial for their effective deployment in real-world applications.

The study questions the validity of AMALIA as a standalone annotator for moral constructs, highlighting its reliance on surface features.

Summary

The paper evaluates the validity of the AMALIA language model as a data annotator for the moral foundation of authority in European Portuguese, comparing its performance to larger open models and examining its reliance on surface correlates rather than theoretical constructs.

Key contributions

  • Evaluation of AMALIA's performance against larger open models for annotating moral constructs.
  • Introduction of the recovery gap method to assess the validity of model annotations.

Notable insights

  • The recovery gap method is used to test whether a model's agreement with human coders is theoretically valid or based on surface shortcuts.
  • Calibration's role in closing the performance gap between holistic and decomposed prompts is explored.

Possible limitations

  • The study is limited to one construct and one corpus.
  • Not stated in the abstract

Abstract

arXiv:2607.08731v1 Announce Type: new Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code by correlated shortcuts. We test this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined by the theory's explicit rule. If calibration closes that gap, some portability should survive across models and languages; where it does not, the construct-model instrument is the likely locus of failure. We ask whether a calibrated English instrument transfers to AMALIA-9B and to European Portuguese. For one construct and one corpus, it does not. Decomposition recovers only about half of AMALIA's holistic performance, and error analysis suggests reliance on surface correlates, especially moral outrage near authority figures. An open multilingual LLM closes the gap on the same Portuguese corpus under the same instructions, pointing away from the corpus as the main explanation. AMALIA can still screen and pre-code at scale, but it cannot yet measure this construct well enough to stand alone. The study is a single counterexample, not a verdict on national models; it argues that sovereign-LLM benchmark batteries should test not only agreement with human coders, but the evidential route by which that agreement is warranted.