Back to today's list

Challenges in annotations by humans and LLMs: A case study of evaluative language

Mirela Imamovic, Aenne Cecilia Kristine Knierim, Khushi Pitroda, Ekaterina Lapshinova-Koltunski

Published Jul 31, 2026Featured #10In the daily list Aug 1, 2026
Daily score61.4
Editorial review6.8
Relevance0.523
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding how LLMs can aid in complex annotation tasks could streamline processes in digital humanities and improve the accuracy of linguistic analyses.

LLMs can effectively assist in annotating complex linguistic phenomena like evaluative language.

Summary

The paper compares human and large language model (LLM) annotations in the context of evaluative language in TED talk transcripts, focusing on Appraisal theory's Attitude subsystem. It evaluates the performance of LLMs against trained linguists and linguists in training, finding that LLMs perform comparably to trained linguists. The study suggests that LLMs can assist in complex annotation tasks.

Key contributions

  • Comparison of LLMs and human annotators in evaluative language annotation.
  • Development and testing of prompts for LLM performance in Appraisal theory classification.
  • Demonstration of LLMs' potential to aid in complex linguistic annotation tasks.

Notable insights

  • LLMs can reach performance levels comparable to trained linguists in complex annotation tasks.
  • The study highlights the potential of LLMs to assist in subjective linguistic annotation tasks, which are traditionally challenging for humans.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.28119v1 Announce Type: new Abstract: In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.