Beyond "AI Language": The case for the idiolectal nature of LLM output
Karolina Rudnicka, Thomas Stephan Juzek
Why It Matters
What makes this one worth your time
Understanding the idiolectal nature of LLM outputs can enhance text detection methods and inform forensic linguistics, which are crucial for applications in security and content verification.
LLM outputs reveal distinct linguistic profiles similar to human idiolects.
Summary
The paper analyzes the outputs of large language models (LLMs) to argue that they exhibit unique, model-specific linguistic signatures, akin to human idiolects, based on datasets from 2024 and 2026.
Key contributions
- Analysis of two distinct datasets of LLM-generated texts across different years.
- Identification of unique linguistic signatures for individual models.
- Demonstration of a generational shift in LLM output styles.
Notable insights
- The study employs stylometric principal component analysis to identify generational shifts in LLM output styles.
- The variation in contraction frequencies among models highlights the nuanced differences in their linguistic profiles.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse two datasets of LLM-generated texts on societal topics: a 2024 corpus of six models (Improta et al. 2024) and a newly generated 2026 corpus using the same prompts featuring six contemporary models. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile. This multi-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models (2026). Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM-generated text detection, forensic linguistics and usage-based approaches to language.