Back to today's list

Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization

Yash Ganpat Sawant

Published Aug 17, 2026Featured #3In the daily list May 1, 2026
Daily score72.9
Editorial review7.5
Relevance0.476
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the limitations of current evaluation methods can guide future research in LLM personalization and improve the reliability of stylistic adaptations.

This research exposes critical gaps in LLM personalization evaluation through a theory-driven lens.

Summary

The paper evaluates stylistic personalization of LLMs using a theory-grounded approach based on authorship verification, revealing a significant authorship gap in existing evaluation metrics.

Key contributions

  • Introduces a theory-grounded evaluation framework for LLM personalization.
  • Demonstrates the inadequacy of existing metrics in capturing authorship differences.
  • Provides empirical results showing the performance of various personalization methods against a calibrated baseline.

Notable insights

  • The use of LUAR as a calibrated baseline provides a more meaningful evaluation compared to ad hoc metrics.
  • The near-zero pairwise correlations among the metrics highlight the importance of theoretical grounding in evaluation.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2604.26460v2 Announce Type: replace Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation grounded in authorship science. We show that grounding evaluation in authorship verification theory transforms what benchmarks can measure. Drawing on three measurement traditions - LUAR (a trained authorship verification model), an LLM-as-judge with decoupled trait matching, and classical function-word stylometrics - we evaluate four inference-time personalization methods across 50 authors and 1,000 generations. The theory-grounded metric (LUAR) provides what ad hoc alternatives cannot: calibrated baselines (human ceiling 0.756, cross-author floor 0.626) that give scores absolute meaning. All methods score below this floor (0.484-0.508), exposing an authorship gap invisible to uncalibrated metrics. The three metrics produce near-zero pairwise correlations (|r| < 0.07), confirming that without theoretical grounding, metric choice determines conclusions - an LLM judge declares a clear winner while LUAR finds no meaningful differentiation. These findings demonstrate the theory-benchmark cycle in action: authorship theory exposes evaluation failures that ad hoc benchmarks miss.