Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization
Yash Ganpat Sawant
Why It Matters
What makes this one worth your time
Understanding the limitations of current evaluation methods can guide future research in LLM personalization and improve the reliability of stylistic adaptations.
This research exposes critical gaps in LLM personalization evaluation through a theory-driven lens.
Summary
The paper evaluates stylistic personalization of LLMs using a theory-grounded approach based on authorship verification, revealing a significant authorship gap in existing evaluation metrics.
Key contributions
- Introduces a theory-grounded evaluation framework for LLM personalization.
- Demonstrates the inadequacy of existing metrics in capturing authorship differences.
- Provides empirical results showing the performance of various personalization methods against a calibrated baseline.
Notable insights
- The use of LUAR as a calibrated baseline provides a more meaningful evaluation compared to ad hoc metrics.
- The near-zero pairwise correlations among the metrics highlight the importance of theoretical grounding in evaluation.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2604.26460v2 Announce Type: replace Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation grounded in authorship science. We show that grounding evaluation in authorship verification theory transforms what benchmarks can measure. Drawing on three measurement traditions - LUAR (a trained authorship verification model), an LLM-as-judge with decoupled trait matching, and classical function-word stylometrics - we evaluate four inference-time personalization methods across 50 authors and 1,000 generations. The theory-grounded metric (LUAR) provides what ad hoc alternatives cannot: calibrated baselines (human ceiling 0.756, cross-author floor 0.626) that give scores absolute meaning. All methods score below this floor (0.484-0.508), exposing an authorship gap invisible to uncalibrated metrics. The three metrics produce near-zero pairwise correlations (|r| < 0.07), confirming that without theoretical grounding, metric choice determines conclusions - an LLM judge declares a clear winner while LUAR finds no meaningful differentiation. These findings demonstrate the theory-benchmark cycle in action: authorship theory exposes evaluation failures that ad hoc benchmarks miss.