From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
Itay Itzhak, Eliya Habba, Gabriel Stanovsky, Yonatan Belinkov
Why It Matters
What makes this one worth your time
This work addresses the gap between traditional benchmark evaluations and practical user experiences, potentially improving LLM assessments in real-world applications.
Formalizing vibe-testing could enhance LLM evaluation by aligning it with real-world user experiences.
Summary
The paper investigates the informal practice of 'vibe-testing' LLMs, formalizes it into a systematic evaluation process, and demonstrates its effectiveness through experiments on coding benchmarks.
Key contributions
- Analysis of user evaluation practices through surveys and social media reports.
- Formalization of vibe-testing as a two-part process involving user personalization.
- Development of an evaluation pipeline that generates personalized prompts and incorporates user-aware evaluation criteria.
Notable insights
- The study reveals that user personalization significantly influences model preference, highlighting the subjective nature of LLM evaluations.
- The introduction of a proof-of-concept evaluation pipeline demonstrates a novel approach to integrating user-aware criteria into model assessments.
Possible limitations
- Potential biases in user feedback from the analyzed resources.
- The proof-of-concept may not generalize across all LLM applications or user contexts.
- Not stated in the abstract.
Abstract
arXiv:2604.14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'': informal experience-based evaluation, such as comparing models on coding tasks related to their own workflow. While prevalent, vibe-testing is often too ad hoc and unstructured to analyze or reproduce at scale. In this work, we study how vibe-testing works in practice and then formalize it to support systematic analysis. We first analyze two empirical resources: (1) a survey of user evaluation practices, and (2) a collection of in-the-wild model comparison reports from blogs and social media. Based on these resources, we formalize vibe-testing as a two-part process: users personalize both what they test and how they judge responses. We then introduce a proof-of-concept evaluation pipeline that follows this formulation by generating personalized prompts and comparing model outputs using user-aware subjective criteria. In experiments on coding benchmarks, we find that combining personalized prompts and user-aware evaluation can change which model is preferred, reflecting the role of vibe-testing in practice. These findings suggest that formalized vibe-testing can serve as a useful approach for bridging benchmark scores and real-world experience.