Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao, Yingyue Luna Luan
Why It Matters
What makes this one worth your time
Understanding how LLMs evaluate creativity compared to humans can guide the selection of appropriate models for tasks requiring creativity assessment, impacting AI's role in creative industries.
LLMs and humans evaluate creativity differently, especially regarding contextual information.
Summary
The paper investigates the alignment between large language models (LLMs) and human evaluations of creativity, identifying differences in evaluation standards across three studies. It finds that LLMs align more closely with humans on intrinsic qualities like novelty but diverge on contextual factors. The study highlights the importance of selecting appropriate LLMs for creativity evaluation based on their standards.
Key contributions
- Identification of differences in creativity evaluation standards between LLMs and humans.
- Empirical evidence showing LLMs' alignment with human evaluations varies by dimension.
- Analysis of how contextual information affects LLM and human creativity ratings differently.
Notable insights
- LLMs tend to rely on a narrower subset of human creativity evaluation standards.
- Different LLMs apply distinct, model-specific standards, affecting their creativity judgments.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.