Back to today's list

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

Kevin Du, Clara K\"umpel, Michelle Wastl, Alex Warstadt

Published Jul 23, 2026Featured #6In the daily list Jul 24, 2026
Daily score63.2
Editorial review7.0
Relevance0.491
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding how linguistic framing affects LLM behavior can inform prompt engineering and improve model robustness.

The paper evaluates how different expressions of belief influence LLMs' reliance on context versus prior knowledge.

Summary

The paper introduces a typology to evaluate how expressions of belief (EoBs) affect large language models' (LLMs) adherence to context versus prior knowledge. It categorizes EoBs based on form, evidentiality, epistemic stance, and tone, and uses this framework to test 16 LLMs of varying architectures, scales, and training stages. The study finds that larger and instruction-tuned models are less context-following and identifies specific EoBs that more effectively persuade LLMs.

Key contributions

  • Introduction of a typology for evaluating expressions of belief in LLMs.
  • Evaluation of 16 LLMs across different architectures, scales, and training stages.
  • Identification of systematic patterns in LLM response behavior based on linguistic framing.

Notable insights

  • Larger and instruction-tuned models tend to be less context-following than smaller and base models.
  • Certain expressions of belief statistically significantly persuade LLMs more consistently.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.18232v2 Announce Type: replace Abstract: Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true. In others, they should stick to their prior knowledge. Notably, users' expressions of belief (EoBs) can take linguistically diverse forms - using presuppositions, evidential and certainty markers, or varied tones - each of which may have a different persuasiveness over the LLMs. We introduce a typology to systematically evaluate how different EoBs affect whether models follow context versus prior knowledge. The typology is grounded in four linguistically motivated dimensions: form, evidentiality, epistemic stance, and tone, spanning 17 fine-grained types. By pairing these EoBs with world knowledge facts, we generate controlled EoB-query pairs that isolate the effect of linguistic variation. Using this benchmark, we evaluate 16 LLMs that differ in architecture (Llama3, Qwen3, Gemma3), scale (1B-30B parameters), and training stages (base vs instruct). We identify meaningful variations in response behavior across these axes, e.g., that bigger models and instruction models tend to be less context-following than smaller models and base models. We further identify specific EoBs that statistically significantly persuade LMs more consistently than others. Our work reveals systematic patterns in how linguistic framing affects LLM context integration, with implications for prompt engineering and model robustness.