Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
Sadia Kamal, Arefa Patwary, Anthony Marchiafava, Atriya Sen, Sagnik Ray Choudhury
Why It Matters
What makes this one worth your time
Understanding prompt robustness is crucial for accurately interpreting LLM outputs, especially in contexts involving subjective or value-laden questions.
Prompt robustness in LLMs varies significantly between objective and subjective questions.
Summary
The paper investigates the robustness of large language models (LLMs) to prompt variations, comparing objective questions with fixed answers to subjective questions that solicit opinions or values. It evaluates four instruction-tuned model families across three objective and three subjective datasets, applying various prompt changes to assess consistency in responses. The study finds significant effects of model, dataset, prompt category, and their interactions on prompt robustness.
Key contributions
- Evaluation of prompt robustness across different question types and datasets.
- Identification of significant interactions between model, dataset, and prompt category on response consistency.
Notable insights
- Prompt robustness is significantly influenced by the type of question and the nature of prompt changes.
- The interaction between dataset type and prompt category has a large effect on response consistency.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.05554v1 Announce Type: cross Abstract: Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.