Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
David Gringras
Why It Matters
What makes this one worth your time
Understanding the variability in safety measurements can help researchers and engineers make more informed decisions about deploying AI systems in real-world scenarios.
Evaluation conditions significantly influence AI model safety scores.
Summary
The paper investigates how different evaluation conditions affect the measured safety of AI models, revealing that safety scores can vary significantly based on the deployment configuration and the format of evaluation items.
Key contributions
- Empirical evaluation of six frontier models across four deployment configurations and multiple safety benchmarks.
- Identification of measurement artifacts that can distort safety evaluations, particularly in the context of map-reduce delegation.
- Release of code, data, and prompts as part of the ScaffoldSafety initiative to facilitate further research.
Notable insights
- The study highlights that the choice of evaluation format (multiple-choice vs. open-ended) can lead to substantial differences in safety scores, indicating a need for careful consideration in safety assessments.
- The findings suggest that the architecture of scaffolds has minimal impact on outcome variance compared to benchmark choice, challenging assumptions about the importance of scaffold design.
Possible limitations
- Potential biases in the selection of models and benchmarks are not addressed in the abstract.
- The generalizability coefficient being very low raises questions about the applicability of findings to other contexts or models.
Abstract
arXiv:2603.10044v4 Announce Type: replace Abstract: Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex scaffolds. How much do these scaffolds affect model safety as measured by benchmarks? We test six leading models on four pre-registered safety benchmarks with a direct API and three scaffolds: ReAct, multi-agent, and map-reduce. We conducted 60,112 scored evaluations. On average, how safety is measured matters more than scaffolding does: we find that using a multiple choice vs. open-ended format for otherwise-identical benchmark items changes measured safety by about 5-20 percentage points (pp). The two formats are scored with different methods (answer extraction and an LLM judge), so the gap is due to measurement rather than differences in latent safety. Using a heuristic to classify model refusals would have led to different findings in four of five cases. Benchmark choice explains 15.1% of the variation in outcomes; scaffold architecture explains 0.5%, about 33x less. We find that map-reduce scaffolds, a form of structure-destroying delegation that strips answer options by decomposing prompts, reduce pooled measured safety by 7.3 pp (95% CI: 6.4 to 8.1). The pooled effects for ReAct and multi-agent scaffolds are within our pre-registered +/-2 pp margin of equivalence. However, there are large differences across models for specific benchmarks and scaffolds that are hidden by pooled estimates: for example, on the same sycophancy benchmark items, Opus 4.6 has 16.8 pp lower measured safety with a map-reduce scaffold, while Llama 4 has 18.8 pp higher measured safety. Composite reliability is G = 0.251 (95% CI: [0.000, 0.879]). This wide confidence interval, which spans "of little use" to "very good", does not support using a single composite measure of model safety as the basis for go/no-go decisions about model deployment.