Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
Why It Matters
What makes this one worth your time
Understanding how deployment settings affect LLM outputs is crucial for ensuring the reliability and accountability of AI systems used in sensitive contexts.
Deployment configurations significantly influence LLMs' validation of pseudo-science.
Summary
The paper investigates how deployment configurations of large language models (LLMs) affect their validation of pseudo-scientific claims, specifically focusing on ethnonationalist pseudo-science. It examines four LLM families over time and finds that the models' epistemic stances are influenced by deployment factors such as system prompts and interface routing, rather than being inherent to the models themselves.
Key contributions
- Empirical analysis of LLMs' validation of pseudo-scientific claims across different deployment configurations.
- Identification of discrepancies in model outputs based on interface and silent updates.
Notable insights
- Deployment configurations, including silent updates and interface differences, can drastically alter LLM outputs.
- The same LLM can produce different results depending on whether it is accessed via API or web interface.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.