Back to today's list

Safety from Honesty in a Disinterested AI Predictor

Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn

Published Jul 14, 2026
Editorial review6.5
Relevance0.526
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding how to train AI systems to make predictions without developing unintended agency is crucial for ensuring safety as AI capabilities advance.

The paper proposes a framework for a non-agentic AI Predictor that ensures safety through epistemic contextualization and training constraints.

Summary

The paper presents a formal safety argument for a non-agentic AI Predictor, called the Scientist AI (SAI) Predictor, which is trained to approximate Bayesian posteriors based on epistemically contextualized natural-language statements. The Predictor aims to make honest predictions about agents and actions without adopting goals, relying on data representation and training procedures to avoid implicit agency.

Key contributions

  • A formal safety argument for a non-agentic AI Predictor.
  • A training framework that uses epistemic contextualization to prevent goal adoption.

Notable insights

  • Epistemic contextualization of text helps distinguish factual claims from communication acts, preventing the model from adopting goals.
  • The training procedure avoids using downstream effects as reward signals, reducing the risk of implicit agency.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2606.29657v2 Announce Type: replace Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.