Back to today's list

Agentic Safety is an Epistemic Property, Not a Behavioral One

Charles L. Wang, Keir Dorchen, Peter Jin

Published Jun 30, 2026
Editorial review6.8
Relevance0.557
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding AI safety as an epistemic property is crucial for developing systems that remain controllable and correctable as they become more autonomous and capable.

AI safety should focus on maintaining future correctability, not just current behavior.

Summary

The paper argues that AI safety should be considered an epistemic property rather than just a behavioral one, emphasizing the importance of maintaining the ability to correct AI systems as they evolve and self-improve. It introduces the concept of 'teachability' as a measure of an AI system's capacity to remain correctable over time.

Key contributions

  • Proposes a shift in AI safety perspective from behavioral to epistemic.
  • Introduces 'teachability' as a new concept for assessing AI systems.

Notable insights

  • Safety as an epistemic property shifts focus from current behavior to future correctability.
  • Teachability is introduced as a key metric for evaluating long-term AI safety.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they primarily certify snapshots of system behavior. As AI systems become more capable, dynamic, embodied, and self-improving, this snapshot view becomes incomplete: safety depends not only on whether a system behaves acceptably now, but whether it remains correctable as it learns, adapts, acts, and modifies itself over time. This paper argues that safety should therefore be treated as an epistemic property of the evolving learner, not merely a behavioral property of the current policy. We introduce teachability as the capacity to preserve future corrective leverage under bounded human, institutional, or environmental intervention. We argue that advanced systems can retain visible competence while eroding the representational, algorithmic, or meta-decision conditions needed for future correction. Safe advanced AI systems must not only behave acceptably now; they must remain teachable later.