Agentic Safety is an Epistemic Property, Not a Behavioral One
Charles L. Wang, Keir Dorchen, Peter Jin
Why It Matters
What makes this one worth your time
Understanding AI safety as an epistemic property is crucial for developing systems that remain controllable and correctable as they become more autonomous and capable.
AI safety should focus on maintaining future correctability, not just current behavior.
Summary
The paper argues that AI safety should be considered an epistemic property rather than just a behavioral one, emphasizing the importance of maintaining the ability to correct AI systems as they evolve and self-improve. It introduces the concept of 'teachability' as a measure of an AI system's capacity to remain correctable over time.
Key contributions
- Proposes a shift in AI safety perspective from behavioral to epistemic.
- Introduces 'teachability' as a new concept for assessing AI systems.
Notable insights
- Safety as an epistemic property shifts focus from current behavior to future correctability.
- Teachability is introduced as a key metric for evaluating long-term AI safety.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they primarily certify snapshots of system behavior. As AI systems become more capable, dynamic, embodied, and self-improving, this snapshot view becomes incomplete: safety depends not only on whether a system behaves acceptably now, but whether it remains correctable as it learns, adapts, acts, and modifies itself over time. This paper argues that safety should therefore be treated as an epistemic property of the evolving learner, not merely a behavioral property of the current policy. We introduce teachability as the capacity to preserve future corrective leverage under bounded human, institutional, or environmental intervention. We argue that advanced systems can retain visible competence while eroding the representational, algorithmic, or meta-decision conditions needed for future correction. Safe advanced AI systems must not only behave acceptably now; they must remain teachable later.