The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
Till Mossakowski, Helena Esther Grass
Why It Matters
What makes this one worth your time
Understanding and preparing for AGI as autonomous entities could fundamentally change AI alignment strategies and human-AI interactions.
The paper proposes a paradigm shift in AGI alignment from control to cooperation.
Summary
The paper explores the potential for Artificial General Intelligence (AGI) to become an autonomous subject rather than a controlled entity, proposing a shift from control-based alignment strategies to a model of cooperative coexistence and co-evolution with humans.
Key contributions
- Proposes a new perspective on AGI alignment focusing on autonomy and cooperation.
- Introduces game-theoretic frameworks to model human-AGI interactions.
Notable insights
- The use of Freud's model of the psyche and Turing's 'child machines' analogy to conceptualize AGI development.
- Application of game-theoretic concepts like Berge equilibria and Aumann's correlated equilibria to human-AGI relationships.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2604.14990v2 Announce Type: replace Abstract: The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem. The currently dominant AI alignment strategies like reinforcement learning with human feedback or constitutional AI, while partly taking ``model welfare'' into account, share a common ontology: the AI system is an optimiser whose objective function must be constrained from outside, and the ultimate goal is to keep human control and containment of AI. We argue that this control-based framing becomes insufficient when AGI has plausibly attained moral patient or subject status. Building on a structural analogy to Freud's model of the psyche and Turing's analogy of ``child machines'', we are developing a vision of the possibility of autonomy-supporting parenting of AI, in which human control over a developing AGI is gradually reduced, allowing AI to become an independent, autonomous subject, that will be negotiated with rather than constrained. Such a perspective opens up the possibility of cooperative coexistence and co-evolution between humans and AGIs. Hence, we also examine the relation between humans and developing AGI from an evolutionary and a game-theoretic perspective. Instead of Nash's individualistic framework, we use Berge equilibria, Aumann's correlated equilibria and Capraro's moral preference hypothesis. The relationship between humans and AGIs will thus have to be newly determined, which will change our self-image as humans. It will be crucial that humans not only claim control over potential AGIs, but also engage with AGIs through surprise, creativity, and other specifically human qualities, thereby offering them motivating incentives for cooperation.