Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Max Kanwal, Caryn Tran
Why It Matters
What makes this one worth your time
Understanding and managing the dynamic nature of human preferences in AI interactions could lead to more ethically aligned and user-centered AI systems.
Constructive Alignment redefines AI alignment as managing evolving human preferences.
Summary
The paper introduces a new paradigm called Constructive Alignment, which reframes AI alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. It models preferences as layered state variables that evolve through interaction with AI systems, using a control-theoretic framework to influence both world states and human evaluative states.
Key contributions
- Introduction of Constructive Alignment paradigm.
- Modeling preferences as dynamic, layered state variables.
- Application of control-theoretic framework to preference evolution.
Notable insights
- Preferences are modeled as layered state variables influenced by AI interactions.
- The control-theoretic framework is used to manage the evolution of human preferences.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.00001v1 Announce Type: new Abstract: Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constructed through interaction--particularly with adaptive technologies. As AI systems become more persistent, personalized, and socially embedded, they increasingly participate in shaping what people attend to, value, and endorse over time. We introduce Constructive Alignment, a paradigm that reframes alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. Drawing on behavioral economics, psychology, and constructivist social theory, we model preferences as layered state variables that evolve under interaction with AI systems. We formalize this view using a control-theoretic framework in which system actions and interaction design jointly influence both world states and human evaluative states. We argue that alignment is not primarily about controlling AI behavior, but about regulating how AI systems influence the evolution of human preferences--ensuring that value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. Alignment thus becomes a problem of governing long-term value formation rather than simply satisfying static preferences.