Back to today's list

Position: Align AI to Our Aspirations, Not Our Flaws

Nikita Kazeev, Bui Nhat Huyen Phan

Published Jun 27, 2026
Editorial review6.5
Relevance0.497
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding how to align AI with human values is crucial for developing systems that are beneficial and safe, making this paper relevant for researchers focused on AI ethics and alignment.

The paper argues for aligning AI with foundational values over pluralistic human preferences.

Summary

The paper critiques the approach of aligning AI with aggregated human preferences, arguing instead for a foundational alignment based on competence, factual accuracy, honesty, and lawfulness. It proposes a framework that respects pluralistic values only when they do not violate these foundational goals and addresses several objections to this approach.

Key contributions

  • Proposes a framework for AI alignment based on foundational values rather than aggregated human preferences.
  • Engages with six objections to the proposed alignment framework, providing a comprehensive discussion on its feasibility and implications.

Notable insights

  • The paper suggests a non-negotiable floor of objective alignment goals for AI, which includes competence and adherence to factual accuracy, honesty, and lawfulness.
  • It highlights the potential dangers of aligning AI with unfiltered pluralistic human values.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2606.13755v2 Announce Type: replace-cross Abstract: We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the values of a Silicon Valley techno-optimist, a degrowth environmentalist, a national-conservative culture warrior, a single-party state cadre, or a devout religious traditionalist. We should not. Human values produce societies that thrive or fail on the merits of those values - from failed states and extreme inequality to declining happiness, political polarization, and government dysfunction in the world's wealthiest democracies. The pluralistic-alignment program correctly diagnoses that there is no single "humanity" to align with, but is dangerous if taken as the main directive. We argue that AI should be trained to a non-negotiable floor of objective alignment goals - competence, bounded by the constraints of factual accuracy, honesty, and lawfulness and that pluralism belongs at the surface (language, register, conventions, missing-context defaults) and across the wide band of legitimate value tradeoffs that respect the floor, but not at the level of values that violate it. We highlight the empirical reality of unfiltered pluralistic values, propose four commitments as a constructive alternative, and engage six credible objections: commercial pressure and practical feasibility, democratic legitimacy, regulatory compliance, over-reliance on institutionalist explanations, the charge that the floor itself is culturally laden, and the limits of Coherent Extrapolated Volition.