Back to today's list

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

Benjamin Minhao Chen, Xinyu Xie

Published Jun 29, 2026Featured #9In the daily list Jun 30, 2026
Daily score59.9
Editorial review6.8
Relevance0.498
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding how people perceive AI actions versus human actions is crucial for developing AI systems that align with societal values, especially in high-stakes scenarios.

Moral judgments of AI actions differ when human design is highlighted, complicating value alignment.

Summary

The paper investigates how moral judgments differ when evaluating human actions, AI actions, and the actions of AI designers, using a runaway mine train scenario with 1,002 U.S. adults. It finds that moral evaluations change significantly when AI actions are attributed to human design, indicating a divergence in moral judgments that complicates the alignment of AI with human values.

Key contributions

  • Empirical study on moral judgments involving AI and human actors.
  • Identification of the 'alignment target problem' in AI ethics.

Notable insights

  • Moral evaluations become more deontological when AI actions are attributed to human design.
  • Visibility of human agency in AI design activates heightened moral constraints.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2604.24155v3 Announce Type: replace-cross Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically.This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adults measured moral judgments in a runaway mine train scenario, varying the subject of evaluation across four conditions: a repairman, a repair robot, a repair robot programmed by company engineers, and company engineers programming a repair robot. We find no significant difference in evaluations of the repairman and the robot. However, judgments shifted substantially when the robot's actions were described as the product of human design. Participants exhibited markedly more deontological, rule-based reasoning when evaluating either the programmed robot or the engineers who programmed it, suggesting that rendering human agency visible activates heightened moral constraints. These findings indicate that people may evaluate humans, AI systems acting in the same situation, and the humans who design them in meaningfully different ways. The fact that these evaluations do not necessarily converge gives rise to the alignment target problem: which normative target should guide the development of artificial moral agents in high-stakes domains, and whether these plural judgments can be reconciled within a coherent account of value alignment.