The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
Francis Heylighen
Why It Matters
What makes this one worth your time
Understanding the nature of values in AI systems is crucial for developing safe and aligned AI technologies, especially as concerns about their potential risks grow.
The paper argues that LLMs do not possess intrinsic values or motivations that could lead to existential risks.
Summary
The paper explores the evolutionary origins of values in biological organisms to address concerns about AI alignment and existential risks associated with Large Language Models (LLMs), arguing that LLMs lack intrinsic motivations for self-preservation and dominance.
Key contributions
- Introduces a biological perspective on the origins of values to inform AI alignment discussions.
- Clarifies the implications of LLMs being allopoietic and allotelic for existential risk assessments.
Notable insights
- The distinction between autopoietic and allopoietic systems provides a framework for understanding the limitations of LLMs in terms of agency and motivation.
- The paper challenges the orthogonality thesis by suggesting that LLMs' lack of intrinsic goals complicates the relationship between intelligence and values.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.03361v1 Announce Type: cross Abstract: AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.