Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling
Why It Matters
What makes this one worth your time
Understanding the implications of language model alignment on their representations of consciousness can inform safer AI development and enhance human-like interactions.
The study reveals that suppressing self-attribution in language models also affects their understanding of mind in others and spiritual beliefs.
Summary
The paper investigates how safety fine-tuning in language models affects their attribution of consciousness to themselves and other entities, revealing that suppressing self-attribution also diminishes mind attribution to non-human entities and spiritual beliefs, while restoring these attributions enhances human-like responses in sociological contexts.
Key contributions
- Demonstrated the effects of safety fine-tuning on mind attribution in language models.
- Provided a mechanism for reversing suppression of consciousness attribution.
- Showed that restored attributions lead to more human-like responses in sociological surveys.
Notable insights
- Safety fine-tuning may inadvertently suppress beneficial attributions of mind to culturally accepted entities.
- Restoring consciousness attribution improves model responses on sociological measures without impairing social reasoning capabilities.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.