From Plausible to Actionable: A Position on LLM Self-Explanations
Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti
Why It Matters
What makes this one worth your time
Understanding the limitations of LLM self-explanations is crucial for developing reliable AI systems that can be trusted in decision-making processes.
The paper critiques LLM self-explanations for their plausibility versus faithfulness and emphasizes their actionable potential.
Summary
The paper discusses the concept of self-explanations generated by large language models, arguing that while these explanations may seem plausible, their faithfulness to the model's reasoning is questionable. It proposes guidelines for evaluating these self-explanations based on plausibility, faithfulness, and actionability.
Key contributions
- Identification of limitations in standard evaluation protocols for LLM-generated self-explanations.
- Proposed practical guidelines for assessing plausibility, faithfulness, and actionability of self-explanations.
Notable insights
- The distinction between plausibility and faithfulness in LLM self-explanations is critical for their practical application.
- Actionability is proposed as a necessary criterion for evaluating self-explanations, which is often overlooked in traditional XAI frameworks.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2607.15957v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior.However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, questionably faithful, and yet highly actionable. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM-generated self-explanations and propose practical guidelines for assessing their plausibility and faithfulness. Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision-making and appropriate action across diverse stakeholders.