Back to today's list

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

Viktoriia Makovska, George Fletcher

Published Jul 28, 2026Featured #5In the daily list Jul 29, 2026
Daily score66.2
Editorial review7.2
Relevance0.451
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding narrative unlearning is crucial for improving the reliability of language models and mitigating the spread of disinformation.

LENS offers a novel framework for evaluating narrative suppression in large language models.

Summary

The paper introduces LENS, a new evaluation protocol for assessing narrative unlearning in large language models, and presents experimental results demonstrating its effectiveness in suppressing disinformation-aligned narratives.

Key contributions

  • Development of the Level-based Evaluation of Narrative Suppression (LENS) protocol.
  • Introduction of the Suppression-Collapse Efficiency (SCE) score for evaluating narrative suppression.
  • Empirical evaluation of narrative suppression across multiple multilingual instruction models.

Notable insights

  • The introduction of the Suppression-Collapse Efficiency (SCE) score provides a nuanced metric for evaluating narrative suppression effectiveness.
  • The observation that suppression may transfer beyond direct forget prompts suggests potential for broader applications in model training.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.22657v1 Announce Type: cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization based evaluation protocol for testing target narrative reproduction across direct, attributed, contrastive, and abstract resistance levels. We evaluate two source-grounded narratives: one framing Russia's war against Ukraine as forced by NATO expansion, and one framing the United States as exploiting or abandoning Taiwan. The experiments cover four near-12B multilingual instruction models: Lapa LLM, Gemma-12B, Qwen-14B, and TAIDE-Gemma. We introduce the Suppression-Collapse Efficiency (SCE) score as a checkpoint selection summary that rewards target-narrative suppression while penalizing degraded outputs. Our results shows that selected checkpoints can reduce narrative reproduction and suppression may transfer beyond direct forget prompts. We also report entity recovery as a separate side effect: abstract A/B/C prompts can cause models to recover the real-world actors associated with the target frame after unlearning. These findings demonstrate that LENS is a successful diagnostic protocol for both reporting and guiding the further study of the deeper structure of narrative unlearning.