Back to today's list

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

Tingyu Li, Le Zhou, Siyuan Li, Yujun Wu, Xinglong Xu, Jingxuan Wei, Conghui He, Cheng Tan

Published Jul 10, 2026Featured #9In the daily list Jun 13, 2026
Daily score64.4
Editorial review7.2
Relevance0.470
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and improving modality transitions can lead to more coherent and effective multimodal models, which are crucial for tasks requiring integrated reasoning across different data types.

MoTiF enhances interleaved reasoning by supervising modality transitions to reduce information loss.

Summary

The paper addresses the issue of modal isolation in interleaved thinking models by introducing a framework called MoTiF, which supervises modality transitions to improve cross-modal coherence and task accuracy.

Key contributions

  • Introduction of modality transition loss to quantify cross-modal issues.
  • Development of the MoTiF framework for supervising modality transitions.
  • Demonstration of improved cross-modal coherence and task accuracy on visual puzzle benchmarks.

Notable insights

  • The concept of modality transition loss quantifies cross-modal hallucination and visual utilization deficit.
  • MoTiF uses transition-level fidelity rather than end-task accuracy for training signals.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2606.12886v2 Announce Type: replace-cross Abstract: Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spatial and physical tasks. However, in complex long-chain scenarios, we identify a fundamental failure mode: generated images diverge from the textual context while subsequent text ignores the visual evidence, causing the two modalities to alternate without genuinely informing each other. We term this Modal Isolation and attribute it to compounding information loss at modality boundaries. We decompose each reasoning cycle into atomic operations and define modality transition loss, quantifying cross-modal hallucination (text-to-image) and visual utilization deficit (image-to-text) at each boundary. We propose MoTiF (Modality Tiransition Fidelity), a two-stage training framework that directly optimizes these transitions: Reflective SFT trains the model to detect and recover from erroneous visual outputs; Flow-GRPO improves image generation fidelity via reinforcement learning. All training signals in MoTiF derive from transition-level fidelity rather than end-task accuracy. Across four visual puzzle benchmarks, this transition-level supervision substantially improves both cross-modal coherence and final task accuracy. The results demonstrate that effective interleaved reasoning requires explicit structural supervision at modality boundaries, not merely scaling or end-task optimization.