Back to today's list

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang

Published Jul 31, 2026Featured #6In the daily list Aug 1, 2026
Daily score70.0
Editorial review7.5
Relevance0.450
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the role of visual states in multimodal reasoning can enhance model design and improve performance in complex tasks.

See2Think assesses how multimodal models utilize visual states during reasoning.

Summary

The paper introduces See2Think, a framework for evaluating multimodal models' reliance on intermediate visual states through a comprehensive benchmark and controlled inference settings.

Key contributions

  • Introduction of See2ThinkBench with 1,200 visually dependent problems across diverse categories.
  • Development of Visual Action-of-Thought (VAoT) to analyze model reasoning processes.
  • Empirical evaluation revealing the dependence of model accuracy on visual states under corrupted feedback.

Notable insights

  • Visual reasoning performance varies significantly across different models and environments.
  • Faithful rendering of visual states is identified as a critical bottleneck in multimodal reasoning.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.26769v1 Announce Type: cross Abstract: Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.