Back to today's list

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

Swapnanil Mukherjee, Agyeya Negi, Tanuja Ganu, Ponnurangam Kumaraguru

Published Aug 7, 2026Featured #4In the daily list Aug 8, 2026
Daily score67.9
Editorial review7.2
Relevance0.481
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and improving cross-modal reasoning in multimodal models is crucial for developing AI systems that can process and integrate information from diverse sensory inputs effectively.

C$^3$PO benchmark reveals the limitations of current multimodal models in cross-modal reasoning.

Summary

The paper introduces C$^3$PO, a benchmark designed to evaluate cross-modal reasoning in multimodal large language models by testing their ability to compose information across modalities and resolve counterfactual conflicts.

Key contributions

  • Introduction of C$^3$PO, a benchmark for evaluating cross-modal reasoning.
  • Analysis of modality dominance in multimodal models using attention probes.
  • Identification of mid-layer attention entropy as a predictor of model success.

Notable insights

  • Attention probes reveal that modality dominance, particularly text, is a major cause of reasoning failures.
  • Mid-layer attention entropy is a predictor of model performance in cross-modal tasks.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.05381v1 Announce Type: new Abstract: Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities' structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature