Back to today's list

Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping

Xuhui Zhan, Tyler Derr

Published Oct 9, 2026
Editorial review—
Relevance0.442
Freshness0.427

Why It Matters

What makes this one worth your time

Write-up coming soon.

Summary

Summary coming soon.

Key contributions

  • Nothing noted yet.

Notable insights

  • Nothing noted yet.

Possible limitations

  • Nothing noted yet.

Abstract

arXiv:2508.12466v3 Announce Type: replace-cross Abstract: Connecting pretrained vision and language models usually involves projecting image features into the language model's input space. Inverse-LLaVA reverses this mapping within decoder attention: language states are projected to the visual feature dimension, and modality-specific maps produce residual query, key, and value updates. Fusion and low-rank adaptation (LoRA) learn jointly from 665K visual instructions, with frozen backbones and no separate alignment stage. Across nine primary benchmark evaluations, the final 7B model approaches two-stage LLaVA-1.5 on several tasks. It scores 78.45% on VQAv2 versus 79.13% for official LLaVA-LoRA, and 50.96% versus 48.56% on VizWiz; TextVQA is lower at 56.96% versus 58.47%. Controlled studies examine fusion components, insertion depth, visual features, and language-model size. Representation analysis shows that the text maps preserve much of the pairwise similarity ordering while changing its geometric spread. Additional paired supervision improves celebrity recognition, while instruction replay repairs caption-induced answer-format failures. Analytical cost expressions and fixed-work profiles separate the additional fusion computation from the omitted alignment stage. These findings establish text-to-vision attention fusion as a practical alternative for instruction-only multimodal adaptation.