Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
Chunyu Ye, Yunhao Zhang, Jingyuan Sun, Chong Li, Yang Zhao, Shaonan Wang
Why It Matters
What makes this one worth your time
This research could advance brain-computer interfaces by enabling more accurate and flexible decoding of brain activity across multiple modalities, potentially impacting fields like neuroprosthetics and communication aids.
A unified BCI framework aligns brain signals with multimodal semantic spaces for improved brain-to-text translation.
Summary
The paper proposes a unified framework for decoding language from the human brain by leveraging Multimodal Large Language Models to align brain signals with a shared semantic space that includes text, images, and audio. It introduces a router module that dynamically selects and fuses modality-specific brain features, demonstrating state-of-the-art performance on fMRI datasets and extending the approach to EEG and MEG data.
Key contributions
- Introduction of a unified BCI architecture for multimodal brain activity decoding.
- Demonstration of state-of-the-art performance on fMRI datasets with various stimuli.
- Extension of the framework to EEG and MEG data, showing flexibility across different resolutions.
Notable insights
- The use of a router module to dynamically select and fuse modality-specific brain features is a clever approach to handle diverse stimuli.
- Aligning brain signals with a shared semantic space that includes text, images, and audio leverages the brain's multimodal processing capabilities.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2505.10356v3 Announce Type: replace Abstract: Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal processing. Inspired by the brain's associative mechanisms, where viewing an image can evoke related sounds and linguistic representations, we propose a unified framework that leverages Multimodal Large Language Models (MLLMs) to align brain signals with a shared semantic space encompassing text, images, and audio. A router module dynamically selects and fuses modality-specific brain features according to the characteristics of each stimulus. Experiments on various fMRI datasets with textual, visual, and auditory stimuli demonstrate state-of-the-art performance, achieving an 8.48% improvement on the most commonly used benchmark. We further extend our framework to EEG and MEG data, demonstrating flexibility and robustness across varying temporal and spatial resolutions. To our knowledge, this is the first unified BCI architecture capable of robustly decoding multimodal brain activity across diverse brain signals and stimulus types, offering a flexible solution for real-world applications.