BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
Yangxuan Zhou, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan, Sha Zhao
Why It Matters
What makes this one worth your time
This work provides a structured evaluation framework that could enhance the integration of LLMs in EEG analysis, potentially improving clinical and research applications.
BrainBench benchmarks LLMs for advanced EEG analysis.
Summary
The paper introduces BrainBench, a benchmark designed to evaluate large language models' capabilities in comprehensive EEG understanding through various analysis tasks and datasets.
Key contributions
- Introduction of a unified benchmark for comprehensive EEG understanding.
- Evaluation of multiple LLMs across various EEG analysis tasks and datasets.
- Provision of a reproducible testbed for future research in LLM-based EEG analysis.
Notable insights
- The benchmark includes diverse subsets that cover a wide range of EEG analysis tasks, highlighting the multifaceted nature of EEG data interpretation.
- The evaluation methodology incorporates multiple validation metrics, ensuring a comprehensive assessment of model performance.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.04156v2 Announce Type: replace Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.