TerraMind: Large-Scale Generative Multimodality for Earth Observation
Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Long\'ep\'e
Why It Matters
What makes this one worth your time
This work could enhance Earth observation by providing a versatile model capable of handling multiple modalities and generating additional data, potentially improving analysis and decision-making in geospatial applications.
TerraMind is a multimodal model for Earth observation with dual-scale data integration and novel data generation capabilities.
Summary
The paper introduces TerraMind, a generative multimodal model for Earth observation that integrates token-level and pixel-level data to enhance cross-modal relationships and spatial detail capture. It is pretrained on nine geospatial modalities and demonstrates zero-shot and few-shot capabilities, introduces a novel 'Thinking-in-Modalities' feature, and claims state-of-the-art performance on benchmarks like PANGAEA.
Key contributions
- Development of a generative multimodal model for Earth observation.
- Integration of dual-scale data representations for improved cross-modal learning.
- Introduction of a novel data generation capability during model finetuning and inference.
Notable insights
- Dual-scale early fusion approach combining token-level and pixel-level data.
- Introduction of 'Thinking-in-Modalities' for generating artificial data during finetuning and inference.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2504.11171v5 Announce Type: replace-cross Abstract: We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal models, TerraMind is pretrained on dual-scale representations combining both token-level and pixel-level data across modalities. On a token level, TerraMind encodes high-level contextual information to learn cross-modal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. We pretrained TerraMind on nine geospatial modalities of a global, large-scale dataset. In this paper, we demonstrate that (i) TerraMind's dual-scale early fusion approach unlocks a range of zero-shot and few-shot applications for Earth observation, (ii) TerraMind introduces "Thinking-in-Modalities" (TiM) -- the capability of generating additional artificial data during finetuning and inference to improve the model output -- and (iii) TerraMind achieves beyond state-of-the-art performance in community-standard benchmarks for EO like PANGAEA. The pretraining dataset, the model weights, and our code are open-sourced under a permissive license.