Back to today's list

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

Yuanyuan Tian, Wenwen Li, Xiao Chen, Michael Brook, Michael Brubaker, Anna Liljedahl, Chitta Baral

Published Jul 2, 2026Featured #7In the daily list Jul 3, 2026
Daily score62.8
Editorial review7.0
Relevance0.473
Freshness0.722

Why It Matters

What makes this one worth your time

This work could improve the ability to link and analyze environmental event reports, aiding in understanding environmental changes and their impacts.

A framework leveraging AI foundation models for enhanced semantic search in geographic information retrieval.

Summary

The paper introduces a framework for semantic search and recommendation of documents with spatial and temporal information using AI foundation models. It proposes two strategies: CAMERA, which combines textual and visual data for richer embeddings, and ASTRA, which enhances similarity ranking by considering spatiotemporal relevance. The framework is tested on a dataset from the Local Environmental Observer Network, showing improved performance over unimodal approaches.

Key contributions

  • Development of the CAMERA algorithm for multimodal event retrieval.
  • Introduction of the ASTRA algorithm for adaptive spatiotemporal re-ranking.

Notable insights

  • Combining textual and visual information can create richer embeddings for document similarity.
  • Incorporating spatiotemporal relevance can improve the effectiveness of similarity ranking.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2606.28369v2 Announce Type: replace-cross Abstract: Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale washed ashore in Alaska) that contain spatial and temporal information, is a critical task in Geographic Information Retrieval (GIR). This work presents a novel framework that leverages AI foundation models, including Large Language Models (LLMs) and Vision-Language Models (VLMs), to enable effective similarity search and ranking for such event documents. To support this goal, we introduce two new strategies: (1) CAMERA (Context-Aware Multimodal Event Retrieval Algorithm), which fuses textual and visual information to generate richer embeddings than those derived from text alone; and (2) ASTRA (Adaptive Spatial and Temporal Re-ranking Algorithm), which improves similarity ranking by incorporating scale-dependent spatiotemporal relevance alongside semantic similarity. Experimental results, using a dataset from the Local Environmental Observer Network, demonstrate that our VLM-enhanced methods outperform unimodal, LLM-based approaches in similarity ranking effectiveness. By automatically linking relevant event reports, the proposed framework helps both data curators and the general public gain deeper insights into environmental change and its localized impacts. These findings highlight the potential of AI foundation models to advance GIR through multifaceted, intelligent analysis that integrates key geographic concepts: space, time, scale, and semantics.