Back to today's list

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

Gabriele La Malfa, Nitay Alon, Emanuele La Malfa, Reuth Mirsky, Stefan Sarkadi

Published Aug 11, 2026
Editorial review6.8
Relevance0.476
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding the limitations of VLMs in safety-critical scenarios is crucial for developing AI systems that can be trusted in real-world applications where human safety is at stake.

The paper critiques VLMs' readiness for safety-critical scenarios, highlighting their reliance on textual priors over spatial evidence.

Summary

The paper evaluates the spatial understanding and decision-making capabilities of Vision-Language Models (VLMs) in safety-critical scenarios using an extended Theory of Space framework called Explore, Map, Remember, and Decide (EMRD). It assesses VLMs' exploration competence, spatial fidelity, memory persistence, and cognitive decision-making, finding that VLMs often rely on textual priors rather than spatial grounding and that their memory diverges from human cognition.

Key contributions

  • Extension of the Theory of Space framework into a safety-critical pipeline called EMRD.
  • Evaluation of VLMs' decision-making capabilities in terms of exploration competence, spatial fidelity, memory persistence, and cognitive decision-making.

Notable insights

  • VLMs frequently select evacuation points based on pre-trained textual priors rather than spatial evidence.
  • Spatial reasoning in VLMs degrades in low-light conditions but is unaffected by texture and color tampering.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.08077v1 Announce Type: new Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly applied to safety-critical scenarios, it is crucial to understand whether VLMs possess robust spatial memory and make reliable decisions. In this paper, we assess whether VLMs' decisions are based on physical evidence or are corrupted by visual-language biases, if their memory processes align with human cognitive patterns, and how they respond to environmental hazards. We extend the ToS framework into a safety-critical, goal-driven pipeline, named Explore, Map, Remember, and Decide (EMRD). We then quantify Exploration Competence (Explore) through metrics of environmental coverage and temporal efficiency, assess Spatial Fidelity (Map), evaluate, with a suite of psychological metrics, Memory Persistence (Remember), and measure, using focal-point metrics, Cognitive Decision-Making (Decide). Our results show that in terms of decision-making capabilities, VLMs frequently select evacuation points based on pre-trained textual priors while lacking the spatial grounding to justify their choices. We also show that spatial reasoning degrades in low-light conditions, but it is not affected by texture and colour tampering. Our findings suggest that VLM memory fundamentally diverges from human cognition, creating unpredictable risks of misalignment.