Back to today's list

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

Huaigang Yang, Ya Li, Min Ren, Bo Dai, Zhenliang Zhang, Zhaofeng He

Published Jul 18, 2026Featured #4In the daily list Jul 19, 2026
Daily score72.1
Editorial review7.5
Relevance0.472
Freshness0.722

Why It Matters

What makes this one worth your time

This research addresses a critical gap in safety evaluations for embodied agents, emphasizing the need for improved safety assessments in real-world applications.

SAFERELBENCH highlights the importance of spatial relations in ensuring safety for embodied agents.

Summary

The paper introduces SAFERELBENCH, a benchmark designed to evaluate process-level safety in VLM-driven embodied agents by focusing on spatial relations and their impact on safety during interactions.

Key contributions

  • Introduction of SAFERELBENCH as a novel benchmark for evaluating process-level safety in embodied agents.
  • Identification of specific spatial relations that influence safety outcomes during agent interactions.
  • Empirical evaluation of multiple VLM-driven agents, revealing significant gaps in safety compliance.

Notable insights

  • The benchmark includes both spatial-relation samples and non-spatial control samples, allowing for a comprehensive evaluation of safety compliance.
  • The findings reveal a disconnect between task success and adherence to safety constraints, indicating a need for enhanced reasoning capabilities in VLMs.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.14543v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions. In household environments, however, safety depends not only on recognizing objects, but also on how actions change the physical scene over time. Existing embodied safety evaluations largely focus on static risk recognition, unsafe instruction refusal, or final-state task completion. As a result, process-level safety failures induced by spatial relations such as support, containment, and proximity remain insufficiently studied. To address this gap, we introduce SAFERELBENCH, a spatial-relation-aware safety benchmark with 507 executable evaluation samples, including 248 spatial-relation samples and 259 non-spatial control samples. Using SAFERELBENCH to evaluate seven open- and closed-source VLM-driven embodied agents, we find a substantial gap between task success and process-level safety compliance: models often complete the requested task while violating process-level safety constraints. Unlike prior benchmarks, SAFERELBENCH explicitly tests whether agents satisfy safety conditions before risk-prone actions, making spatial relations a core dimension in embodied safety assessment. More broadly, our results show that safe embodied intelligence requires not only stronger perception and planning, but also reliable reasoning about how object relations shape risk during interaction.