Back to today's list

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu, Stephanie Lukin, Cynthia Matuszek

Published Jul 22, 2026Featured #4In the daily list Jul 23, 2026
Daily score71.4
Editorial review7.5
Relevance0.454
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the difference between hazards and anomalies is crucial for improving the reliability of VLMs in safety-critical applications, which can directly impact emergency decision-making.

This research clarifies the distinction between hazards and anomalies in VLM evaluations for enhanced safety reasoning.

Summary

The paper evaluates Vision-Language Models (VLMs) in the context of safety-critical systems, distinguishing between hazards and anomalies to improve safety reasoning assessments.

Key contributions

  • Introduced a clear distinction between hazard and anomaly in VLM evaluations.
  • Evaluated state-of-the-art VLMs across multiple datasets and prompting strategies.
  • Provided a public dataset for further research in anomaly and hazard detection.

Notable insights

  • VLMs often conflate anomalous elements with hazards, indicating a potential flaw in their safety reasoning capabilities.
  • The introduction of a separate evaluation framework for hazards and anomalies could lead to more nuanced assessments of model performance.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.18325v1 Announce Type: cross Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds.