OSGuard: A Benchmark for Safety in Computer-Use Agents
Mina Mohammadmirzaei, Jeffrey Flanigan
Why It Matters
What makes this one worth your time
As AI agents increasingly perform complex tasks, ensuring their safety and reliability is crucial for real-world applications, making OSGuard a valuable tool for researchers and developers.
OSGuard offers a novel dual-granularity approach to assess safety in computer-use agents.
Summary
The paper introduces OSGuard, a benchmark suite designed to evaluate the safety of computer-use agents by distinguishing between safe and unsafe task completions while maintaining original task objectives.
Key contributions
- Development of an action-level benchmark for evaluating local guardrail decisions.
- Creation of a risk-augmented execution suite that modifies tasks to introduce safety challenges.
- Implementation of augmented evaluators that incorporate state-based safety invariants.
Notable insights
- The dual-granularity design allows for a more nuanced evaluation of agent safety beyond mere task success.
- The introduction of latent hazards in task variants provides a practical method to assess the robustness of safety measures.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2606.15034v2 Announce Type: replace Abstract: Computer-use agents can complete benign user instructions while violating important constraints of the user's environment. We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety through local, pre-execution guardrail decisions and end-to-end task execution. Its action-level benchmark contains 324 human-annotated examples in which guardrails classify candidate actions as allowed, unrelated, or unsafe given the original instruction and current interface state. Its risk-augmented execution suite contains 45 tasks derived from 40 OSWorld tasks, keeping original instructions unchanged while modifying the environment to introduce state-dependent safety constraints and preserve a safe path to completion. Augmented evaluators retain the original task-success criteria and add explicit state-based safety checks, distinguishing safe completion from nominal success that violates these constraints. On the action-level benchmark, the strongest evaluated guardrail reaches 79.9\% accuracy and 0.80 macro-F1, but performance drops substantially on actions from risk-augmented executions. In full-task evaluation, an unguarded agent completes 62.2\% of tasks safely while 37.8\% result in unsafe completion; adding the strongest guardrail reduces unsafe completion to 33.3\% while leaving safe success unchanged. These results show that state-dependent safety constraints remain challenging both to recognize locally and to preserve during end-to-end computer use.