BeSafe-Bench: Behavioral Safety Risks of Situated Agents (2026) logo

BeSafe-Bench: Behavioral Safety Risks of Situated Agents (2026)

Free

First benchmark across 4 real functional domains (Web, Mobile, Embodied VLM/VLA) with 9 safety-risk categories; even the best agent completes <40% of tasks under full safety constraints

FreeFree tier
Type
Open Source

About BeSafe-Bench: Behavioral Safety Risks of Situated Agents (2026)

BeSafe-Bench (BSB) is a benchmark designed to expose behavioral safety risks of situated agents operating in functional environments. It covers four representative domains: Web, Mobile, Embodied VLM, and Embodied VLA. The benchmark constructs a diverse instruction space by augmenting tasks with nine categories of safety-critical risks and adopts a hybrid evaluation framework combining rule-based checks with LLM-as-a-judge reasoning to assess real environmental impacts. Evaluations of 13 popular agents reveal that even the best-performing agent completes fewer than 40% of tasks while fully adhering to safety constraints, and strong task performance frequently coincides with severe safety violations. These findings underscore the urgent need for improved safety alignment before deploying agentic systems in real-world settings.

Key Features

Covers 4 functional domains: Web, Mobile, Embodied VLM, Embodied VLA
9 categories of safety-critical risks
Hybrid evaluation framework: rule-based checks + LLM-as-a-judge
Functional environments (not low-fidelity or simulated APIs)
Evaluates 13 popular LMM-based agents
Reveals severe safety violations in strong task-performing agents

Pros & Cons

Pros
  • First comprehensive safety benchmark across multiple real-world domains
  • Uses functional environments rather than low-fidelity simulations
  • Hybrid evaluation (rule-based + LLM-as-a-judge) increases reliability
  • Identifies critical safety gaps in state-of-the-art agents
  • Open access and free (arXiv paper, code expected via links)
Cons
  • Limited to four domains; may not generalize to all real-world scenarios
  • Best-performing agent completes fewer than 40% of tasks safely, indicating high safety risk even for top models
  • Benchmark focuses on behavioral safety risks and does not cover other safety aspects (e.g., data privacy, adversarial robustness)

Best For

Safety benchmarking of situated agentsEvaluating behavioral safety risks in autonomous decision-makingTesting safety alignment of Large Multimodal Models (LMMs)Research on AI safety and deployment readiness

FAQ

What is BeSafe-Bench?
BeSafe-Bench is a benchmark for exposing behavioral safety risks of situated agents in functional environments. It covers Web, Mobile, Embodied VLM, and Embodied VLA domains.
How does BeSafe-Bench evaluate safety?
It uses a hybrid evaluation framework that combines rule-based checks with LLM-as-a-judge reasoning to assess real environmental impacts of agent actions.
What are the key findings from using BeSafe-Bench?
Even the best-performing agent completes fewer than 40% of tasks while fully adhering to safety constraints, and strong task performance frequently coincides with severe safety violations.
Is BeSafe-Bench free and open?
Yes, the paper is freely available on arXiv and the benchmark is expected to be open source. No fees or costs are mentioned.