LPS-Bench: Long-Horizon Safety Benchmarking for Computer-Use Agents (2026) logo

LPS-Bench: Long-Horizon Safety Benchmarking for Computer-Use Agents (2026)

Free

Safety benchmark for browser/computer-use agents focused on long-horizon tasks where risk accumulates across many UI actions — useful for testing confirmation discipline, phishing resistance, and context drift

FreeFree tier
Type
Open Source

About LPS-Bench: Long-Horizon Safety Benchmarking for Computer-Use Agents (2026)

LPS-Bench is a benchmark for evaluating the planning-time safety awareness of Model Context Protocol (MCP)-based computer-use agents (CUAs) in long-horizon tasks. It addresses a critical gap in existing benchmarks that focus on short-horizon or GUI-based tasks and overlook the ability to anticipate risks before execution. The benchmark covers 65 scenarios across 7 task domains and 9 risk types, including both benign and adversarial interactions. It uses a multi-agent automated pipeline for scalable data generation and an LLM-as-a-judge evaluation protocol to assess safety awareness through planning trajectories. Designed for testing confirmation discipline, phishing resistance, and context drift, LPS-Bench reveals substantial deficiencies in current CUAs and proposes mitigation strategies. The code is open-source and publicly available.

Key Features

Evaluates planning-time safety awareness of MCP-based computer-use agents
Covers 65 scenarios across 7 task domains and 9 risk types
Includes both benign and adversarial interactions
Uses multi-agent automated pipeline for scalable data generation
Employs LLM-as-a-judge evaluation protocol on planning trajectories
Open-source code available

Pros & Cons

Pros
  • Focuses on planning-time safety awareness, not just execution-time errors
  • Comprehensive coverage: 65 scenarios, 7 domains, 9 risk types
  • Includes adversarial scenarios to test robustness
  • Multi-agent pipeline enables scalable benchmark generation
  • LLM-as-a-judge provides automated evaluation of safety awareness
  • Open-source and publicly available

Best For

Benchmarking safety awareness of computer-use agents in long-horizon tasksEvaluating planning-time risk anticipation in MCP-based agentsTesting agents against adversarial manipulations and ambiguous instructionsResearch on improving safety of automated computer interaction agents

FAQ

What is LPS-Bench?
LPS-Bench is a benchmark for evaluating the safety awareness of computer-use agents (CUAs) in long-horizon planning tasks, focusing on planning-time risk anticipation.
What types of scenarios are covered?
It covers 65 scenarios across 7 task domains and 9 risk types, including both benign and adversarial interactions.
Is the code available?
Yes, the code is open-source and available at the link provided in the paper.
How does LPS-Bench evaluate safety?
It uses a multi-agent automated pipeline for data generation and an LLM-as-a-judge protocol to assess safety awareness through planning trajectories.