RiskWebWorld: GUI Agents in E-commerce Risk Management (April 2026) logo

RiskWebWorld: GUI Agents in E-commerce Risk Management (April 2026)

Free

Realistic interactive benchmark for GUI agents in high-stakes professional workflows — 100 real-world e-commerce risk scenarios testing sequential decision-making under uncertainty

FreeFree tier
Type
Open Source

About RiskWebWorld: GUI Agents in E-commerce Risk Management (April 2026)

RiskWebWorld is the first highly realistic interactive benchmark for evaluating GUI agents in e-commerce risk management. It features 1,513 tasks sourced from production risk-control pipelines across 8 core domains, capturing authentic challenges such as uncooperative websites and partial environmental hijackments. The benchmark includes a Gymnasium-compliant infrastructure that decouples policy planning from environment mechanics, supporting scalable evaluation and agentic reinforcement learning. Evaluations reveal a significant capability gap, with top generalist models achieving 49.1% success while specialized open-weights models fail nearly entirely. Agentic RL training on this infrastructure improves open-source models by 16.2%, positioning RiskWebWorld as a practical testbed for developing robust digital workers.

Key Features

First highly realistic interactive benchmark for GUI agents in e-commerce risk management
1,513 tasks sourced from production risk-control pipelines across 8 core domains
Captures authentic challenges of risk operations on uncooperative websites and partial environmental hijackments
Gymnasium-compliant infrastructure decoupling policy planning from environment mechanics
Supports scalable evaluation and agentic reinforcement learning (RL)
Evaluated across diverse models; top generalist models achieve 49.1% success rate
Agentic RL improves open-source models by 16.2%

Pros & Cons

Pros
  • Realistic and production-sourced tasks covering diverse risk domains
  • Includes uncooperative websites and environmental hijackments not seen in consumer benchmarks
  • Open-source and freely available (arXiv paper and associated code)
  • Supports agentic RL for iterative improvement of agents
  • Provides a standardized infrastructure for reproducible evaluation
Cons
  • High difficulty – even top generalist models only achieve 49.1% success
  • Specialized open-weights GUI agents perform near total failure without additional tuning
  • Focused specifically on e-commerce risk management, limiting general applicability
  • Requires substantial compute for agentic RL training

Best For

Evaluating GUI agents in high-stakes e-commerce risk managementBenchmarking sequential decision-making under uncertainty in professional workflowsTraining and improving open-source GUI agents via reinforcement learningDeveloping robust digital workers for risk operations on uncooperative websites

FAQ

What is RiskWebWorld?
RiskWebWorld is a realistic interactive benchmark for evaluating GUI agents in e-commerce risk management, featuring 1,513 tasks from production pipelines across 8 core domains.
How many tasks does RiskWebWorld contain?
It contains 1,513 tasks sourced from production risk-control pipelines.
What infrastructure does RiskWebWorld use for agent training?
It provides a Gymnasium-compliant infrastructure that decouples policy planning from environment mechanics, supporting scalable evaluation and agentic reinforcement learning.
What are the key findings from evaluating models on RiskWebWorld?
Top generalist models achieve 49.1% success, while specialized open-weights GUI agents fail nearly entirely. Agentic RL on this infrastructure improves open-source models by 16.2%.