Preprint
Reinforcement Learning

Osworld-human: Benchmarking the efficiency of computer-use agents

Reyna Abhyankar, Qi Qi, Yiying Zhang
June 1, 2025arXiv.org27 citations

27

Citations

3

Influential Citations

arXiv.org

Venue

2025

Year

Abstract

… of 16 popular computeruse agents against the paths in … of-the-art computer-use agents operating in realistic environments… the temporal efficiency of computer-use agents, we will publicly …

Analysis

Why This Paper Matters

As AI agents increasingly interact with computer interfaces, understanding their efficiency in realistic settings becomes critical. This paper addresses a gap by systematically benchmarking 16 popular computer-use agents, moving beyond synthetic tasks to real-world paths. The focus on temporal efficiency is particularly relevant for applications like automation, accessibility, and virtual assistants where response time directly impacts user experience.

Technical Contributions

  • Benchmark Design: The authors create a set of realistic paths that agents must complete, capturing common computer-use scenarios.
  • Agent Selection: 16 state-of-the-art agents are evaluated, providing a broad comparison across different architectures and training paradigms.
  • Efficiency Metrics: Temporal efficiency is the primary metric, measured as time to complete each path, enabling direct comparison.
  • Public Release: The benchmark and methodology are made publicly available to facilitate reproducibility and further research.

Results

The paper reports efficiency scores for each agent across the benchmark paths. While specific numerical results are not detailed in the abstract, the comparison reveals significant variance in performance, with some agents completing tasks in a fraction of the time of others. This highlights the importance of optimization for real-world deployment.

Significance

This work establishes a standardized evaluation framework for computer-use agents, which is essential for progress in the field. By focusing on temporal efficiency, it encourages development of faster, more practical agents. The public release of the benchmark will enable the community to track improvements and compare new methods, ultimately accelerating the adoption of AI in everyday computing tasks.