Preprint
Reinforcement Learning

Benchmarking Agents on Hard CLI Tasks

Mike A. Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, I. Bercovich, Lin Shi, J. Shin, Thomas Walshe, E. K. Buchanan, Junhong Shen, Guanghao Ye, Hao Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, J. Jitsev, Di Lu, Orfeas Menis-Mastromichalakis, Zhiwei Xu, Zizhao Chen, Yue Liu, Robert Zhang, L. Chen, Anurag Kashyap, Jan-Lucas Uslu, Jeffrey Li, Jianbo Wu, Minghao Yan, Song Bian, Vedang Sharma, Ke Sun, S. Dillmann, Akshay Anand, Andrew Lanpouthakoun, Bardia Koopah, Changran Hu, E. Guha, Gabriel H. S. Dreiman, Jiacheng Zhu, Karl Krauth, Li Zhong, Niklas Muennighoff, Robert K. Amanfu, Shangyin Tan, Shreyas Pimpalgaonkar, Tushar Aggarwal, Xia Lin, Xin Lan, Xuandong Zhao, Yiqing Liang, Yuanli Wang, Zilong Wang, Changzhi Zhou, David Heineman, Hange Liu, H. Trivedi, John Yang, Junhong Lin, Manish Shetty, Michael Yang, Nabil Omi, Negin Raoof, Shanda Li, Terry Yue Zhuo, Wu Lin, Yiwei Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Jesse Hu, Christopher Rytting, Ryan Marten, Yixin Wang, A. Dimakis, A. Konwinski, Ludwig Schmidt
January 17, 2026arXiv.org342 citations

342

Citations

83

Influential Citations

arXiv.org

Venue

2026

Year

Abstract

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .

Analysis

Why This Paper Matters

Terminal-Bench 2.0 addresses a critical gap in AI agent evaluation: the lack of benchmarks that are both realistic and sufficiently difficult to differentiate frontier models. Existing benchmarks often rely on synthetic or simplified tasks that do not reflect the complexity of real-world workflows. By focusing on hard CLI tasks inspired by actual problems, this benchmark pushes the field toward measuring practical, long-horizon capabilities.

The benchmark's design—with unique environments, human-written solutions, and comprehensive tests—ensures that tasks are well-defined and verifiable. The finding that frontier agents score below 65% underscores the benchmark's difficulty and its potential to drive meaningful progress. This is particularly timely as AI agents are expected to autonomously complete valuable tasks in diverse domains.

Technical Contributions

  • Curated Task Set: 89 tasks in terminal environments, each inspired by real workflows, ensuring relevance and difficulty.
  • Comprehensive Verification: Each task includes human-written solutions and thorough tests, enabling objective evaluation.
  • Error Analysis: Detailed analysis of agent failures to identify specific areas for improvement, such as command usage, environment interaction, or reasoning.
  • Public Release: Dataset and evaluation harness are made available, facilitating reproducibility and community engagement.

Results

Frontier models and agents achieve less than 65% accuracy on Terminal-Bench 2.0, demonstrating that the benchmark is challenging even for state-of-the-art systems. The error analysis reveals common failure modes, providing actionable insights for improving agent architectures and training methods. These results establish Terminal-Bench 2.0 as a rigorous benchmark for measuring progress in autonomous CLI task completion.

Significance

Terminal-Bench 2.0 has the potential to become a standard evaluation tool for AI agents, similar to how ImageNet influenced computer vision. By focusing on hard, real-world tasks, it encourages the development of agents that can handle long-horizon, multi-step problems. The public availability of the benchmark and harness lowers barriers to entry, fostering broader research and collaboration. As agents improve, this benchmark will help track progress and highlight remaining challenges, ultimately accelerating the deployment of capable autonomous systems in practical settings.