ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
342
Citations
83
Influential Citations
arXiv.org
Venue
2026
Year
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .
Terminal-Bench 2.0 addresses a critical gap in AI agent evaluation: the lack of benchmarks that are both realistic and sufficiently difficult to differentiate frontier models. Existing benchmarks often rely on synthetic or simplified tasks that do not reflect the complexity of real-world workflows. By focusing on hard CLI tasks inspired by actual problems, this benchmark pushes the field toward measuring practical, long-horizon capabilities.
The benchmark's design—with unique environments, human-written solutions, and comprehensive tests—ensures that tasks are well-defined and verifiable. The finding that frontier agents score below 65% underscores the benchmark's difficulty and its potential to drive meaningful progress. This is particularly timely as AI agents are expected to autonomously complete valuable tasks in diverse domains.
Frontier models and agents achieve less than 65% accuracy on Terminal-Bench 2.0, demonstrating that the benchmark is challenging even for state-of-the-art systems. The error analysis reveals common failure modes, providing actionable insights for improving agent architectures and training methods. These results establish Terminal-Bench 2.0 as a rigorous benchmark for measuring progress in autonomous CLI task completion.
Terminal-Bench 2.0 has the potential to become a standard evaluation tool for AI agents, similar to how ImageNet influenced computer vision. By focusing on hard, real-world tasks, it encourages the development of agents that can handle long-horizon, multi-step problems. The public availability of the benchmark and harness lowers barriers to entry, fostering broader research and collaboration. As agents improve, this benchmark will help track progress and highlight remaining challenges, ultimately accelerating the deployment of capable autonomous systems in practical settings.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba