PreprintarXiv.org2026
Benchmarking Agents on Hard CLI Tasks
Mike A. Merrill, Alexander G Shaw, Nicholas Carlini, et al.
Terminal-Bench 2.0 is a hard benchmark of 89 real-world-inspired CLI tasks where frontier agents score below 65%, with error analysis and public dataset/harness.
342Jan 17, 2026Benchmarks
arXiv