Benchmarking Agents on Hard CLI Tasks
Mike A. Merrill, Alexander G Shaw, Nicholas Carlini, et al.
Terminal-Bench 2.0 is a hard benchmark of 89 real-world-inspired CLI tasks where frontier agents score below 65%, with error analysis and public dataset/harness.