AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
Yizhe Chi, Wenyi Li, Deyao Hong, et al.
AI4AI-Bench benchmarks LLM agents on rewriting training algorithms across 10 repositories, finding best agents close only a fifth of the gap to optimal.