LLMs Still Can’t Plan
Karthik Valmeekam, Kaya Stechly, Subbarao Kambhampati
This paper evaluates current LLMs and new Large Reasoning Models on PlanBench, finding that while OpenAI's o1 shows quantum improvement, it still falls far short of saturating the benchmark.