LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
FreeEvaluating planning abilities of LLMs and LRMs on PlanBench
FreeFree tier
About LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
This paper presents a preliminary evaluation of OpenAI's o1 (Strawberry) model on PlanBench, an extensible benchmark for assessing the planning abilities of large language models (LLMs). It investigates whether Large Reasoning Models (LRMs), such as o1, overcome the planning limitations of autoregressive LLMs. The study finds that while o1 shows a quantum improvement over prior models, it still falls far short of saturating the benchmark, raising important questions about accuracy, efficiency, and guarantees before deployment.
Key Features
PlanBench: An extensible benchmark for planning abilities of LLMs
Evaluation of OpenAI's o1 (Strawberry) model
Comparison between LLMs and LRMs (Large Reasoning Models)
Analysis of planning capabilities with focus on accuracy, efficiency, and guarantees
Pros & Cons
Pros
- Provides a comprehensive evaluation of planning abilities
- Highlights the quantum improvement of o1 over previous models
- Extensible benchmark that can be used for further research
Cons
- o1's performance still far from saturating the benchmark
- Raises questions about accuracy, efficiency, and guarantees before deployment
- Preliminary evaluation, not definitive
Best For
Evaluating planning abilities of large language modelsBenchmarking AI reasoning modelsResearch on AI planning and reasoning
FAQ
What is PlanBench?
PlanBench is an extensible benchmark developed in 2022 for evaluating the planning abilities of large language models.
How does OpenAI's o1 perform on PlanBench?
o1 shows a quantum improvement over previous models but still falls far short of saturating the benchmark, raising questions about accuracy, efficiency, and guarantees.