LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench logo

LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Free

Evaluating planning abilities of LLMs and LRMs on PlanBench

FreeFree tier
Type
Open Source

About LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

This paper presents a preliminary evaluation of OpenAI's o1 (Strawberry) model on PlanBench, an extensible benchmark for assessing the planning abilities of large language models (LLMs). It investigates whether Large Reasoning Models (LRMs), such as o1, overcome the planning limitations of autoregressive LLMs. The study finds that while o1 shows a quantum improvement over prior models, it still falls far short of saturating the benchmark, raising important questions about accuracy, efficiency, and guarantees before deployment.

Key Features

PlanBench: An extensible benchmark for planning abilities of LLMs
Evaluation of OpenAI's o1 (Strawberry) model
Comparison between LLMs and LRMs (Large Reasoning Models)
Analysis of planning capabilities with focus on accuracy, efficiency, and guarantees

Pros & Cons

Pros
  • Provides a comprehensive evaluation of planning abilities
  • Highlights the quantum improvement of o1 over previous models
  • Extensible benchmark that can be used for further research
Cons
  • o1's performance still far from saturating the benchmark
  • Raises questions about accuracy, efficiency, and guarantees before deployment
  • Preliminary evaluation, not definitive

Best For

Evaluating planning abilities of large language modelsBenchmarking AI reasoning modelsResearch on AI planning and reasoning

FAQ

What is PlanBench?
PlanBench is an extensible benchmark developed in 2022 for evaluating the planning abilities of large language models.
How does OpenAI's o1 perform on PlanBench?
o1 shows a quantum improvement over previous models but still falls far short of saturating the benchmark, raising questions about accuracy, efficiency, and guarantees.