Scalable Online Planning via Reinforcement Learning Fine-Tuning logo

Scalable Online Planning via Reinforcement Learning Fine-Tuning

Free
FreeFree tier
Type
Open Source

About Scalable Online Planning via Reinforcement Learning Fine-Tuning

Replaces tabular search methods with online model-based fine-tuning of a policy neural network via reinforcement learning. Demonstrates state-of-the-art results in self-play Hanabi and outperforms tabular search in Atari's Ms. Pacman. Aims to scale planning in stochastic and partially observable environments.

Key Features

Replaces tabular search with online RL fine-tuning of policy networks
Scalable to large search spaces in stochastic and partially observable environments
Achieves state-of-the-art results in self-play Hanabi
Outperforms tabular search in Atari Ms. Pacman
General approach applicable to various planning domains

Pros & Cons

Pros
  • Achieves state-of-the-art results in benchmark games
  • More scalable than traditional tabular search methods
  • General methodology that can be applied to different domains

Best For

Game AI (Hanabi, Ms. Pacman)Planning under stochasticity and partial observabilityScalable online planning in complex environments

FAQ

What is the main contribution of this paper?
The paper proposes replacing tabular search with online model-based fine-tuning of a policy network via reinforcement learning to achieve scalable planning in stochastic and partially observable environments.
What games were tested?
The method was tested on self-play Hanabi and the Atari game Ms. Pacman.
Does this method outperform existing search algorithms?
Yes, it achieves a new state-of-the-art result in self-play Hanabi and outperforms tabular search in Ms. Pacman.