Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm logo

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Free

A general RL algorithm mastering chess, shogi, and Go through self-play

FreeFree tier
Type
Open Source

About Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

AlphaZero is a general reinforcement learning algorithm that achieves superhuman performance in the games of chess, shogi (Japanese chess), and Go entirely through self-play, without any domain-specific knowledge beyond the game rules. Starting from random play, it learns tabula rasa within 24 hours and convincingly defeats world-champion programs in each domain. The algorithm generalizes the AlphaGo Zero approach into a single architecture that can master multiple challenging board games.

Key Features

General reinforcement learning algorithm applicable to multiple board games
Tabula rasa learning from self-play with no human knowledge except game rules
Achieves superhuman performance within 24 hours starting from random play
Defeats world-champion programs in chess, shogi, and Go
Single architecture for multiple challenging domains

Pros & Cons

Pros
  • Learns without any human-crafted features or domain expertise
  • Quickly achieves superhuman performance (within 24 hours)
  • General algorithm that works across multiple games without modification
  • Defeats state-of-the-art specialized programs in chess, shogi, and Go
Cons
  • Requires substantial computational resources for self-play and training
  • Currently limited to deterministic, perfect-information board games
  • Not yet demonstrated for real-world or partially observable domains

Best For

Game AI research and developmentBenchmark for reinforcement learning algorithmsStudy of self-play and tabula rasa learningDemonstration of general-purpose game-playing agents

FAQ

What games does AlphaZero master?
AlphaZero masters chess, shogi (Japanese chess), and Go.
How does AlphaZero learn?
It learns entirely through self-play reinforcement learning, starting from random play with no domain knowledge except the game rules.
How long does it take to achieve superhuman performance?
AlphaZero achieves superhuman performance within 24 hours of training.
Does AlphaZero use any human-designed heuristics or opening books?
No, AlphaZero learns tabula rasa using only the game rules and neural network self-play.