Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
FreeA general RL algorithm mastering chess, shogi, and Go through self-play
FreeFree tier
About Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
AlphaZero is a general reinforcement learning algorithm that achieves superhuman performance in the games of chess, shogi (Japanese chess), and Go entirely through self-play, without any domain-specific knowledge beyond the game rules. Starting from random play, it learns tabula rasa within 24 hours and convincingly defeats world-champion programs in each domain. The algorithm generalizes the AlphaGo Zero approach into a single architecture that can master multiple challenging board games.
Key Features
General reinforcement learning algorithm applicable to multiple board games
Tabula rasa learning from self-play with no human knowledge except game rules
Achieves superhuman performance within 24 hours starting from random play
Defeats world-champion programs in chess, shogi, and Go
Single architecture for multiple challenging domains
Pros & Cons
Pros
- Learns without any human-crafted features or domain expertise
- Quickly achieves superhuman performance (within 24 hours)
- General algorithm that works across multiple games without modification
- Defeats state-of-the-art specialized programs in chess, shogi, and Go
Cons
- Requires substantial computational resources for self-play and training
- Currently limited to deterministic, perfect-information board games
- Not yet demonstrated for real-world or partially observable domains
Best For
Game AI research and developmentBenchmark for reinforcement learning algorithmsStudy of self-play and tabula rasa learningDemonstration of general-purpose game-playing agents
FAQ
What games does AlphaZero master?
AlphaZero masters chess, shogi (Japanese chess), and Go.
How does AlphaZero learn?
It learns entirely through self-play reinforcement learning, starting from random play with no domain knowledge except the game rules.
How long does it take to achieve superhuman performance?
AlphaZero achieves superhuman performance within 24 hours of training.
Does AlphaZero use any human-designed heuristics or opening books?
No, AlphaZero learns tabula rasa using only the game rules and neural network self-play.