Eric Jang – Building AlphaGo from scratch
May 15, 2026
Summary
Eric Zhang explains how to rebuild AlphaGo from scratch, detailing the game of Go’s rules and the immense computational complexity of its search tree. He breaks down Monte Carlo Tree Search, including action selection algorithms like PUCT, and demonstrates how neural networks for policy and value prediction make this search tractable. The discussion covers architectural choices, such as ResNets versus Transformers, and the importance of initializing models with human expert data before self-play.
Topics discussed
Introduction and Eric Zhang's background
Rules of Go and game mechanics
The complexity of the Go search space
Monte Carlo Tree Search and UCB1 formula
Value networks and pruning the search tree
Neural network architectures for Go
MCTS selection, expansion, and evaluation steps
Backup process and policy refinement
Self-play training and the Bitter Lesson
Comparison with LLM training and RL variance
Advantage functions and Q-learning concepts
Scaling laws and test-time compute tradeoffs
Robustness, error correction, and robotics
Supervised vs. Reinforcement Learning signals
Automated AI research and future directions
Listen ad-free on Castria