Dwarkesh Podcast Dwarkesh Podcast

Eric Jang – Building AlphaGo from scratch

May 15, 2026

Summary

Eric Zhang explains how to rebuild AlphaGo from scratch, detailing the game of Go’s rules and the immense computational complexity of its search tree. He breaks down Monte Carlo Tree Search, including action selection algorithms like PUCT, and demonstrates how neural networks for policy and value prediction make this search tractable. The discussion covers architectural choices, such as ResNets versus Transformers, and the importance of initializing models with human expert data before self-play.

Topics discussed

Introduction and Eric Zhang's background Rules of Go and game mechanics The complexity of the Go search space Monte Carlo Tree Search and UCB1 formula Value networks and pruning the search tree Neural network architectures for Go MCTS selection, expansion, and evaluation steps Backup process and policy refinement Self-play training and the Bitter Lesson Comparison with LLM training and RL variance Advantage functions and Q-learning concepts Scaling laws and test-time compute tradeoffs Robustness, error correction, and robotics Supervised vs. Reinforcement Learning signals Automated AI research and future directions
Listen ad-free on Castria