The next big breakthrough will be AIs learning on the job
Jun 26, 2026 · 19m
Summary
Dwarkesh Patel analyzes the industry bet that scaling RLVR on verifiable tasks will yield AGI, contrasting this with the slow progress in complex domains like computer use. He argues that true generalization requires continual learning to distill real-world experience into model weights, proposing techniques like on-policy self-distillation and "dreaming" simulations. The episode envisions a future where AI improves primarily through on-the-job learning from broad deployment rather than pre-release training.
Topics discussed
The RLVR bet: Scaling verifiable tasks to achieve AGI
Why computer use lags behind coding and math
The challenge of non-replayable real-world domains
Limits of context windows and the need for weight updates
Current limitations of online learning and sample efficiency
On-policy self-distillation (OPSD) for continual learning
Dreaming: AI rehearsing skills in simulated environments
Future scenario: Broad deployment and on-the-job learning
Listen ad-free on Castria