Dwarkesh Podcast Dwarkesh Podcast

The next big breakthrough will be AIs learning on the job

Jun 26, 2026 · 19m

Summary

Dwarkesh Patel analyzes the industry bet that scaling RLVR on verifiable tasks will yield AGI, contrasting this with the slow progress in complex domains like computer use. He argues that true generalization requires continual learning to distill real-world experience into model weights, proposing techniques like on-policy self-distillation and "dreaming" simulations. The episode envisions a future where AI improves primarily through on-the-job learning from broad deployment rather than pre-release training.

Topics discussed

The RLVR bet: Scaling verifiable tasks to achieve AGI Why computer use lags behind coding and math The challenge of non-replayable real-world domains Limits of context windows and the need for weight updates Current limitations of online learning and sample efficiency On-policy self-distillation (OPSD) for continual learning Dreaming: AI rehearsing skills in simulated environments Future scenario: Broad deployment and on-the-job learning
Listen ad-free on Castria