Dwarkesh Podcast Dwarkesh Podcast

The next big breakthrough will be AIs learning on the job

Jun 26, 2026 · 19m

Summary

The episode critiques the current AI reliance on Reinforcement Learning from Verifiable Rewards (RLVR), arguing it fails in complex, non-simulatable real-world tasks due to sample inefficiency. It advocates for continual learning via techniques like On-Policy Self-Distillation (OPSD) and "dreaming" to distill in-context insights into model weights. This shift would enable AIs to learn from diverse, unstructured real-world interactions, moving beyond static training toward dynamic, on-the-job improvement.

Topics discussed

The RLVR bet: Scaling verifiable tasks to achieve AGI Why computer use lags: The need for replayable simulators The challenge of non-stationary real-world domains Limits of context windows and the need for weight updates Sample efficiency in online learning and human analogy On-policy self-distillation (OPSD) for continual learning Dreaming: Test-time training via internal simulation Future scenario: AI improving through broad deployment
Listen ad-free on Castria