The next big breakthrough will be AIs learning on the job
Jun 26, 2026 · 19m
Summary
The episode critiques the current AI reliance on Reinforcement Learning from Verifiable Rewards (RLVR), arguing it fails in complex, non-simulatable real-world tasks due to sample inefficiency. It advocates for continual learning via techniques like On-Policy Self-Distillation (OPSD) and "dreaming" to distill in-context insights into model weights. This shift would enable AIs to learn from diverse, unstructured real-world interactions, moving beyond static training toward dynamic, on-the-job improvement.
Topics discussed
The RLVR bet: Scaling verifiable tasks to achieve AGI
Why computer use lags: The need for replayable simulators
The challenge of non-stationary real-world domains
Limits of context windows and the need for weight updates
Sample efficiency in online learning and human analogy
On-policy self-distillation (OPSD) for continual learning
Dreaming: Test-time training via internal simulation
Future scenario: AI improving through broad deployment
Listen ad-free on Castria