Dwarkesh Podcast Dwarkesh Podcast

The data black hole at the center of AI

Jun 19, 2026 · 11m

Summary

Dwarkesh argues that AI progress relies on massive data scaling rather than improved sample efficiency, as models require millions of times more data than humans to learn. He dismisses objections regarding evolutionary pre-training and scaling laws, noting that current architectures cannot bridge this efficiency gap. Despite this, he suggests AI can still automate white-collar work by amortizing training costs, though solving the remaining sample efficiency problem may require automated AI research.

Topics discussed

Defining intelligence as sample efficiency The massive scale of expert data and RL training Why open source models catch up to frontier AI Comparing human vs AI data consumption Rebutting the evolution pre-training argument Addressing the multimodal sensory data objection Why scaling model size won't fix sample efficiency Sponsor: Mercury Command AI banking tool Why sample efficiency matters for automating work AI research automation and future intelligence
Listen ad-free on Castria