The data black hole at the center of AI
Jun 19, 2026 · 11m
Summary
Dwarkesh argues that AI progress relies on massive data scaling rather than improved sample efficiency, as models require millions of times more data than humans to learn. He dismisses objections regarding evolutionary pre-training and scaling laws, noting that current architectures cannot bridge this efficiency gap. Despite this, he suggests AI can still automate white-collar work by amortizing training costs, though solving the remaining sample efficiency problem may require automated AI research.
Topics discussed
Defining intelligence as sample efficiency
The massive scale of expert data and RL training
Why open source models catch up to frontier AI
Comparing human vs AI data consumption
Rebutting the evolution pre-training argument
Addressing the multimodal sensory data objection
Why scaling model size won't fix sample efficiency
Sponsor: Mercury Command AI banking tool
Why sample efficiency matters for automating work
AI research automation and future intelligence
Listen ad-free on Castria