How to train your data
Jun 25, 2026 · 26m
Summary
David Pierce interviews Atlantic’s Alex Reisner about the critical role of training data in AI, exploring how models are built from scraped books, articles, and music. They discuss the secrecy surrounding data sources, the reliance on platforms like YouTube and Common Crawl, and the ethical implications of using human-created content without consent. The conversation also debunks the viability of synthetic data due to model collapse and highlights the emerging industry of paid creators making content specifically for AI training.
Topics discussed
Intro: Training data and AI creative expression
90 Seconds: Apple price hikes, Disney settlement, AI ads
Sponsors: ServiceNow and Geico
Why training data defines AI model capabilities
Why AI companies keep training data secret
Reverse engineering AI data sources via research papers
The role of Common Crawl and data curation challenges
Data laundering: Nonprofits scraping for AI companies
Sponsors: Geico, Fetch Pet Insurance, Monday.com, Thumbtack
YouTube as a primary source for AI training data
The 'manifest destiny' attitude of the AI industry
Synthetic data and the problem of model collapse
The emerging industry of content created for AI
Outro, listener feedback, and final sponsors
Listen ad-free on Castria