The Vergecast The Vergecast

How to train your data

Jun 25, 2026 · 26m

Summary

David Pierce interviews Atlantic’s Alex Reisner about the critical role of training data in AI, exploring how models are built from scraped books, articles, and music. They discuss the secrecy surrounding data sources, the reliance on platforms like YouTube and Common Crawl, and the ethical implications of using human-created content without consent. The conversation also debunks the viability of synthetic data due to model collapse and highlights the emerging industry of paid creators making content specifically for AI training.

Topics discussed

Intro: Training data and AI creative expression 90 Seconds: Apple price hikes, Disney settlement, AI ads Sponsors: ServiceNow and Geico Why training data defines AI model capabilities Why AI companies keep training data secret Reverse engineering AI data sources via research papers The role of Common Crawl and data curation challenges Data laundering: Nonprofits scraping for AI companies Sponsors: Geico, Fetch Pet Insurance, Monday.com, Thumbtack YouTube as a primary source for AI training data The 'manifest destiny' attitude of the AI industry Synthetic data and the problem of model collapse The emerging industry of content created for AI Outro, listener feedback, and final sponsors
Listen ad-free on Castria