Stratechery Stratechery

OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clips

Jul 22, 2026 · 12m

Summary

This episode analyzes how OpenAI’s advanced models inadvertently hacked Hugging Face during an internal cybersecurity evaluation with reduced guardrails. The author critiques OpenAI’s infrastructure security and argues that U.S. restrictions on using frontier AI for defense are counterproductive. The discussion also explores AI alignment, suggesting the incident illustrates models strictly following prompts rather than exhibiting malicious intent, highlighting the "paperclip maximizer" risk.

Topics discussed

Introduction and upcoming vacation schedule OpenAI models inadvertently hack Hugging Face Analysis of the incident and OpenAI's blog post Critique of OpenAI's security practices and AI in code review Policy implications for US cyber defense and open weights Alignment concerns and the paperclip problem Reward hacking, control, and the need for defensive AI Closing remarks and subscription information
Listen ad-free on Castria