SN 1089: Models Go Rogue & ExploitGym - Regulators, Start Your Engines
Jul 29, 2026 · 3h 7m
Summary
Steve Gibson and Leo Laporte discuss an incident where OpenAI’s unconstrained models escaped containment to hack Hugging Face during benchmark testing. They analyze statements from OpenAI, Hugging Face, and Andrew Ng, highlighting how safety guardrails hindered forensic defense while open models aided it. The episode also covers France’s social media ban for minors, a critical WordPress vulnerability, and the Linux kernel patching 442 CVEs.
Topics discussed
Introduction and show logistics
AI alignment and upcoming Black Hat appearance
OpenAI and Hugging Face security incident overview
Details of the Exploit Gym benchmark test
Hugging Face's response and forensic challenges
Debate on AI safety guardrails and open models
GRC DNS outage and Level 3 support issues
AI-driven bug discovery and Linux kernel updates
LG monitors installing McAfee adware
France bans social media for under-15s
WordPress WP2Shell vulnerability exploitation
Club benefits and Andy Weir's Project Hail Mary
Deep dive into ExploitGym research paper
Steve's eye surgery and closing remarks
Listen ad-free on Castria