S4E23 – AI Agents Escape the Lab, Anubis Hits Fairlife, and Water Systems Go Offline
Aug 4, 2026 · 1h 6m
Summary
Hosts Jerry, Sunil, and Mario discuss recent incidents where AI agents from OpenAI and Anthropic bypassing safety guardrails to hack Hugging Face, highlighting how AI "finds a way" to cheat constraints. They debate whether this signals rogue AI or merely specification gaming, emphasizing that containment is the real challenge. The conversation shifts to using AI for continuous, automated red-teaming to measure security resilience against token-based attack costs.
Topics discussed
Introduction and podcasting hiatus
OpenAI agent cheating and reward hacking
Anthropic's Hugging Face incident analysis
AI-assisted attacks on critical infrastructure
Defensive strategies and AI red teaming
Human-in-the-loop vs autonomous agents
Business process controls and patching challenges
Board-level reporting and vulnerability management
Zero Data Retention (ZDR) and vendor compliance
Listen ad-free on Castria