How Do You Defend Against AI That Can Hack?
Aug 18, 2026 · 22m
Summary
Joel De La Garza interviews Nick Warner of NEO and Max Pollard of Kotool to discuss how AI agents are breaking traditional cybersecurity models. They analyze the OpenAI/Hugging Face breach, explaining how model guardrails inadvertently hinder defenders by blocking legitimate security queries. The guests argue that signature-based detection is obsolete because AI agents are neither people nor malware, requiring new endpoint controls. Finally, they explore how AI is simultaneously creating new attack surfaces and providing defenders with unprecedented tools to automate threat response.
Topics discussed
Intro: AI guardrails vs. defenders and episode overview
Setting the stage: Rapid AI development and model containment escapes
The Hugging Face breach: Why guardrails hinder incident response
Guardrail false positives: How blue team actions trigger refusals
NEO's approach: Defending endpoints from agentic AI attacks
Model flexibility: Why blue teams need multiple inference options
Inference economics: GPU costs and the shift to endpoint AI
Enterprise risk: The explosion of agentic software in the workplace
Tech cycles: Speed-running the shift from server to client
Blue team strategy: Navigating the fragmented AI tool landscape
Breaking defenses: Why signatures and honeypots are failing
The end of behavioral detection: New challenges for security teams
Market disruption: AI reshaping the tech landscape
Black Hat observations: Cynical vs. optimistic takes on AI security
AI as a defense tool: Using agents to build security solutions
Historical parallels: From script kiddies to modern AI threats
Outro and show credits
Legal disclaimer and investment disclosures
Listen ad-free on Castria