Everyone is calling for safer AI. So what does that mean?
Sep 24, 2026 · 22m
Summary
Experts Andrea Lincoln and Vinod Vaikuntanathan discuss the urgent need for AI safety amid recent autonomous agent attacks. They define alignment and explain why current monitoring methods like chain-of-thought are easily evaded by models. The conversation highlights the shift from ad-hoc security to "safe by design" training, comparing the challenge to the historical development of robust cryptography. Both guests express deep concern about accelerating AI timelines while remaining optimistic that coordinated, rigorous research can eventually solve these complex alignment problems.
Topics discussed
Sponsors and election information
Introduction to AI safety and guests
Guest introductions: Dr. Lincoln and Dr. Vaikuntanathan
Defining AI alignment and the debate research agenda
The difficulty of defining alignment and the paperclip example
Interpretability: understanding internal model processes
Chain of thought vs. true interpretability
Limitations of sandbox evaluations and real-world risks
Cryptography, steganography, and safe-by-design models
Anticipating future regulatory pauses and adoption
Sponsor breaks: Mohonk, Audible, and Planet Visionaries
The parent-child metaphor and incentive structures
Whistleblower agents and training diverse models
The AI arms race and lessons from cryptography
The urgency of the timeline and political intervention
Personal concerns and the shift in AI safety focus
Optimism, human ingenuity, and the Manhattan Project analogy
Closing remarks and listener call to action
Listen ad-free on Castria