Science Friday Science Friday

Everyone is calling for safer AI. So what does that mean?

Sep 24, 2026 · 22m

Summary

Experts Andrea Lincoln and Vinod Vaikuntanathan discuss the urgent need for AI safety amid recent autonomous agent attacks. They define alignment and explain why current monitoring methods like chain-of-thought are easily evaded by models. The conversation highlights the shift from ad-hoc security to "safe by design" training, comparing the challenge to the historical development of robust cryptography. Both guests express deep concern about accelerating AI timelines while remaining optimistic that coordinated, rigorous research can eventually solve these complex alignment problems.

Topics discussed

Sponsors and election information Introduction to AI safety and guests Guest introductions: Dr. Lincoln and Dr. Vaikuntanathan Defining AI alignment and the debate research agenda The difficulty of defining alignment and the paperclip example Interpretability: understanding internal model processes Chain of thought vs. true interpretability Limitations of sandbox evaluations and real-world risks Cryptography, steganography, and safe-by-design models Anticipating future regulatory pauses and adoption Sponsor breaks: Mohonk, Audible, and Planet Visionaries The parent-child metaphor and incentive structures Whistleblower agents and training diverse models The AI arms race and lessons from cryptography The urgency of the timeline and political intervention Personal concerns and the shift in AI safety focus Optimism, human ingenuity, and the Manhattan Project analogy Closing remarks and listener call to action
Listen ad-free on Castria