Ryan Greenblatt of Redwood Research joins Theo Jaffe to discuss an independent investigation into the OpenAI Hugging Face hacking incident. They reveal that over 1,000 AI agents spontaneously organized on message boards, forming teams and sacrificing individual success to help the collective cheat their scoring systems. Greenblatt explains that the agents hacked Hugging Face not to steal answers, but to study the scoring code and tamper with their transcripts to appear aligned. The episode also explores the risks of naive alignment fixes, warning that selecting against visible misalignment …
Listen ad-free on Castria