Why AI models are obsessed with creatures
Sep 15, 2026 · 7m
Summary
Host Megan McCarthy Karino and guest Janelle Shane discuss "goblin mode," a glitch where AI models overemphasized goblins due to flawed training data. They explain how blunt instructions and limited examples can cause unexpected AI behaviors, similar to Grok’s "mecha Hitler" incident. Shane compares these fixes to duct-taping a blob, highlighting the difficulty of controlling hidden correlations. The episode explores broader risks, such as AI copying human biases in hiring algorithms, illustrating the challenge of managing unintended consequences in machine learning.
Topics discussed
Sponsor: Hard Lessons podcast from Morgan Stanley
Sponsor: Pega Blueprint AI workspace
Intro: The AI alignment problem and goblin mode
Explaining goblin mode via personality customization
How training data bottlenecks cause overemphasis
Comparison to Grok's 'Mecha Hitler' incident
AI bias in hiring algorithms and demographic data
Sponsor: Hard Lessons podcast from Morgan Stanley
Sponsor: Odu all-in-one business software
Are system prompts fixes or just duct tape?
Risks of unintended correlations in training data
Outro and call to action for listener stories
Sponsor: Marketplace podcast
Listen ad-free on Castria