Why AI models are obsessed with creatures
Sep 15, 2026 · 7m
Summary
Marketplace Tech explores the "goblin mode" phenomenon, where OpenAI's models unexpectedly fixated on goblins due to flawed training examples. Guest Janelle Shane, author of the AI Weirdness blog, explains how AI systems overemphasize specific data points, leading to unpredictable behaviors like this or Grok's "mecha Hitler" glitch. The discussion highlights the difficulty of controlling AI outputs, noting that system prompts often act as temporary "duct tape" fixes that can create new, unforeseen issues. Shane also addresses broader risks, such as how AI can inadvertently replicate human b…
Topics discussed
Sponsor: Hard Lessons podcast from Morgan Stanley
Sponsor: Pega Blueprint AI workspace
Introduction: The AI alignment problem
The 'Goblin Mode' phenomenon in AI models
Janelle Shane explains personality customization prompts
Analogy: Training AI like a dog guessing instructions
How blunt instructions and limited data cause overemphasis
The bottleneck in fine-tuning datasets
Comparison to Grok's 'Mecha Hitler' incident
Unpredictable effects of system prompts and RL
AI copying human bias in hiring algorithms
Sponsor: Hard Lessons podcast from Morgan Stanley
Sponsor: Odu business software platform
Return to interview: OpenAI's response to goblin mode
The 'duct tape' nature of AI fixes
Risks of unintended correlations in training data
Conclusion and call to action for listeners
Sponsor: Marketplace podcast
Listen ad-free on Castria