What I Learned Testing GPT-5.5
Apr 24, 2026
Summary
This episode analyzes the launch of OpenAI’s GPT-5.5, evaluating its performance against Anthropic’s Opus 4.7 through benchmarks, community reactions, and the host’s personal tests. The discussion highlights GPT-5.5’s strengths in coding, long-running agentic tasks, and knowledge work, while noting trade-offs in design and planning. It also covers OpenAI’s strategic shift toward democratization and efficiency, contrasting it with Anthropic’s recent quality issues.
Topics discussed
Intro, sponsors, and Operator's Bonus announcement
GPT-5.5 launch context and benchmark dominance
Benchmark nuances, pricing, and efficiency debates
Community reactions: Hype vs. reality and Mythos comparison
Nuanced user views on workflow impact and value
Coding capabilities: Reliability, speed, and planning
Design, presentations, and spreadsheet accuracy tests
Sponsors: KPMG, Blitzy, Granola, and Mercury
OpenAI's communication strategy and democratization
Host tests: Script prep and meta-planning workflows
Host tests: Creative kit and web app execution
Host tests: Art book design and data analysis
Competitive landscape: OpenAI vs. Anthropic narrative
Future outlook: O3, rapid releases, and conclusion
Listen ad-free on Castria