Opus 5.5 is the best model ever made, Grok 4.7 exists, GPT 6 Sol & Luna Impress, and why Jev isn't the solution to everything
Sep 30, 2026 · 2h 8m
Summary
Theo and Ben discuss recent AI model releases, including GPT-6 Soul, Luna, Grok 4.7, and Opus 5.5. They analyze Grok 4.7’s disappointing performance, high token usage, and misleading pricing claims, contrasting it with the superior value of Opus 5.5. The hosts also explore the "middle child" syndrome of mid-tier models and introduce Jev, a fast, cheap "System 1" model designed for rapid pattern matching and classification tasks.
Topics discussed
Intro: Best model drop of the week and Grok 4.7 debate
Sponsors and deciding the order of model discussion
OpenAI GPT-6 Sol release and ChatGPT product focus
Twitter autoplay issues and X Money APY discussion
GPT-6 Sol positioning as the new 'middle child' model
Grok 4.7 release timeline and Elon Musk's predictions
Grok 4.7 pricing, marketing, and SpaceX employee hype
Grok 4.7 performance in coding and agent tasks
Grok 4.7 benchmark scores and token efficiency claims
Competitive advantages of non-lab companies like Cursor
Cursor's strategy and rebuilding sites with Paper
Introduction to Jeff: A classification-focused model
Jeff or No Jeff: Use cases for classification tasks
Criticism of using Jeff for compaction and routing
Jeff vs Go: Comparing niche model archetypes
Jeff Router popularity and token usage statistics
Anthropic Opus 5.5: Bias disclosure and initial impressions
Opus 5.5 pricing, reasoning levels, and benchmark analysis
Opus 5.5 speed, efficiency, and comparison to Fable
Opus 5.5 cost structure and cache read pricing
Opus 5.5 in practice: Fish slot port and game dev
Opus 5.5 on complex tasks: Persona 3 and Rust rewrite
Cost per task, workflow friction, and verification tooling
Experimenting with banning unit tests in agent workflows
Local model setup with DGX Sparks and outro
Listen ad-free on Castria