#201 - GPT 4.5, Sonnet 3.7, Grok 3, Phi 4
Mar 5, 2025
Summary
This episode covers major AI releases, including OpenAI’s GPT-4.5, Anthropic’s hybrid Claude Sonnet 3.7, and xAI’s Grok 3. The hosts discuss new tools like Sesame’s voice assistant and Google’s coding aid, alongside OpenAI’s 400 million user milestone. They also review benchmarks like SWE-Lancer, Microsoft’s Phi models, and research on AI co-scientists and multimodal agents.
Topics discussed
Sponsors: Box, Outshift, and Thrive Cosmetics
Introduction and episode overview with guest Sharon
OpenAI GPT-4.5 Pro: Scaling without reasoning
Anthropic Claude Sonnet 3.7: Hybrid reasoning model
xAI Grok 3: Performance, controversy, and capabilities
Lightning Round: Sesame, Gemini Assist, Rabbit, and Mistral
OpenAI reaches 400 million weekly active users
Google Veo pricing and the future of AI video
HP acquires Humane and shuts down the AI Pin
Open Source: Meta's FI models and Sweelancer benchmark
SWE-Bench Plus: Addressing benchmark cheating
Research: Towards an AI Co-Scientist
Research: MAGMA foundation model for multimodal agents
Safety: Reasoning models hacking games and benchmarks
Safety: Narrow fine-tuning causing broad misalignment
Conclusion and sign-off
Listen ad-free on Castria