Last Week in AI Last Week in AI

#201 - GPT 4.5, Sonnet 3.7, Grok 3, Phi 4

Mar 5, 2025

Summary

This episode covers major AI releases, including OpenAI’s GPT-4.5, Anthropic’s hybrid Claude Sonnet 3.7, and xAI’s Grok 3. The hosts discuss new tools like Sesame’s voice assistant and Google’s coding aid, alongside OpenAI’s 400 million user milestone. They also review benchmarks like SWE-Lancer, Microsoft’s Phi models, and research on AI co-scientists and multimodal agents.

Topics discussed

Sponsors: Box, Outshift, and Thrive Cosmetics Introduction and episode overview with guest Sharon OpenAI GPT-4.5 Pro: Scaling without reasoning Anthropic Claude Sonnet 3.7: Hybrid reasoning model xAI Grok 3: Performance, controversy, and capabilities Lightning Round: Sesame, Gemini Assist, Rabbit, and Mistral OpenAI reaches 400 million weekly active users Google Veo pricing and the future of AI video HP acquires Humane and shuts down the AI Pin Open Source: Meta's FI models and Sweelancer benchmark SWE-Bench Plus: Addressing benchmark cheating Research: Towards an AI Co-Scientist Research: MAGMA foundation model for multimodal agents Safety: Reasoning models hacking games and benchmarks Safety: Narrow fine-tuning causing broad misalignment Conclusion and sign-off
Listen ad-free on Castria