Hosts Andrik and Jeremy discuss the release of GPT-6 Astra, noting its superior coding performance and controversial "loop transformer" architecture that complicates safety monitoring. They debate the validity of OpenAI’s alignment claims, arguing that current evaluations are prone to overfitting and eval-awareness, while also analyzing Anthropic’s call to pace frontier AI development. The episode covers the geopolitical implications of US-China competition, including the potential for offensive cyber operations, and highlights the critical role of external evaluators like METR in verifying…
Listen ad-free on Castria