ADSP: Algorithms + Data Structures = Programs ADSP: Algorithms + Data Structures = Programs

Episode 303: Open Models, GPU Kernels & autoresearch with Mark Saroufim

Sep 11, 2026 · 34m

Summary

In ADSP episode 303, hosts Connor and Bryce interview Mark Seraphim of Core Auto about the rapid convergence of open and closed LLMs and the complexities of inference serving. They discuss how AI agents are democratizing GPU kernel optimization, allowing non-experts to achieve top-tier performance in competitions like NVFP4. The conversation highlights the critical need for rigorous property-based testing over traditional tolerance checks to ensure numerical correctness in automated code generation. Seraphim also shares insights on the high token costs of auto-research workflows and the evo…

Topics discussed

Intro: AI-generated kernels and the NVFP4 competition Episode 303 introduction and agenda Open vs closed models and the state of Chinese AI Inference economics: managed deployments vs usage-based How inference providers impact model quality perception The iceberg of inference: quantization and numerics Non-determinism and versioning in LLM serving Kernel testing: tolerances vs property-based testing Physics simulations and correctness checks analogy Historical correctness testing in diffusion models Core Auto's mission: automating the research lab GPU Mode journey and the shift to AI education Auto-research tools for researchers and code optimization Costs of AI kernel generation and open model efficiency Economies of scale in inference and local clusters Token spend breakdown and efficiency improvements Addressing AI code style, reuse, and design limitations Challenges in search orchestration and preventing stalls Unsteered AI search and the value of literature Impact of public examples on AI kernel performance Outro and show notes
Listen ad-free on Castria