Episode 303: Open Models, GPU Kernels & autoresearch with Mark Saroufim
Sep 11, 2026 · 34m
Summary
In ADSP episode 303, hosts Connor and Bryce interview Mark Seraphim of Core Auto about the rapid convergence of open and closed LLMs and the complexities of inference serving. They discuss how AI agents are democratizing GPU kernel optimization, allowing non-experts to achieve top-tier performance in competitions like NVFP4. The conversation highlights the critical need for rigorous property-based testing over traditional tolerance checks to ensure numerical correctness in automated code generation. Seraphim also shares insights on the high token costs of auto-research workflows and the evo…
Topics discussed
Intro: AI-generated kernels and the NVFP4 competition
Episode 303 introduction and agenda
Open vs closed models and the state of Chinese AI
Inference economics: managed deployments vs usage-based
How inference providers impact model quality perception
The iceberg of inference: quantization and numerics
Non-determinism and versioning in LLM serving
Kernel testing: tolerances vs property-based testing
Physics simulations and correctness checks analogy
Historical correctness testing in diffusion models
Core Auto's mission: automating the research lab
GPU Mode journey and the shift to AI education
Auto-research tools for researchers and code optimization
Costs of AI kernel generation and open model efficiency
Economies of scale in inference and local clusters
Token spend breakdown and efficiency improvements
Addressing AI code style, reuse, and design limitations
Challenges in search orchestration and preventing stalls
Unsteered AI search and the value of literature
Impact of public examples on AI kernel performance
Outro and show notes
Listen ad-free on Castria