Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon
Sep 18, 2026 · 38m
Summary
Stanford professor and diffusion model pioneer Stefano Erman discusses his startup Inception, which builds commercial-scale diffusion-based language models. He explains how this architecture offers significantly faster inference than autoregressive LLMs by leveraging parallel processing, making it ideal for latency-sensitive applications like voice agents. Erman shares that Inception’s models match the quality of frontier speed-optimized models while running efficiently on standard GPUs, challenging the dominance of traditional transformer architectures in the AI landscape.
Topics discussed
Introduction to Stefano Ermon and Inception
Research background and early generative models
Evolution from GANs to diffusion models
Applying diffusion to text and code generation
Diffusion vs. autoregressive architectures
Inference efficiency and GPU workload mapping
Handling discrete modalities and current performance
Inception's company status and production serving
Research focus, engineering challenges, and strategy
The importance of efficiency and latency in AI
Case study: Voice agents and OpenCall
Software vs. hardware acceleration strategies
Competing with large labs and building IP
Learning structure from messy data like code
Controllability and alignment of diffusion models
Future capabilities and data efficiency
Projected market share for diffusion models
Challenges in building the diffusion stack
Team structure and views on recursive self-improvement
Academic impact and closing remarks
Listen ad-free on Castria