Inside vLLM: The Engine Powering Open-Source AI
Aug 6, 2026 · 46m
Summary
Alena Berger and Matt Bornstein interview Simon Mo, CEO of Inferable and lead maintainer of vLLM, on how open-source inference has become critical AI infrastructure. They discuss why enterprises are shifting to open-weight models for control and cost efficiency, the evolving licensing landscape, and the technical challenges of serving large language models. The conversation also covers the future of open-source AI versus proprietary systems and the importance of community-driven optimization.
Topics discussed
Introduction: Simon Moe, vLLM, and the open source AI landscape
The evolution of AI inference from BERT to modern LLMs
Open source models becoming critical infrastructure for startups
vLLM's role in the stack and supporting model releases
The partnership behind model releases and the Open Weights pledge
Cost vs. Control: Why enterprises choose open source models
Licensing changes and the economics of funding model training
The hidden complexity and cost of training frontier models
Thought experiment: What if GPUs were 99% cheaper?
Hugging Face cyber attack and the need for controllable guardrails
Moderation challenges and the social media analogy for AI
Differentiating open weight models through environment and data
Distillation, data policies, and the future of AI progress
Listen ad-free on Castria