AI + a16z AI + a16z

Inside vLLM: The Engine Powering Open-Source AI

Aug 6, 2026 · 46m

Summary

Alena Berger and Matt Bornstein interview Simon Mo, CEO of Inferable and lead maintainer of vLLM, on how open-source inference has become critical AI infrastructure. They discuss why enterprises are shifting to open-weight models for control and cost efficiency, the evolving licensing landscape, and the technical challenges of serving large language models. The conversation also covers the future of open-source AI versus proprietary systems and the importance of community-driven optimization.

Topics discussed

Introduction: Simon Moe, vLLM, and the open source AI landscape The evolution of AI inference from BERT to modern LLMs Open source models becoming critical infrastructure for startups vLLM's role in the stack and supporting model releases The partnership behind model releases and the Open Weights pledge Cost vs. Control: Why enterprises choose open source models Licensing changes and the economics of funding model training The hidden complexity and cost of training frontier models Thought experiment: What if GPUs were 99% cheaper? Hugging Face cyber attack and the need for controllable guardrails Moderation challenges and the social media analogy for AI Differentiating open weight models through environment and data Distillation, data policies, and the future of AI progress
Listen ad-free on Castria