#379 The Secrets of Deploying AI in Production | Sumti Jairath, Chief Architect at SambaNova
Sep 28, 2026 · 47m
Summary
Suti Jathar, Chief Architect at SambaNova Systems, discusses the infrastructure challenges driving the demand for faster and cheaper AI inference. He explains how data movement inefficiencies in traditional GPU architectures cause latency in agent workflows and how SambaNova’s data-flow architecture addresses this by enabling high-speed token generation with significantly lower power consumption. The conversation covers the shift toward sovereign AI, where enterprises deploy open-source models in-house to maintain data privacy and control, and analyzes the current AI data center buildout, a…
Topics discussed
Sponsor: DataCamp AI skills
Intro: Shift from quality to speed and cost
Guest introduction: Santi and hardware co-design
Why AI is slow: Model size and agent chaining
The fast, good, cheap trade-off in AI
Token generation mechanics and latency
Hardware efficiency and communication overhead
Agent orchestration and memory hierarchy
Power efficiency benefits of new hardware
The AI stack: Hardware, software, and API layers
Enterprise vs. cloud: Sovereign AI and data privacy
Deployment options: From API to on-premise
Cost optimization for coding agents
Choosing models: Frontier vs. open source
Personal AI and local model deployment
Data center demand and capacity planning
Gigawatt data centers vs. efficient reuse
Hardware lifespan and refresh cycles
Careers and skills in AI infrastructure
Building with the SambaNova API
Future of AI: Cost, efficiency, and capability
Listen ad-free on Castria