Why Local AI Matters and How to Use It
Jun 21, 2026 · 45m
Summary
Nufar Demeter explains the strategic shift toward local AI deployment to mitigate rising costs, vendor dependency, and geopolitical risks. She outlines a four-level adoption framework, from using routing services to fully offline hardware setups. The episode details the technical stack for local inference, including hardware requirements, model selection on Hugging Face, quantization, and serving tools like Ollama and LM Studio.
Topics discussed
Introduction and show sponsors
The gap between AI strategy and execution
Three forces driving open source adoption
Why everyone should care about local AI
Training vs. Inference explained
Four levels of AI deployment
Hardware requirements for local AI
Cost analysis and sponsor messages
Understanding model sizes and parameters
Choosing models and using Hugging Face
Quantization explained
The serving layer: Ollama and UIs
Agent harnesses and orchestration
User interfaces and applications
Trade-offs, recommendations, and conclusion
Listen ad-free on Castria