Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Aug 25, 2026 · 1h 18m
Summary
Patrick O'Shaughnessy interviews Neil Nova, founder of SAIL Research, who is building a "token factory" focused on ultra-low-cost, long-running AI agents rather than low-latency chatbots. They discuss the technical trade-offs between throughput and latency, the architecture of GPUs versus specialized chips like Cerebras, and the future of background AI tasks in cybersecurity and deep research.
Topics discussed
Sponsors: Ramp and Felix by Rogo
SAIL Research: The Token Factory for Long-Running Agents
Human Time Scales and the Rise of Background AI Agents
Deep Research and Autonomous Cybersecurity Agents
Verifiable Tasks vs. Human Taste in AI
SAIL's Software Stack and GPU Efficiency Strategy
Neil's Experience at NVIDIA and GPU Culture
GPU Batching, Latency, and Hardware Interconnects
Memory Hierarchy: SRAM vs. DRAM and Cerebras
Context Windows and the Transformer Architecture
Scaling Laws, Attention, and Data Quality
Recursive Self-Improvement and Kernel Engineering
Data Center Infrastructure and Chip Arbitrage
Market Dynamics: Training vs. Inference Spend
Distributed Compute and Power Constraints
Compute Orchestration and Company Culture
Open Source, Distillation, and the Future of Tokens
Advice for Hardware Entrepreneurs and Closing Thoughts
Listen ad-free on Castria