Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Aug 25, 2026 · 1h 23m
Summary
Sale Research’s founder explains their strategy to become the cheapest provider of AI inference by optimizing for throughput rather than latency, targeting long-running background agents. The episode details how they leverage open-source models and heterogeneous hardware stacks, including specialized chips like Cerebras, to maximize cost efficiency. Key topics include the shift from interactive chatbots to proactive agents, the technical tradeoffs between SRAM and DRAM, and the future of verifiable tasks in cybersecurity and deep research.
Topics discussed
Introduction: Mission to make intelligence tokens cheap
The shift from chatbots to long-running agents
Applications: Cybersecurity and background agents
Vision: Abundant intelligence and personal agents
Sponsor segments and enterprise adoption
Hardware deep dive: NVIDIA history and Tensor Cores
GPU efficiency: Latency vs. throughput trade-offs
Memory architectures: SRAM, DRAM, and HBM
KV Cache and long-context inference challenges
Transformer architecture and the future of data
Kernel engineering and software optimization
Chip strategy: Heterogeneous serving and arbitrage
Market analysis: AI bubble vs. real demand
Data center infrastructure and power constraints
The 'Scavenger' strategy for chips and power
Supply chain bottlenecks and manufacturing
Open vs. Closed models and distillation
Divergent views on hardware and future tech
Listen ad-free on Castria