Invest Like The Best Invest Like The Best

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

Aug 25, 2026 · 1h 23m

Summary

Sale Research’s founder explains their strategy to become the cheapest provider of AI inference by optimizing for throughput rather than latency, targeting long-running background agents. The episode details how they leverage open-source models and heterogeneous hardware stacks, including specialized chips like Cerebras, to maximize cost efficiency. Key topics include the shift from interactive chatbots to proactive agents, the technical tradeoffs between SRAM and DRAM, and the future of verifiable tasks in cybersecurity and deep research.

Topics discussed

Introduction: Mission to make intelligence tokens cheap The shift from chatbots to long-running agents Applications: Cybersecurity and background agents Vision: Abundant intelligence and personal agents Sponsor segments and enterprise adoption Hardware deep dive: NVIDIA history and Tensor Cores GPU efficiency: Latency vs. throughput trade-offs Memory architectures: SRAM, DRAM, and HBM KV Cache and long-context inference challenges Transformer architecture and the future of data Kernel engineering and software optimization Chip strategy: Heterogeneous serving and arbitrage Market analysis: AI bubble vs. real demand Data center infrastructure and power constraints The 'Scavenger' strategy for chips and power Supply chain bottlenecks and manufacturing Open vs. Closed models and distillation Divergent views on hardware and future tech
Listen ad-free on Castria