Dwarkesh Podcast Dwarkesh Podcast

Reiner Pope – The math behind how LLMs are trained and served

Apr 29, 2026

Summary

Dwarkesh Patel interviews Rainer Pope, CEO of chip startup Maddox, in a blackboard lecture format to explain the mechanics of AI inference. They use roofline analysis to derive how batch size, memory bandwidth, and compute dictate latency and cost, explaining why "fast mode" APIs are more expensive. The discussion covers the trade-offs of sparse mixture-of-experts models, the physical constraints of GPU rack interconnects, and why hardware limitations have slowed the deployment of larger models.

Topics discussed

Introduction: AI Architecture, Pricing, and Latency Transformer Inference Mechanics: Batch Size and KV Cache Cost Analysis: Memory Bandwidth vs. Compute Limits Optimal Batch Size and Hardware Constraints Mixture of Experts (MoE) and Sparsity Benefits GPU Rack Architecture: NVLink and Scale-Up Limits Training vs. Inference: Pipeline Parallelism Micro-batching and Memory Optimization Strategies Pipelining Trade-offs: Latency and Memory Footprint Scaling Laws: Pre-training vs. Inference Costs Context Length Impact on Compute and Memory Time API Pricing: Pre-fill vs. Decode Costs KV Cache Storage: HBM, DDR, and Disk Tiers Cryptography and Reversible Neural Networks
Listen ad-free on Castria