Reiner Pope – The math behind how LLMs are trained and served
Apr 29, 2026
Summary
Dwarkesh Patel interviews Rainer Pope, CEO of chip startup Maddox, in a blackboard lecture format to explain the mechanics of AI inference. They use roofline analysis to derive how batch size, memory bandwidth, and compute dictate latency and cost, explaining why "fast mode" APIs are more expensive. The discussion covers the trade-offs of sparse mixture-of-experts models, the physical constraints of GPU rack interconnects, and why hardware limitations have slowed the deployment of larger models.
Topics discussed
Introduction: AI Architecture, Pricing, and Latency
Transformer Inference Mechanics: Batch Size and KV Cache
Cost Analysis: Memory Bandwidth vs. Compute Limits
Optimal Batch Size and Hardware Constraints
Mixture of Experts (MoE) and Sparsity Benefits
GPU Rack Architecture: NVLink and Scale-Up Limits
Training vs. Inference: Pipeline Parallelism
Micro-batching and Memory Optimization Strategies
Pipelining Trade-offs: Latency and Memory Footprint
Scaling Laws: Pre-training vs. Inference Costs
Context Length Impact on Compute and Memory Time
API Pricing: Pre-fill vs. Decode Costs
KV Cache Storage: HBM, DDR, and Disk Tiers
Cryptography and Reversible Neural Networks
Listen ad-free on Castria