The Next Frontier of AI Video Is Control
Sep 17, 2026 · 39m
Summary
a16z General Partner Jennifer Lee discusses H3 Max with Fal co-founders Gorka Mirtsevin and Buwan Tashkaya, highlighting how post-training and systems optimization made video generation significantly faster and cheaper. The episode explores how these advancements enable real-time, continuous video streams with memory and action control, unlocking new consumer and professional experiences. They also cover the integration of Blender workflows for Hollywood, emphasizing the shift toward greater controllability over camera angles, lighting, and character motion to meet industry needs.
Topics discussed
Introduction: H3 Max and the generative media market
Benchmarking H3 Max against competitors
Why post-training an open-source model like Minimax H3
Token market fit and compute constraints in AI
System optimizations and breaking the roofline
Cost, speed, and quality trade-offs
Post-training for efficiency and hardware utilization
Optimizing the multi-component video pipeline
Hardware implications: Hopper vs Blackwell
H3 Max Turbo: Real-time generation capabilities
Future focus: Quality and controllability
Surprise launch and viral real-time experiments
Internal creative explosion and parallel projects
Viral live streams: Twitch, Levelio, and Fall Live
Technical deep dive: Continuous video and memory
H3 Max Director: Action-controlled continuous video
Market dynamics: Bursting innovation cycles
Creator usage and voice-directed workflows
Memory architecture: Raw video vs system prompts
Economics: Serving costs and hardware diversity
Hollywood workflows: Blender and GPT Astra integration
Professional controls: Lip sync, motion, and camera
Unified post-training infrastructure for any model
Hollywood adoption and specific studio needs
Bridging the gap between consumer and pro expectations
Legal, data residency, and US-hosted models
Generative Media Conference and industry shift
Closing remarks and podcast outro
Listen ad-free on Castria