The world of voice AI, with Mati Staniszewski of ElevenLabs
Apr 14, 2026 · 1h 0m
Summary
Matty Stanishevsky, co-founder of ElevenLabs, explains how their AI audio models use neural networks to predict phonemes and capture human-like emotional inflection. He details the company’s expansion from text-to-speech into voice agents, speech-to-text, and music, highlighting the technical challenges of real-time conversational orchestration. Stanishevsky also discusses the "deployment gap" in consumer voice interfaces, upcoming features like personalized transcription, and the business economics of scaling high-fidelity audio AI.
Topics discussed
Introduction to Matty Stanishevsky and 11 Labs
How audio models work: from analog to neural
Defining phonemes and voice model representations
Data strategy and annotation innovations
Historical context: The Mechanical Turk
11 Labs business breakdown and product suite
Platform strategy vs. vertical applications
The gap between LLMs and voice AI deployment
Consumer voice assistants and PDF reading
Challenges in real-time voice agents and Turing test
Voice in subscription checkout flows
Personalized transcription and speaker-specific models
Speech enhancement and emotional control
Cascaded vs. speech-to-speech architectures
User behavior changes with voice interfaces
Dubbing, translation, and restoring lost voices
Examples of persistent voice agents
Business economics, CapEx, and model scaling
Future model sizes and architecture trends
Deployment priorities and popular use cases
Revenue growth and customer expansion strategies
Company structure and team dynamics
Self-serve motion and billing models
AI-native organization and internal tooling
Government and healthcare voice agent deployments
Culture, agency, and closing remarks
Listen ad-free on Castria