Cheeky Pint Cheeky Pint

The world of voice AI, with Mati Staniszewski of ElevenLabs

Apr 14, 2026 · 1h 0m

Summary

Matty Stanishevsky, co-founder of ElevenLabs, explains how their AI audio models use neural networks to predict phonemes and capture human-like emotional inflection. He details the company’s expansion from text-to-speech into voice agents, speech-to-text, and music, highlighting the technical challenges of real-time conversational orchestration. Stanishevsky also discusses the "deployment gap" in consumer voice interfaces, upcoming features like personalized transcription, and the business economics of scaling high-fidelity audio AI.

Topics discussed

Introduction to Matty Stanishevsky and 11 Labs How audio models work: from analog to neural Defining phonemes and voice model representations Data strategy and annotation innovations Historical context: The Mechanical Turk 11 Labs business breakdown and product suite Platform strategy vs. vertical applications The gap between LLMs and voice AI deployment Consumer voice assistants and PDF reading Challenges in real-time voice agents and Turing test Voice in subscription checkout flows Personalized transcription and speaker-specific models Speech enhancement and emotional control Cascaded vs. speech-to-speech architectures User behavior changes with voice interfaces Dubbing, translation, and restoring lost voices Examples of persistent voice agents Business economics, CapEx, and model scaling Future model sizes and architecture trends Deployment priorities and popular use cases Revenue growth and customer expansion strategies Company structure and team dynamics Self-serve motion and billing models AI-native organization and internal tooling Government and healthcare voice agent deployments Culture, agency, and closing remarks
Listen ad-free on Castria