Builders Free to read

Only Build Where Audio Is the Bottleneck

ElevenLabs started because Polish TV still dubs every foreign film with one flat voice reading all the parts.

· 4 min read

Mati Staniszewski on David Senra's podcast

Mati Staniszewski is the co-founder and CEO of ElevenLabs, the voice AI company he started in 2022 with Piotr Dabkowski, his best friend of 15 years. Before ElevenLabs he was a deployment strategist at Palantir, where he brought optimization models to customers like the NHS during the COVID vaccine rollout, and before that he built risk models at BlackRock. Dabkowski came from Google, where he worked on text models for the knowledge graph. ElevenLabs was the first company either of them had started, and it launched before the first version of ChatGPT. These signals are from Staniszewski's interview with David Senra on David Senra's podcast.

  1. The Polish Dubbing Problem. ElevenLabs started because foreign movies in Poland are still voiced by a single narrator who reads every part, male and female, with the emotion stripped out. Mati Staniszewski and his co-founder grew up with it and noticed it was still happening in 2021. The first plan was to dub that content with the original voices and intonation. When they broke dubbing into its 3 steps (transcription, translation, and regenerating speech in the new language), the existing research for each step sounded robotic.
  2. Creators Redirected the Roadmap. The creators they pitched on dubbing said they'd want it one day but had more pressing problems: fixing a badly recorded line in post-production, hearing a script read aloud before recording, or having AI narrate over a video. ElevenLabs parked the language problem and built speech generation first. That put research and product on the same task, since a convincing text-to-speech model was the missing piece for both.
  3. Audio-Only Research. ElevenLabs keeps its research team on audio and nothing else, and Staniszewski says the remaining gains in audio models come mostly from architecture, ahead of data and compute. Audio also mixes science with art, because every voice is subjective and listeners have strong preferences. He says the company explicitly won't touch general intelligence, knowledge work, or coding, which he calls "not our strength, not our domain."
  4. The Bottleneck Test. Before any new product, the team asks whether its audio models give a unique advantage there, and if audio isn't the bottleneck, it's out of scope. Two years ago they paused work on lip-sync and avatars because video quality was the weak link, and the best audio couldn't fix that experience. Open-source video models have since improved, so they are revisiting it. They still won't build text-to-video from scratch; they want the case where a customer already has an asset and needs audio added or changed.
  5. A Marketplace for the Long Run. Staniszewski expects the pure research lead to shrink over time, so he is building product moats around it. One is a voice marketplace: people record and authenticate their voice, share it, and earn money every time it's used. It has 20,000 voices today. "George," the ElevenReader voice Senra listens to, belongs to a voice actor who gets paid on every listen.
  6. Land and Expand at Deutsche Telekom. Deutsche Telekom started with marketing because 2 years ago voice agents were too slow and unreliable, so the first use was daily AI-voiced news podcasts in its Magenta app. After ElevenLabs proved impact (more than a proof of concept), it moved into call-center agents wired into Telekom's CRM, then an in-network agent that T-Mobile subscribers can pull into a live call to book appointments or translate in real time. Agents and creative tools now bring in most of ElevenLabs' revenue. Fintech (Revolut, Klarna, Customers Bank) adopts fastest, followed by healthcare and telcos, with retail and e-commerce picking up this year.
  7. Forward-Deployed Engineers on the Product Team. ElevenLabs' FDEs report into product, not go-to-market, because their second job is to bring what they learn back into the platform for the next customer. Skip that step, Staniszewski says, and the work is "effectively just services." The customer keeps its domain expertise, while the product patterns built around it become reusable. He picked this up at Palantir, which flew him to the North Sea in his first month; at BlackRock his team had to review his emails before he could send them externally.
  8. Ten Reports, Five Layers. Most ElevenLabs managers have close to 10 direct reports, teams usually have fewer than 10 people, there are no titles, and the org chart is capped at 5 layers. Nearly every internal doc is open to everyone. That transparency lets AI summarize what's happening across teams, so Staniszewski can spot problems himself instead of waiting for someone to raise them. He hopes AI will let him remove layers over time.
  9. Engineers in Every Function. Talent, ops, and legal teams at ElevenLabs each include engineers who automate work and help colleagues use AI. Small, independent teams also mean AI adoption happens bottom-up, with no mandate from the top. The company runs a central revenue engineering group whose AI SDR lets website visitors talk to a voice agent instead of filling out a drop-down form. People who speak leave far more detail about their problems and use cases than people who type.
  10. Imperfection Sounds Human. ElevenLabs first tried to build a voice agent that made no filler sounds, and it didn't sound human. When they added "ums," approval jumped and people said they were happy to talk to it. Staniszewski says voice carries emotion, intonation, pauses, and imperfections that text lacks. That is also why benchmarking text-to-speech is so hard: different models use different voices, which makes them hard to compare.
  11. No Price for the Best Idea. ElevenLabs has had 3 or 4 concrete acquisition offers, the most recent in June of last year, and turned all of them down. Asked if he will sell, Staniszewski said no, adding that the offers "should never be an interesting proposition" even though they would have made life easy. He is 31 and sees this as his best idea at the best possible timing. Dabkowski, he says, is just as committed.
  12. Technology on Human Terms. For decades people learned the language of machines (keyboards, screens, programming languages), and Staniszewski wants to reverse that so technology works through voice, the most basic way humans communicate. Within 12 months he expects AI conversations to combine IQ and EQ: reading how you feel, pausing, thinking, and stepping back into the conversation. As intelligence improves, he argues, the next bottleneck is how people talk and work with it. The founders' early pitch slides pointed to the Babel fish from The Hitchhiker's Guide to the Galaxy, built into the devices people already own.

Watch the full video at https://www.youtube.com/watch?v=RFccAuyPPOg.

sig·nal·ful /ˈsɪɡ.nəl.fəl/ adjective — full of signal.

Get Signalful in your inbox.

One story free among every issue. Members unlock all, plus access the full archive.