ElevenLabs AI voice technology works by training models on raw audio data rather than text tokens, allowing the system to capture not just what is said but how it is said — the emotion, intonation, and subtle human qualities that make a voice feel real. Co-founder Mati recently explained in a conversation with Andreessen Horowitz that humanity has been trying to synthesize human voice since the 1700s, and ElevenLabs believes it is finally close to crossing the threshold that separates a voice that sounds plausible from one that genuinely makes you feel something.

How Does ElevenLabs AI Voice Technology Work?

Most AI language models are trained on text — breaking language into tokens that humans created. ElevenLabs takes a fundamentally different approach. By training on raw audio, their models learn to capture everything that text simply cannot encode: the warmth of a whisper, the weight of a dramatic pause, the subtle cultural flavor baked into the way a sentence rises at the end.

Early timeline of voice synthesis — from 1700s mechanical voice boxes to early digital synthesizers to Siri 00:45 Early timeline of voice synthesis — from 1700s mechanical voice boxes to early digital synthesizers to Siri Watch at 00:45 →

The company currently has specialized models for voice, sound effects, and music, but the long-term vision is a single unified model capable of generating any kind of audio. Imagine a voice that seamlessly transitions into a musical score, or spoken words that morph into cinematic sound effects. Mati describes this as one of the most exciting technical frontiers ElevenLabs is chasing: a model smart in audio could become a blueprint for a model smart in any raw data domain.

The research and product teams are deliberately kept in close feedback loops. Researchers test their models directly on the product; the product surface immediately surfaces what users actually need. It is a flywheel that Mati believes gives ElevenLabs a meaningful edge over companies that are strong on one side but not both.

Will Voice AI Replace Screens as Our Main Interface?

Mati argues that voice is poised to become the next fundamental interface for humans interacting with computers — as significant a shift as the move from keyboards to touchscreens. The vision is striking in its specificity: imagine a student wearing headphones in a classroom, with access to the world's smartest physicist, mathematician, or historian, all speaking naturally and responsively in real time.

Today, most technology is screen-first. You carry a laptop and a phone. Mati sees much of that moving into the background as voice handles more of the cognitive load, freeing people to be more present in the physical world around them. The screen does not disappear, but it stops being the center of gravity.

Mati describing the Polish film dubbing problem that sparked the ElevenLabs idea 03:10 Mati describing the Polish film dubbing problem that sparked the ElevenLabs idea Watch at 03:10 →

There is also a business case that is already playing out. ElevenLabs launched to a few thousand users on its waitlist in early January — and within weeks, those few thousand had become hundreds of thousands. The demand signal was, by Mati's own admission, a magnitude higher than first-order expectations.

How Did ElevenLabs Grow From 2 Founders to 300 Employees?

The origin story starts, fittingly, in Poland. When Mati and his childhood best friend Piotr watched foreign films growing up, every character — male or female, hero or villain — was dubbed by a single monotone narrator. All the emotion, all the intonation, simply disappeared. In 2021, the two realized that this problem, embarrassingly, had not been solved. Mati was at Palantir. Piotr was at Google. They started exploring the problem on weekends.

They invited an early group of users, iterated on feedback, and refined the use cases that genuinely resonated. By the time they raised their Series A, they had seven people. A year later, they had a few dozen. Today, ElevenLabs operates in over 11 cities with more than 300 employees and is doubling every six months.

Team members sharing their unconventional backgrounds — astrophysics, White House, competitive gaming 06:20 Team members sharing their unconventional backgrounds — astrophysics, White House, competitive gaming Watch at 06:20 →

The hires themselves came from unconventional places:

  • An astrophysics and applied physics graduate who had been quietly building a text-to-speech project during his master's degree rather than attending lectures
  • A former White House staffer who an investor told to do whatever it took to join
  • A music generation researcher whose work Piotr discovered online and cold-contacted immediately
  • A competitive Dota 2 player ranked in the top 250 in Europe who channeled that obsessive ambition into building product

The throughline: proof of excellence in something. It did not have to be a traditional resume credential. It had to demonstrate that this person goes all in.

Why Did ElevenLabs Remove All Job Titles?

One of the most counterintuitive decisions ElevenLabs made as it scaled was to eliminate job titles entirely. The reasoning is both practical and philosophical. Titles, Mati explains, create implicit hierarchies that discourage people from asking questions, proposing ideas, or admitting they need help. Remove the titles and you remove the social friction.

It also functions as a hiring filter. If a candidate's opening move is negotiating for a VP title, they are probably not the right fit for a flat, high-autonomy environment. When Mati first spoke publicly about dropping titles, a former colleague reached out saying she had heard about it and wanted in immediately. She now leads hiring.

The broader culture is described by investors and employees alike as low-bureaucracy, high-autonomy, and almost deliberately fuzzy in its hierarchy. Teams are small. Ownership is real. Researchers can get access to a training cluster and start building on an idea without navigating layers of approval. Mati and Piotr are described as the yin and yang of the company — Mati the relationship builder, Piotr the technical genius who is, in the words of one colleague, so far ahead that the second-smartest person they know is significantly less smart than him.

What Is the Vocal Turing Test and Can AI Pass It?

ElevenLabs has given itself a specific and ambitious north star: become the first company to pass what Mati calls the vocal Turing test. The idea is simple to state and fiendishly hard to achieve — build an AI voice that sounds so genuinely human, so emotionally present and contextually intelligent, that you cannot tell you are talking to a machine.

Current voice AI, including the Siri-era synthesizers that represented a real leap forward, still does not cross that threshold. It sounds realistic enough. It does not make you feel anything. Mati believes the gap is closable, and that the key is training models on raw audio rather than text-derived representations of speech.

An AI that passes the vocal Turing test is not just a party trick. It is the foundation for everything from real-time language translation that preserves cultural nuance, to emotionally intelligent tutoring systems, to audio interfaces that are more expressive and efficient than any screen.

Can AI Voice Finally Break Language Barriers?

One of the most compelling use cases Mati describes is the dissolution of language and cultural barriers. Today, traveling to another country means navigating not just vocabulary but an entire cultural register — the way things are said, the emotional subtext carried in tone and rhythm. Learning a language gets you partway there. Truly immersing yourself takes years.

Voice AI, Mati argues, can collapse that timeline. Speak any language in the world. Understand not just the words but how they are delivered. Feel like a closer participant in another culture, not an outsider reading subtitles. The technology to do this is not theoretical — it is the direct extension of what ElevenLabs is already building.

Why Is Voice the Only AI Modality That Makes You Feel Something?

Text can move you. A poem, a story, a well-crafted argument — these things matter. But Mati draws a sharp distinction: text engages the mind. Voice reaches somewhere deeper. An ASMR whisper, a deep cinematic boom, a voice cracking with genuine feeling — these can transport you in a way that no string of tokens can replicate.

This is why Mati believes voice will eventually carry the majority of human-machine communication. It is faster than typing. It is richer in information. And it is the only AI output that can make you feel, in his words, truly alive. For a company whose founding insight was that a single monotone dubbing voice was stealing the soul from foreign films, that is a fitting mission to be chasing.