Voice AI Went From Seven-Year Leaps to Seven-Month Leaps. What Happens by 2028?
Human-sounding speech is becoming the baseline. The next voice AI race is about interruption, action, multilingual reliability and owning an outcome.
In 2016, getting a neural network to produce human-ish raw audio was a research breakthrough.
In 2026, I can interrupt an AI mid-sentence, switch languages, ask it to check a system, and expect it to continue without forgetting why we started talking.
Ten years on paper.
It feels more like two different industries.
I have been going deep on voice because we are building in it at Praxy, and because the progress is currently easy to misunderstand. Most people notice that voices sound more natural. That is the least interesting part now.
The actual shift is this:
Voice AI is moving from generating speech to participating in a live system.
It has to listen while speaking, recognise when a user is done, handle an interruption, carry emotion without becoming creepy, use tools, survive background noise, switch languages and recover when it gets something wrong.
That is why I think the next two years will change the voice market more than the previous ten.
2016: make the waveform believable
Google DeepMind published WaveNet in 2016.
Instead of stitching together recorded fragments, it generated audio one sample at a time. On US English, DeepMind reported a mean opinion score of 4.21, compared with 4.55 for human speech, while reducing the quality gap to previous systems by more than 50%.
It also needed enormous computation. Beautiful research. Not something you casually put inside a support call.
Tacotron and Tacotron 2 then simplified the stack. Tacotron 2 reported a score of 4.53 versus 4.58 for human recordings.
By 2018, the research question was already shifting from "can a machine speak?" to "how do we make this simpler and more controllable?"
Still, these were synthesis systems. Text went in. Audio came out.
No turn-taking. No reasoning. No action.
2023: three seconds becomes an identity
Microsoft's VALL-E showed in early 2023 that a three-second recording could condition a new voice. It was trained on 60,000 hours of English speech and treated audio tokens more like language-model tokens.
That changed the product imagination.
Voice was no longer a small library of carefully recorded speakers. It could be a property sampled from a person.
Meta's Voicebox expanded the task again. One model could generate speech, edit a segment, remove noise and transfer style across languages. Meta reported it as 20x faster than VALL-E and substantially better on word error rate in its comparison.
The model was becoming a general audio editor, not a narrator.
But the user experience still mostly behaved like a walkie-talkie:
- You speak.
- The system transcribes.
- The language model thinks.
- A speech model reads the answer.
That pipeline can sound wonderful and still feel dead.
2024: the voice learns to interrupt
Humans do not take clean turns.
We overlap. We say "hmm" to show we are following. We stop when someone cuts in. We change a sentence because we hear confusion. A useful phone conversation contains lots of tiny timing decisions that never appear in a transcript.
Kyutai's Moshi was an important milestone because it treated conversation as full-duplex audio. The system could listen and speak at the same time, handle overlap and respond with roughly 200ms practical latency.
For context, the paper notes that human response latency averages around 230ms and overlap occupies roughly 10% to 20% of speaking time.
That is the difference between "audio attached to a chatbot" and a spoken system.
Once the system hears continuously, timing becomes part of intelligence.
2025-26: the stack starts collapsing
OpenAI's Realtime API brought speech-to-speech systems into a developer product. Google added native audio dialogue with affective conversation, proactive audio and tool use. Qwen2.5-Omni gave open-source builders a real-time multimodal Thinker-Talker architecture.
The pipeline did not disappear overnight. Cascaded systems remain extremely useful because each piece can be swapped, measured and controlled.
But the direction is clear.
Speech recognition, reasoning and speech generation are becoming one continuous model interaction, especially in consumer experiences.
The gains are not only aesthetic:
- The model can hear emotion and non-verbal cues that a transcript discards.
- It can start responding before a sentence fully ends.
- It can update its response while listening.
- It can preserve timing and tone through translation.
- It can decide that a sound is relevant without turning it into text first.
Voice is becoming a native input-output space for intelligence.
Naturalness is heading toward demo saturation
I do not mean voice quality is solved.
Prosody still breaks. Long-form consistency can drift. Emotion can sound performed. Background audio and accents expose systems quickly. Cloned voices create obvious consent problems.
But for English product demos, "it sounds human" is no longer a durable differentiator.
This is the same thing that happens in every model market. A capability goes from impossible -> magical -> expected -> checkbox.
The value moves one layer up.
Can it finish the job?
The voice AI market has four layers
Market maps can make voice look like 150 unrelated logos. I see four actual jobs.
1. Foundation and speech models
OpenAI, Google, Meta, Qwen, ElevenLabs, Cartesia and others are pushing speech generation, recognition and native audio reasoning.
This layer is capital- and data-heavy. Quality moves quickly and API pricing keeps compressing.
2. Speech infrastructure
Deepgram and similar companies optimise transcription, synthesis, streaming, deployment, observability and enterprise reliability.
This layer wins when "works in a demo" has to become "works through noise, telephony, scale and compliance".
3. Horizontal voice-agent platforms
Vapi, Retell and peers make it easier to assemble calls, models, tools, routing, testing and providers.
This layer grew because the raw stack was painful. It will remain useful, but the pricing pressure is real as model providers add orchestration and vertical apps build more of the stack themselves.
4. Vertical outcome owners
Healthcare scheduling. Restaurant calls. Insurance intake. Collections. Recruiting screens. Sales qualification. Support.
These products can price against a resolved task, booked appointment or collected payment. They own workflow context and usually integrate into the system where the outcome is recorded.
This is where I expect the most durable companies to emerge.
a16z's 2025 voice-agent update found that 22% of the YC class it examined was building in voice, with 69% of the voice companies targeting B2B. Sierra Ventures mapped more than 150 companies across the stack.
That much attention means the generic layers will get crowded fast.
Follow the money: voice is becoming an operating layer
The funding tells the same story.
ElevenLabs raised $500 million at an $11 billion valuation in February 2026. It later said ARR passed $500 million in the first four months of 2026. Those are company-reported numbers, but even with the usual discount, this is not a niche text-to-speech market.
Deepgram raised $130 million at a $1.3 billion valuation and acquired a restaurant voice-agent company, moving from infrastructure toward an applied vertical.
Vapi raised a $50 million Series B and says enterprise revenue grew 10x. Cartesia raised a $64 million Series A around low-latency foundation models.
Capital is funding all four layers, but the stories are converging:
- models want enterprise workflows
- infrastructure wants applications
- orchestration wants evaluation and governance
- vertical apps want to own more of the stack
The voice alone is becoming the least defensible piece of a voice company.
What English benchmarks hide
This is the part I care about most from our work on Praxy Voice.
An English demo can feel solved while the global product is nowhere close.
Indic languages introduce different phonology, scripts, code-mixing, named entities and data availability. A person may start a sentence in Hindi, insert an English product name, use a local place name and end in a regional dialect. The transcript can look almost right while the interaction feels obviously wrong.
We built Praxy Voice with 1,220 hours of licensed Indic audio because evaluation on clean English speech was not telling us enough. Our related work on phonological similarity, language-aware speech evaluation and a TTS-STT improvement flywheel kept pointing to the same product lesson:
Speech quality is local.
A system can top a broad benchmark and fail on the exact family name, district, medicine or code-mixed phrase that determines whether the user trusts it.
By 2028, I expect enterprise buyers outside English-first markets to stop treating multilinguality as a language dropdown. They will evaluate the actual acoustic and cultural distribution of their calls.
That creates room for local data, routing, pronunciation systems and evaluation layers even as foundation models improve.
Six predictions for 2028
Predictions are easy when they cannot be checked. So I am attaching a confidence level and a failure condition to each one.
1. Consumer voice moves to native audio by default
Confidence: high
For consumer assistants, the emotional and timing information lost in transcription is too valuable. Native speech-to-speech systems will become the default experience, with text appearing as an optional artifact.
This is wrong if the top consumer assistants in 2028 still visibly behave as STT -> LLM -> TTS systems with turn-based latency.
2. Enterprise voice stays modular for longer
Confidence: high
Banks, healthcare providers and large support organisations need to inspect what was heard, what policy was applied, what the model said and what tool ran. Modular systems make substitution, audit and failure analysis easier.
Native audio will enter the stack, but often behind a control plane that exposes transcripts, policies and actions.
This is wrong if regulated companies broadly adopt opaque end-to-end voice models without separate controls or evaluation.
3. Per-minute pricing starts dying
Confidence: medium-high
Minutes are an infrastructure cost, not a customer outcome.
As model and telephony costs fall, buyers will push toward booked appointments, qualified leads, resolved calls and collected revenue. Platform margins based only on marking up minutes will compress.
This is wrong if per-minute pricing remains the dominant unit in enterprise contracts in 2028.
4. Voice becomes a wedge, not the whole product
Confidence: high
The most valuable "voice AI" companies will gradually stop describing themselves that way.
They will be restaurant operating systems, patient-access platforms, collections products or support systems. Voice will be their fastest interface into a workflow.
This is wrong if broad horizontal voice layers keep stronger pricing and retention than workflow owners.
5. Non-English evaluation becomes a procurement feature
Confidence: high
Buyers will ask for performance by language, accent, code-mix, domain and noise condition. One aggregate word error rate or a polished demo will not be enough.
This is wrong if global deployments are still selected mainly through clean English WER and MOS scores.
6. On-device voice becomes a real tier
Confidence: medium
Wake-word detection and transcription already run locally in many products. Better small models and personal hardware will move more listening, memory retrieval and low-risk action on-device.
The benefit is not only cost. It is latency, availability and the ability to keep sensitive ambient data off a server.
This is wrong if local voice remains marginal outside basic transcription and wake words.
What I would build around
If I were starting a voice company now, I would not begin with "we have a very natural voice."
I would begin with a workflow where:
- voice is the easiest interface
- the user already has a reason to call
- the outcome can be measured
- domain context improves with use
- errors can be recovered, not hidden
- multilingual reality is a product advantage
- a system of record sits on the other side
Then I would make evaluation part of the product from day one.
Did the system understand the right entity? Did it follow policy? Did it transfer at the correct moment? Did the outcome happen? What does failure look like for this specific customer?
The seductive voice demo is the start of the sale.
The boring evaluation dashboard is often the start of the company.
By 2028, voice will stop feeling like a feature
We are used to thinking of software as something with a screen.
Voice makes software present in a different way. It can sit inside a phone call, a car, a headset, a home, a field worker's day. It can notice urgency, ask a follow-up and act without asking the user to find the right menu.
That will make voice much larger than "talking to an AI".
It will also make the failures more consequential.
By 2028, the winners will not be the companies with the most human-sounding demo. Human-sounding will be the admission ticket.
The winners will be the ones that can listen in the real world, earn permission, complete a job and prove they did it correctly.
The industry spent a decade teaching machines to speak.
The next two years are about teaching them when to listen, when to act and when to shut up.
Research note: technical milestones link to original papers or first-party releases. Funding and ARR figures are company-reported unless otherwise stated. Predictions are my own and include explicit conditions under which I would consider them wrong.