#18: Into ElevenLabs' research bet
Deepgram releases a new multilingual speech recognition model for real-time agents
Top News
ElevenLabs acquires a team from a Polish voice AI startup to bolster research
ElevenLabs continues to grow in revenue and valuation, but it also wants to remain, at its core, a research lab around sound and audio. The company’s value came from its ability to create unique, expressive voices. Now, it wants to accelerate research in voice generation by acquiring a team behind a Polish startup called Papla.
Papla was founded by Hubert Siuzdak, Jakub Abramczyk, Dominik Stachura, and Tomasz Dentko.
The startup largely focuses on voice generation with an API that provides realistic voices in English in different accents. The company claims on its website that you can clone a voice in just 10 seconds of audio. The company’s model ranked higher than models from ElevenLabs and Cartesia on the TTS Arena V2 leaderboard.
At the core, research, and the Polish connection, which possibly attracted ElevenLabs to hire the team.
The key achievement of the startup comes from Siuzdak’s research. He wrote a paper on solving for natural-sounding voice generation with better pronunciation through a two-staged architecture. His second research was around making a fourier-based vocoder output high quality while using fewer resources. His third research was focused on using low bitrates such as 0.98 kbps for speech and 2.6 kbps for music for high-fidelity output.
All this means that ElevenLabs can chase top-end research problems that could eventually result in products.
Deepgram’s new speech recognition model focuses on multilingual code-switching use cases
Fresh off the $130 million funding in January, Deepgram released a new speech recognition model, or rather, a multilingual version. The company first introduced its Flux model for English last October. This new line of models has built-in turn detection and interruption handling.
The company has now introduced a version that supports 10 languages and can handle mid-sentence code switching (people switching from English to Hindi or Spanish). The model focuses entirely on real-time conversations, which is helpful for voice agents.
Deepgram’s CEO, Scott Stephenson, said that Flux’s technical design allows the model to pass the context around better.
“In order to pass the audio Turing test, you have to pass context around. For instance, the end-of-turn detection. That’s context being passed around, noticing that somebody is about to end their sentence or about to switch to a different language when they speak, and then coming back. And this is extra context, which is useful downstream in the system,” he told me.
Stephenson thinks that as Flux is well-suited for real-time voice agents on call, the company’s Net Promoter Score (NPS) will improve. He also thinks that baked-in turn detection is unique to Deepgram’s model and will help the company stand out in the market.
The new Flux model is priced at $0.47 per hour of streaming, as compared to $0.45 per hour of AssemblyAI Universal-3 Pro and $0.96 of Google Cloud Chirp.
(Full interview with Deepgram CEO Scott Stephenson coming this week)
Signals & Experiments
Pronunciation learning has been one of the prime use cases for voice in translation tools. This week, Google Translate rolled out a new feature that lets you practice your pronunciation and enunciation on the tool’s 20th anniversary.
The feature tells you the level of your pronunciation, listens to the ideal pronunciation, and compares both versions as well. At launch, the feature is available in India and the U.S. with support for English, Hindi, and Spanish. This is a great tool for language learning, and I personally believe that hearing and listening to words in a language and then speaking it makes a lot of difference.
In a way, this feature is something out of the Duolingo book, but without structured sessions. I think it would also be useful for Google (or some other tool) to list common words and phrases a traveller can use when they are visiting a particular country. They can have fun while practicing and also preparing for the trip.
Quick Bytes
Service AI company Avoca, which handles calls for businesses, lets them design campaigns and coaches them on the sales process, reached unicorn valuation with its latest raise. The company said it has raised $125 million across three rounds from investors such as General Catalyst, Meritech Capital Partners, and Kleiner Perkins. But it hasn’t really broken down how much it raised in seed, Series A, and Series B, respectively
India’s Bajaj Finance handled 52 million calls through AI in FY 2026, the company said during its latest earnings calls. It also said thanks to AI, it is spending a third less on call center calls, as 30% of its outbound is voice agents
Otter rolled out its enterprise connection suite and a way to search across the entire knowledge base. The company is finally catching up to other players in the market, such as Read AI and Fireflies.ai
Wispr Flow launched its marketing campaign in the Indian city of Bengaluru on auto rickshaws, but there were mixed reactions as some folks said the language should have been a mix of English and Kannada (the state language). Some others said that the name reminded them of sanitary products. Either way, more people now know about the dictation app
Taylor Swift joins a horde of celebrities to trademark her voice and protect it from deepfakes. Reuters noted that the pop star’s voice has been used in fake advertisements and political campaigns
A new survey from Luminate indicates that Gen Z’s interest in AI-generated music has dropped significantly. As we reported last week, despite the rise in the number of AI tracks, Deezer said that overall listenership remains lower than 3%
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
This newsletter is by Ivan Mehta, a freelance tech reporter at TechCrunch. This newsletter is an attempt at covering what is happening in the industry of voice, audio, and music.
Thank you for tuning in. Keep listening.
Email: voiceaiweek@gmail.com or im@ivanmehta.com








