Week in voice AI #8
Eleven billion for ElevenLabs
I am Ivan Mehta, a consumer tech reporter at TechCrunch. For more than a year, I have covered different aspects of voice AI. This newsletter is an experimental attempt at covering what is happening in this industry, which is growing at a rapid pace.
ElevenLabs nabs $500M in funding at $11 billion valuation
ElevenLabs was founded in 2022. In 2023, when it raised its series A funding, it was valued roughly at $100 million. In 2024, it got to the unicorn status, and in early 2025, it was valued at $3.3 billion after a $180 million round. Later in 2025, the valuation doubled to $6.6 billion through a secondary round. Now, with its latest $500 million round, it is $11 billion. Essentially, in the last three years, the valuation has increased by 110x.
The latest round saw Sequioa Capital, which participated in the secondary round, leading the charge. a16z quadrupled its investment in ICONIQ, which led the last equity round, and tripled its investment. There might be one or two other investor names that we might see later this month, as ElevenLabs indicated in its blog post. It is likely that NVIDIA also increased its investment in the voice AI company.
Besides investment growth, the company has had a good revenue run rate as well. It ended the year at $330 million in ARR and took only five months to reach there from $200 million in ARR.
The company’s own blog, along with Lightspeed’s blog, suggested that ElevenLabs would focus more on agents and generation beyond voice. Earlier this year, the company announced a partnership with LTX to create videos from audio, and we might see more announcements in that direction.
Sarvam releases a bouquet of India-focused models
Sarvam, one of the most well-funded Indian AI companies, released a bunch of models as part of its 14-day drop. India is a voice-first country as there are many nuances and differences between spoken and written language. Many folks largely “talk” using WhatsApp through voice notes, and know their way around smartphones and apps by learning a few keywords.
Savarm Audio is a tuned version of Sarvam 3B, which can handle 22 Indian languages along with English. One key part of Indian speech is that people in India often mix English with their native language, and most tools are yet to account for that. The new audio model claims to take care of mix-coded language.
Sarvam Dub is a real-time dubbing model that claims to preserve speaker's voice characteristics. Indian TV news channel Republic used the model to translate India’s union budget speech for 2026 in real time in early February.
Sarvam BulbulV3 is a new text-to-speech model with more expressiveness. The company said that in its test, the model matched or outperformed other competitors like ElevenLabs and Cartesia in Indian languages and handled complex scenarios well.
While these models clock some impressive benchmarks on paper, for Sarvam, the challenge might be to convince Indian companies to ditch models from OpenAI and Google. Plus, these models need to be cost-effective and perform well in real-life scenarios.
Quick Bytes
Sarah McLachlan and Mac DeMarco are supporting a campaign in Canada to stop the spread of unlicensed AI music.
Apple plans to allow third-party voice chatbots that could control CarPlay. This means that CarPlay can work with the likes of ChatGPT or Gemini.
Speechify adds its app to ChatGPT and introduces celebrity voices like Snoop Dogg, MrBeast, or Gwyneth Paltrow.
DeepL makes its voice API generally available with the capability for real-time translation and transcription. The company claims that through this API, call centers can hire customer service experts even if they don’t know a particular language.
Amazon’s Alexa+ conversational assistant is now available to all users in the U.S.
Spotify is adding more limits to its developer mode, requiring them to get a premium account, and limiting test users to five users. The company also reduced the number of API endpoints devs can use.
Must Read
Sky News detailed how fraudsters are using AI in the music world. First, they are creating AI tracks, and then they are building bots to “listen to those tracks” to earn royalties. Thibault Roucou, Deezer’s head of revenue, told Sky that more than 85% of all listens on AI-generated music is fraudulent.
Signals & Experiments
I love testing productivity tools. So as a part of that, I have tested many meeting notetakers. But for more than a year, Circleback was a permanent meeting bot that I used. I also use Granola for a lot of meetings where I don’t necessarily need the audio recording of the meeting.
Recently, I was looking to switch the apps merely because I wanted the new experience, and my Circleback discount code that gave me the app subscription for almost $12 per month was running out. Since technically I am a team of one, I didn’t need a $25 per month subscription, and I started looking for alternatives.
Fireflies is a good notetaker that I have used in the past, and it was giving me a deal to get a year-long subscription for $90. So I caved in.
This is not about one tool being better than another. But it is about switching costs. I didn’t particularly find that there was an easy tool for me to transfer all my meetings to Fireflies. I might be able to use the MCP server connections and transfer recordings, but for an average user, that might be a headache. Another switching cost is lost AI context or insight.
One tool’s AI might give you different summaries and potential action items than others. Plus, it is not clear if you have set up automation in these tools; it is easy to transfer them to a new one.
We are still early in the process of figuring out AI tools that work for us. This applies in both personal and enterprise areas. There might still be tools and expertise for enterprises to move their data from one platform to another. However, for personal users or small teams, these tools don’t seem to be apparent.
Thank you for tuning in. Keep listening.
Email: voiceaiweek@gmail.com or im@ivanmehta.com


