#16: ElevenLabs' growth gets louder with $430M ARR milestone
The company added $100 million in net new ARR this quarter
Top News
ElevenLabs clocks in $100 million in net new ARR for Q1 2026
ElevenLabs seems to be growing louder with its growth story. The company has had a stellar quarter. It closed the biggest round in voice AI with a $500 million Series D round in February at a $11 billion valuation. But apart from investors pumping in money, the company seems to be on an upward trajectory for adding new revenue.
CEO Mati Staniszewski said that the company has added $100 million in net new ARR this quarter. Earlier this year, he mentioned that ElevenLabs ended 2025 with $330 million ARR.
ElevenLabs took 20 months to cross $100 million in ARR, 10 months to cross $200 million, five months to reach the current figure of $330 millon, and three months to add $100 million and reach $430 million, or $450 million if we go by what Staniszewski said on the Cheeky Pint podcast with Stripe CEO Patrick Collison.
If the growth rate persists, the company could end up with over $550 to $600 million ARR by the end of Q2 and possibly reach $1 billion ARR by the end of the year.
The company signed important contracts this quarter, including Deutsche Telekom, Revolut, and Klarna.
Apart from signing enterprise contracts, it also teamed up with consultancy firms like Boston Consulting Group (BCG) and Deloitte, so they can suggest ElevenLabs to their clients or use the voice AI company in its own customer service deployments. A key point Staniszewski mentioned on the Cheeky Pint is that the company has over 50% of ARR through sales-led enterprise.
Staniszewski said that growth was possible because voice agents powering enterprise conversations (for sales, marketing, and support) became more reliable in recent times, attracting buy-in from large organizations. Plus, he hinted that existing accounts also expanded through cross-department pollination or increased usage of the voice startup’s tech.
Its geographical expansion also seems to be going well. India, which is one of the key markets for the company, now has an expanded team for sales and go-to-market. The market is exceeding sales expectations.
ElevenLabs is also building other capabilities apart from its core speech generation and enterprise agent business. Two moves from this quarter stand out. First, as I broke the news about this first, the company launched an ElevenMusic app and website for consumers to create AI-powered music and discover other tracks. At the moment, this platform might be too early to generate any revenue, but it could be a move to warm up to the prosumer audience.
The second and more crucial part of the business is ElevenCreative, its toolkit for marketing and creation. In February, the company launched Flows, a node-based creation tool. This is where design tools like Visual Electric, Figma-owned Weavy, and Flora operate in. ElevenLabs’ move is to involve designers in a creative process.
Staniszewski hinted at a future where ElevenLabs gets more involved in ad creation and the marketing process during The Cheeky Pint podcast.
“Then you can have all the way to the marketing use cases where we are your partner for working even outside of the conversational agent space of how you create a great marketing campaign,” he said.
This is a lucrative market where companies like Adobe and Canva operate. While ElevenLabs won’t directly compete with them for static campaigns, the company will jostle for ad and creative dollars where there is more audio and video. But regardless, ElevenLabs seems to be on track for a high-growth year.
Signals & Experiments
As I live in India, WhatsApp is one of the most used applications for me to converse with friends, family, and professional contacts. For the last year or so, I have been using apps like Wispr Flow, Monlouge, and Typeless for replying to messages on WhatsApp and have largely ditched voice notes.
Some of my friends send me voice notes if they have something to say that is more than a few sentences. Or even delivery or service folks send voice notes to convey information or ask questions. Because I work from home, I am able to listen to those voice notes, but at times, I just want to read the gist and reply if needed. Though when I try to transcribe voice notes on WhatsApp, I often see so many blanks even if the other person spoke in English, and the words were easily legible.
I don’t know what model Meta is using, but it is a terrible one. They might be using a less capable model that works within the confines of end-to-end encryption. But we do have models that work well on edge devices like phones. As people use voice more to interact with phones through apps like dictation and assistants, they will demand better models on both ends of the speech spectrum.
This will also increase demand for apps and models that can work with low or no connectivity. We already saw interest in Google’s dictation app that can work offline. We need more of that on phones and laptops.
Model Behaviour
Google Gemini 3.1 Flash speech model gets expressive with audio tags
ElevenLabs made its name for high-quality voices. Google now wants to move in the same direction with its new Gemini 3.1 Flash text-to-speech model. Besides high-performance benchmarks, the new model has tags that you can write in the script and direct the model to be expressive. For instance, you can add [excited] or [bored] as speech tags to make the AI-generated speech sound like that.
ElevenLabs has had this kind of tech with its v3 Alpha release last year. But more companies could follow suit as they are evolving speech models beyond monotonous agent use cases. Expressive controls are especially important as voice and video companies are looking to be more involved in the creative process, including ads and fictional video clips.
xAI jumps into the enterprise voice game with new individual speech recognition and generation APIs
Last year, The Information reported that xAI is struggling to sell its technology to enterprises. Voice is one such tech where there is a growing enterprise demand. xAI seems to be aligned with that thesis with the launch of individual text-to-speech and speech-to-text APIs that also power Grok’s voice mode.
The company is positioning itself as a cheaper alternative to competitors by pricing these APIs lower. For instance, the speech-to-text API is priced at $0.10 for an hour of batch processing as compared to Assembly AI’s latest model’s API priced at $0.21 for an hour of processing.
Grok’s tex-to-speech API is also significantly cheaper at $4.20 per million characters (they couldn’t resist it) as compared to $50 per million characters for ElevenLabs’ business plan.
Quick Bytes
Voice-based security threats can go beyond cloning. Bleeping Computer noted that a new platform called ATHR uses voice calls to carry out phishing attacks and get login details of platforms like Google, Microsoft, and Coinbase. The report said that on dark web forums, the software is advertised at fees of $4,000 and 10% comission from profits
Regulators are looking at voice-based scams now. The U.S. Federal Trade Commission (FTC) said that AI is playing a key part in the increased threat of Robocalls. Meanwhile, Sen. Maggie Hassan sent a letter to ElevenLabs, LOVO, Speechify, and VEED for them to explain what kind of guardrails they are taking to prevent malicious usage. This could get tougher as a study in Nature suggested that people often misidentify an AI voice as a real one
In India, parties are using deepfakes on social platforms in political campaigns in the state of Tamil Nadu. Election bodies all over the world will need to keep a watch on how different political entities use AI to reach voters
Zoom partnered with Sam Altman’s World to prove that participants in a specific meeting are humans. World will verify the image based on the device capture, original capture that was registered in an Orb, and a video frame that other participants see. After verifications, participants will see a badge next to a verified person’s tile
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
Email: voiceaiweek@gmail.com or im@ivanmehta.com
This newsletter is by Ivan Mehta, a consumer tech reporter at TechCrunch. For more than a year, I have covered different aspects of voice AI. This newsletter is an experimental attempt at covering what is happening in this industry, which is growing at a rapid pace.





