Week in Voice AI#9: Value of calls
Echos of India AI summit
I am Ivan Mehta, a consumer tech reporter at TechCrunch. For more than a year, I have covered different aspects of voice AI. This newsletter is an experimental attempt at covering what is happening in this industry, which is growing at a rapid pace.
Top news
Voice AI was central at India’s AI summit as the call volumes ramped up
There is a reason why India is called a voice-first country. At India’s mega AI summit, while there was a lot of deal signing around enterprise contracts and data centers, a lot of the talk surrounded voice AI. Google said during its events that Indian users were among the highest global adopters of voice and visual search. OpenAI also said that it saw a high volume of voice conversations.
There is a rising number of calls that are handled by AI in India. Bajaj Finance CEO recently said that the company expects listen to 100 million calls in a year. Travel company Ixigo said more than 76% of its calls were handled by AI in the quarter that ended in December.
Companies powering these voice calls are also seeing high volumes. Enterprise AI company Kore.ai CEO told me that they handle roughly 200 million calls through AI in India. India’s AI company Sarvam said that its voice platform currently handles 2 million minutes of calls every day. The company’s co-founder, Partyush Kuma,r said that this volume is expected to cross 100 million minutes a day by the end of this year.
Separately, tons of startup founders are trying to build pipelines for voice calls to cater to different industries. Some of them, in stealth mode, are toying with use cases beyond sales and support to try to make a name in different industries like real estate.
It’s hard to estimate India’s voice AI market. One way to gauge is to take a part of the call center market, which is in billions. But voice companies I have talked to said that AI is enabling use cases that go beyond customer support, such as marketing team follow-ups.
Besides business talk, there were model releases too. Sarvam launched its speech-to-text and text-to-speech models along with a dubbing platform. Gnani, which focuses on BFSI, launched a zero-shot text-to-speech model with voice cloning capacity. The Indian government-backed Bhashini platform released an open-source voice agent stack called VoicERA.
For model makers, India is not just a high-volume but a high-value market. Both ElevenLabs and Deepgram count this as their second-highest-grossing market. ElevenLabs CEO Mati Staniszewski told MoneyControl that the company’s goal is to reach $100 million in annual recurring revenue in India. Cartesia is also offering local data residency for enterprises by teaming up with Blue Machines.
Voice companies that offer a diversity of solutions, in terms of accents and voices, will find success in India.
Wispr Flow launches on Android
Dictation app Wispr Flow launches on Android today after being on Mac, Windows, and iOS for a while. Besides Typless, there aren’t as many AI-powered voice dictation apps on Android as on iOS, so this is a good launch to expand Wispr Flow’s reach.
The new app has a floating bubble interface, which is different than having a separate keyboard for iOS. This also allows you to type in certain words if the dictation goes awry.
The company also noted that it has rewritten its core infrastructure, and that would mean that dictation is 30% faster than before.
Quick bytes
Twilio said that voice revenue growth accelerated to “high-teens” and has been best since 2022. The company also said that voice AI revenue growth is expected to grow by 60% year-over-year in Q4.
ElevenLabs insured its voice AI agents AIUC-1 certification, which is an equivanelt to “SOC 2” certification but for AI agents. For this certification, AI systems go through over 5,000 simulated tests to measure security and reliability.
Former NPR host David Greene sued Google for cloning his voice to be used in its knowledge tool NotebookLM.
Google launched music generation in Gemini through a new Lyrica 3 model. You can generated 30 second clips by using text, images, and videos.
Speechify launched its SIMBA 3.0 model for all kinds of voice applications.
Apple Music now lets you build playlists using AI prompts in iOS 26.4 beta.
Signals & Experiments
I really like using voice dictation to reply to emails and messages. While my emails are largely in English, I have had a good success rate in using apps like Wispr Flow and Monlouge; messages are a different beast.
I, like many Indians, speak in a mixed-code language. This means I’d often mix English with Hindi or Gujarati (the primary languages I speak). This creates a conundrum for dictation apps, as either they will output the whole thing in Devanagari (Hindi’s script) or Gujarati script, or get the output wrong. This deters me from using voice AI in messaging apps unless I am replying to someone in English.
However, some releases that I have seen in the last few weeks have given me some hope. First, Sarvam released a few voice models and tools that cater to code-mixed languages with output in both native and romanized script. This is very useful for everyday conversations. There is a chance that I prefer Roman script, and the person on the other end prefers the native script. By using these tools, both of us can look at the same conversation and understand the context better.
I played around with some of the tools on Sarvam Playground and in its new ChatGPT competitor app, Indus, and the output was consistent. I hope someone builds consumer voice tools using these APIs.
Second, along with its Android app, Wispr Flow is also launching a dedicated Hinglish model. This will roll out across apps over the next few weeks, and in my early testing, I was satisfied with the initial results.
Thank you for tuning in. Keep listening.
Email: voiceaiweek@gmail.com or im@ivanmehta.com


