Signal
Meeting notetakers are still a hot and crucial tool in the productivity stack. Wispr released its own notetaker a few weeks ago, and then scheduling tool Calendly followed suit with its own notetaker. There is no dearth of meeting notetaking tools, yet more companies want to own this part, because transcripts contain action items and hidden prompts that could be used to trigger automation processes, which is the core promise of AI tools.
In Focus
In a survey conducted by caller ID company Truecaller, 76% of respondents in India said that they prefer communicating with businesses through calls. This is a small indicator of why voice AI could be an important business in India, given the potential of automating processes ranging from customer service to appointment booking and sales.
Separately, in a conversation with Pratyush Kumar, co-founder of one of India’s marquee AI startups, Sarvam, VC firm PeakXV’s Rajan Anandan said that voice AI can be a $1 billion category in the country, with calls costing less than ₹5 per minute (5.24 cents), which is below the cost of a human agent in Business Process Outsourcing (BPO) offices.
Depending on the cost of calls, India’s domestic BPO market is estimated to be around $5 billion. The opportunity to serve customers is large, but the race to the bottom on pricing might not be the way. Several voice AI startups in India offer rates lower than ₹5 per minute for automated calls. The challenge with that game is that there is always a bigger fish, which can offer similar capabilities at a lower price.
The argument for voice AI is a bit deeper than cheaper calls. Since AI agents don’t need rest, they can be available at any time. Plus, they take up low-priority tasks that human agents might not get to. Observability platform Last9’s co-founder Kuldeep Dhankar said that agents can take up tasks that humans don’t have time for.
The initial promise of voice AI was handling calls at volume. Then it shifted to providing context-aided customer service. Now, the aim is to automate complex tasks such as user onboarding through voice.
A large market in India could be defined by the crude term “form filling.” Sites and apps in India are hard to navigate. People often find the process of buying something or booking an appointment cumbersome. So they call the helpline numbers and keep waiting until a human agent picks up.
In the same vein, there might be an opportunity to provide government services through voice calls or make it easier for people to access public services. But only a handful of companies will be able to sell their services to governments.
The argument people working on voice, and companies that have adopted voice, make is that there is a need to create new workflows that AI can handle, rather than using existing human workflows, because AI will fail at it.
The bull case is that voice actually becomes an interface, and tasks that took many clicks/taps across many pages can become a simple voice command. This command could be something within the app or could also be a call.
The bear case is that startups in India don’t innovate enough on the technology layer, create differentiation with models, and play the sales game of getting pricing driven down.
Chaitanya Chokkareddy, CTO of CX platform Ozonetel, said that despite startups claiming to offer various capabilities, most customers ask for lower pricing. Plus, enterprise customers are not looking to replace human support and outreach teams just yet. His argument, made in another post, was that a voice agent’s capability to replace a human is not there yet.
One narrative I have heard from voice AI startups is that they are not looking to replace human teams either. They want to deploy voice AI to let agents add to the tally of outbound calls for available leads and incoming support calls.
India is also a tough market because of language barriers. Two people speaking the same language can speak differently. They also mix languages mid-sentence in an unpredictable way. On top of that, users are often in a noisy or unstable network environment.
These are hard technical challenges, and the return on investment might not be that high when you solve for these problems. As Last9’s Dhankar said, voice startups will need to sell at a very high volume to get good margins.
In India, there are 50-60 startups working on voice AI. Some of them are established CX players, and some of them are newer startups working across sectors. For voice startups catering to every industry, there is a chance of losing the price war and burning a lot of cash, or losing contracts to bigger players offering better service packages.
In voice, there are two types of startups that are typically acquisition targets. First, startups that have solved technical problems for voice models, noise cancellation, evals, or similar other areas. Second are companies that focus on a particular vertical, like finance, healthcare, or food, and have had great success in the industry. Current model makers or agent platforms can gobble them up with established sales motion and integrate their products.
India is a potentially large voice market because people like to talk and use voice to connect with businesses and even navigate apps. The country is in the top five revenue-generating markets for voice AI startups like ElevenLabs and Deepgram. If businesses can make it easier to interact with them through voice, there is a big opportunity. But if startups stick to just enterprise calling and price wars, they might be run into the ground.
Model behaviour
Daily CEO and Pipecat creator Kwindla Kramer launched PhoneLLM, a fine-tuned NVIDIA Nemotron Nano 30B for voice agents. Kramer’s argument is that often voice stacks use models that use thinking mode to get answers or context, which increases latency. The new PhoneLLM model turns off the thinking.
Google released Gemini Transcribe 3.5 in both live and batch mode. This model powers dictation on the Gemini Mac app and Rambler tool on Pixel and other Android devices.
Quick Bytes
Artists including Nicola Coughlan, Hugh Bonneville and Matt Lucas have asked the UK government to take action against unauthorized voice cloning and give rights to their voices. This is an important one to track as more countries across the world are looking at providing protection against unauthorized voice cloning.
Australia and Finland ban AI music from charts
Australia’s music body said that it is barring fully AI-generated tracks from its official charts. Finland also made a similar move with its new rules. To be clear, these changes are not against using AI, as they still accept songs where there is significant human involvement.
The Indian AI company launched a new 30B large language model, along with a platform to quickly build voice agents. The model is not specifically tuned for voice use cases, but rather a generic language model.
Plaud and Boat bet on earphones as a voice AI interface
Plaud launched a new device called Plaud One, which is a pair of earbuds that also have an eSIM in its case to trigger agents. Separately, India’s Boat released a Crest AI platform, powered by the Gemini model, to let users converse with AI assistants. These two are separate roads meeting at one point. One is a meeting note-taking company taking a stab at earphones; another is an audio company taking a stab at AI conversations. It is not clear, though, if having AI capabilities is a pull for consumers.
Wispr is prepping for an AI assistant and offline support
Wispr is preparing to release an AI assistant that allows users to query all meetings. It also seems to be preparing offline capability.
Call for Sponsors
I am looking for sponsors for this newsletter to support my reporting. I spend hours every week researching different topics and reporting on various slants. I am based out of India, so getting a Stripe account is a very tough task, and that prevents me from turning on subscriptions for the newsletter.
If you want to support independent writing and reporting, get in touch with me at im@ivanmehta.com
Deals Corner
Ringg ($10 million): Indian voice AI company catering to enterprises raised an extension round. The startup also wants to get into other modalities like WhatsApp and chat.
Investors: PeakXV.
Stability ($76 million): Stability produces a lot of models for image, video, and music generation. But given the investor group, the round seems to be tilting towards sound and music generation.
Investors: Electronic Arts, Sony Music Group, Universal Music Group, Warner Music Group, and investment firms AMD Ventures and Pacific Alliance Ventures.
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com








