Week in voice AI #6
Pipelines are getting bigger
I am Ivan Mehta, a consumer tech reporter at TechCrunch. For more than a year, I have covered different aspects of voice AI. This newsletter is an experimental attempt at covering what is happening in this industry, which is growing at a rapid pace.
Top News
Google gets Hume AI’s team in another acqui-hire
Google is notorious at acqui-hiring a startup’s top talent. We have seen this with Character.AI and Windsurf. This week, Google grabbed the top brass of voice AI startup Hume, including CEO Alan Cowen.
The main reason behind this acquisition is to improve Gemini’s voice capabilities. This is not surprising given that there is a rise in voice-based conversational features with assistants.
Hume’s core strength is making “emotionally intelligent” models. The company claims that its data and models are tagged with over 200 emotions and over 400 voice characteristics. This give pontential model users a way to make their assistants sound more human and realistic. The company also focuses on measuring different successful outcomes based on different voice traits and nuances, for instances how the meaning of a sentence changes with a switch in tone.
The company’s new CEO, Andrew Ettinger, who is also an investor in the company, said Hume AI will release new models in the coming months and is on track to bring over $100 million in revenue this year.
OpenAI partner LiveKit gets $100 million and turns into a unicorn
Voice infrastructure companies are becoming more important as the use of the medium by both enterprises and consumer companies is rising. LiveKit is one startup that has an open-source infrastructure for voice used by top companies like OpenAI, xAI, Salesforce, and Tesla. The company raised $100 million in funding from Index along with Salesforce Ventures, Altimeter Capital Management, Hanabi Capital, and Redpoint Ventures.
LiveKit’s key offering is network pipelines that enable low latency for voice agents. Plus, it provides an agent builder, a telephony integration, and hosts this infrastructure on its cloud. The toolkit also handles things like turn detection for developers.
The numbers related to LiveKit usage are impressive. The startup was handling 30 to 40 million daily sessions last June. That number would have likely gone up. The startup claims that its LiveKit Agent framework is being downloaded 1 million times per month.
Two Indian voice AI companies pick up funding
India is an important voice AI market, because it has a massive userbase that relies on voice as a communication medium with each other, digital tools, and companies. This means organizations ranging from enterprises to FMCG companies are ready to pay money to use voice AI for customer support, hiring, and sales.
Banking on this momentum, two Indian startups, Bolna and Ring,g picked up funding. Bolna is an orchestration platform, akin to a Vapi or LiveKit, that raised $6.3 million from General Catalyst, Blume Ventures, and Y Combinator. The company is already handling 200,000 calls per day and is on track for $700k in ARR.
Ringg AI is a competitor that raised $5.5 million from Arkam Ventures, with participation from Groww’s Founder Fund, and Cred founder Kunal Shah. The platform is handling 1.5 million conversations per month, serving companies like publicly listed trading platform Groww, logistics company Shiprocket, and telemedicine platform Practo.
A key insight is that for both platforms, there are plenty of self-serve clients paying them to run pilots or deploy voice agents without getting in touch with the sales or support team. Though companies might need some customization at one point, they are fine to get started with a platform that is good enough to serve base-level requirements.
Quick bytes
ElevenLabs appointed former Elastic GM Karthik Rajaram as India country head. Krisp also hired Vimal Nair as growth head for the country. Global companies are likely make more growth appointments in the Indian market.
Qwen open-sourced its text-to-speech models of 1.6B and 0.7B size.
Shunya Labs launched a new speech recognition model that understands Indian code-mixing and multilinguality.
Sweden’s music industry body banned a song from official charts because it was AI-generated. The song had been streamed over 5 million times on Spotify.
ElevenLabs created an album with Grammy-nominated artists, who used the startup’s music-generating AI.
OpenAI is likely to ship its first device in 2026, and it might be a pair of earbuds.
Steve Downes, a voice actor behind Halo’s Master Chief, urged not to clone his voice. On his YouTube channel, he said that while voice cloning is generally harmless, it could deprive an actor of their work. We have previously seen instances of gaming voice actors raising concerns around voice cloning.
Stat
In India’s voice market, English and Hindi are still the biggest languages served. For an international model provider, Deepgram, Hindi is the second biggest use case. However, other languages are catching up. For instance, Bolna said that its calls in Tamil are steadily rising, followed by other languages like Marathi and Telugu
Signals & experiments
NVIDIA’s new speech-to-speech model, PersonaPlex-7B, is gaining attention. The key part about this model is that it can listen to input while talking. What this essentially means is that, in theory, you don’t need any turn-taking or special interruption handling.
Here is how it works according to NVIDIA:
The model operates on continuous audio encoded with a neural codec and predicts both text tokens and audio tokens autoregressively to produce its spoken responses. Incoming user audio is incrementally encoded and fed to the model while Personaplex simultaneously generates its own outgoing speech, enabling natural conversational dynamics such as interruptions, barge-ins, overlaps, and rapid turn-taking. Personaplex runs in a dual-stream configuration in which listening and speaking occur concurrently. This design allows the model to update its internal state based on the user’s ongoing speech while still producing fluent output audio, supporting highly interactive conversations
While the conversational part is impressive, there are a few lines of thought when it comes to the intelligence of the model. Users/developers might be able run base voice-based tasks locally with these kinds of models. But they would need external models or tools to access data or business logic, as pointed out by Vinod Khosla.
A few people pointed out that it would be good to run small models with memory context rather than rely on large language models for context. A lot of good voice consumer applications still bank on cloud processing, even if it doesn’t involve any kind of knowledge fetching. If local models get more powerful, a lot of these tasks are offloaded to the device and become available systemwide.
Thank you for tuning in. Keep listening.
Email: voiceaiweek@gmail.com or im@ivanmehta.com






