Signal
Large Language Models (LLMs) might be the trickiest and most costly part of the voice stack. Companies are now focusing on this part, in addition to developing voice models. Recently, Pipecat’s Kwindala Kramer released PhoneLLM that turns “thinking” mode off for better latency. Last week, Phonely released an LLM trained on over 10 million phone conversations. We might see more LLMs tuned for the voice stack to reduce latency and give more use-case-specific answers.
In Focus
Microsoft and Meta join Google in transcription model wars
In the past few weeks, there has been a barrage of transcription models. Google released its Gemini 3.5 Transcribe a few weeks ago. Wispr released its ASR model, Canto, along with its fundraising announcement last month.
Last week, two big tech companies showed off their new bets in the transcription model kerfuffle: Meta and Microsoft.
Meta released its Muse Voice Transcribe model, which is the company’s first real-time offering. The model was “extensively verified” for over 25 languages and can identify up to 20 speakers, the company claims.
One interesting technological tidbit was that the model spends a bit more time transcribing hard words and speeds up processing for the words it is confident about. We don’t know how that will play out in real-life performance and if it would increase accuracy at all.
On the other hand, Microsoft released its MAI-Transcribe-2 model with a claimed Word Error Rate (WER) of 5.3% over the FLEURS benchmark. It said that the new model processes the audio faster, at a rate of 1 hour of recording in 10 seconds.
Why is there a flurry of new models on the market? There are a few reasons. All companies now have a ton of audio processing to do internally as well as on their consumer-facing surfaces where it offers dictation. For instance, Meta and Google both offer system-wide dictation through their Meta AI and Gemini apps.
Plus, all these companies are integrating more voice features in their coding and productivity apps. All speech recognition models are helpful if you have a ton of meetings and video content to transcribe and analyze.
The biggest market for these companies would be enterprise calling and call centers. Microsoft also dabbles in legal speech capture and powering clinical scribes.
Meta and Wispr are the newer entrants in the market, and the latter doesn’t even offer an API yet. Wispr clearly wants to focus on becoming a transcription layer for hardware and consumer apps. For Meta, the idea would be to power its small and medium business suite for support and consumer outreach.
Benchmarks are interesting from an academic standpoint, but price is what will attract customers. Google is promising a blended price of $0.54 per hour, Meta’s pricing is at $0.18 per hour, and Microsoft is the lowest at $0.1 per hour.
There is an ongoing notion that “voice models are being commoditized” in the voice AI world. But amid this belief, we have seen more models than ever. Low-hanging fruit for developers is driving prices of these models down. On a macro level, companies still want to own model infrastructure even if prices are getting commoditized. Speech recognition is still not perfect in any sense, especially for non-English languages. There is a lot of work to be done.
Quick Bytes
Google wants you to talk to its apps. In Gmail, you can search for emails through voice and ask follow-up questions. In Keep, you can refer to notes or add a new note by speaking. Docs is the most interesting one, as its voice assistant can fetch information from Gmail and Drive to create a draft.
ElevenLabs deploys its AI to Indian appliance company Havells
I am all for voice commands, but this one seems like a tech searching for a problem. Voice commands to turn on/off lights or change the speed of the fan are good, and maybe in India you need regional language support. One feature in this suite is error resolution, which might be better than “agentic commands.” Through this, users can ask what a particular error code shown on the appliance screen means.
Music company Roland created a new tool called Melody Flip, which has over 250 gen AI loops based on different music themes. Musicians can import these files into their digital workstations (DAWs) and work on them like editable elements.
Suno defines its download limits
Suno will start curbing downloads to reduce spam on music streaming services. Free tier users get seven downloads for life, Pro users get 20 downloads per month, and Premier users get 60 downloads per month.
IIT Madras-based lab Bodhan AI launched open-weight models across speech and vision, supporting over 23 languages. This continues the trend of more open-weight models being released to cater to Indian languages.
Numbers Game
Indian startup Navana AI said it crossed 100 million minutes in automated voice calls
Deals Corner
Wonderful ($550 million): CX funding doesn’t seem to be slowing down as Wonderful hits $5 billion in valuation.
Investors: Insight Partners (lead), Index Ventures, IVP, Vine Ventures, 9Yards, and Bessemer Venture Partners
Apate.ai ($8.5 million): The company focuses on fraud prevention. It deploys voice and text agents to waste scammers’ time.
Investors: Lobby Capital (lead), OIF Ventures, Investible, Concept Ventures, and Baobab Ventures.
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com



