Signal
India is a very complex country, and it’s a big market for voice AI because a lot of business is conducted over calls. Yet, in enterprise deployment and model release charts, there are more companies from abroad. That might be changing as Indian companies are beginning to release deployable models.
Last week, Murf AI released the Falcon 2 speech generation model that costs $0.01 per minute. Separately, Blue Machines released the Floe speech detection model, with support for 11 languages and contextual detection. In India, given that there are so many languages, more models with specific language slants and lower cost will find their footing long-term.
In Focus
Why Wispr scored a $2B valuation round
When a certain category of people in tech, namely founders and investors, start using a tool and find it indispensable, the startup is often valued at a high price. We have seen this happen with Granola, the meeting note-taking tool people in tech can’t live without because they fear that they will forget an important thing mentioned in the meeting.
Wispr is in the same boat with its dictation tool, Flow. Voice is the next big interface and is a running thesis within Silicon Valley. While a lot of money in voice AI is concentrated on automating enterprise sales and customer support, meeting note-taking and dictation are two consumer-facing use cases.
The startup’s $280 million round includes tons of new investors such as Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund, and PLUS Capital. The company also got money from artists and athletes, which is a big marketing play.
How much a company should be worth is a conversation, a strange combination of math and hope. I want to get into the latter part to understand why investors might think that Wispr is worth $2 billion.
Is good dictation enough?
Dictation itself is not a unique or defensible feature. Wispr Flow has competitors like Willow, Aqua, Superwhisper, Typeless, and Monologue, followed by a swath of indie apps and open-source projects vibe-coded over a weekend. People in tech may move around these tools and might use an open-source project for their own use, but not all products are maintained in the same way, and not all of them have similar features.
Wispr Flow has built a recognizable brand, and a lot of its users are experiencing AI-powered dictation for the first time and, in turn, getting wowed by it. Other tools might have some features a set of users will like, but they don’t have the marketing power of Wispr.
In the past few weeks, many people have complained about Wispr’s transcription quality dwindling. But only a few will completely ditch it to use another tool. The company has accepted that it needs to improve the quality, and as I reported earlier in the month, a new model rollout is on the way.
The company said that the new model, called Canto, will reduce the number of daily edits users have to make to their dictated output by 30-35%. But a speech recognition model unlocks new business avenues for Wispr.
A model for new business models
A few weeks ago, I broke the news of Wispr launching its Granola-like notetaker. Features aside, starting a meeting note-taking tool means that you would need more speech recognition prowess than before. Wispr might still use external models for some speech recognition, but banking on its own model will let it save massive costs.
Powering speech recognition on hardware is another way for Wispr to make money. The company’s CEO Tanay Kothari told me in an earlier conversation that the company is not yet thinking about opening up its API to developers. But the startup is experimenting with this model with a few pilots.
For instance, Wispr Flow partnered with ring maker Oasis to let users directly dictate to their devices like laptops without needing an external mic. Also, one unannounced partnership is Wispr powering speech recognition for Pebble’s Index 01 notetaking ring.
Chasing Jarvis
Wispr started as an idea to build hardware that lets you, well, Whisper. Kothari has talked about his fascination with Iron Man’s intelligent assistant, Jarvis, and chasing a dream to build something similar to that.
The company recently established an Interfaces Lab to build its own models and also explore new interaction modalities with computers. The startup hired Ariya Rastrow, a researcher who was on an early team developing Alexa at Amazon, to head the new lab’s research efforts.
A long-term goal would be to become an all-around assistant while powering speech or voice interfaces in other devices and apps.
Investor speak
Most new investors in Wispr’s round are talking about habit-forming and user needs. Customer-focused fund Forerunner’s Eurie Kim focused on Wispr building tech that reduces keyboard usage.
“When software becomes abundant, what's scarce — what's always been scarce — is understanding the human on the other end well enough that they trust you with something intimate: our thoughts and our voice,” she said.
Meanwhile, Goodwater and Acrew Capital talked about Flow becoming a daily driver, and Together Fund’s Lakshmi Shankar mentioned Wispr having a large swathe of proprietary data for model training.
“Wispr has one of the largest purpose-built datasets of context-labeled audio anywhere, the kind of proprietary, edit-informed data that would cost peers an inordinate amount to recreate,” Shankar said.
Experiments
More people are building personal tools as voice model makers are allowing generous usage for tinkerers. Voice companies also want to see different use cases that could be deployed commercially.
A person builds a voice agent that calls their friend at 7 am and has a conversation with them to motivate them to go to the gym. This could be great for reminder calls in hotels or appointment booking.
Numbers Game
Wispr’s funding was the biggest story of the week, so let’s dive into some of its numbers provided by the company and investors:
Users have dictated over 60 billion words with Flow.
People using Flow cross a 50% voice-to-keyboard usage ratio by month four and are at a 72% voice and 28% keyboard ratio by month five.
Wispr Flow has retained over 70% of its users after 12 months.
Roughly 60% of dictation is in non-English languages.
Over 10,000 enterprises are using Wispr, and users in over 125,000 companies are using the dictation product.
The company’s revenue has grown by 150% in each of the last four quarters.
Quick Bytes
As the volume of AI-generated songs increases, streaming platforms are thinking more about marking these tracks better. Apple Music, which has adopted a voluntary approach to labeling, will actively start monitoring AI-generated songs later this year, as per Variety. The report said that music labels will need to mark songs with “a material portion of the content” created with AI.
Wispr Flow got another competitor in Meta in the dictation apps market. The company released a Mac app that can dictate speech in all apps using its own Muse Spark model.
India’s Sarvam starts deploying in consumer apps
Sarvam, one of India’s biggest AI startups, has claimed to have the best Indic language models for a long time, but its focus has been on enterprise deployment until now. In the past week, the company announced a partnership with HP for a dictation app and also said that it will power call screening on Equal AI.
The Rolling Stones’ Mick Jagger said that he is open to artists using AI as a tool to create music as long as it is original. He said artists should have their own thoughts and composition ideas going into it. “I don’t want people just putting stuff out there that can sound exactly like the Rolling Stones — I think that’s obviously wrong,” he said.
Adobe has been working on making Firefly a complete AI-powered creative studio to compete with different companies ranging from ElevenLabs to Canva. Firefly now has the ability to generate music, speech, sounds, and effects within its studio using Adobe’s own and third-party models. Sound and music generation is taking off in marketing; Adobe doesn’t want to lose its place as a creative suite leader.
Deals Corner
Wispr ($280 million): One of the biggest voice deals in recent times to take Wispr beyond transcription
Investors: Menlo Capital (lead); Acrew, Forerunner, Goodwater, Peak XV, Together Fund, and PLUS Capital (New investors).
Incredible (formerly VISS.AI) ($2.7 million): The startup is developing an agentic assistant that can complete tasks across Mac and Windows
Investors: Spintop Ventures (lead); Starbright Invest, Magnus Emilson.
HeyBreez ($2.5 million): It’s an enterprise voice agent design and deployment startup with a focus on the Middle East
Investors: Lunara Partners (lead); Jabbar Group, DASH
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com



