#31: Wispr Flow to overhaul quality with new models
Investors are pouring money into speech generation startups
Signal
There seems to be a massive demand for AI-generated voices. Companies are trying to sound more “human” than ever, and everyone has a different definition of that. When you try to program “humanness” into AI-generated voices, you can generate a lot of variety, which is important for both creative and enterprise use cases. Investors are ready to back startups in this area as there are a lot of opportunities for both research and market expansion. For companies, there is room to play around with accuracy, expression, and tone of these voices and master one or more categories.
In focus
Wispr Flow is rolling out a new model for quality improvement
Wispr Flow is one of the most popular transcription apps out there. However, the drawback of this popularity is that when something goes even slightly wrong, people start to take notice. Over the last few weeks, many people have complained about Wispr Flow’s accuracy and model quality dropping significantly.
In June, co-founder and CEO Tanay Kothari asked users to express their biggest gripes with the tool. One of the top requests was to improve accuracy.
In a chat with me last week, Kothari said that a year ago, the company decided that it couldn’t just put a band-aid on quality issues with quick fixes. The right way was to introduce a better model that solves current issues.
“We started realizing about a year ago that for these issues, the right answer is not to kind of put a Band-Aid and make those small changes. It comes back to why the model is making these kinds of mistakes, and realizing it hasn’t been trained with more intelligence. To solve this problem, we needed to train a bigger model,” Kothari said.
He said that the new model, which will likely improve quality significantly, is already rolled out to some users, and in the coming weeks will be available to all. The company also plans to better communicate what model changes can and can’t do for users and why there are quality drop-offs for some users and use cases.
New research lab and company direction
Wispr Flow is also taking steps not to be just a transcription company. It hired Ariya Rastrow, a researcher who worked on Alexa, as chief scientist to establish Wispr Advanced Interfaces Labs. Under this project, the company will research new models and new modalities of interaction with AI.
Kothari said that he and his co-founder Sahaj Garg have been thinking about cutting-edge research and new interfaces for a long time. But the challenge was to hire and assemble a team of people capable of working on these problems.
The new lab team comprises 13 people now and will grow bigger. The Wispr CEO said that the company is able to attract top talent as researchers want to work on cutting-edge product problems.
The aim with this lab is to build a new generation of models with reasoning models or intelligence behind them to serve new interaction paradigms.
“I think the whole generations of models that we're going to build are primarily going to be in the human-to-computer interaction space. Right now, it's just voice in, text out, and then you want to add a lot more modalities on both sides. We will use our own architecture and data for that,” Kothari said.
The company is not immediately thinking of building an API business out of its own models. However, it plans to publish more about research findings in the coming months.
As I reported last week, the company is also set to introduce a meeting note-taking feature and take on the likes of Granola, Firlies and Read AI. With this launch, the company could be taking the direction of being an executive assistant.
However, getting into meeting note-taking means the company is set to collect more data from you. Kothari told me that Wispr wants to be respectful and clear about how much context users are ready to give to the app, and if need be, take a step back.
With these moves, Wispr Flow is going to look like a different company, competing in various domains rather than being just a dictation startup.
Experiments
I recently started testing the Vocci ring. My initial impression was that it might be designed for personal note-taking, just like Sandbar’s Stream or Pebble 01 Index ring. However, strangely, it is built more like a meeting notetaker, as you start and stop recording with a button.
The ring is good at capturing audio and transcribing conversations. But the software is not ideal for when you need to take short notes. Second, I like meeting notetakers for in-person conversations to be visible to the other person. With this ring, the recording light is barely visible to me. Plus, it is easy for other attendees not tof notice a ring if a user doesn’t inform them that they are recording the meeting.
While this could count as a feature for some, as we form our meeting etiquette in the AI world, people would be wary of devices like these.
Numbers game
ElevenLabs reached the mark of $600 million ARR in June. This marks another quarter with the company progressing towards that elusive $1 billion ARR mark. The company’s time to reach the next $100 million milestone has constantly been reducing.
Quick Bytes
A few weeks ago, I said OpenAI is going to think more about being a strong enterprise player in voice AI. The company released two transcription models in the week called: GPT-Live Transcribe for real-time transcription, and GPT-Transcribe for batch mode.
RHEI debuted a new agentic platform called Made for Music, which is designed for marketing a new track or an album release. The platform can create the press release, playlist pitch, cover art, and short-form promo assets.
Friend released a new version of its pendant, which comes with a speaker and a pre-loaded voice that you can’t change. The company told Wired it is using OpenAI’s model to power the companion. We could see more gadgets in similar lines.
A report by 404 Media highlighted that since Spotify’s AI music labeling is not working as intended, alternative sources such as SoullessMusic.com and Sloptracker have cropped up for users to check for AI-generated tracks. Increasingly, users are demanding more accountability from streaming platforms.
There is a new class action lawsuit against Granola alleging the app is secretly recording meetings without users’ consent. Granola typically asks users to inform other parties that they are transcribing the meeting. The case highlights the lack of tools that can do this on behalf of users.
India’s Sarvam released two new models for voice AI: Bulbul v4 for speech generation with better vocal range and Saaras v4 for automatic speech recognition with support for multi-speaker diarization.
Deals corner
Fish Audio ($50 million): The company started as an individual project from an NVIDIA researcher to generate better voices. Largely working in the text-to-speech domain, it has 8 million users and $21 million in ARR
Investors: Coreline Ventures and Capital Today (lead), 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
Smallest AI ($13 million): A text-to-speech startup that closed its funding this week. Smallest concentrates on smaller models for specific languages and dialects.
Investors: Seligman Ventures (lead), Sierra Ventures, and 3one4 Capital.
Encore AI ($30 million): The startup observes voice calls and other customer conversations to train voice agents that can work with call center executives.
Investors: Team8 (lead), Planven, Lukatz, and Garage.
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com





