#20: Are you ready to ramble?
The battle for native mobile dictation, real-time interaction models, and Fish Audio's copyright issue
Top News
Google makes an AI dictation move as Wispr’s valuation balloons
Google announced a new AI-powered dictation feature called Rambler at its Android Show: IO Edition event last week. The feature lives in Gboard, and users can access it in any app to dictate speech and get clean text free of filter words.
Google already released an experimental app called AI Edge Eloquent last month on iOS with a focus on offline models. However, there was no keyboard, and you had to manually paste the dictated text into another app. Rambler is a more polished and productized version that is integrated with Gboard and will be accessible to Pixel and Samsung Galaxy users to start with.
The first thought in the minds of many users would be that this is bad for apps like Wispr Flow and Typeless. But I have a few counterpoints. Dictation apps are still a desktop-heavy category. These apps have their mobile counterparts, but people use them because they use the desktop version, and possibly want to carry over the same experience, along with their custom dictionary.
Rambler might help users get acquainted with AI-powered dictation. But if anyone seeks advanced features, they will move to other apps.
This is where the discussion about AI-powered dictation being a feature, product, or company starts. Google saw that many people are using apps in this category and decided to build a feature, but that doesn’t mean power users will automatically gravitate towards it.
The trend is evident in another news story this week: Wispr AI is seeking to raise $260 million in funding at a $2 billion valuation, up from $700 million in the last round, as reported by Bloomberg. Menlo Ventures, who are admirers of the product and led its Series A, will reportedly lead the round again, as per the report.
The app is popular with venture capitalists and folks working at companies like Nvidia, Amazon, Groupon, and Mercury. And as we have seen in the case of Granola, if certain people in tech feel that they can’t do without a particular product, the valuation rises.
Fish Audio vs voice artists
For the last few days, many voice-over artists on LinkedIn have been complaining about an AI platform called Fish Audio, saying that their voices ended up on the platform as models without permission, with many of them being used thousands of times.
Fish Audio lets users upload voices and have others use them for voice cloning and AI-powered reading/narration projects. Artists mentioned that this model has allowed others to upload their voices without their consent. The platform has a takedown process for requesting the removal of audio. But several artists complained that the process is tedious, and it took them days to complete the takedown request.
“I’m sorry this happened. The reported community-uploaded model has been removed. Fish Audio is a UGC platform, and we take reports of impersonation or rights violations seriously. In the past month alone, we’ve removed over 100 reported or policy-violating community voice models as part of our ongoing enforcement efforts,” Fish Audio co-founder and CEO Rissa Cao said in a reply to a voice-over artist’s post.
Copyright is a massive issue in the voice world. Several artists like Taylor Swift and Indian singer Arijit Singh have moved to protect their voices. Lawmakers are also looking at this problem. Denmark passed a law this year that protected the voices of its citizens from AI cloning. In 2024, the state of Tennessee passed the ELVIS (Ensuring Likeness Voice and Image Security) Act with similar protections.
Signals & Experiments
The mobile dictation dilemma
Apps are realizing that voice is a growing input method for a lot of users. Voice input is particularly useful when users need to dump in a big chunk of text. Email is one target use case, and increasingly, the second one is AI apps for prompts.
I have had mixed experiences with these tools. I very regularly use this tool called Littlebird, an assistant that knows my work context by “reading” my screen. I have tried using its own transcription tool for prompting, but it often falls flat. Plus, on desktop, I have the option of using apps like Wispr Flow or Willow in any app easily.
It’s a different ballgame when it comes to mobile experience. As I have written before, Apple is breaking the AI dictation experience with a new flow in iOS that requires you to swipe back to the app you were typing in after activating a voice session. So when email apps like Avec use their own dictation engine, it might make life easier for people. Plus, as dictation is one of its core features, the company will keep working on improving quality and personalization.
Model Behaviour
Solving for Duplex and Expression
Labs are edging towards making voice interactions more human. AI labs are focusing on two main fronts: turn detection and interruptions, and voices sounding more expressive.
Mira Murati’s Thinking Machines Lab demoed a full duplex multimodal model that can listen and react at the same time. The demo showed real-time interaction, real-time translation, and a model triggering a response when a person walked into the frame.
Resemble AI released a new model called Dramabox, which lets you describe a character and then have the model enact the voice using a script. The model is doing this without any specific expression tags, which is kind of impressive.
Quick Bytes
OpenClaw creator Peter Steinberger said that he has built agents that listen to his (and his team’s) meetings and proactively start working on tasks like creating pull requests based on the discussions. This is a goal for a lot of meeting note-taking companies, but they might not have the level of access to the tool they would ideally hope for.
A study in the UK suggested that 40% of general physicians used ambient AI scribes to transcribe patient sessions and take notes. This study comes after the National Health Service (NHS) laid out guidelines around using the technology.
Musician and producer Jack Antonoff, who has worked with artists like Taylor Swift, Lorde, Lana Del Rey, and Kendrick Lamar, lambasted the use of AI in music creation.
“So to everyone who is gassed up about the new ways you can fake making art, by all means, drive right off that cliff. We’re genuinely happy to see you go,” he said in a post on his Instagram account.Meta is finally on the conversational assistant train. The company demoed a voice chat with Meta AI, powered by its Muse Spark model. Users can point the camera at different objects and ask questions about what is in the frame using voice.
Indian investors had a good week. PeakXV led voice orchestrator Vapi’s $50 million round, which took the startup to a $500 million valuation. Meanwhile, growth stage firm Activate invested in ElevenLabs through a special vehicle. The firm will help the voice AI giant with India expansion.
Off Topic
I started testing Project Mirage’s Dune gadget, which has three buttons and can change its function based on the app that is open. I haven’t quite gotten to making custom controls for it, but right now, the biggest use case for me is to paste URLs into Slack, and also use mic and video controls.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance tech reporter at TechCrunch. This newsletter is an attempt at covering what is happening in the industry of voice, audio, and music.
Email: voiceaiweek@gmail.com or im@ivanmehta.com






