Week In Voice AI #2
I am Ivan Mehta, a consumer tech reporter at TechCrunch. For more than a year, I have covered different aspects of voice AI. From voice generating to voice dictation and assistants, I have seen multimillion-dollar funding rounds to new app launches, and startups concentrating on making voice a primary interface for interaction with users.
This newsletter is an experimental attempt at covering what is happening in this industry, which is growing at a rapid pace. While the title says “voice,” I consider all kinds of AI x Audio as a valid scope for the newsletter.
I don’t yet have a set format for this. So bear with me as I play around and find out what resonates the best.
Users of this dating app are spending 26 minutes on average on voice AI onboarding
Typically, dating apps either just let you fill in basic details and get started with swiping or have you fill out long forms to know more about you. But now, since both speech-to-text and large language models are getting good, apps are trying out voice as an onboarding method.
Known, a company backed by Forerunner Ventures, wanted to replicate the old-school onboarding of apps like eHarmony, which took time and intent in terms of knowing users. But Known didn’t want to make users fill out forms, so instead, it created an AI-powered voice agent for onboarding. This makes it easier for the app to capture information about a user, ask dynamic questions based on user responses, and also get more data for potential matches.
The startup told TechCrunch that users spend 26 minutes on average on its onboarding process, with the longest call being 1 hour 38 minutes in its test phase.
In the next year, we might see more startups using voice or avatars as an onboarding mechanism to gather more context from users. While people want to add more context while using dating apps, a longer onboarding process could turn people off, as they just want to get started with the app.
Top News
Meta releases a new sound segmentation model
Meta released a new model called SAM Audio (Segment Anything Model), which lets you type in prompts to separate the voice, music, or sounds from the original audio or video clip.
What’s more, users can use visual cues like clicking on a person in a video to isolate their voice. Plus, you can also prompt on timestamps to find the audio that you might want to separate. Users can combine these three types of prompts to get the audio they want.
The company also released an audio judge model to determine the quality of audio segmentation. The company said that it has designed its evaluation framework based on how humans perceive sounds. There is also a new audio separation benchmark called SAM Audio-Bench that allows you to assess how different prompts separate different sounds.
You can try out the model by uploading your own clip and prompting it to separate sounds using this demo site.
Wispr Flow acquires its first company: a voice email replying tool called Yapify
AI-powered dictation tool Wispr Flow has made it clear during its fundraising announcements that it wants to do more than dictation. The company wants to build assistants that help users complete their tasks using voice.
The Menlo Ventures-backed company made its first acquisition in Yapify, a Chrome extension that helps users reply to emails with voice.
Dan Meier, Yapify’s co-founder, said that the company started with email, but wanted to build support for other tools as well. Now, it will continue doing that as part of Wispr Flow.
Voice is becoming the primary way people interact with technology.
It's not just faster input. It's a fundamentally new interface for everything.
For Yapify, email was just the beginning. My co-founder, David, built the foundation for intelligent voice input: a voice AI agent that went far beyond dictation, turning rambling into ready-to-send drafts by pulling in previous threads, recipient details, and context.
Our users kept asking for that same intelligence across their entire workflow: In Slack, Gong, WhatsApp, Message, LinkedIn, Github, Docs, …
That's what Wispr Flow built - voice that works everywhere, across every app.
Wispr Flow is already present on desktop and iPhone. While the desktop application works in the browser, Yapify’s acquisition can possibly allow it to let users try the Wispr Flow app without downloading it using the extension.
The company is also offering a year's worth of free subscription for Yapify users.
Quick hits
The voice agent market for enterprises is not slowing down anytime soon. UK-based Poly AI raised a $86 million Series D round co-led by Khosla Ventures, Georgian, and Hedosophia.
ChatGPT is removing voice mode from its macOS app on January 15, 2026. The voice mode will be available on the web app if you still want to use it from your Mac. The company said this will allow it focus on “ more unified and improved voice experiences across our apps.”
Google is adding a new version of real-time translate, based on Gemini’s speech-to-speech upgrades, to the Translate app.
Krisp launched its enterprise SDK for accent conversion that companies can run on their own servers.
ElevenLabs Agents now support WhatsApp deployment. This can be helpful in getting the company’s revenue up in countries like India and Brazil.
Thank you for tuning in. Keep listening.
Email: at voiceaiweek@gmail.com or im@ivanmehta.com


