Signal
Over the past few years, productivity platforms like Notion and ClickUp have moved into the meeting notetaking space. Now, we are seeing a reverse trend. In the past week, Fireflies announced an email assistant that organizes emails by priority and also drafts replies. Separately, Fathom integrated with Superhuman Go to let users pull up summaries or action items from meetings.
Meeting notetakers already listed drafting follow-up emails as a feature. But with these newer launches, they are moving up the productivity stack. Given how many meeting notetakers and email clients are around, we might see many of them being acquired by productivity companies trying to build a full stack of solutions.
In focus
Deepgram releases a new speech model, and why AI agents need better marketing
Every speech company I speak with says it wants to make its AI voices sound human. Their reasoning is that human-sounding voices can have better conversations with other humans on a call.
Deepgram, which released its Flux text-to-speech model last week, wants to focus on solving customer problems rather than thinking about speech modalities of an agent.
The company said its new speech model, paired with its own Speech-to-text model and context pipeline, can handle user interruptions, changing directions, pauses, and sharing complex information. CEO Scott Stephenson talked about building a robust content pipeline in an interview with Week in Voice AI in May.
“The next phase of this market will not be won by the model that sounds best in a demo. It will be won by systems enterprises can trust to complete business-critical work when conversations become unpredictable,” Scott Stephenson, CEO and Co-Founder of Deepgram, said in a statement.
The company’s new model can respond to a person with just 80ms of latency while retaining context from prior conversations.
While all the model talk is fine, I feel that labs are sometimes exaggerating the need for human tonality in calls to airlines. As a customer, I wouldn’t think, “Oh, that bot sounded so human,” if my query is not resolved. Most customers wouldn’t think about AI’s voices if they get answers. However, many conversations with AI bots are as broken as the ones with human agents. The difference is that people might give other people some grace and expect perfect results from bots.
I am tired of reading marketing material that sounds like “Our AI just works in enterprise, and others suck.” There has to be a better way to convey what your model can do in enterprise settings and for customers. Because if AI was so good, more people would be asking to chat with these bots instead of humans.
Experiments
I wanted to highlight some voice experiments I spotted in the wild. The core idea in both projects is to solve a personal need and have voice be the layer of interaction for people who are not technically sophisticated. The people who made these projects did the technical heavy lifting and hid the complexities in an efficient way.
Dan Peguine built a box for his children to press a button and send a voice message to grandparents on WhatsApp. When grandparents reply to them, the box plays a notification tone, and kids can play the messages at the press of the button. The design is like an old answering machine, which is by design not to involve any screens.
Yashaswini Singh made a WhatsApp-based grocery ordering system for her grandmother, who can send voice notes of items she would want to order to a bot. The bot, powered by ElevenLabs voice, will confirm the order and place it. The payment method is cash on delivery, which means that no payment authentication is required.
Both experiments are deeply personal but can resonate with a lot of people while solving for narrow use cases. Voice-only interaction gadgets can reduce screen time for people who are trying to have a digital detox. On the flip side, even if you are interacting with one app (WhatsApp in this case), voice commands can act as prompts to trigger tasks.
The second example is tough to execute at scale given the number of technical layers and APIs involved. If there is no proper error handling, the end user will get confused and/or frustrated. The shopping use case only works when I don’t care about specific prices or brands. When I ask the voice agent to buy rice, it can either pick a brand and quality based on the factors defined on the back end, or it has to know what specific rice I order every time. The moment there is any unreliability, I would switch to ordering directly on the grocery app and ditch the voice agent.
Numbers game
Deepgram said that it now has $100 million in ARR
Google said that Gemini now has 1 billion monthly active users, with 63% using voice to chat with the assistant.
A survey from NextPhone says that AI bots have over 90% resolution rates for callback requests, but have over 55% resolution rate for booking appointments. That means AI agents aren’t yet great at longer conversations with complex turns.
Quick Bytes
As AI proliferates in music-making tools, we might see more artists using them in the creative process to different degrees. Rapper Tyga, who first said the album was human-made, later admitted that he used AI in the process. It’s not clear if it was AI or just bad music; the album got a 0.0 score on Pitchfork.
Spotify will label AI personas of artists on the app and will keep the songs away from recommendation algorithms. We might just see the rise of “virtual artists,” just like virtual influencers and a whole lot of business around getting users to stream this music to earn money.
Wispr Flow is thrown into a privacy debate
A week ago, transcription tool Wispr Flow’s employee posted data about how Indians speak as compared to other countries. This sparked a debate about privacy and user data retention. CEO Tanay Kothari mentioned that the data was based on people who opt in to contributing to model training, and 90% of users opt out during onboarding. It is not clear what data the company retains when they use the data for model training, and whether it removes any PII-related data.
Google’s Pixel 11 gets useful voice features
Google’s Pixel 11 series gets useful voice features including live translation and transcription for content on the screen. The new phone also has the company’s Wispr Flow competitor Rambler built in.
The company launched a conversational voice agent, which lets users ask questions like “How much did I spend on food?” and get visual charts as answers. More apps should let us navigate reports by asking voice-based unstructured queries.
Deals corner
Wippi ($1.2 million): The Indian company is making gadgets for children that focus on screen-free engagement
Investors: 12 Flags Capital (lead), Warmup Ventures, Ventana Ventures, Ruvento Ventures
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com





