#30: ChatGPT and Claude voice get promising, but clunky
Voice AI assistants are good, but not good enough for a workout
Editorial note: For this week’s edition, I am trying a new format to include more analysis and reporting rather than aggregation. At the top, there will be Signal of the week, with a short analysis of a trend. I am moving my tests and experiments to their own section. I am also introducing a numbers section to highlight reports, milestones, and stats.
Signal
There are hundreds of hardware and software tools that are ready to “hear” what you say, but they are not yet fully ready to “do” what they heard from you. Voice updates from OpenAI and Anthropic, coupled with Genspark’s meeting notetaker hardware, suggest that companies are busy trying to acquire customers at a voice capture layer. Most of these companies claim users “get work done” through their voice layer, but integration with other software and commands results in final tasks not being smooth yet.
In focus
OpenAI and Anthropic update their voice stack with an agentic slant
Silicon Valley has been chasing the vision of an ever-present and efficient assistant like Jarvis in the Iron Man movies. The assistant listens to you perfectly and executes tasks. We do not have power suit-fetching voice assistants yet, but they can definitely fetch information and reschedule meetings.
Last week, Anthropic updated Claude’s voice mode, which now allows users to choose from Opus, Sonnet, and Haiku. This means that the company updated the intelligence aspect of Claude. Notably, there is no change in the voice model itself, so if you had complaints about how it talks to you, there won’t be any difference.
Another big difference is that the voice assistant can access app integrations like Gmail, Calendar, Notion, and Canva to execute tasks. It can rearrange your meeting or create a design document.
OpenAI, which had updated its conversational chops a few weeks ago, decided to update the ChatGPT desktop app, so the voice assistant can access Codex and Work to trigger agents and have them work in the background without you having to stare at the screen.
The promise of these updates is big, as it allows you to complete work without typing anything or staring at the screen. A startup founder said that he talked to Codex remote on mobile while he was on a hike.
There are a few hindrances, though; you will need to check on agents periodically to see if they are working as intended. For Codex, you will need to remember thread names, as voice runs in a different thread and needs you to call out other agents explicitly.
I partially blame this on OpenAI’s very confusing ChatGPT desktop app, where you don’t know the difference between Codex, Chat, and Work unless you explicitly remember what task you want to complete in what section.
As compared to OpenAI, Anthropic’s voice agent is useful for more general tasks, but there are some teething and latency issues as people switch between models.
AI companies are finally getting their voice features out of the conversation mode, which was useful for brainstorming and saving a transcript. However, with the new features, users can at least begin to do some tasks reliably.
My experiments
I have had brainstorming chats with ChatGPT’s new voice mode on my phone, but I wanted to put it through a different test of longer conversation and memory retention. My gym usually gets crowded during the evening, and I missed my usual time window because of work. I didn’t want to go to the gym, so I asked ChatGPT voice mode to prepare a mixed upper body and cardio workout.
Some exercises were repetitive, and I asked ChatGPT to swap them out. After a bit of back-and-forth, we settled on a routine.
The next task for ChatGPT was to stay with me during the workout and guide me. The first pass I tried was asking it to set an internal timer of 30 seconds per exercise. This went downhill fast. The timer failed multiple times, and ChatGPT told me that 30 seconds were over within 10-15 seconds of starting an exercise block. I pointed out the mistake a few times, and after that, ChatGPT said that it couldn’t do timers.
The alternative I found was that I would run my own timer. Then, once an exercise block was finished, I would ask ChatGPT to tell me the next exercise. This was good during the upper body block. However, after that, ChatGPT moved straight to cooldown rather than cardio, and I had to remind it of that.
The voice mode is good for conversation, but it needs to have a better memory and agentic tools to help with tasks on the phone. I’ll test this once again in a few weeks.
Feature that everyone should use
I recently started re-testing Willow’s dictation app on desktop. One neat feature in the latest update is that when you edit a particular word repeatedly, the app can detect the change and will ask you to add that to the dictionary.
Numbers game
Deezer said that for the first time AI-generated music crossed 50% of daily uploads on the platform. The company noted the volume of songs has gone up to over 90,000 AI tracks every day. Here is how the volume of AI songs has increased on Deezer”
Quick Bytes
Suno’s troubles continued last week as TechCrunch reported that a November breach contained records of 55.3 million users, according to the Have I Been Pwned website. Separately, thousands of musicians joined a class action lawsuit against Suno and Udio.
Amazon aims to switch models powering Alexa+ conversations as AWS costs could go up to $1.7 billion this year, according to Business Insider. Anthropic models are just not sustainable for Amazon, and it could use its own models.
LG and Ericsson announced a partnership to bring voice AI directly to the network layer. We have seen Indian telco Reliance Jio make such announcements too. While the tech sounds interesting, we need to see what privacy challenges it can bring.
Most Meeting notetakers focus on post-conversation knowledge and tips based on the transcript. Now, Otter wants to coach people during the conversation with its Live Assist feature. In the past, we have seen the likes of Cluely use a similar interface for real-time coaching.
At a $700M valuation, you can’t just be a dictation company. Wispr Flow announced a new Interfaces Lab led by new chief science officer Ariya Rastrow, who worked on Amazon Alexa. He wrote that Wispr will focus on how users can interact with AI better, starting with the voice modality.
Deal corner
Cast Insights ($4.5 million): A platform that provides audio insights from television, radio, podcasts, and livestreams for newsrooms and consultancies
Investors: Abstract Ventures (lead), SRI Ventures (lead), LocalGlobe, Martin Audio, Syndicate (VC)
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com





