#28: Chatting with ChatGPT
OpenAI's big bet on voice models
In Focus
OpenAI releases new conversational models to boost voice usage
The headline number from OpenAI this week was that more than 150 million people use ChatGPT’s voice features — voice mode and dictation. Assuming over 900 million people use ChatGPT weekly, that means roughly 16.7% are engaging with its voice features, giving us a clear sense of the massive consumer demand for audio-first tools. This gives us a rough idea of how large the consumer market for voice tools is right now.
OpenAI released a new pair of models last week, named GPT-Live-1 and GPT-Live-1-mini, to enhance ChatGPT’s voice mode for live conversations.
Architectural changes
The company’s naming of the voice model is a bit of a mess because the voice model was previously powered by a model named “Advanced voice model” (AVM). AVM was a pipeline of speech-to-text, a large language model, and text-to-speech. This meant that turn detection and interruption handling were clunky. Plus, you didn’t always get the answers that you wanted.
The company has adopted a new architecture to solve these problems. First, the GPT-Live series is a pair of full-duplex models. This means they can listen and talk at the same time without having to wait for someone’s turn. Second, reasoning, search, and agentic capabilities are separate to keep the conversation going and get access to the best knowledge models like GPT-5.5 for more accurate answers.
App usage
OpenAI has big ambitions with the model. The company said during a media briefing that with this kind of change to voice mode, it thinks of voice “as a kind of primary interface to computing.” With improved conversational skills, OpenAI is betting that its weekly active user base for voice will surge well past 150 million
While OpenAI didn’t provide any data about average session time, the team that built the feature said they had many conversations in voice mode spanning 30-40 minutes. With the new and improved voice mode, the company would like increased session times.
How will OpenAI use these models?
Apple, Google, and Amazon have tried to make their assistants sound more natural, conversational, and knowledgeable. While all these companies have their own devices, OpenAI has to rely on existing platforms for distribution. The new voice mode could set the stage for the company's rumoured devices like earbuds and a phone. Notably, OpenAI was sued by Apple this week over issues of trade theft and breach of contract. It might not derail the AI company’s hardware plans just yet, but we have to keep watching the case to know more.
What’s more, OpenAI plans to bring these models to its API for developers and enterprises. The call center and customer support market is hot. With models that can better handle interruptions, OpenAI could make itself a serious player in the market to compete with the likes of ElevenLabs, Deepgram, and Cartesia.
Signals and Experiments
I got access to the new ChatGPT Voice mode with the rollout and decided to put it to the test. The voice sounds more expressive, and you have the option to switch intelligence (Instant, High, and Medium) and languages directly from a quick settings menu present while using voice mode.
The model still occasionally talks over you, but when it realizes its mistake, it pauses much more naturally than the previous version. Overall, the conversation flowed much better
Earlier in the week, I had a conversation with ChatGPT voice mode about the impact of voice as a modality on app engagement, and the older model wasn’t able to fetch any data. The newer model cited some studies for me.
OpenAI claimed that you can get instant translations while speaking because of the full-duplex nature of the model. In practice, the latency works, but I wish I could say the same about the quality of voice and translation. I tried conversing with the new voice mode in English, and it was speaking Hindi; the pronunciation was that of a foreign-born Indian with a heavy accent, and the assistant spoke bookish Hindi at times that people don’t really use in real life.
The company said that the new voice mode supports “most spoken languages” but didn’t list them. I wish that multilingual coverage were better because I would like to use the assistant for translation when I travel.
The second issue for me was basic work that an assistant can do in voice mode. I asked it to list some points we discussed about newsletter topics, and it started narrating the points without actually listing them. A few times, it told me it could make a framework for graphics for certain stories, but it never ended up making one because it can’t do so.
OpenAI did show off the ability to give visual answers, but those are limited to sports schedules, maps, and weather. I wish the company would give the voice mode the ability to do more things, including creating documents. ChatGPT already has a way to transcribe your prompts, but it would be great if I could keep talking about tasks and the voice mode performs them in the background.
This week, OpenAI launched OpenAI Work, a new agent for project work, and also updated the ChatGPT desktop app to merge Work and Codex. With plugins and coding ability, I would really like to chat with the agent while I might be checking email or doing other tasks.
OpenAI has hinted that it wants people to use their voice for work-related tasks, and the new voice mode could just be the beginning of that path.
Quick Bytes
Paris-based company Gradium extended its seed round of $70 million in December to another $30 million, bringing the total to $100 million. The company is focusing on building real-time speech models and solutions.
Music industry bodies have come together to create a program for AI track labeling on streaming services with markers for “AI-assisted” and “AI-generated” tracks. Streaming services like Apple Music and Tidal have adopted their own labeling, while Spotify shows AI involvement in song credits.
Speechify launched its Simba 3.2 model, which is currently topping the artificial analysis leaderboard. The model also has a competitive base pricing of $10 per million characters, which reduces to $6 per million characters at scale.
Character.AI is thinking about audio drama with AI characters through its c.ai FM product. The company is currently testing these serialized dramas in a limited way, with professional writers penning stories and AI-generated characters narrating the script.
Music companies are now trying ethical ways to train models. Music production platform Landr said that more than 30,000 artists have opted into its AI revenue share program. The company has also created a $1 million pool to fund advances for eligible tracks joining the program.
Google Voice added transcription and note-taking capabilities that already exist within Google Meet. Users can press the “Notes” button during the call to get a summary, transcript, and action items.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com


