#29: ChatGPT, turn on my light
OpenAI is building a smart speaker
In Focus
Into OpenAI’s hardware play with voice at center
We have seen many rumors about OpenAI’s hardware play with a pair of earbuds and a phone reportedly in the works. Last week, Bloomberg reported that the company’s first device could be a smart portable speaker with a camera that acts like an AI companion.
The speaker will have different sensors to understand your environment, along with a camera. Just like many AI assistants, OpenAI’s idea sounds like a ploy to collect as much context as possible about you and your surroundings in order to personalize its outputs.
More than 900 million people use ChatGPT for work, productivity, travel planning, shopping, and more. OpenAI would hope that the speaker can carry that context over and assist users in all kinds of tasks.
Here is what Bloomberg said about how the device would work:
OpenAI envisions the device anticipating needs, surfacing information proactively and serving as an expert on its user, they said. Though the speaker is designed to stay in the home, it will be easy to move around the house.
The weird bit Bloomberg reported is that the speaker will have mechanical components that move while responding, so users feel that it is alive. That description makes the device sound creepy rather than helpful. I would want my smart speaker to be efficient rather than emotive.
The speaker will directly go up against the likes of Amazon Echo, Sonos, and Google Nest, which have been around for a long time. OpenAI’s bet would be that with more sensors and an assistant that knows more about you, the speaker would attract many users.
The company has already deployed a new voice model, called GPT-Live, which sounds more expressive and human-like. The model is also better at holding conversations with improved turn-taking and interruption handling. The speaker can use a model from that family.
We have heard “AI is the new interface paradigm” many times. For OpenAI, voice might be the modality that unlocks it. The company is working with Jony Ive’s LoveFrom, hoping that the former Apple design head’s experience of creating an iconic new category of devices can wave his magic wand once again.
There are a few hurdles, though. On paper, the ChatGPT app on your phone can do many of the tasks a speaker can do, and there is no clear advantage to having an OpenAI speaker. Most smart speakers don’t have a camera, and in a world where users are wary of AI’s privacy implications, they might not want to have a device that can see what you do all the time from a company that has been accused of privacy infringement.
Smart speakers are good at a limited number of tasks, like setting timers or alarms, adding reminders, or playing songs. They can also control your smart home devices, but that part is unreliable at times. That is because either the speaker doesn’t understand what you said, or there is a breakdown in many pieces of software that connect your bulb to your speaker. AI or no AI, the basic expectation from any smart speaker would be that it can do all of the above without failing.
A more recent hindrance is Apple’s case against OpenAI, alleging that former employees now working at the AI giant stole trade secrets. The court procedure could take a long time, and it might not prevent OpenAI from releasing devices just yet, but there could be future roadblocks, thanks to the case.
AI devices haven’t really taken off in the mainstream yet. OpenAI would hope that it could breach this barrier.
Signals and Experiments
Listeners prefer multi-character AI voices over a human narration: Survey
Spoken conducted a survey of over 1,000 listeners in the U.S. about audiobooks narrated by its AI. The company said that 61% of people mistook AI narration for human narration. Spoken has a Multi-Cast tech with which it can power various characters for book enactments. The company pitted the tech against a single-person human narration.
Key observations from this survey:
The AI narration had a higher percentage favorability at 61% vs 53% for human narration.
31% of listeners said that they would give AI narration a go. The number jumped to 65% after they heard a sample.
46% of people who heard the AI version showed an intent to purchase a title, but human narration rated higher, with 49%.
The audio publishing industry reached $2.43 billion in 2025. For smaller publishers or authors who might not have a budget for voice actors, AI could be helpful.
Platforms are offering tools for users to make creation easier. Apple started a digital narration program in 2023. Last year, Audible partnered with publishers by offering over 100 virtual voices to increase its catalog with rapid audiobook creation.
ElevenLabs offers a product for audiobook creation, along with incentives for distribution on its platform. Spotify started accepting books created on ElevenLabs in 2025 and launched a creation platform powered by the voice AI company’s model this year.
Book creation platform Inkfluence AI noted that audiobook creation increased 5x from March to May 2026. As the speed of creation and the variety of voices increase, many authors and publishers might not opt for human narration, impacting voice actor jobs.
Quick Bytes
Another day, another Suno controversy. 404 Media reported that a hack revealed the company scraped services like YouTube, Deezer, and lyrics site Genius to train its models. Suno previously admitted that it used public sources to get data for its model, but maintains that it is within fair use rights.
Avatar company Synthesia launched the latest version of its dubbing product, which lets teams refine transcripts and perform edits without spending tokens. The company said that its improved lip-syncing is better at syncing small mouth movements.
Dictation company Willow made its transcription plan unlimited for base use. The company launched two new models, Willow Frontier Mini and Willow Frontier Pro. The mini model will power the free tier, and the Pro model is for faster and more accurate transcription for paid tiers.
Local transcription is really useful when your internet is choppy. Communication coaching app Hedy said that it has improved its local dictation abilities on both iOS and macOS.
Digital audio workstation (DAW) company BandLab acquired an AI music creation studio, Aiode, to expand its licensed audio-to-audio offerings. Aiode claims that its models are created in collaboration with musicians.
Deal Corner
Rime ($24 million): The voice AI model company raised a Series A
Investors: M13 Ventures (lead), Twilio Ventures, Corazon Capital, Unusual Ventures
Overtone ($18 million): Hinge founder’s new voice-based dating app raised a seed round
Investors: FirstMark Capital, Pace Capital, and Match Group.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com


