#21: Google wants you to talk to AI, Spotify wants you to listen to it
Big Tech's voice bets
Top News
Google’s voice interaction push and Spotify’s voice generation push
I. Google IO was packed with announcements about AI agents. However, a sub-theme was voice interaction across different apps and modalities. These features might excite people who already use voice for dictation and prompts. In isolation, they look good, but for a company whose services are so widespread, the new voice announcement seemed fragmented. I’ll break down what I mean by that as I go along.
One core idea Google showed off was that now you can talk to its Workspace apps like Gmail, Keep, and Docs. But the function of voice in each app was different. In Gmail, it was more about providing a fuzzy context to the assistant and letting you find an email; in Keep, it was to polish up transcription and post a formatted chunk of text as a note; in Docs, it was about interacting with an assistant that can take multistep commands, including fetching information from sources like email and files in Drive. But are these all one assistant? Do they carry the context from one place to another? It is not clear.
WSJ reporter Nicole Nguyen tested an early version of the Docs Live product, and let’s say there are still some things that need to be ironed out before the feature goes into production.
Google’s product strategy has always been confusing, and the company often duplicates features and products. Earlier this month, it released an AI-powered dictation feature called Rambler, which cleans up speech. Now, a similar feature exists in Keep. Maybe the idea is that if I don’t use Gboard, I still get to use the built-in dictation power. But what takes precedence when there are two features?
There are some places where Google didn’t talk about the voice feature out loud. For instance, its agentic coding platform Antigravity now has voice support. Plus, the company also showed off voice commands on its Gemini Mac app that followed instructions and got tasks done. We don’t know if the latter is going to make it to a wider release.
One surprise was Google announcing its “Audio glasses” that will arrive this year. These glasses isn’t designed just for music listening, but they are akin to Meta Ray-Bans and don’t have a screen as a display. The company is developing these glasses in partnership with Warby Parker and Gentle Monster.
With these glasses, users can fire up Gemini to get information or complete a task like ordering a coffee. The glasses can use cameras for visual context to translate text or speech, give you directions based on what you are seeing, and read out details about the object in front of you. Plus, you can snap pictures and ask Gemini to edit them for you. A lot of people use Meta’s glasses to take photos, but a few of the AI features Google listed could be useful if they work in real-life conditions.
II. Spotify also had a big voice AI-focused announcement week with a bunch of new products and features unveiled at its investor day event. Google’s whole spiel was about talking to AI, and Spotify’s concentration was on generating voice content through AI.
The company’s marquee announcement in this area was personal podcasts, a way for users to use a prompt for a topic to generate a new podcast using AI. You can prompt to create a daily or a weekly briefing about a current affairs topic or a sports team that you are interested in, or generate a one-time podcast to understand a concept.
Other products like Google NotebookLM, ElevenLabs Reader, and Adobe Acrobat allow for a similar podcast creation for personal usage. There are apps in this area, too. Former NotebookLM developers tried building an app called Huxe, which is shutting down. Another app called Sun is working in the same modality under a16z Speedrun. Spotify understands that more people are trying to generate AI audio personal entertainment and information use cases, and wants to be at the center of this.
The company also released an experimental desktop app called Studio by Spotify Labs, which will connect to your email, calendar, and notes to create audio briefings about your day or a trip. Ideally, you don’t need a separate app for that, but Spotify seems to have a larger ambition to be in the productivity space. The company hints at this in the app description:
With your permission, it can take action on your behalf: researching topics, using a web browser, organizing information, and helping complete tasks. It can also connect to the tools you use every day, like your calendar, inbox, and notes.
Every company wants to build agents that take action on users’ behalf. Spotify might not be there yet, but it can test the waters with its new app. Given its desktop presence, Spotify could easily venture into other modalities like meeting recordings.
Spotify is not stopping at personal audio generation, though. The company released two other products for public consumption. The company is partnering with ElevenLabs to let authors create audiobooks using AI voices. It also struck a deal with Universal Music Group (UMG) to let users produce AI remixes.
Then there are discovery features to get to the podcasts and audiobooks using AI. If you have questions about an episode or a concept, you can also ask the assistant about them rather than going to Google or ChatGPT. Spotify is loading up the app and tools with AI to get to every kind of audio possible, but this might end up confusing users.
Signals and Experiments
April was an interesting month for voice AI funding. The dealflow amount was higher than that of March, but there was only one deal over $100 million, signaling mid-sized cheques.
Deals in the customer service vertical keep on flowing, but we saw that HR/recruiting, along with healthcare, are strong sectors for investors to fund startups.
April was also the first month in the year when no new unicorns were minted.
While April was good, May seems to be shaping up as the blockbuster month with two $700 million plus deals.
Quick Bytes
As the market for AI-generated personal podcasts heats up, former NotebookLM developers’ app Huxe is shutting down. The company didn’t provide core reasons, but said that the team will move on to other projects.
The creator of the newly open-sourced transcription model, Mega ASR, claims that it performs better than Gemini 3 Pro and Qwen 3 ASR.
LG has partnered with Kardome to add voice tech in its TVs that allows users to control the device with voice, even when audio is playing loudly. We definitely need more tech that works in loud environments.
ElevenLabs said that voice creators have earned $22 million through its products, up from $11 million cited last year. The company is adding more than 200,000 premium titles to its ElevenReaders Library. It said that subscribers who pay $11 per month can get 20 more hours of listening.
Stability AI released a new family of music creation models, called Stable Audio 3.0. The small models are designed to make sound effects and music of up to two minutes locally on a device. The larger models are designed for professional use cases with the ability to conjure up a six-minute-long track.
Suno is in more legal trouble as the indie band The American Dollar filed a lawsuit against the startup, alleging that it used 236 of their songs to train AI models. The band said it has suffered a nearly 80% loss in licensing revenue due to AI copying its style.
Brett Adcock’s new venture, Hark, a personal AI platform including software and hardware, got a massive $700 million in funding. The company said it is developing an assistant that uses speech, vision, and memory to interface with users.
ElevenLabs said that it will now use Google’s SynthID to watermark AI-generated audio. Meanwhile, TikTok also said in context of its UMG deal that it plans to ramp up its efforts to remove unauthorized AI music from its platforms.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance tech reporter at TechCrunch. This newsletter is an attempt at covering what is happening in the industry of voice, audio, and music.
Email: voiceaiweek@gmail.com or im@ivanmehta.com







Really like your curation of news in Voice and AI.