Signals
I. The music industry is trying to work out new business models in AI amid lawsuits
The music industry is still battling with AI companies over copyright cases with regard to using data in model training. But it is also thinking about how to make the most of the situation and earn money. Last week, Suno released a new model, which is trained from scratch with licensed data from Warner Music Group, BMG, and Believe.
On the other hand, ElevenLabs signed a partnership with Universal Music Group, which is in an active legal battle with Suno. The voice AI company also released a new v2.5 model and enabled lossless downloads. We will see more “industry-approved” models in the future, and their sound signature will depend on the catalog of data from labels and other sources. Some models might be great at background scoring, and some might be great for ads.
As Musixmatch co-president Rio Caraeff told Week in Voice AI in an interview, we will see more collaborations as it’s monetarily beneficial for platforms, rights holders, and model makers to work together.
II. India’s Global Fintech Fest becomes a voice AI fest
The finance industry is one of the best targets for voice AI companies. This tweet possibly sums it all up.
There were new partnerships and products announced, including Sarvam’s partnership with Mahindra Finance, Cashfree’s voice-powered payment recovery stack, and Arrowhead’s own pipeline for BFSI clients. In the coming months, we might get to know about more deals signed between finance institutions and voice AI companies. Right now, the mood in India around voice AI is buoyant, and startups are taking advantage of that by signing new deals.
In Focus
Apple tries to weave privacy into its new Audio Intelligence features
Apple released two new features with its new Apple Watch in an attempt to indicate that we don’t yet need separate AI gadgets for productivity and accessibility. First is a Live Rewind feature, which transcribes the last 15 seconds of audio, in case you missed anything a person said. The second Siri Recap, a Granola-like meeting notetaker summary without any transcript, along with action items and key points of a conversation.
Both those features sound like always-on features on the face of it. Apple did specify that you can turn them off at any point, and users have full control. But it’s also interesting to understand the processing pipeline of how the company handles privacy for these features when they are on.
Apple Watch 12 and Apple Watch Ultra 4 have a secure enclave within the S11 chip for initial secure processing of audio. For iPhone 16 and above, there is a Secure Enclave in the chip for raw audio processing.
Apple uses a new audio-verified pairing besides Bluetooth to ensure that audio data is sent only between two secure enclaves while keeping the communication encrypted. Backed-up Live Rewind or Siri Recap is encrypted end-to-end.
Live Rewind
Live Rewind works when you double-press the crown on your Apple Watch and shows the last 15 seconds of transcript. Apple Watch keeps a rolling buffer within its chip, so there is no continuous conversation transcript. There is only a backup if you ask Siri a follow-up question or ask it to save it in the Siri app.
Apple said it is also alert people around you with a sound and an on-screen message that you are using the Live Rewind feature:
When activated, an audible chime plays from the speaker on your Apple Watch, even if your Apple Watch is on silent or you have headphones connected. A full display animation and microphone indicator appear on the Apple Watch display, and a distinct double-press gesture is required to activate Live Rewind. This is to alert those nearby that the feature has been activated.
Siri Recap
There have been calls for Apple to buy Granola, but I think for Apple, storing transcripts and all kinds of recording-related data could be privacy trouble in waiting. The company’s new feature is the closest iteration of a meeting notetaker or a conversation capture.
With this feature, Apple doesn’t keep any transcript and also removes financial data, government-assigned identifiers, authentication data, and certain personal identifiers. Personally, I think meeting notetakers should have this toggle for their apps too. Users will get a summary, action items, and highlights from a conversation.
What’s more, users can set a schedule to have this feature on “At work” or turn it off at night.
When the feature is on, Apple runs a conversation detection model on the Apple Watch. Once it detects the conversation, it encrypts the audio and starts transferring it to the iPhone. A model on the Secure Enclave starts transcribing it, and the final product is roughly half of the original transcript, removing tone markers, non-essential language, filler words, and redundancies while preserving the topics.
Then, to further summarize this block of text, it is sent to Apple’s Private Cloud Compute (PCC) for further processing. Apple collates data such as the title of a meeting, if any, in the calendar, rough location like a school, cafe, or a park, and uses it to assign a title to the final summary. Apple uses its own foundation model to generate a summary and key points, and discards the audio and the transcript.
The company said that it doesn’t notify anyone if the Siri Recap feature is on because the summary is “comparable to notes a person might write after a conversation” and there is no saved transcript. Though its support page reads:
While audio is not recorded and stored, and no speakers are attributed, consider those around you where conversations may be private or sensitive.
What does it mean for meeting notetakers?
Like I wrote last week, I don’t think meeting notetakers will be impacted by this. This might work in their favor if people like getting meeting notes and would want advanced features.
A lot of users value transcripts and recordings, which meeting notetakers offer. They also use these transcripts with other AI tools to create custom workflows.
However, meeting notetakers do need to provide some of the privacy protections offered by Apple as optional features, such as private data filtering, secure processing, or discarding transcripts.
How will this change user behavior?
This is a tricky question given people can record conversations all the time, and despite Apple’s filters, sensitive information can still leak through.
There are also hanging legal questions in two-party or all-party consent regions as to how lawmakers will interpret this feature if there is a lawsuit.
I have heard people say “I am not running Granola” when they are at a dinner meeting because they assume the other person is always recording.
Yes, apps like Granola and Circleback are present on Apple Watch, but their scale is a lot less than the number of Apple Watches sold every year. Months from now, we are likely to see stories about people sneakily recording conversations and some sensitive information leaking into summary notes.
Numbers Game
A survey of 550 decision-makers found that AI usage dropped from 95% to 64% year-over-year.
India’s Sarvam said that it has handled 325 million voice AI minutes in the preceding 12 months.
Another Indian startup Navana.ai, said it has handled 100 million voice AI minutes to date.
Quick Bytes
Countries worldwide are looking at preventing deepfakes of people to prevent scams. China is also looking into these issues for its population of over a billion people. The country’s top court is issuing guidelines around how to handle cases related to impersonation or improper usage of voices and identity.
Reports mentioned that Yemeni forces received an audio resembling Lieutenant General Tareq Saleh’s voice giving an order to retreat. The forces accused Iran-backed Houthi fighters of using AI. We will see more usage of AI in warfare and politics to deflect and deceive.
OpenAI has made its first full duplex model available through API. The company also demoed use cases with its early partners, including Yelp for reservations, Hatch service bookings, Cognition for Devin AI agent, Speak for its AI-powered tutor, HeyGen for its live avatars, and PicsArt to access creative tools.
Last week, GamersNexus reported that hackers can take advantage of LG’s ever-listening capabilities for the wake word. After that, LG defended its stance, saying that its mics are only listening for the assistant hotword and the mics are not meant to record conversations. But as ambient devices around us grow, we will see more concerns around privacy.
Deals Corner
Navana.ai ($4.2 million): Indian startup that builds calling agents, especially for the BFSI industry
Investors: Ronnie Screwvala (lead), Antler India, Sharad Sanghi, Sandeep Singhal, and Paula Mariwala.
Tellia ($5 million): A voice-first interface for agricultural field teams
Investors: Revent (lead), Grey Silo Ventures, Jeriko, and Fund F.
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com




