#24: Apple gets Siri-ous about its assistant
Top News
Apple’s new Siri AI looks promising, but needs speech improvements
Apple announced the much-anticipated Siri AI (why not just Siri?) at its Worldwide Developer Conference (WWDC). The new Siri is a working iteration of what Apple promised two years ago — an assistant that can better understand you, get insights from the data on the device, look at the screen, and use that context to provide answers.
The company also said that users will be able to customize Siri’s voice by controlling pace and expressiveness. All voice companies and Assistant makers have been trying to get AI to sound more human. The demo of Siri Apple showed in the WWDC keynote video didn’t sound expressive, but this may change as the new betas roll out.
Another small but important announcement from Apple was an improvement in dictation with an on-device model. It’s not clear if the dictation features would equate to using an app like Wispr Flow, Willow, or Monolouge. However, as I have reported before, Apple broke how these apps worked with its iOS 26.4 upgrade. With its own dictation feature built into the default keyboard, users won’t have to switch apps to activate a dictation session.
There is also a dedicated app for the new assistant, which you can use as an alternative to ChatGPT or Gemini. But I haven’t got a chance to test it long enough.
After being on a weeklong waitlist, I got access to the new Siri just today. Here are my very early observations from a few hours of testing:
Siri can understand me better, even with filler words and complex sentences. However, as soon as I finish one sentence, Siri fires up requests to fetch knowledge even when I am completing my request. When I said, “What are the World Cup matches for today, and where can I watch them?” Siri started searching for matches the moment I said “today.”
For the above request, for some reason, Siri fetched matches between “Last Friday and Thursday.”
The voice is not very expressive, but it is better than previous versions of Siri. Right now, the expressiveness dials are in preview, and there is no way to adjust them.
I turned on the advanced dictation preview, but there is no specific indicator to tell me if I am using the new dictation or the old one.
I am going to keep testing the new Siri for the coming weeks and report my observations.
Model Behavior
Google releases a new live translation model with support for 70 languages
Google dropped a new live translation model powered by Gemini 3.5 with support for over 70 languages. In last week’s newsletter, I talked about using translation apps in China. The new model works on Google Translate, which might not work in countries like China, unless I download the models.
On the enterprise side, the company is debuting the model in Google Meet for enterprises next month. It noted that companies like Grab are experimenting with the new model for customer calls already. This will create a competition between Google and translation providers like Krisp, DeepL, and Palabra
Apple’s new on-device model for powering speech and dictation
Apple’s new Siri is powered by a series of models, including Apple Foundation Model (AFM) 3 and AFM 3 Core Advanced. It’s the second model that is powering the new expressive voice and dictation. The company said that the Core Advanced model is a 20-billion-parameter model, but it acts like a mixture of expert model with 1-4 billion active parameters.
Signals & Experiments
Robot’s slip of tongue
During my China tour, I visited a robotics company called Keenon. They had demoed various robots, and one of them was a humanoid with talking capabilities. When someone asked about a tech blog, it messed up the pronunciation and made it into profanity.
Tech companies might perfect all the robot movements, but numerous examples of chatbots going rogue have taught us that it is highly likely that robots might slip up while conversing with humans at any time. Companies will need to be extra careful as their robots will be out in the wild, and a customer who will experience this mishap might not be as forgiving as tech journalists looking at demos.
Quick Bytes
Wispr Flow admits that the app’s performance hasn’t been up to date. The company said it is taking steps like scaling its infrastructure to tackle outages and rolling back changes to its autocorrect algorithm. It claimed to have bolstered internal tooling to catch issues quickly.
India’s Equal AI raises $30 million in funding from Prosus for its AI call screening assistant that supports 10 languages. The app currently handles calls from unknown callers, but it has ambitions to handle calls from known contacts to know the reason. The startup also wants to make outbound calls with its app to book appointments. All of this is very ambitious in India’s multilingual landscape, where not all things are digitized.
Rylo (formerly Nagish), the platform that helps people with hearing impairments, raised $85 million in funding led by Canaan. The round also has contributions from General Catalyst’s Customer Value Fund (CVF) in the form of growth equity. With the new funding, the company aims to add workplace accessibility, network improvements, and real-time fraud detection to its app.
Krisp dropped its Voice Translation v3 model with the ability to audit calls live for admins. Plus, the new model has industry-specific vocabulary built in, and it is easier for admins to add custom phrases.
Deezer has been battling against AI music hard. The company has been offering a tool to other streaming platforms to identify AI-generated music. Last week, the company released a tool for users to identify such music using their playlists imported from other platforms like Spotify and Apple Music. The tool supports detection in 27 languages across 20 platforms.
Warner Music Group acquired Sureel AI, a tech company that can track how AI is using existing tracks by artists. For labels, AI has presented a new question about how to determine whether the tracks it owns are being used. AI’s usage of data fed for training is fuzzy. There isn’t a standard way to say that a specific percentage of a track was used to generate a new one. But any way for labels to measure usage could be critical for business.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com



