Week in voice AI #13: Open-source models galore
And Granola funding
I am Ivan Mehta, a consumer tech reporter at TechCrunch. For more than a year, I have covered different aspects of voice AI. This newsletter is an experimental attempt at covering what is happening in this industry, which is growing at a rapid pace.
Top News
Open source models from Mistral and Cohere hit the shelves
French AI company Mistral is trying to become a platform to create end-to-end voice agents with the launch of its new text-to-speech model after releasing Voxtral Transcribe 2 a few months ago. The lightweight 4B model can run on edge devices such as phones, laptops, and smartwatches, the company claimed.
The model has a time-to-first-audio (TTFA) — a measure of when the model starts “speaking” after receiving input — of 90 ms for a 10-second sample of 500 characters. The model can render a 10-second clip within 1.6 seconds.
Now that Mistral has both speech-to-text and text-to-speech models, the company aims to support end-to-end agents.
On the other hand, Cohere launched its first voice model in Cohere Transcribe, a 2B model that supports 14 languages. The company claimed that it outperforms competitors such as Zoom Scribe v1, ElevenLabs Scribe v2, and OpenAI Whisper Large v3 on average Word Error Rate with WER of 5.42.
Both companies are thinking that releasing an open source model will attract a lot more developer interest, and also enterprise interest, as they can tune the model according to their needs easily.
Granola raises $125 million at $1.5 billion valuation
Earlier this month, Granola was caught in a bit of a storm when users found out that their data was locked after their plan was being downgraded. What’s more, the company also locked down its local database, breaking some agent workflows as I wrote last week.
This didn’t deter Granola from securing $125 million in Series C funding from Danny Rimer at Index Ventures. The company’s valuation rose 6x from $250 million in the last round to $1.5 billion this year.
For years, Granola has acted as a “Single player” app, which sits on your computer, transcribes meetings, and takes notes. But with this fundraiser and feature release, the company wants to become more useful for enterprises and teams. Granola announced a new feature called Spaces, a way to organize meeting notes and folders within Granola for a team.
Granola already had folders, but with this release, they move within Spaces. Users can create folders within folders and control access. You can also filter folders by a person or a company. Plus, you can ask questions to Granola’s AI assistant with the context of Spaces or Folders.
The company’s founder, Chris Pedregal, told Bloomberg that agentic AI features are the next in the pipeline for Granola. One thing that stood out was Pedregal agreeing to the fact that meetings are commoditized. ElevenLabs CEO Mati Staniszewski also said a similar thing about voice models last year.
Quickly growing companies are understanding that models and features are replicable. When Granola was caught in the data skirmish, multiple people spun up open-source alternatives. Granola is still popular among users, but it recognizes that people can move to other alternatives, but companies and enterprises will treasure the shared context and AI workflows that could potentially stem from it. This also means that other players will step in to address the prosumer market with advanced features, given that Granola is charging $14 per month per user.
Quick bytes
Indian voice AI company Gnani.ai, which focuses heavily on voice agents for the financial sector, raises $10 million
As I reported earlier this week, ElevenLabs is set to launch its music generation app soon
Dell and Deepgram are partnering to build a dedicated voice AI infrastructure
Google expanded its live search feature, which lets you converse with search in real-time and share a video feed globally
Google Translate’s real-time headphone translation expands to iOS. The feature will now be available in more countries, such as Germany, Spain, France, the UK, and Japan.
Music AI platform Suno launches v5.5 of its platform with features like voice cloning and custom models based on their tracks
Google launches Lyria 3 Pro model that understands song structures such as intros, verses, choruses, and bridges. The company is also making the model available for developers through Vertex Cloud, Gemini API, and Google AI Studio
Smallest AI releases a new text-to-speech model with a word error rate (WER) 5.38%
Signals and Experiments
Last week, 404 Media covered a company called WebinarTV, which secretly scans for open Zoom links all across the internet, gets into these Zoom meetings, records them, and hosts them as webinars on its platform. The report noted that a lot of people got to know about the site when Webinar TV “featured” them, and they got an email.
One person said that they specifically didn’t record the webinar because it was about politics. They got an email from the site stating they were featured on an AI-generated show by WebniarTV.
There are other reports stating the notorious activity of the site. Stanford warned its employees and students that sites like WebinarTV could end up on Stanford.
A report from the cybersecurity collective CyberAlberta said that WebinarTv tries to get access to meetings using third-party browser extensions, which can gain access to webinar links through calendar linking or when a user submits a link for services like transcriptions.
This is very creepy for me as I attend a lot of AI meetings. Largely, I only have a few people in a meeting, so if a bot or another user joins, it is easy to spot. WebinarTV likely targets public meetings where a lot of people join and unjoin. But meeting sabotage is a real issue, and we would need to know more about how different meeting AI notetaking tools are handling security as they are getting more common and prevalent.
Thank you for tuning in. Keep listening.
Email: voiceaiweek@gmail.com or im@ivanmehta.com


