#26: Heidi's health hardware
Why is it important to own voice stack in the health world
In focus
Why Heidi is building specialized hardware for health
Health-focused voice AI companies have raised over $280 million in funding this year, including Assort Health’s recent $120 million Series C raise to scale voice agents that have handled over 190 million calls. Of that total, nearly $100 million has gone to companies that have a scribe product for notetaking and follow-up actions.
Hardware meeting notetakers are a rising category. Plaud, Anker, Vibe, and Pocket are focusing on professions like doctors and healthcare providers to rely on their gadgets to take notes or record patient sessions. These companies are also targeting all professions where in-person meetings and conversations are taking precedence.
In contrast, Australia-based Heidi Health is taking matters into its own hands by using its specialized models — trained to understand medical terminology better — and approaching clinics directly to deploy its scribe. The company also decided to build its own hardware scribe to improve transcription quality.
I spoke with the company’s CEO, Tom Kelly, about its approach, which has now scaled to handling 3 million weekly visits. Heidi began as a tool for a patient’s history and a chatbot built on top of that.
Building its own models
The company realized that healthcare providers spend a lot of work in admin-related tasks such as follow-ups and filling forms based on information patients provided. That is why Heidi started focusing on making its scribe top-notch. It also wanted to sell directly to clinics to get direct feedback and improve the product.
Kelly said that the company tuned models for medical use and claimed they are better than what frontier labs have to offer for speech-to-text and language summarization.
“In 2025, we struggled a lot with third parties having outages or variations in performance. So we would notice, for example, when a Frontier Lab launched a brand new model, they would be very excited about it. Everyone at Heidi would be disappointed because we would experience a degradation in the old model,” Kelly said.
He added that another problem was that Heidi didn’t have a full view of what happens to requests underneath. That is why the company started building and relying on its own fine-tuned models.
Heidi said that often, generalized models try to “guess” the medical term, which could be dangerous for patients. That is why the startup tries to ground its models in specialized medical terminology for different sectors.
Hardware play
Besides models, Heidi wants to own other parts of its voice stack, including hardware. Rather than relying on just the phone mic or third-party gadgets, the company built its own notetaking device for better transcription quality and, in turn, better follow-up actions.
“The most common cause of bad transcripts is having a bad microphone or a phone that’s put in a bad position. So it’s really important for us to never miss a doctor’s session. We want it to be high quality, and we want to be able to hear what they’re doing,” Kelly reasoned behind building its own device.
Heidi said that healthcare providers often move around a lot, and when the startup tested third-party mics, they weren’t accurate. When it started to test its own device, the results were much better, and it went on to develop its own hardware.
More than just a scribe
The company is also starting to test and deploy its own Wispr Flow-like dictation tool, but for clinical use. The idea is similar: users press a key on their desktop and transcribe what they need across apps.
The company’s aim is to run all of its processing offline in the future. Kelly said that it is possible to run reliable dictation or transcription on a device, but for running LLMs, you either need a dedicated device for a hospital with on-prem deployment or buy a machine with large RAM and storage. It also wants to tap into its origins and reason from a patient’s history to provide a correct diagnosis.
A Menlo Ventures report from 2025 suggested that AI note-taking in healthcare saw a $600 million spend. Investors are very interested in this category as clinics might not switch scribes easily, and startups that build a context graph for health over time will have a dominant market positioning.
Signals & Experiments
Amazon’s Alexa assistant has existed for nearly a decade in India. After its introduction in 2017 with support for English, it later added Hindi. Yes, it understands both languages and sometimes a mix of them. Alexa needs drastic improvements in understanding users.
I have a fan and a table lamp connected to the smart home ecosystem that I can control through Alexa. And often Alexa switches on one instead of another. And at times, it just plays a song instead of switching off the light. I barely ask Alexa for facts or news because it often takes me three to four times to just get my request right.
Last week, I reported for TechCrunch that Amazon is testing Alexa+ in Hindi in India. With this experiment, I hope Amazon is focusing more on accurately understanding Indian speech rather than just adding a “conversational” assistant.
Amazon and Alexa still command a large share of India’s smart home market. However, if you want an assistant to chat with for news, facts, music, and more, there are dozens of options out there with new ones launching every week. Thanks to newer AI models, some of them are good at understanding varied speech. Amazon will need to bring better models to create a stickier experience for consumers.
Quick Bytes
Coval, a startup built by former Waymo engineer Brooke Hopkins, raised $28 million in a round led by Norwest. Voice AI is scaling up, and enterprises are in need of an infrastructure that tells them how agents are performing and if their money is being used effectively.
Voice startup Speechify launched its dictation tools for iOS and Mac. With people building their own tools for dictation, the market is heating up. People in the tech industry might choose open-source or niche tools or create one for themselves. However, enterprise customers will tilt towards the likes of Wispr Flow, Willow, and Superwhisper for compliance reasons.
Zoom is leaning into the ethos of agents working based on meeting notes and discussions with Zoommates. These new agents can create first-draft presentations, reports, project plans, proposals, and spreadsheets from conversations.
Voice AI company Modulate released a new API to catch AI-generated music. There is a strong wave of AI-generated music making it to streaming services, and companies like Deezer have been fighting it hard with their own tools. Labs like ElevenLabs are adopting Google’s SynthID for generated audio watermarking. But it is still not a widely accepted standard across the board.
Off Topic
I love testing new apps, and I have been obsessed with this iOS weather app called Good Air. It is beautiful and has great transitions as you scroll to different times. Plus, it also tells me the AQI level near me. As someone who lives in one of the most polluted cities. That is important information, and at least I get to look at the horrible weather with a good interface.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com




