Signals
HumanX transcribed all sessions, and all conferences should do the same
When I go to conferences, one of my top pieces of feedback is to make different sessions available to the press as soon as possible. At HumanX, held in Amsterdam last week, I didn’t have to do that because the organizers created a press portal with sessions uploaded quickly after they finished, with Deepgram powering transcriptions. The press portal also had notable quotes and numbers that speakers said on stage, which were extracted from the transcriptions.
People who were not part of the press and still wanted to see key points discussed in a session could use the HumanX app, where they would see a summary generated by Granola. In large conferences, you might only be able to attend a few talks depending on your role. You might be giving or conducting interviews, attending pre-scheduled or spontaneous meetings. And even if you can’t rewatch all sessions, summaries can at least give you an outline of them.
More conferences should be providing ways for attendees to engage with sessions even when they are at the event.
In Focus
Meta’s new hardware gives Muse a voice, eyes, and ears
For more than a decade, Meta has been trying to become a platform that could reach users independent of iOS or Android. The Quest series and Metaverse were some of the early gambits at achieving this. Meta Ray-Ban has been fairly successful, and its Muse AI agent is showing early signs of high-engagement usage; the combination of smart glasses and an AI agent could become a torchbearer for the company’s bid to find a new platform.
Meta has used voice as an input mechanism for the Ray-Ban smart glasses, with the camera providing additional context. However, people have had privacy concerns over recording. Meta’s defense has been its recording light, which is on when you record. In a conversation with Joanna Stern, Mark Zuckerberg argued that people are recording with their phone held up all the time; other users don’t get to know about it because there is no recording light.
The counterargument is that one is a camera on your head in the form of glasses, and the other is a visible device held up by a person. The privacy debate might not be deterring enthusiasts and creators from buying the smart glasses, but there is a certain negative connotation about them.
That’s why when Meta released the Ray-Ban Meta Audio during the keynote, it almost felt like a response to privacy concerns. However, products like these might have been in development for months, if not years. The new device runs purely on a voice interface. Users can chat with Muse through voice and get answers through speakers. The glasses also act like a pair of earphones to listen to music and audio. If you are someone who hates the idea of wearing headphones, earphones, or ear clips, these could be a good alternative to consider.
Meta also released its Tamagotchi-like device, called Muse Charm, to control the Muse agent. This is a consumer version of countless devices on the market used for agent orchestration. Meta has given Muse a face in the form of the mascot Jolly and a voice as well. Users can interact with Muse through “a state-of-the-art real-time voice model,” as per Meta. The keychain-like device has a camera for additional context, too. Pendants like Looki L1 have tried this trick.
Meta is betting big on voice as both input and output with its new set of devices. It’s not always a good thing, given that answer engines and large language models behind the agents spit out long answers that you might not want to hear. People do like to see some visual output even if an AI assistant is reading it out in parallel. The stickiness of the devices and voice-based agents will also depend on the accuracy of their answers.
Numbers Game
UMG said that over 50% of the catalog in a streaming service comes from indie DistroKid, which could cause more AI songs to surface on that service.
Voice AI company Kardome’s survey revealed that 83% of respondents mainly or only use a physical remote, indicating that voice hasn’t taken over TV controls.
Quick Bytes
ChatGPT’s mobile app now gets a voice agent that gets work done
OpenAI enabled voice-based work on the ChatGPT desktop app that lets users use Codex and Work through voice. The company has now brought similar capability to the mobile app where Plus and Pro users last week can go to the Work tab and use their voice to create documents or draft emails on the go.
ElevenLabs releases new studio update
ElevenLabs released a new Studio 4.0 update to tie in sound, music, and video creation under one interface. The company has added frame-level zoom control for better cuts, and users can pick from over 10,000 AI voices.
Jujutsu Kaisen voice actor takes TikTok to court over AI-generated cloning
Japanese voice-over artist Kenjiro Tsuda sued TikTok over apparent AI clones of his voice appearing in short videos. The platform maintained that the voice was a “generic male voice” that might be similar to Tsuda’s voice.
Deals Corner
Heidi Health ($340 million): Melbourne-based medical AI company that has both a scribe and dictation product for clinicians. Out of $340 million, $240 million is from General Catalyst’s Customer Value Fund (CVF).
Investors: Blackbird (lead, first check into the company), Phoenix Court, P72 VC
Sela ($21 million): San Francisco voice AI company whose agents handle mortgage sales, helping lenders originate roughly $1B in loans per month.
Investors: Costanoa (lead), Emergence Capital
Dextr AI ($6.7 million): Startup using voice agents to handle hotel reservations, guest requests, and staff coordination.
Investors: Elevation Capital (lead), Foundation Capital
InTouchNow.ai (£2.3 million): UK voice AI for NHS primary care. It makes AI call handlers that answer GP practice calls 24/7, book appointments, and do triage, live at 150+ practices.
Investors: Ada Ventures (lead), Exceptional Ventures, Kadmos Capital








