Hume AI's life after the Google deal
And its new Voice EQ benchmark
Earlier this year, Google poached Hume AI’s team of engineers, including CEO Alan Cowen. The startup, which had raised more than $80 million in funding, worked on models that understand users’ emotions.
Before the deal, the company focused on building emotionally intelligent voice and language models and a pipeline for emotional labeling and metadata from audio. After the deal, under new CEO Andrew Ettinger’s leadership, the company has taken a new direction.
In a chat with me, Ettinger said the company now positions itself as a neutral player by becoming an emotionally intelligent voice infrastructure company.
Hume developed models around semantic space theory, measuring 48 emotions. While the models and pipeline were robust, the company realized that there was high value in its valuation and the data processing framework infrastructure.
“Our infrastructure was something that no one else in the world had because of the level at which we measured and got to that granularity. And so we started licensing that infrastructure out, and have built a fantastic business on top of that in the last five or six months,” Ettinger said.
“You could look at us as Switzerland, where we’re trying to provide infrastructure. The fact that we built this infrastructure with our research for our own models gives us a really unique perspective from first principles to explain and show people how we did this across hundreds of millions of hours of data and human judgments and emotion collections along with an evaluation infrastructure,” he added.
Leaning more into being a model-agnostic infrastructure company, Hume AI is releasing a new benchmark, called Real World VoiceEQ, to evaluate models on emotional intelligence and how “human” they are. The company tested more than 40 models for 60 metrics like Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Speech-to-Speech (S2S), and Speech Understanding. It grounded its results in over a million evaluations, made by 10,000 participants, across the world using its Kairos platform, which is in private preview.
Here are some of the top observations from the study:
Models can specialize in one ability, like technical accuracy, and falter in another, like emotional intelligence. This means we will see more specialized models. Plus, one-size-fits-all benchmarks are not enough to evaluate different features of a model.
Voice models have become better at speaking and understanding the emotions of other speakers. However, they still don’t understand nuances like tonality and pacing.
As we have covered in my story about Treble x Hugging Face benchmarks to measure model performance for real-life scenarios, evaluation metrics often ignore situations like noisy environments, overlapping speakers, and emotional expression.
One of the biggest challenges for the voice AI industry is speech-to-speech models that listen and speak at the same time.
The study also revealed that even with accent variation in English, models can struggle to understand speech. Here is a passage from the research paper that suggests the error rate goes up 1.5x–2x for South Asian languages:
The accent results show that the native versus non-native gap is the most informative fairness axis. Across selected systems, the native-to-non-native WER gap ranges from +2.15 percentage points for ElevenLabs Scribe to +8.04 percentage points for AssemblyAI Universal-3 Pro. This gap is larger and more diagnostic than a standard-American versus other-regional-accents comparison, because the latter mixes easier native accents with harder non-native speech and hides the underlying fairness issue. Under Kachru’s Three Circles taxonomy, the second-language (Outer Circle) group — Indian, South Asian, Nigerian, Filipino, and related varieties — is the hardest condition for many models. In this subgroup, the top ten systems fall between 6.58% (Scribe) and 12.96% (Modulate Velma-2) WER, while the same systems score between 4.38% and 6.96% on native English.
Hume’s Olya Ossipova said that while studying benchmarks, the company observed that the current way of measuring voice AI models is often just to look at a transcript.
“When you think about expressiveness, what we found is that the model that performs best on expressiveness only ranks fourth in terms of taking action on that expressiveness. So there was actually like a very interesting gap that even if a model understands emotion, it doesn't really know what to do with it,” Ossipova said.
She also mentioned that there is an issue in real-life testing of models where LLM judges use objectivity rather than subjectivity, so you need to perform human evaluations as well. Ossipova said that peering into these problems will give model builders a way to make voice AI models more efficient.
Hume is currently running Kairos, which lets you automate evaluation models for this index, in a private preview. The company is also planning a new partnership with Hugging Face for benchmarks. There is a rising trend of evaluating voice models in different real-life situations, and also assessing models based on their emotional understanding.



