Deepgram CEO says it is harder to build a good open-source speech-to-text model
Inside Deepgram's new Flux model with in-built turn detection and why context passing will be important for voice agents
In a world where new companies like ElevenLabs and Cartesia have taken over a lot of headlines, Deepgram has operated in the voice AI space for more than a decade. The company is having a solid year with $130 million in Series C funding in January, which made it a unicorn. The company also struck a partnership with IBM and acquired a quick-serve voice company, OfOne, with the intent to get into verticalized spaces.
While Deepgram offers full-stack services, it is revered for its speech recognition models. The company says companies have built over 2,000 products on top of its voice stack. I sat down with the company’s CEO, Scott Stephenson, to talk about its new Flux multilingual model release, context carrying for voice agents, and why it is difficult to build open-source, speech-to-text models.
Flux multilingual
Last year, the company made a shift in its model with a new Flux series, focusing on powering real-time conversation instead of batch processing.
The earlier Nova series of models is still active, but it wasn’t designed from the ground up with live conversation in mind. As more enterprises shift to AI-aided consumer service and experience, demand for voice agents is rising, and voice AI providers are racing to build models and infrastructure that solve problems in real-time conversations, such as interruptions and audio cleaning.
In multiple countries, people speak in a mixed code manner. That means they switch between languages mid-sentence. In April, Deepgram made its Flux series of models multilingual, with a focus on ten languages. The company has built conversational nuances such as turn detection and code switching in the Flux model.
Stephenson thinks that more companies are going to switch to Flux given the high demand in the consumer support world.
“Nova was a prerecorded and batch-first architecture that we also made real-time. It was more for real-time, closed captions, and that type of thing. So less for hyper-performant voice agents with the lowest latency and that type of thing,” Stephenson said.
“Flux is probably more performant, is probably faster, and it certainly is more performant in the task of if you’re switching languages and noticing that’s happening and then having a higher quality, higher accuracy when that’s happening.”
He said that typically models looked for silence to detect turns, which wasn’t the best method because the user might have gotten distracted or was trying to remember something. If I were recalling my account number, I don’t need a voice agent to interrupt me.
Deepgram said that the company needed to include different filter words and cadences of different languages to make the model truly multilingual.
Passing context in voice models
Deepgram CEO noted that the way you make voice agents’ voice context will matter as much as having high-performing models. For instance, if your voice agent can detect that someone is in a loud environment, and it speaks more slowly and louder, the person on the other end of the call would be able to understand the voice AI bot better.
Stephenson said that in the current voice agent paradigm, where there is an architecture of speech-to-text, text-to-text, and text-to-speech models, the more context-like turn detection and pauses that you pass to models, the more accurate and “human-like” the agents become. He said that cues like a change in tone are critical to detecting if someone is ending their turn.
“This is one of the biggest changes that is going to happen in this next evolution and revolution in voice agents, where you add context everywhere. And this is what the flux line is about. Flux conversational speech recognition adds more context to perception. And later on, with more flux releases, we’ll talk about adding context to the rest of the stack,” he said.
“I think about this is that there are sort of two things that you’re fighting when you’re building a voice agent. Between turns, you don’t want to wait 10 seconds of silence until you finally kick off and start responding. Then you have the other side, which is if you look for just a little bit of silence, and any little bit of silence you find, you jump right in. Then everybody is like, ‘This thing is interrupting me all the time,” he said.
It’s hard and expensive to do open-source speech-to-text models. There are way fewer people doing that.
Future of the open-source model
Stephenson says when the cost to get started in voice drops, we might see less activity on the model front in open-source. But we might see more activity in people building code for adding context or orchestrating voice agents.
He thinks that it is hard to build a good open speech-to-text model.
“It’s hard and expensive to do open-source speech-to-text models. There are way fewer people doing that. Whisper was launched, I don’t know, three years ago, three and a half years ago, and that probably still is the best open source speech-to-text engine. So there isn’t a lot of activity in that area,” he said.
He thinks the reason for that is that someone building a speech-to-text model has to get millions of hours of speech data to work with and have money to train models.
Notably, labs like Mistral and Cohere have released open-source transcription models recently, but they do have financial backing to spend on research.
He thinks that it is easier for builders to get into open-source voice generation and orchestration than transcription.
Way forward for DeepGram
The company has different lines of models for users, but it wants to give them the option for “good, better, best” to choose from in different parts of the voice stack. Just like ChatGPT offers Instant, or Google offers Flash, Deepgram wants to offer a cheaper and faster version of its top models.
In an interview with Week in Voice AI, Deepgram investor Zach DeWitt called “Google of voice.” I asked Stephenson about the comparison; he leaned into the cultural parallel, noting that Deepgram’s work environment allows people to expand their knowledge.
“Google has an amazing work environment. So when you’re at Deepgram, there are sort of infinite possibilities to expand your knowledge and move from one area of the company to another and engage with different folks and have ample time that’s outside of your scheduled work that you’re supposed to be doing in order to expand and explore,” Stephenson said.




