An Interview with Wing VC's Zach DeWitt on how voice can unlock trapped knowledge
The death of the unrecorded meeting
Last month, Wing VC released a list of prominent enterprise startups across different stages. The list featured many voice companies, including ElevenLabs, Deepgram, Decagon, Granola, Gong, Listen Labs, and Retell.
I sat down with Wing VC’s Zach DeWitt to understand voice AI’s opportunity and challenges. DeWitt has been at the firm for over eight years and looks after its application investing.
Wing VC has been an early investor in unicorns like Gong and Deepgram and thinks that voice is a prominent opportunity.
Voice as an interface
When I asked DeWitt about what he thinks Wing VC’s list represents, like many other people in tech, he feels that voice is the new interface.
“Yeah, I think the thesis that we have, and this will be expressed in different ways, is that voice is just going to be the dominant interface. I mean, I think voice is going to take over everything we do. And it’s uncomfortable today because we really like having a graphic user interface, a GUI. We really like clicking on things. And that’s still going to exist in certain use cases,” he said.
He referenced apps like Wispr Flow and Granola, which he uses every day. However, dictation apps are currently an input mechanism that works well on desktop, and for Granola, speech is a modality to capture.
For years, consumers have interfaced with assistants like Siri and Alexa, but they haven’t had a satisfactory result in getting the results they want. There is a way to go before users talk to their gadgets more than they type. DeWitt agreed with this and said that we will slowly see the rise of voice-first interfaces in apps.
“Right now, voice is a second-class citizen. It fits into the existing interface layer we have. As you said, you use voice to input a block of text into Slack. That’s voice working around the existing interface. A lot of interfaces today, almost all of them, are not built for a voice-first feature. And so what does that [a voice -first interface] look like? I think we’re going to have much more fluid interfaces, much more just-in-time adaptive interfaces based on the command, and so I think voice is going to be more fluid between desktop and mobile,” he said.
DeWitt thinks that we don’t have interfaces or mechanisms that sync across devices seamlessly. For instance, there aren’t efficient assistants that can switch context between devices such as mobile, desktop, and car while keeping the context alive.
“The interfaces are going to change a lot, and as voice becomes more dominant, people get more addicted to granola and whisper flow, which I can’t live without either product right now. I think we’re going to see new types of interfaces that are crafted around voice, and I think this is really the great opportunity for application founders and teams to reimagine interfaces,” he said.
What’s more, on the consumer end, DeWitt feels that in the previous generation, there have been attempts to build AI speech coaches, tutors, or therapists. With the current generation of models and tech, there could be more successful companies in this area.
I think every conversation will be recorded in the enterprise to gain insight.
Enterprise opportunity
Areas like customer support and sales are clear examples of voice AI’s opportunity. I asked DeWitt what other enterprise trends he foresees in voice. One of DeWitt’s most striking predictions was that enterprise operations will get more transparency and insights due to voice.
“I think every conversation will be recorded in the enterprise. It’s easy to dismiss this as, oh, that sounds like a Big Brother initiative. But actually, it’s way better for people inside the company if you have a manager who maybe is not following conduct or is aggressive or whatnot. It’s way better to have this stuff recorded, and the business knows about it. But more importantly, just to really understand all of the trapped knowledge that’s happening inside these conversations. So certainly customer support, certainly sales, right? Anything that has external conversations,” he said.
He also thinks that hiring is a big opportunity in voice.
“You can also add hiring [to the list of verticals where voice AI can shine]. Hiring is going to be a big opportunity for voice AI. There are a lot of businesses that are starting to screen candidates using voice AI. And then there are questions businesses and enterprises can’t even answer: We had an interview with this candidate. Did our side show up on time? Were there two people in the meeting or three? What was covered? Does the candidate have redundant experience where you know that the candidate asked 10 questions? In the next interview, they’re asked eight of those same 10 questions again. So I think that we’re going to see in hiring, we’re certainly going to see internal meetings too,” he said.
I think there’s really a limitless opportunity in voice in vertical applications.
Besides this, he thought that there were specialized voice AI companies that work in a specific vertical, such as a research assistant for scientists or a voice system for restaurants, that could shine through. DeWitt gave me an example of Slang AI, which recently raised $36 million in Series B funding:
One example is a company that we’re invested in called Slang. It’s A really good founding team that came out of Spotify, and they’re the leader in restaurant voice and hospitality voice. If you’re a restaurant and you’re calling to change a reservation, ask if there’s space, ask about dietary restrictions, or ask about group dining, you’re probably going to be on hold, it’s going to be noisy, and it’s going to be hard to hear. They use AI to handle the conversations end-to-end, like Gong.
Earlier this year, Deepgram acquired a restaurant-focused voice AI company, OfOne. DeWitt thinks that while the likes of OpenAI and Anthropic are not currently looking to dominate voice verticals, voice AI companies could go deeper in industry-specific slants and compete with verticalized agent providers. He said that, for instance, Deepgram is focusing on the medical sector as the company understands the vocabulary and intricacies of working in the health sector.
“And so I think these vertical application businesses are going to be competing to understand the customer better, to deliver the best workflow, to give the best customer experience with the highest NPS (Net Promoter Score). And they’re going to be relying on these foundation models,” he said.
Specifically on Deepgram, he said that the company has ambitions to be the “Google of voice” and it aims to power and search every conversation inside companies.
Future and opportunities
Companies and research labs around the world are releasing more voice models. DeWitt thinks that’s a good thing as it helps more startups and ideas to sustain at a lower cost. However, large model makers like ElevenLabs or Deepgram would have an advantage in the ability and infrastructure to scale and reduce the cost along the way, which might not be easy for an open-source model.
DeWitt sees that the future and opportunity in voice is tremendous, and we can see companies that are worth billions of dollars.
“My belief is that the voice market itself is going to be, you know, probably one of the biggest markets we’re going to see and in our generation. And so just being elite at that is certainly enough to build a multi-billion dollar business,” he said.


