#27: Local processing (Acoustic version)
The post-tokenmaxing era: Local AI via sound waves
In Focus
India’s ToneTag bets on its acoustic tech for local processing
After the short era of “Tokenmaxing”, companies using AI are being more cautious about their token budgets and trying to minimize costs. One way that has emerged in this frugal era is to offload part of the processing to local models and chips.
Even if companies can infer some queries without having to go to the cloud, they can start to save on costs. Indian fintech company ToneTag is doing this by processing acoustic signals and answering merchant queries about recent payments locally.
After the Unified Payments Interface (UPI) gained prominence as a digital payment medium in India, banks, payment providers, and fintech startups started deploying small speaker boxes with QR codes. These QR codes “announced” the payments merchants received, which was helpful for sellers who were new to digital payments or might not have a smartphone handy to check the status of a payment.
Using acoustic signals for AI processing
Until now, these boxes have been basic, with some functions like recalling recent payments and volume control. ToneTag, backed by Amazon, MasterCard, and Qualcomm, produces these boxes. The startup recently revealed a new project called eKosha, which uses AI to answer merchant queries.
The startup has a patent for using audio to process payments. At the core, the patent is about encoding and decoding data by embedding it into sound waves. While the original use of this was to facilitate payments between two devices, such as a feature phone and a payment device, the technology could be used to decode requests for language models.
ToneTag’s CEO, Kumar Abhishek, told me that phone processors are capable enough to convert all speech into digital signals like text and hand over the request to models behind Siri or Google Assistant. But for payment boxes, there needs to be a different strategy.
“When you start adding AI to IoT devices, where you have less compute, less power, and less memory, it's very difficult for you to keep listening to requests and process them. We were able to do this because we do analog filtering to make local processing possible,” Abhishek said in an interview.
What can this tech do?
With this implementation, merchants can ask anything about the box itself and troubleshoot problems. Plus, the AI assistant can answer questions about cash flow and recent transactions. This saves merchants from frantically checking bank apps to see the status of a payment or calling customer care to get basic information. The box supports 10 Indic languages and mixed-coded speech with plans for expansion.
ToneTag is also working with banks to put AI to work through cloud queries. With the device, merchants can ask banking-related queries to the device, and also get suggestions for financial products. With all these requests, the bank can gauge who might be interested in lending-related products.
“Because these boxes are conversational, banks could extend their MSME banking right on the box. A merchant can request a checkbook or a credit through voice commands. All tasks that needed a call to the relationship manager or a bank visit could be done directly from the box,” Abhishek said.
ToneTag’s new box has a design quirk for shopkeepers in a dedicated query button. This button serves two purposes: you don’t have to go across the room to start a query, ensuring queries are intentional rather than accidental activations or casual questions about the weather. Given that there is a mandatory button press for the query, merchants have to be close to the device (and, in turn, the mic), resulting in better quality of audio capture. The company said that the merchant's perception of the device and query volume saw an uptick after this implementation.
India has more than 23 million soundboxes deployed across merchants. The problem for fintech companies is merchants retaining soundboxes. Often, distributors offer a box from a company to merchants at a cheaper rate, but as subscription prices increase, merchants would move to another box because the switching cost is low. Abhishek and ToneTag, which aims to reach 500,000 merchants with the new eKosha box in the coming month, think that AI and embedded services will get merchants to stick around with the same box.
Quick Bytes
Investors are bullish that meeting capture is one of the biggest unlocks of AI. After pouring millions into digital notetakers, they are looking at the physical space where the likes of Plaud and Viaim are competing. Now, another competitor, Pocket, which has sold over 130k devices, got $11 million in funding from Accel, Y Combinator, Peak XV, and ElevenLabs CEO and co-founder Mati Staniszewski.
ElevenLabs, which has multiple celebrity voice partnerships, will power Gene Wilder’s voice for Netflix’s upcoming reality show “Wonka’s The Golden Ticket.” The company has gotten permission to use the voice from the Gene Wilder Estate, as per Variety.
Suno is prepping to launch its API for music generation with licensed data. The end goal is to get blessings from the label for partnership and provide other toolmakers with the ability to create sounds or tracks. But legal trouble shows no signs of abating for the company as Winamp Group sued the AI company for unauthorized use of its data for model training.
ElevenLabs’ value continues to soar. In a fresh tender offer, the voice AI company is looking at a $22 billion valuation, up from $11 billion during its fundraise in February.
Some reports have suggested that a good chunk of AI-generated music is to game streaming service systems and generate revenue. Tidal is taking a stance against that by demonetizing AI music revenue. The company will tag fully AI-generated tracks, but it didn’t specify what kind of tech it might use to detect AI in songs.
Signals & Experiments
I wanted to escape Delhi’s heat, so I took a two-day trip to Dehradun. While I was there, I stayed in the Taj hotel, and the hospitality was generally great. On the first night, I asked the front desk to send me a toothbrush and a few water bottles. I didn’t get them for 40 minutes. I called them again. This time, I got water bottles, but no toothbrush.
First, I missed my robot friends from hotels in China. Second, can taking requests through voice or chat solve this? I have seen hotels ask you to contact them over WhatsApp through a QR code on the TV screen. But often they just have hotel room information in the chat.
If I can leave a voice note or type my request out, there is less of a chance of missing out on items in my request. This use case doesn’t particularly need large language models. Just a way to get my entry to the request queue to the front desk.
Partner Spotlight with Atomik Growth
In a chat with a16z, ElevenLabs CEO Mati Staniszewski talks about how he and his co-founder, Piotr Dabkowski, got inspired to solve for badly dubbed movies and created a company that thinks that voice is the new interface for human-computer interaction.
Sponsored content
Thank you for tuning in. Keep listening.
This newsletter is by Ivan Mehta, a freelance reporter at TechCrunch. It covers AI and technology in voice, audio, and music. Email: voiceaiweek@gmail.com or im@ivanmehta.com




Interested to see what success for ToneTag 6 months from now.