Modulate, the frontier audio AI firm, has secured $25m in fresh capital as organisations deploying voice agents discover that a transcript alone cannot reveal fraud, emotion or synthetic speech.
Future Ventures led the round, with Hyperplane and Lakestar also taking part. Modulate will direct the capital towards AI and machine learning research, product development, engineering, developer relations and partnerships.
It also intends to widen the range of APIs, models and deployment options it offers to developers, and to scale its developer and partner ecosystems in response to accelerating demand for AI that grasps the full context of spoken communication.
The raise arrives after a stretch of strong technical and commercial progress. Modulate’s models currently process upwards of 10m hours of audio every month, and the cumulative total has recently passed 600m hours.
The company took first place on Hugging Face’s Open ASR Leaderboard for transcription and presently holds the top spot on the platform’s deepfake speech benchmark. Batch transcription through its API costs $0.03 per hour, while its deepfake detection achieves 98.9% accuracy on public benchmark data.
Central to the business is Velma, a flagship platform that examines audio directly rather than relying on text. It picks up signals such as tone, emotion, intent, emphasis, synthetic speech and conversational behaviour.
These can be applied individually or combined to flag broader events, including attempted fraud, AI agent breakdowns, harassment, customer frustration and policy breaches. Because Velma works in real time, applications can step in while a conversation is still under way. According to the company, Velma delivers twice the accuracy of conventional LLMs in identifying true positives and produces seven times fewer false positives.
Velma runs on Modulate’s Ensemble Listening Model (ELM) architecture. Instead of depending on one enormous foundation model, ELM coordinates over 100 specialised audio models, choosing and blending them to produce precise outputs with far lower inference compute. Modulate reports efficiency gains of up to 1,000x compared with a single large model, cutting the cost, energy and memory needed for large-scale audio analysis.
Modulate builds audio-native AI that helps machines interpret what is actually taking place in a voice interaction. Its technology is deployed daily to shield healthcare institutions from deepfake attackers, strengthen emotional awareness in voice agents, curb extremism and harassment on social platforms, monitor voice agent performance, identify child grooming in voice conversations and conceal the identities of agents operating in high-risk situations through voice masking. Organisations also use its models to spot suspicious behaviour in high-risk calls and to recognise when a voice agent is misreading or irritating a customer.
The company is also enlarging its team and infrastructure to serve developers and partners across security, customer experience, communications, AI agent supervision and trust and safety. Planned work includes new industry models, additional SDKs and APIs, more developer relations resources, partner integrations and support for new deployment environments. The goal is to let firms building voice agents, communications platforms and security tools embed advanced audio understanding without creating specialised models in-house.
Modulate CEO and co-founder Carter Huffman said, “Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript. We’re already using audio-native AI to protect organizations from deepfake attacks, help voice agents understand emotion and respond with more empathy, identify dangerous behavior in online conversations, and monitor whether voice agents are actually performing the way they’re supposed to.
“Underneath all of that are more than a hundred specialized models working together to understand what’s really happening across audio, with dramatically less cost and compute than traditional large models.”
Huffman added, “Developers shouldn’t have to rebuild the audio intelligence layer every time they create a new voice experience. Our mission is to build the models and infrastructure that let them focus on the application they want to create. The opportunity facing audio-native AI is expanding incredibly quickly. We’ve built the technology and proven it at scale, and this investment lets us grow the team and move faster to meet that demand.”
Copyright © 2026 RegTech Analyst
Copyright © 2026 RegTech Analyst





