Modulate Inc., a voice AI innovator specializing in audio-native models that analyze raw conversation audio for emotion, intent, and fraud signals, has raised $25 million to broaden developer access and industry integration.
- Raised $25M to scale developer access to specialized audio AI models
- Velma platform uses multiple small models for efficient real-time audio analysis
- Applications span gaming, healthcare security, and AI agent evaluation
Market signal
Modulate’s recent $25 million funding round underscores growing demand for voice AI that processes nuanced audio signals beyond transcription. By focusing on emotion, tone, intent, and deepfake detection, its technology meets emerging needs for more intelligent voice interfaces in sectors like gaming, healthcare, and AI-driven customer service.
The company’s distinctive multi-model approach enables up to 1,000 times greater processing efficiency compared to monolithic models, making it commercially viable for large-scale real-time audio applications. With over 10 million audio hours processed monthly and top rankings on Hugging Face leaderboards, Modulate is positioned as a leader in audio-native AI.
Operator impact
Operators running interactive voice services or monitoring voice content in real time can leverage Modulate’s models to enhance fraud detection, caller sentiment analysis, and moderation of harmful or inappropriate speech. Healthcare providers benefit from voice deepfake detection to protect against impersonation attacks, while gaming platforms continue to use the models for harassment prevention.
Developers and operators gain from the upcoming software development kits (SDKs) and APIs aimed at simplifying integration of audio intelligence layers into varied voice experiences. The company’s focus on industry-specific models and partner ecosystems will offer operators tailored solutions and more deployment flexibility.
What to watch next
In the near term, monitoring the rollout and adoption of Modulate’s new SDKs and APIs will be key to understanding how rapidly audio-native AI integrates into developer workflows and voice applications. Expansion into vertical-specific models could further accelerate deployment in regulated sectors like healthcare and security.
Additionally, growth in use cases for AI agents and live conversation analytics will test the scalability and adaptability of Modulate’s Ensemble Listening Model architecture. Upticks in fraud and deepfake voice threats may drive stronger demand for real-time voice authentication and behavioral signal analysis capabilities.