Company profile
Gradium
Voice layer for modern applications and AI agents with real-time, scalable voice APIs.
- Category
- NLP & speech
- Headquarters
- Paris, Île-de-France
- Sells to
- Developers
- Business model
- Freemium, Usage-based API
- Deployment
- API
- Pricing
- Usage based · from $13/mo · free tier
- Builds own models
- Yes
- Modalities
- speech, Audio
What Gradium does
Gradium provides real-time, scalable voice APIs for modern applications and AI agents, built on cutting-edge research. They offer Text-to-Speech (TTS), Speech-to-Text (STT), smart turn-taking, and instant voice cloning. Gradium builds the technological backbone, models, and infrastructure to support all voice applications, focusing on natural, expressive, real-time voice interactions at scale. Their founders have invented and published methods and algorithms behind many existing voice and audio models, translating this research into production-ready systems for developers and enterprises while continually pushing research frontiers.
Products
- Text-to-Speech (TTS)Turn text into natural sounding voice with low-latency and high-quality output.
- Speech-to-Text (STT)Transcribe audio into text instantly with best-in-class accuracy, including keyword boosting for specific vocabulary.
- Live TranslationConvert voice to voice in real time.
- Voice CloningOffers instant or pro voice cloning capabilities.
- On-Device TTSOffline real-time natural voice generation.
- Speech-to-Speech TranslationReal-time translation of spoken language.
- GradbotTool to build voice agents in a single prompt.
Key capabilities
- Real-time voice APIs
- Scalable voice APIs
- Smart turn-taking
- Instant voice cloning
- On-device TTS
- Keyword boosting for Speech-to-Text
- Multilingual support (English, French, German, Spanish, Portuguese)
- Low-latency (below 300ms time-to-first-token when streaming)
- Voice library with multiple voices
- Semantic VAD (Voice Activity Detection) in STT streams
- Adaptive delay control
- Studio access
- API access
- Commercial use
- Engineering Support (for grant recipients)
- Early Access to Newest Models and Research Previews (for grant recipients)
Use cases
- Building production voice systems
- Creating voice agents
- Integrating voice capabilities into modern applications
- Translating spoken language in real-time
AI approach
Gradium provides real-time, scalable voice APIs for Text-to-Speech (TTS), Speech-to-Text (STT), smart turn-taking, and instant voice cloning. They build their own models and infrastructure based on cutting-edge research, translating over a decade of open research into production-ready systems. They also push the frontier of research to invent next-generation voice algorithms.
Tech named: TTS, STT, smart turn-taking, instant voice cloning, real-time voice AI, Keyword boosting, neural audio codecs, audio language models, Semantic VAD, Adaptive delay control
What it says sets it apart
- Built on cutting-edge research
- Founders have invented and published methods and algorithms behind most voice and audio models existing today
- Translating over a decade of open research into production-ready systems
- Continually pushing the frontier of research to invent the next generation of voice algorithms
- Low-latency and high-quality output
- Best-in-class accuracy for Speech-to-Text
- Keyword boosting for STT to recognize specific vocabulary
Funding rounds we track
Nvidia, NVIDIA, FirstMark Capital, Eurazeo, DST Global Partners, Eric Schmidt, Xavier Niel
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from Gradium's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.