Company profile
Deepgram
Real-time API platform for speech-to-text, text-to-speech, and voice agents.
- Category
- NLP & speech
- Headquarters
- San Francisco, California
- Sells to
- Mixed
- Business model
- Usage-based API, Freemium
- Deployment
- Cloud / SaaS, On-premise, Self-hosted, API, Hybrid
- Pricing
- Pay As You Go, Growth (pre-paid credits), Enterprise (custom) · free tier
- Builds own models
- Yes
- Modalities
- speech, Audio, Text
What Deepgram does
Deepgram is a foundational AI company specializing in voice technology. It provides a real-time API platform for speech-to-text (STT), text-to-speech (TTS), and voice agents, enabling human-machine interactions. The platform is built on proprietary voice-native foundation models and runtime infrastructure, designed for high accuracy, low latency, and enterprise reliability. Deepgram's offerings include advanced speech recognition models like Nova-3 and Flux, text-to-speech models like Aura-2, and a unified Voice Agent API that combines STT, TTS, and LLM orchestration. The company also offers Audio Intelligence models for extracting insights from conversational audio, such as sentiment analysis, intent recognition, and topic detection. Deepgram supports various deployment options, including cloud, dedicated single-tenant, in-VPC, and self-hosted environments, catering to diverse business needs and compliance requirements.
Products
- Nova-3Deepgram's highest performing and most accurate real-time speech-to-text model, supporting 45+ languages with advanced capabilities like Speaker Diarization, Smart Formatting, Keyterm Prompting, and Automatic Language Detection. It is recommended for most use cases, especially audio with multiple languages, background noise, crosstalk, and far-field audio. Nova-3 Medical is a specialized version for clinical environments.
- FluxThe first Conversational Speech Recognition model designed to handle interruptions, with built-in turn detection, natural interruption handling, and ultra-low latency. Flux Multilingual supports 10 languages and handles multiple languages within a single conversation.
- Aura-2Deepgram's professional, enterprise-grade text-to-speech model, optimized for high-throughput applications at enterprise scale. It produces natural, human-like voices (40+ English voices with localized accents) with sub-200ms latency, designed for real-time conversational AI applications and voice agents.
- Voice Agent APIA unified API that combines speech-to-text, LLM orchestration, and text-to-speech into a single solution for building enterprise-ready, real-time conversational AI agents. It features built-in barge-in detection, turn-taking prediction, function calling, and mid-session control.
- SagaThe Voice OS, pioneering the Voice OS for productivity.
- Audio IntelligenceModels that extract actionable insights from conversational audio using real-time APIs powered by task-specific language models. Features include summarization, sentiment analysis, intent recognition, and topic detection.
- Nova-2Deepgram's previous generation speech-to-text model, supporting 30+ languages and optimized for handling background noise, cross-talk, and alphanumerics common in call center interactions. It was also available as Nova-2 Medical.
- Aura-1Deepgram's previous generation text-to-speech model.
- Employee AssistAn in-ear AI assistant for restaurant operations, providing hands-free, real-time support for tasks like inventory tracking, recipe management, QA compliance, and multilingual assistance.
- Voice AI OrderingAutomates drive-thru, phone, and self-ordering kiosk channels with a purpose-built AI phone assistant that handles orders in noisy, fast-paced restaurant environments, integrating with POS, CRM, and existing ordering systems.
Key capabilities
- Real-time and batch processing
- Cloud and self-hosted deployment options
- Ultra-low latency (under 300 milliseconds for STT, sub-200ms for TTS)
- High accuracy speech recognition (Nova-3, Flux)
- Natural, human-like text-to-speech (Aura-2)
- Multilingual support (45+ languages for Nova-3, 10 languages for Flux)
- Speaker Diarization (multi-speaker detection)
- Smart Formatting (punctuation, capitalization, paragraphing, dates, currency)
- Keyterm Prompting (boost accuracy for specific jargon)
- Redaction (automatically identify and remove sensitive PII)
- Filler words transcription
- Numerals conversion to digits
- Built-in turn detection and natural interruption handling (Flux, Voice Agent API)
- Unified Voice Agent API (combines STT, TTS, LLM orchestration)
- Conversational control (barge-in detection, turn-taking prediction, function calling)
- Full model ownership for optimized latency and tuning
- BYO LLM & TTS integration options
- Scalable cost optimization with flat-rate pricing and volume discounts
- Audio Intelligence (summarization, sentiment analysis, intent recognition, topic detection)
- Domain-tuned pronunciation for industry-specific terminology (TTS)
- Context-aware delivery (adjusts pacing, tone, expression for TTS)
- Enterprise-ready AI voices (40+ English voices for Aura-2)
- Noise robustness and handling of crosstalk
- HIPAA compliance (for medical models and deployments)
- GDPR and regional data residency support
- Real-time POS sync (for restaurant solutions)
- Noise filtering
- Observability and evals
- Telemetry
- Model routing
Use cases
- Building Voice AI products, platforms, and autonomous agents
- Transcription and analytics
- Real-time, human-like voice agents
- Customer support
- Order taking
- Medical transcription
- Conversational AI applications
- Speech analytics
- Media transcription
- Call tracking and analytics
- Call driver monitoring
- Agent performance coaching
- Agent assist with real-time guidance
- Quality assurance and compliance
- Sentiment analysis
- Intelligent routing of dissatisfied callers
- Building conversational chatbots
- Automating drive-thru, phone, and kiosk ordering
- Employee assistance in restaurant operations (inventory tracking, recipe management, QA, multilingual assistance)
- Documenting patient encounters
- Processing medication orders
- Powering next generation space tech
- Transcribing calls and meetings with background noise and crosstalk
- Captioning podcasts and recorded/live video streaming
- Building voicebots and AI agents
AI approach
Deepgram is a research-driven, foundational AI company focused on building voice technology that transforms human-to-machine interactions. They develop proprietary voice-native foundation models and runtime infrastructure for speech-to-text, text-to-speech, and voice agents. Their models are trained on extensive audio data (50,000+ years of audio, 1 trillion+ words) and are optimized for accuracy, low latency, and enterprise reliability. They offer various models like Nova-3 (STT), Aura-2 (TTS), Flux (Conversational STT), Voice Agent API, and Saga (Voice OS), with options for customization and deployment flexibility (cloud, self-hosted, dedicated single-tenant, in VPC).
Tech named: Deep Learning, Natural Language Processing, LLMs, task-specific language models, domain-specific language models (DSLMs), end-to-end deep learning, waveform analysis
Industries served
- Software Development
- Contact Centers
- Healthcare
- Legal
- Finance
- Media
- Restaurants
- Customer Service
- Sales
What it says sets it apart
- Industry's voice AI leader with 50,000+ years of audio processed and over 1 trillion words
- Lowest latency and highest accuracy in real-time APIs
- Enterprise reliability and scalability
- Unified Voice Agent API reduces complexity, latency, and cost by combining STT, TTS, and LLM orchestration
- Proprietary voice-native foundation models and runtime infrastructure
- Flexible deployment options: cloud, dedicated single-tenant, in-VPC, self-hosted, on-premises
- Cost-effective with transparent pricing and volume discounts
- Superior accuracy in noisy, accented, or overlapping speech
- Ultra-low latency for real-time applications (under 300ms for STT, sub-200ms for TTS)
- Go global with a single API supporting 50+ languages
- Domain-specific accuracy, especially for medical terminology (Nova-3 Medical)
- Full model ownership across the voice stack for optimized performance
- Ability to integrate own LLM or TTS while retaining Deepgram's orchestration
- Lightweight, purpose-driven, and fine-tuned Audio Intelligence models for specialized topics
- Laser-focused on building Voice AI technology for restaurants, with custom models based on menu data and real-time POS sync
- Faster transcription creation of pre-recorded audio (up to 40x faster than alternatives)
- Research-driven, foundational AI company with end-to-end deep learning approach
Funding rounds we track
AVP, Alumni Ventures, Princeville Capital, University of Michigan, Columbia University, Twilio, ServiceNow Ventures, SAP, Citi Ventures
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from Deepgram's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.