AI News Today.

Artificial intelligence, professionally covered

Company profile

Inworld AI

Realtime Voice AI for human-like conversational experiences.

inworld.aiProfile compiled July 20266 source pages read
Category
NLP & speech
Headquarters
Mountain View, California
Sells to
Developers
Business model
Usage-based API, SaaS subscription
Deployment
Cloud / SaaS, API, On-premise
Pricing
Tiered subscription with usage-based credits · from $25/mo · free tier
Builds own models
Yes
Modalities
speech, Audio, Text

Inworld AI is a realtime AI research lab and model provider, offering Voice AI built for realtime conversation that feels as human as it sounds. They provide top-ranked text-to-speech and speech-to-speech, along with user-aware LLM routing and speech recognition that understands user context and emotions. Their platform is built with enterprise-grade security and compliance, including a zero-trust framework and continuous monitoring.

  • Realtime TTS-2Flagship, top-ranked model for ultra-realistic, context-aware speech synthesis with steerability, supporting over 200 languages and optimized for real-time use.
  • Realtime TTS 1.5 MaxRich, expressive speech with maximum stability, supporting 15 languages and optimized for real-time use (<200ms median latency).
  • Realtime TTS 1.5 MiniUltra-fast model for when latency is the top priority (~120ms median latency).
  • Realtime STT APIUnified integration point for industry-leading transcription providers, supporting synchronous transcription and real-time bidirectional streaming over WebSocket. Includes Inworld's first-party STT model and integrations with Groq, AssemblyAI, Soniox, and Deepgram.
  • Realtime API (Speech-to-Speech)Enables low-latency, speech-to-speech interactions with an optimized STT, LLM, and TTS pipeline, following the OpenAI Realtime protocol.
  • Realtime RouterPowerful LLM routing to optimize for every user and context, enabling a single agent to dynamically handle different user cohorts or facilitate A/B tests.
  • Realtime Text-to-Speech (TTS)
  • Realtime Speech-to-Text (STT)
  • Realtime Speech-to-Speech API
  • Realtime LLM Routing
  • Emotionally expressive voice synthesis
  • Context-aware speech synthesis
  • Zero data retention (ZDR)
  • Precise voice cloning
  • Voice design
  • Natural language steering for TTS-2
  • Multilingual support (200+ languages for TTS-2, 30 for Inworld STT, 100+ for Whisper STT)
  • Optimized for real-time use
  • Ultra-low latency (120ms median for TTS 1.5 Mini)
  • Instant voice cloning
  • Professional voice cloning (add-on)
  • Custom pronunciation
  • Speaking rate control
  • Temperature control
  • Timestamps with phonetic details and visemes
  • API access (streaming and non-streaming)
  • TTS Playground
  • WebSocket and WebRTC transports for Realtime API
  • Automatic interruption-handling and turn-taking
  • Conversational awareness (Realtime TTS-2 conditions on prior turns)
  • OpenAI compatibility for Realtime API
  • Unified STT API for multiple providers
  • Voice Profile (age, pitch, emotion, vocal style, accent) for Inworld STT
  • Configurable turn-taking for Inworld STT
  • Automatic end-of-turn detection for STT
  • Manual mode for client-controlled turn boundaries for STT
  • Enterprise-grade security and compliance
  • Zero-trust framework
  • Continuous monitoring
  • Real-time encryption and isolation
  • End-to-end encryption (AES for data in transit and at rest)
  • Microsegmentation with automatic policy enforcement
  • Multi-cloud and on-premises deployment flexibility
  • Enterprise SSO (SAML/OIDC)
  • Social login support
  • Role-based access controls
  • Automated provisioning
  • Continuous threat detection
  • High-priority alerting
  • Globally redundant infrastructure
  • Built-in compliance controls
  • Building natural and engaging experiences with human-like speech quality
  • Voice agents and character-driven apps
  • Content creation
  • Small projects
  • Growing projects and small teams
  • Production applications
  • Large deployments
  • Evaluation and prototyping
  • Synthesize speech
  • Clone and design voices
  • Generate audiobooks
  • Transcribe audio
  • Chat through the LLM Router
  • Building voice agents with WebSocket, mic input, and audio playback
  • Building voice agents with browser-native WebRTC
  • General-purpose transcription for recorded audio
  • Multilingual streaming transcription
  • English-optimized streaming transcription
  • High-accuracy, sub-300ms latency streaming transcription
  • Real-time Whisper transcription
  • English conversational voice agents with built-in turn detection, interruption handling, and barge-in awareness
  • Multilingual conversational voice agents with automatic language switching mid-conversation

Inworld AI is a realtime AI research lab and model provider specializing in voice AI. They offer proprietary text-to-speech (TTS-2, TTS 1.5 Max, TTS 1.5 Mini) and speech-to-text (Inworld STT-1) models, alongside a Realtime API for speech-to-speech interactions and a Realtime Router for LLM orchestration. They also integrate with and provide access to third-party STT models (Groq, AssemblyAI, Soniox, Deepgram) and 220+ LLM models via their Router.

Tech named: AI, Voice AI, TTS, LLMs, Runtime, AI Agent Builders, AI Orchestration, Generative AI Infrastructure, Realtime API, STT, Speech to Speech AI, Voice AI Agent, TTS-2, LiveKit agents, user-aware LLM routing, speech recognition, Realtime TTS 1.5 Max, Realtime TTS 1.5 Mini, streaming API, non-streaming API, TTS Playground, instant voice cloning, professional voice cloning, natural language steering, phonetic details, visemes, custom pronunciation, timestamps, pause controls, speaking rate control, temperature control, multilingual support, Realtime API (Speech-to-Speech), OpenAI Realtime protocol, WebSocket, WebRTC, automatic interruption-handling, turn-taking, conversational awareness, Realtime Router, Realtime Speech-to-Text (STT) API, Inworld (first-party) STT-1, Voice Profile (age, pitch, emotion, vocal style, accent), configurable turn-taking, Groq/Whisper-large-v3, AssemblyAI/universal-streaming-multilingual, AssemblyAI/universal-streaming-english, AssemblyAI/u3-rt-pro, AssemblyAI/whisper-rt, Soniox/stt-rt-v4, Soniox/stt-rt-v5, Deepgram/flux-general-en, Deepgram/flux-general-multi, automatic language switching

  • #1 ranked text-to-speech and speech-to-speech
  • User-aware LLM routing
  • Speech recognition that understands user context and emotions
  • Ultra-realistic, context-aware speech synthesis
  • Zero data retention (ZDR) policy
  • Precise voice cloning capabilities
  • Accessible price point
  • Top-ranked flagship TTS-2 model with steerability
  • Ultra-low latency options
  • Unified STT API for multiple industry-leading transcription providers
  • OpenAI Realtime API compatibility and migration path
  • Enterprise-grade security and compliance built into the AI platform
  • SOC 2 Type II certified
  • HIPAA & GDPR compliant
  • Multi-cloud and on-premises deployment flexibility
  • Dedicated account management and Slack channel for Enterprise clients
$50MSeries A2022-08-24thesaasnews.com

Section 32, Intel Capital, Founders Fund, Accelerator Investments LLC, First Spark Ventures, Kleiner Perkins, BITKRAFT Ventures, CRV, Microsoft’s M12 fund, Micron Ventures, LG Technology Ventures, SK Telecom Venture Capital, NTT Docomo Ventures, The Venture Reality Fund, M12

$7MSeed2021-11-10citybiz.co

Kleiner Perkins, CRV, Meta

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from Inworld AI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.