AI News Today.

Artificial intelligence, professionally covered

Company profile

Together AI

Full-stack AI platform for inference, model shaping, and pre-training.

together.aiProfile compiled July 202613 source pages read
Category
AI infrastructure
Headquarters
San Francisco, California
Sells to
Developers
Business model
Usage-based API
Deployment
Cloud / SaaS, API
Pricing
Usage-based (per 1M tokens, per image, per 1M characters, per video, per audio minute, per GPU per hour)
Builds own models
Yes
Modalities
Text, Image, Video, Audio, speech, Code

Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. It serves as a full-stack AI platform, offering a high-performance inference engine for reliable and fast scaling, on-demand GPU clusters, and massive-scale AI factories. The company continuously advances the field by productizing cutting-edge research from its world-leading AI systems research team, combining research velocity with production-grade infrastructure to enable companies to reliably scale AI-native applications.

  • Serverless InferenceThe fastest way to run open-source models on demand, powered by cutting-edge inference research. No infrastructure to manage, no long-term commitments.
  • Batch InferenceCost-effectively process massive workloads asynchronously. Scales to 30 billion tokens per model with any serverless model or private deployment.
  • Provisioned ThroughputCommitted inference capacity with token-based pricing, reserved throughput, and a 99% uptime SLA. Offers drop-in API compatibility for production workloads with no infrastructure to manage.
  • Dedicated Model InferenceDeploys models on dedicated infrastructure, purpose-built for teams needing speed, control, and optimal economics.
  • Dedicated Container InferenceGPU infrastructure purpose-built for generative media workloads, deploying video, audio, and image models with performance acceleration powered by Together Research.
  • Accelerated ComputeScales from self-serve instant clusters to thousands of GPUs, optimized for better performance with Together Kernel Collection.
  • SandboxProvides fast, secure code sandboxes at scale to set up full-scale development environments for AI apps and agents.
  • Managed StorageHigh-performance managed storage for AI-native workloads, including object storage and parallel filesystems optimized for AI, with zero egress fees.
  • Fine-TuningFine-tunes open-source models for production workloads using the latest research techniques to improve accuracy, reduce hallucinations, and control behavior without managing training infrastructure.
  • Voice SolutionA complete voice stack for real-time production use, enabling deployment of real-time voice agents with ultra-low latency and production-scale reliability. Integrates various STT, LLM, and TTS models through a single API.
  • Faster inference powered by cutting-edge research
  • Lower cost with workload-specific optimization
  • Faster pre-training with Together Kernel Collection
  • Full-stack cloud for the entire AI development journey
  • Serverless, Batch, Provisioned Throughput, Dedicated Model, and Dedicated Container Inference options
  • Accelerated Compute with GPU clusters
  • Code Sandboxes for AI app and agent development
  • Managed Storage with zero egress fees
  • Fine-tuning for open-source models
  • Real-time voice agent deployment with ultra-low latency
  • Access to a comprehensive voice model library (open-source and proprietary)
  • Scalable infrastructure with dynamic autoscaling across global regions
  • Support for custom models and containers
  • Guaranteed performance on single-tenant GPU instances
  • Autoscaling and traffic spike handling
  • OpenAI-compatible API for running AI models
  • Launch H100 and B200 GPU clusters with attached storage
  • Running open-source AI models on demand
  • Processing massive AI workloads asynchronously
  • Deploying production workloads with committed inference capacity
  • Deploying models on dedicated infrastructure for speed and control
  • Deploying video, audio, and image models for generative media workloads
  • Scaling GPU clusters for AI workloads
  • Setting up development environments for AI apps and agents
  • Storing high-performance data for AI-native workloads
  • Fine-tuning open-source models for production
  • Building and deploying real-time voice agents
  • Developing AI-powered coding platforms with real-time in-editor agents
  • Training and deploying frontier reasoning models
  • Running browser-use AI agents at production scale
  • Building customer-specific EOB parsers
  • Training APT-1 models
  • Scaling generative video and image APIs
  • Building clinical AI
  • Achieving AI independence with reliable inference
  • Accelerating mental health AI
  • Prototyping AI platforms
  • Building world-class Thai language models
  • Building AI customer support bots
  • Scaling AI companions

Together AI provides a full-stack AI platform for engineers and researchers, focusing on inference, model shaping, and pre-training. They offer serverless inference, batch inference, provisioned throughput, and dedicated inference options for various open-source models. The platform also includes accelerated compute (GPU clusters), sandboxes for development, and managed storage. A key offering is fine-tuning open-source models. Together AI emphasizes cutting-edge research, including kernel optimization, and supports both open-source and proprietary models across modalities like text, image, video, and audio. They also provide infrastructure for real-time voice agents, integrating various STT, LLM, and TTS models.

Tech named: Together Kernel Collection, ParallelKernelBench, Violin, Parcae, EinsteinArena, AI for Systems, Aurora, DeepSeek V4 Pro, MiniMax M3, Kimi K2.7 Code, GLM-5.2, PrismML Ternary Bonsai 27B, Inkling, Gemma 4 31B, NVIDIA Nemotron 3 Ultra, Qwen3.7-Plus, Kimi K2.6, Qwen3.7-Max, gpt-oss-120B, Qwen3.5-397B-A17B, Qwen3.5 9B, Gemma-4-31B-it-Pearl, Cogito v2.1 671B, Rnj-1 Instruct, Llama 3.3 70B, Gemma 3n E4B Instruct, gpt-oss-20B, GLM-5.1, MiniMax M2.7, Qwen3.6-Plus, Qwen2.5 7B Instruct Turbo, Llama 3 8B Instruct Lite, Qwen3 235B A22B Instruct 2507 FP8 Throughput, GPT Image 2, Wan 2.6 Image, Nano Banana Pro (Gemini 3 Pro Image), FLUX.2 [pro], Ideogram 4.0, Gemini 3.1 Flash Image (Nano Banana 2), Qwen Image 2.0 Pro, Qwen Image 2.0, FLUX.2 [dev], FLUX.2 [flex], FLUX.2 [max], FLUX.1 Kontext [pro], FLUX1.1 [pro], Juggernaut Pro Flux, GPT Image 1.5, FLUX.1 Kontext [max], FLUX.1 [schnell], SD XL, Ideogram 3.0, HiDream-I1-Full, Juggernaut Lightning Flux, Qwen Image, Google Imagen 4.0 Fast, ByteDance Seedream 4.0, Google Imagen 4.0 Preview, Gemini Flash Image 2.5 (Nano Banana), Google Imagen 4.0 Ultra, ByteDance Seedream 3.0, NVIDIA Parakeet TDT 0.6B V3 Realtime, NVIDIA Nemotron 3 ASR Streaming 0.6B, Cartesia Sonic-3, Orpheus TTS, Kokoro-82M TTS, Cartesia Sonic-2, ByteDance Seedance 2.0, Google Veo 3.0, Kling 1.6 Standard, Kling 2.1 Master, Kling 2.1 Pro, Kling 2.1 Standard, Vidu 2.0, Vidu Q1, Wan 2.2 I2V, Wan 2.2 T2V, Sora 2, PixVerse v5, ByteDance Seedance 1.0 Lite, ByteDance Seedance 1.0 Pro, Google Veo 3.0 Fast + Audio, Google Veo 3.0 Fast, Google Veo 3.0 + Audio, Google Veo 2.0, MiniMax Hailuo 02, MiniMax 01 Director, NVIDIA Parakeet TDT 0.6B v3, NVIDIA Nemotron 3.5 ASR, Whisper Large v3, Whisper Large v3 (Streaming), Multilingual e5 large instruct, Llama Guard 4 12B, NVIDIA HGX H100, NVIDIA HGX H200, NVIDIA HGX B200, NVIDIA HGX B300, GB200 NVL72, GB300 NVL72, MiniMax, Rime, Deepgram, OpenAI, Cartesia, Mist v3 Omni, Deepgram Flux, Deepgram Aura-2, Deepgram Nova-3, Deepgram Nova-3 Multilingual, Arcana V3 Turbo, Blackwell architecture, ARM host optimization, NVFP4 quantization pipeline, TensorRT-LLM, NVIDIA Grace CPU, Celery queues, Slurm clusters

  • Software Development
  • AI-native developer tools
  • Enterprise Software
  • Video Generation & AI Content Creation
  • Healthcare
  • Media
  • Customer Service
  • AI Native Cloud purpose-built for AI engineers and researchers
  • Full-stack AI platform covering inference, model shaping, and pre-training
  • Powered by cutting-edge research from a world-leading AI systems research team
  • Combines research velocity with production-grade infrastructure
  • Offers significant cost savings (e.g., 60% lower cost, 60% total cost savings for Hedra)
  • Achieves faster inference (e.g., 3x faster inference on Blackwell, 95% faster TTFT)
  • Provides faster pre-training (e.g., 90% faster)
  • Ultra-low latency for voice agents (sub-second STT-to-TTS latency, under 500ms end-to-end)
  • Reliable production partner with proactive risk management and fast issue resolution
  • Optimized performance at scale with lower cost than closed-source providers
  • Strict privacy standards
  • Expert support and ability to scale to meet rapid growth
  • Innovation partner enabling creative boundaries without compromise
  • Infrastructure capacity to handle viral moments and traffic surges
  • Joint engineering collaboration for kernel optimization
  • Early access and deployment on frontier infrastructure like NVIDIA Blackwell (GB200 NVL72)
  • ARM host optimization and custom kernels for Blackwell Tensor Core
  • Efficient parallelism across NVIDIA GB200 NVL72
  • Shortens weights-to-production cycle
  • Supports rapid iteration with quick deployment of new model variants

Aramco Ventures, Sequoia Capital, Coatue Management, Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, March Capital, Pegatron, SE Ventures, S Ventures, Vista Equity, Salesforce Ventures, DTCP Growth, Lux Capital, Geodesic, PSP Partners, SentinelOne’s S Ventures, Emergence, SentinelOne's S Ventures, SentinelOne S Ventures, Schneider Electric, Khosla Ventures, Tiger Global Management

General Catalyst, Prosperity7

General Catalyst, Prosperity7, Salesforce Ventures, DAMAC Capital, Nvidia, Kleiner Perkins, March Capital, Emergence Capital, Lux Capital, SE Ventures, Greycroft, Coatue, Definition, Cadenza Ventures, Long Journey Ventures, Brave Capital, SK Telecom

$102.5MSeries A2023-12-11maginative.com

Kleiner Perkins, NVIDIA, Emergence Capital, New Enterprise Associates, Prosperity 7, Greycroft, 137 Ventures

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from Together AI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.