Company profile
Modular
Unified AI inference platform for high-performance, portable compute from kernel to cloud.
- Category
- AI infrastructure
- Headquarters
- Silicon Valley
- Sells to
- Developers
- Business model
- Freemium, Usage-based API, Licensing
- Deployment
- Cloud / SaaS, On-premise, Hybrid, Edge, API, Self-hosted
- Pricing
- Per token for shared endpoints, per minute for dedicated endpoints, per minute for Your Cloud deployment. Free self-hosted container. · free tier
- Builds own models
- No — builds on existing models
- Modalities
- Text, Image, Video, Audio, Code
What Modular does
Modular is building AI's unified compute layer, offering a platform that simplifies AI development and deployment. It provides modular and composable infrastructure, open-sourcing its language and engine. The Modular Platform unifies AI under a single framework, offering text, audio, and image inference with state-of-the-art performance. It supports deployment with shared endpoints, dedicated endpoints, in Modular's cloud or the customer's VPC, and with custom models. The platform is designed to address the high costs, fragmentation, and hardware lock-in prevalent in the AI industry, making AI's compute layer efficient and accessible to all.
Products
- The Modular PlatformA unified AI inference platform for high-performance, portable compute, enabling full optimizations from GPU kernel to API endpoint. It offers a unified stack from kernels to cloud, built from the ground up for heterogeneous compute, scaling AI from cloud to edge across CPUs, GPUs, and ASICs.
- MAX (Modular Accelerated eXecution)A high-performance, hardware-agnostic serving framework that automatically optimizes kernels and request execution across accelerators. It provides 2x performance improvement over vLLM on diverse hardware through a single container and OpenAI-compatible API. MAX also includes PyTorch-like model APIs and AI coding skills for easy custom model porting, and supports 1000+ models out of the box.
- MojoA high-performance systems language used to write 100s of SOTA, composable GPU kernels. It allows extending or writing custom GPU kernels for maximum performance across accelerators and simplifies GPU programming with modular kernel architecture, compile-time abstractions, and zero-cost performance across modern GPU hardware. Mojo 1.0 is officially in beta.
- Modular CloudA hosted cloud environment for running AI workloads at production scale with SOTA performance on NVIDIA and AMD GPUs. It offers full-stack workload customization, performance tuning, and deep observability. It also provides shared and dedicated endpoints with per-token or per-minute pricing.
- Self-Hosted (MAX and Mojo in a container)A deployment option allowing users to run MAX and Mojo in a container on their own infrastructure, supporting NVIDIA, AMD, and Apple Silicon. It's free for all developers and provides SOTA inference performance on any supported GPU vendor.
- Your CloudA deployment option built on Modular's production-hardened BYOC infrastructure, where inference runs in the customer's VPC. Modular manages the control plane, while the customer owns hardware, data, and cloud credits. It includes everything in Dedicated Endpoint plus deployment in your cloud or on-premise, data never leaving your VPC, performance optimization of specific pipelines, and custom APIs.
Key capabilities
- Unified AI inference platform
- High-performance, portable compute
- Full optimizations from GPU kernel to API endpoint
- 2x Performance on a unified stack
- AI on any GPU (NVIDIA, AMD, Intel, ARM, Apple Silicon)
- 50% cost saving
- Unified stack from kernels to cloud
- Scales AI from cloud to edge (CPUs, GPUs, ASICs)
- High-performance, hardware-agnostic serving framework (MAX)
- Automatically optimizes kernels and request execution
- OpenAI-compatible API
- Supports 1000+ models out of the box (e.g., DeepSeek, Kimi)
- PyTorch-like model APIs
- AI coding skills for custom model porting
- 100s of SOTA, composable kernels written in Mojo
- Extend or write custom GPU kernels
- Natively heterogeneous hardware compatibility
- Flexible deployment: Modular Cloud, Your Cloud, Self Hosted
- Shared endpoints (per-token pricing, no infrastructure to manage)
- Dedicated endpoints (reserved GPUs, per-minute pricing)
- Custom model deployment
- Compiler-native speculative decoding
- GPU vendor flexibility at scale
- 90% smaller serving footprint (<700MB runtime)
- Custom attention and decoding strategies in Mojo
- Full-stack programmability in Mojo
- Compiler-optimized inference for image generation (up to 4x PyTorch performance)
- Compiler-fused diffusion pipelines
- Hardware-portable diffusion
- Ultra-low latency voice serving
- Compiler-aware auto-scaling for voice
- Custom voice model deployment with Mojo kernels
- Dual GPU vendor support for voice models
- On-device voice with Apple Silicon
- MLIR compiler fuses entire inference path
- Prefix caching for improved latency
- Intelligent batching and memory management
- Inflight batching without requiring chunked prefill
- Paged KVCache strategy for efficient memory usage
- SOC 2 Type 2 certified
Use cases
- Running AI across GPUs and CPUs for demanding inference workloads
- Generating images and text (and videos soon)
- Deploying top open models or custom models with flexible deployment
- Testing and prototyping AI models with shared endpoints
- Deploying custom or fine-tuned models on optimized infrastructure
- Running AI models and pipelines on any supported hardware
- Writing custom kernels for novel architectures
- Optimizing AI pipelines across text, image, and video
- Real-time conversational voice with sub-100ms latency
- Character voice synthesis at scale for interactive gaming and virtual worlds
- Deploying proprietary TTS, voice cloning, or speech-to-speech models
- Privacy-sensitive applications, offline use cases, and edge deployments for TTS
- Generating marketing visuals, product mockups, and social content at sub-second speeds
- Avatar and character generation for gaming, virtual worlds, and social platforms
- Inline code completion for AI-powered code editors
- Code chat, Q&A, and code review with long-context models
- Agentic coding workflows for multi-step code generation, testing, and iteration
- On-device developer tools for local code completion
- Deploying GenAI models from Hugging Face with an OpenAI-compatible endpoint
- Customizing models and tuning GPU kernels
- Building LLMs from scratch
- Real-time patient conversations for AI health agents
- Generating production-quality images with compiler-optimized inference
- Text-to-video and image-to-video generation
AI approach
Modular provides a unified AI inference platform for high-performance, portable compute, optimizing AI pipelines from GPU kernel to API endpoint. They offer a proprietary systems language, Mojo, for writing custom GPU kernels and an inference serving framework, MAX, that optimizes kernels and request execution across various accelerators. The platform supports a wide range of open-source models and allows for custom model deployment, focusing on hardware portability across NVIDIA, AMD, Intel, ARM, and Apple Silicon.
Tech named: Mojo, MAX (Modular Accelerated eXecution), MLIR compiler, GPU kernels, PyTorch-like model APIs, OpenAI-compatible API, vLLM, TensorRT, SGLang, Continuous batching, KV-cache optimization, Streaming-aware scheduler, Speculative decoding, Prefix caching, Intelligent batching, Memory management, Inflight batching, Paged KVCache strategy, MoE serving, Agentic programming tools, Neural audio codec, LLM backbone, Speech-Language Model (SpeechLM), UNet/DiT, VAE, Text encoder, Scheduler, LoRA fine-tuning, CUDA, ROCm
Industries served
- Gaming
- Healthcare
- Software Development
- Marketing
- Design
- Media and Entertainment
What it says sets it apart
- One unified stack from kernels to cloud for heterogeneous compute
- True hardware portability across NVIDIA, AMD, Intel, ARM, and Apple Silicon
- Significant cost savings through higher GPU utilization and faster model compilation/runtime
- MAX serving framework offers 2x performance improvement over vLLM
- Mojo language for high-performance, composable GPU kernels
- Flexible deployment options: Modular Cloud, Your Cloud, Self Hosted
- Compiler-optimized performance, not just wrapper-optimized
- 90% smaller runtime for faster scaling and lower storage costs
- Compiler-native speculative decoding for code generation
- Ultra-low latency and compiler-aware auto-scaling for voice synthesis
- Ability to run proprietary models with custom Mojo kernels
- SOC 2 Type 2 certified for security and compliance
- Forward-deployed engineers for workload optimization and custom kernel development
Funding rounds we track
Thomas Tull’s US Innovative Technology fund, DFJ Growth
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from Modular's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.