Company profile
Modal
High-performance AI infrastructure for developers to run inference, training, and batch processing.
- Category
- AI infrastructure
- Headquarters
- New York City, New York
- Sells to
- Developers
- Business model
- Usage-based API, Freemium
- Deployment
- Cloud / SaaS, API
- Pricing
- Usage-based API · from $250/mo · free tier
- Builds own models
- No — builds on existing models
- Modalities
- Text, Image, Video, Audio, speech, Code, Multimodal
What Modal does
Modal provides high-performance AI infrastructure that developers love, enabling them to run inference, training, batch processing, and sandboxes with sub-second cold starts, instant autoscaling, and a developer experience that feels local. It is described as the production cloud for AI, offering an AI-native runtime built for speed at any scale, elastic cloud capacity to autoscale from 0 to 1000+ GPUs instantly, and production-ready features with out-of-the-box observability. Modal's infrastructure is engineered from the ground up for heavy AI workloads, optimizing for inference behavior and supporting the full training loop from fine-tuning to multi-node runs. It also provides an execution layer for AI agents through isolated, flexible, and scalable sandboxes. The platform is designed to make it easier to iterate and ship applications for data, AI, and machine learning, by building its own infrastructure including a custom file system, container runtime, scheduler, and container image builder.
Products
- Modal SDKYour cloud environment, in code. Stay in Python, ship to the cloud. Composable primitives that specify everything from logic to hardware in one place.
- AI-native runtimeEngineered from the ground up for heavy AI workloads, with super-fast autoscaling and containers that boot instantly.
- Elastic cloud capacityAutoscale from 0 to 1000+ GPUs, instantly. Modal routes workloads across clouds and regions in real time. Get the GPUs you need in seconds, with no commitments or capacity planning.
- InferenceEngineered for inference, from the proxy layer to the GPU scheduler, every part of Modal's stack is optimized for how inference workloads actually behave. Supports LLM inference, multi-modal inference, batch and async inference, and online inference.
- TrainingBuilt for the full training loop, from single-GPU fine-tuning to parallel hyperparameter sweeps to multi-node runs. Supports fine-tuning, reinforcement learning, multi-node training, and parallel hyperparameter sweeps.
- SandboxesDesigned to scale agents, providing isolated containers for running AI code and RL rollouts with sub-second scheduling and support for 100k+ concurrent sandboxes. Used for coding agents, background agents, and RL rollouts.
- BatchProduct for batch processing.
- NotebooksHigh-performance GPU notebooks combining an intuitive interface with near-instant cold starts, offering flexibility for lightweight experiments to large-scale multi-GPU training jobs.
- Core PlatformCloud infrastructure designed for AI workloads, with memory snapshotting, smarter filesystem, deep GPU capacity pool, efficient batching and scheduling, and observability features.
- VolumesDistributed storage for persisting data and sharing artifacts across sandbox runs.
- BucketsStorage primitive.
- QueuesData structure primitive for coordinating workloads.
- DictsData structure primitive for coordinating workloads and storing lightweight state.
- TunnelsNetworking primitive to expose ports from running sandboxes.
- ProxiesNetworking primitive.
- GLM 5.2 FP8Large model good for coding, agents, and tool use, deployable via Modal endpoint.
Key capabilities
- Sub-second cold starts
- Instant autoscaling
- Developer experience that feels local
- Composable primitives
- AI-native runtime
- Elastic cloud capacity
- Autoscale from 0 to 1000+ GPUs
- Real-time workload routing across clouds and regions
- Out-of-the-box observability
- Integrated logging
- Full visibility into every function, sandbox, and container
- Optimized for inference workloads
- Support for LLM Inference on various GPUs (H100s, A100s, A10Gs)
- Scale to zero between requests, burst to handle demand
- Multi-modal Inference (image generation, video, audio, embeddings)
- Batch and Async Inference (evals, embeddings, re-ranking, dataset generation)
- Online inference with sub-10ms overhead latency
- Globally distributed compute
- Support for token streaming, WebRTC, WebSocket
- Full training loop support (fine-tuning, RL, multi-node, hyperparameter sweeps)
- SFT, LoRA, full fine-tunes on B200s, H100s, A100s
- Thousands of concurrent RL trajectories
- Access to up to 128 B200s with 3200 Gbps Infiniband networking
- Gang-scheduled multi-node training with single line of code
- Launch hundreds of experiments simultaneously
- Isolated and flexible sandboxes for AI agents
- Programmatic spin-up of fresh, isolated sandboxes
- Custom images and dependencies in sandboxes
- Autonomous agents with tools, context, and credentials in sandboxes
- Spin up hundreds of thousands of concurrent RL rollout environments
- GPU-accelerated research with H100s, A100s, A10Gs on demand
- Pay-by-the-second pricing with no reserved capacity
- Memory snapshotting for fast model loading
- Optimized filesystem for fast startup
- Deep GPU capacity pool across multiple clouds
- Near-max GPU utilization (2-3x higher throughput)
- Real-time visibility with rich dashboard
- Granular metrics and insights for debugging
- First-party integrations with telemetry providers
- Seamless integration of data pre-processing, training, and serving
- Instant cold-starts for notebooks (under 5 seconds)
- Fast iteration in notebooks
- Swap GPUs on the fly in notebooks (up to 8 H100s or B200s)
- Petabyte-scale storage with Modal Volumes
- Real-time collaboration in notebooks
- Automatic idle shutdown for notebooks
- Modern AI code editing (Pyright type checking, rich outputs, AI-powered suggestions)
- Customer-supplied encryption keys
- Audit logs
- SOC 2 compliance
- HIPAA compatibility
- Okta SSO
- Custom SAML SSO
- Role-Based Access Control (RBAC)
- HTTPS for secure connections (TLS 1.3)
- All user data encrypted in transit and at rest
- Memory-safe programming languages (Rust, Python)
- Automated synthetic monitoring for network and application isolation
Use cases
- LLM Inference
- LLM Fine-Tuning
- Generative Model Inference
- Generative Model Training
- Computational Biology
- Audio Generation
- Image Generation
- Video Generation
- Web Scraping
- Batch Jobs
- Batch Embeddings
- Scaling Out
- AI Agents
- Reinforcement Learning
- Sandboxes
- Background Agents
- Multi-modal Inference
- Batch and Async Inference
- Online inference
- Single-GPU fine-tuning
- Parallel hyperparameter sweeps
- Multi-node training
- Coding agents
- RL rollouts
- GPU-accelerated research
- Language Models
- Image, Video, 3D processing
- Audio Processing
- Sandboxed Code
- ML-driven molecular design
- AI app generation
- Document processing
- Podcast transcription
- Deploying OpenAI-compatible LLM services
- Optimizing tokens per second in batch LLM processing
- Deploying OpenCode agents
- Designing protein binders with ESMFold2
- Transcribing speech in batches with Whisper
- Voice chat with LLMs
- Building AI coding platforms
- Custom pet art from Flux with Hugging Face and Gradio
- Deploying really big language models
- Editing images with Flux Kontext
- Folding proteins with Boltz-2
- Serverless WebRTC for YOLO detections
- Sandboxing LangGraph agents' code
- Serving diffusion models
- Low latency SGLang for interactive language models
- Streaming transcripts with Kyutai STT
- Creating custom music videos
- Making music with ACE-Step
- RAG Chat with PDFs
- Animating images with generative video models
- Building protein folding dashboards
- Deploying Hacker News Slackbots
- Retrieval-Augmented Generation (RAG) for Q&A
- Document OCR job queues
- Parallel processing of Parquet files on S3
- Deploying a text-to-speech (TTS) API with Chatterbox Turbo
- Running vLLM server in OpenAI-compatible mode
- Training models from scratch
- Hosting popular libraries (YOLO, Blender, Streamlit, SQLite, Algolia)
- Connecting to other APIs (Discord, Google Sheets, OpenAI, Tailscale, Prometheus)
- Managing data (S3 buckets, DuckDB, DBT, LoRA Playground)
- Running untrusted code in Functions
- Running commands in Sandboxes
- Networking and security in Sandboxes
- File access in Sandboxes
- Snapshots in Sandboxes
- VM Sandboxes
- Secrets and environment variables management
- Scheduling and cron jobs
- HTTP Applications (Servers, Web Functions, Streaming endpoints)
- Modal Auto Endpoints
- Tunnels and Proxies for networking
- Cluster networking
- Passing local data
- Storing model weights
- Cloud bucket mounts
- Dataset ingestion
- Optimizing cold start performance
- Memory Snapshots for performance
- High-performance LLM inference
- Handling failures and retries
- Preemption
- Timeouts
- GPU health monitoring
- Connecting Modal to Datadog
- Connecting Modal to OpenTelemetry
- Slack notifications
- Workspace and account settings
- Service users
- Billing management
- Developing and debugging with Jupyter notebooks
- Asynchronous API usage
- Global variables
- Region selection
- Container lifecycle hooks
- Parametrized functions
- Dynamic function configuration
- S3 Gateway endpoints
- GPU Metrics
AI approach
Modal provides high-performance AI infrastructure for developers, offering serverless GPUs, sub-second container starts, and native storage. It's designed for low-latency inference, model fine-tuning, and production-ready sandboxes at scale. The platform supports various AI workloads including LLM inference, multi-modal inference, batch and async inference, fine-tuning, reinforcement learning, multi-node training, and hyperparameter sweeps. It also provides sandboxes for coding agents, background agents, and RL rollouts. Modal builds its own custom infrastructure including a file system, container runtime, scheduler, and container image builder, optimized for AI workloads.
Tech named: Python, H100s, A100s, A10Gs, B200s, Infiniband, PyTorch, Axolotl, Unsloth, Hugging Face TRL, Weights and Biases, TensorBoard, Qwen, Flux, Whisper, LoRA, vLLM, SGLang, ESMFold2, Chai-1, Boltz-2, YOLO, Blender, Streamlit, SQLite, Datasette, Algolia, OpenAI API, Tailscale, Prometheus, Pushgateway, S3, DuckDB, DBT, Gradio, Chatterbox Turbo, FastAPI, peft, debian_slim, CUDA, LangGraph, Node.js, Ruby, PHP, GRPO, verl, TRL, Liquid AI embeddings, TEI, MongoDB, Parquet, ESM3, Molstar, ColBERT, Vision-Language Model, Flux Kontext, LTX-Video, Stable Diffusion, Moshi, Kyutai STT, ACE-Step, FastRTC, OpenCode, Claude Agent SDK, Ministral 3, Nemotron 3, FastMCP, Gemma, UMAP, dots.ocr
What it says sets it apart
- Instant GPU access
- Sub-second container starts
- Native storage
- Simple to serve low-latency inference
- Easy model fine-tuning
- Access to production-ready sandboxes at scale
- Rebuilt infrastructure layer for AI workloads
- Developer experience that feels local
- Composable primitives for logic and hardware specification
- AI-native runtime engineered for heavy AI workloads
- Super-fast autoscaling
- Containers that boot instantly
- Autoscale from 0 to 1000+ GPUs instantly
- Routes workloads across clouds and regions in real time
- No commitments or capacity planning for GPUs
- Out-of-the-box observability with integrated logging and full visibility
- Optimized stack for inference workloads
- Supports any model or inference engine on various GPUs
- Scales to zero between requests, bursts to handle demand
- Handles full training loop in a single code file
- Thousands of concurrent RL trajectories in sandboxes
- Access to high-end GPUs (B200s, H100s, A100s) with Infiniband networking
- Sandboxes are native to the same stack as training infrastructure
- Isolated, flexible, and scalable sandboxes for AI systems
- Programmatic spin-up of fresh, isolated sandboxes with custom images
- Autonomous agents running securely in full, isolated dev environments
- Fast enough RL rollouts to saturate GPU inference resources
- Pay-by-the-second pricing with no reserved capacity for GPUs
- Memory snapshotting for loading large models in seconds
- Optimized filesystem for fastest startup
- Deep GPU capacity pool across multiple clouds without quotas or reservations
- Efficient batching and scheduling for near-max GPU utilization (2-3x higher throughput)
- Robust primitives to connect services, persist data, and coordinate workloads
- Real-time visibility and granular metrics for debugging
- Seamless integration of data pre-processing, training, and serving
- Intuitive interface for notebooks with near-instant cold starts
- Fast iteration and real-time results in notebooks
- Ability to swap GPUs on the fly in notebooks
- Petabyte-scale storage with global distributed file system (Volumes)
- Real-time collaboration in notebooks (multiple cursors, live edits)
- Automatic idle shutdown for notebooks to save costs
- Modern AI code editing features in notebooks
- Strong security and privacy commitments (SOC 2, HIPAA, encryption, memory-safe languages)
- Proprietary custom file system, container runtime, scheduler, and container image builder
- Usage-based pricing, only paying for actual compute time
- Serverless architecture for cost-effectiveness with spiky workloads
This profile was compiled from Modal's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.