Company profile
TrueFoundry
Enterprise-grade AI Gateway for secure, scalable, and governed AI systems.
- Category
- AI infrastructure
- Headquarters
- San Francisco, California
- Sells to
- Enterprise
- Business model
- SaaS subscription, Usage-based API
- Deployment
- Cloud / SaaS, On-premise, Hybrid, Self-hosted
- Pricing
- Tiered · from $499/mo · free tier
- Builds own models
- No — builds on existing models
- Modalities
- Text, Multimodal
What TrueFoundry does
TrueFoundry provides an enterprise-grade AI Gateway that encompasses an LLM Gateway, MCP Gateway, and Agent Gateway, enabling enterprises to securely connect, observe, and govern access to models, tools, guardrails, and agents from a single control plane. Beyond the gateway layer, TrueFoundry enables organizations to deploy and train custom LLMs on GPUs, host MCP servers, and run custom agents—all through a Kubernetes-native interface. It supports on-premise and VPC installations for both AI Gateway and deployment environments. TrueFoundry ensures enterprise-grade compliance with SOC 2, HIPAA, and ITAR standards. With built-in autoscaling, caching, and resource optimization, TrueFoundry empowers organizations to build, deploy, and govern AI systems securely, efficiently, and on a future-safe stack. The company's mission is to eliminate the complexity of AI infrastructure by creating self-sustaining systems where AI manages AI, enabling businesses to focus on innovation.
Products
- AI GatewayA unified layer to manage model access, routing, guardrails, and cost controls across teams. It provides a universal API to route over 250 LLMs, with features like load balancing, fallbacks, budget limits, security, compliance, request logging, and observability. It also supports RBAC on models, virtual models, self-hosted models, and multiple gateway endpoints.
- LLM GatewayPart of the AI Gateway, specifically designed for managing and governing access to Large Language Models. It includes features like virtual models, Claude Code hooks for guardrails, unified provider caching, and budget limiting.
- MCP GatewayCentralizes all MCP servers with authentication, allowing secure connection between LLMs and tools. It includes an MCP Registry, virtual MCP servers, and supports RBAC on MCPs.
- Agent GatewayA platform to build, register, and govern AI agents. It includes an Agents Registry, support for TrueFoundry agent harness, remote agents, and a centralized, versioned, and governed Skills Registry.
- Unified AI DeploymentsA single control plane to deploy, scale, and operate LLMs, agents, MCP servers, workflows, jobs, and ML models across cloud, VPC, and on-prem environments. It supports GPU acceleration, autoscaling, auto-shutdown, and provides a consistent developer experience.
- Prompt Lifecycle ManagementVersions, manages, and monitors prompts to ensure high-quality, repeatable behavior across agents and use cases.
- GuardrailsSecures agents with guardrails, including a Guardrails Registry, Model Guardrails, and MCP Guardrails. It supports partner guardrails integration and custom guardrail hooks.
- ObservabilityMonitors and audits the entire AI stack, providing monitoring metrics, request traces, monitoring data access control, trace routing, and logging and export functionalities.
Key capabilities
- Enterprise-grade AI Gateway
- LLM Gateway
- MCP Gateway
- Agent Gateway
- Kubernetes-native interface
- On-premise and VPC installations
- Autoscaling
- Caching
- Resource optimization
- Compliance with SOC 2, HIPAA, GDPR, GLBA, FFIEC, SOX, CCPA/CPRA, VPPA, FERPA, COPPA, HITECH
- Unified API for 250+ models
- RBAC (Role-Based Access Control)
- Virtual models
- Self-hosted models
- Model Playground
- Budget limiting
- Rate limiting
- Semantic Caching
- Control Center (Weight-based Routing, Latency-based Routing, Priority-based Routing, Fallbacks)
- Logs with custom retention
- Traces with export to custom storage buckets
- Feedback on traces
- Custom Metadata
- Cost per team/user/model/application
- Metadata filtering
- Alerts
- MCP Servers (up to 200 tools per server)
- Tool calls per month
- Comprehensive metrics for MCPs
- Advanced authentication for MCPs
- Self-hosted MCPs
- Prompt Management (versioning & variables)
- Partner Guardrails integration
- Custom Guardrail Hooks
- SSO (Single Sign-On)
- Org management
- Audit logs
- Deployment modes (SaaS, VPC, On-prem, Air-gapped)
- Deployment customization
- Data Lake Export
- Connect multiple storage buckets
- Multiple gateway planes
- Gitops (Infrastructure as code)
- Community, Production, and Priority Support
- Dedicated Onboarding
- Standard and Enterprise-Grade SLA
- Deploy and serve open-source or proprietary LLMs with GPU acceleration
- Run long-running AI agents with memory and tool execution
- Deploy MCP servers to securely expose tools, APIs, and enterprise systems
- Orchestrate multi-step AI workflows
- Run batch jobs, training workloads, and scheduled AI tasks
- Deploy and serve traditional machine learning models
- Autoscaling for inference endpoints and agent services
- Auto-Shutdown to control costs
- Unified deployment experience across AWS, Azure, GCP, and on-prem
- Integrated logs, metrics, and events for every deployment
- Native monitoring and alerting
- Production-ready deployment features (health checks, rollout strategies)
- Secure secret management
- Seamless CI/CD integrations
- Framework-agnostic tracing (OpenTelemetry-compliant)
- Full Agent Observability
- Infra Observability (GPU, CPU, Cluster)
- Granular Role-Based Access Control (RBAC)
- Immutable Audit Logging
- Compliance-Ready Architecture
- Unified Monitoring and Alerting
- Real-Time Policy Enforcement
- Automated Resource Optimization
- LLM fine-tuning
- High-performance backends (vLLM, TGI, Triton)
- Support for agent frameworks (Langgraph, CrewAI, AutoGen)
- Airgapped installs can load templates and catalogue assets from local disk
- Model deployments can inherit default pod security settings
- EKS cluster discovery picks up Karpenter NodeOverlays
- Kubernetes proxy access requires manage permissions
- Spark tasks image pull credentials and ephemeral storage fixes
- Scheduled workflows service account fix
- Noma Security guardrail provider integration
- Custom base_url for Anthropic, AWS Bedrock, Bedrock Mantle, AWS Claude
- Budget Limiting V2 with independent rules, scope filters, provider account filtering, email alerts
- Secret group search matches individual secret names
- HashiCorp Vault AppRole auth supports custom role path
- Spark jobs can run on schedule
- Reranking models support in Azure OpenAI and Azure Foundry
- Virtual models can route image requests (generate, edit, variation)
- Wafer model provider integration
Use cases
- Securely connect, observe, and govern access to models, tools, guardrails, and agents
- Deploy and train custom LLMs on GPUs
- Host MCP servers
- Run custom agents
- Orchestrate Agentic AI
- Manage agent memory, tool orchestration, and action planning
- Maintain a structured, discoverable registry of tools and APIs for agents
- Version, manage, and monitor prompts
- Host any AI Model (LLM, embedding model, custom models)
- Finetune any model and deploy updated checkpoints
- Provision dedicated Model Control Protocol (MCP) servers
- Serve agents built with Langgraph, CrewAI, AutoGen, or custom orchestration
- Observe agents and underlying infrastructure (prompt execution to GPU performance)
- Govern and enforce compliance across enterprise-grade AI
- Establish granular Role-Based Access Control (RBAC)
- Record all activity for audit readiness
- Track latency, throughput, token usage, costs, and GPU utilization
- Enforce policies related to data residency, usage quotas, rate limits, and cost control
- Prototyping and testing AI workflows
- Shipping AI features for small teams
- Advanced account management and priority SLAs for teams with strict data controls
- Running AI at scale with strict compliance for medium and large organizations
- Managing model access, routing, guardrails, and cost controls across teams
- Providing a consistent interface for all model providers, policies, and telemetry
- Standardizing how teams interact with LLMs, embeddings, and RAG components
- Optimizing for cost or latency without changing applications
- Accessing and governing all models
- Exploring model features in a playground
- Using models within frameworks
- Load balancing with fallbacks for virtual models
- Cost governance with budget limiting
- Usage governance with rate limiting
- Connecting LLMs securely to MCP tools
- Registering MCP servers with RBAC
- Configuring security of MCP servers
- Tool-level RBAC across virtual MCP servers
- Securing agents with guardrails
- Enforcing guardrails for models and MCP servers
- Registering agents with RBAC
- Building agents with TrueFoundry agent harness
- Adding and using remote agents in code
- Centralizing, versioning, and governing skills
- Optimizing infrastructure costs
- Automating management of reliability and costs
- Monitoring and auditing the entire AI stack
- Inspecting traces and spans across requests
- Controlling who can view metrics and traces
- Defining where traces and metrics are stored
- Defining logging and external system exports
- Fraud Detection & Investigation Agent
- Regulatory & Compliance Monitoring Agent
- Credit & Loan Decisioning Assistant
- Customer Support & Servicing Agent
- Loan Portfolio Monitoring Agent
- Risk & Market Intelligence Agent
- Incident remediation agent
- Cloud resource provisioning agent
- Release governance agent
- Identity & access management Agent
- API key hygiene agent
- Partner deployment safety agent
- Content personalization agent
- Content moderation agent
- Churn prediction
- Ad optimization
- Lifecycle optimization agent
- Clinical documentation agent
- Patient triage and Support automation
- Sales representative copilot
- Patient scheduling automation
- Defect tracking agent
- Claims processing
- Enrollment support agent
- Attendance management agent
- Helpdesk triage agent
- Device management agent
- LMS support agent
- Library management agent
- Automating security and compliance workflows
- AI-native development
- Agent governance
- Cost management for AI workloads
- Scaling AI solutions for large user bases
- Processing large volumes of tokens and model requests
- Unifying multiple model providers
- Evaluating guardrails
- Achieving per-feature cost visibility
- Centralizing GenAI and accelerating Deep Learning deployment
- Managing multi-cloud LLMs
- Scaling agentic AI for IVR calls
- Scaling multi-model agents across modern and legacy systems
- Improving GPU cluster utilization with LLM agents
- Modernizing deployments and shortening software release lifecycles
- Deploying proprietary LLM Models
- Shipping LLM use cases in regulated environments
- Personalizing gaming with AI at massive scales
- Scaling Oral Reading Fluency (ORF) solutions
- Saving cloud costs and serving machine learning at scale
- Simplifying customer journeys for online marketplaces
AI approach
TrueFoundry provides an enterprise-grade AI Gateway that encompasses an LLM Gateway, MCP Gateway, and Agent Gateway, enabling enterprises to securely connect, observe, and govern access to models, tools, guardrails, and agents from a single control plane. It supports deploying and training custom LLMs on GPUs, hosting MCP servers, and running custom agents through a Kubernetes-native interface. The platform focuses on managing, scaling, and securing AI workloads, including LLMs, embedding models, and traditional ML models, with features like autoscaling, caching, and resource optimization. They also offer finetuning capabilities for models.
Tech named: LLM Gateway, MCP Gateway, Agent Gateway, Kubernetes, vLLM, TGI, Triton, Langgraph, CrewAI, AutoGen, OpenTelemetry, Grafana, Datadog, Prometheus, Noma Security, Anthropic, AWS Bedrock, Bedrock Mantle, AWS Claude Platform, Vertex AI, Gemini CLI, HashiCorp Vault, Spark, Azure OpenAI, Azure Foundry, Wafer, KServe
Industries served
- Software Development
- Banking & Investment Services
- Financial Services
- Technology Companies
- Media & Communication Services
- Healthcare & Life Sciences
- Education
- Gaming
- Digital Health
- Online Pharmacy Marketplaces
What it says sets it apart
- Enterprise-grade AI Gateway encompassing LLM, MCP, and Agent Gateways
- Securely connect, observe, and govern access to models, tools, guardrails, and agents from a single control plane
- Deploy and train custom LLMs on GPUs
- Kubernetes-native interface for deployment and training
- Supports on-premise, VPC, hybrid, and air-gapped installations for complete data sovereignty and isolation
- Ensures enterprise-grade compliance with SOC 2, HIPAA, GDPR, GLBA, FFIEC, SOX, CCPA/CPRA, VPPA, FERPA, COPPA, HITECH
- Built-in autoscaling, caching, and resource optimization
- Automated resource optimization without manual intervention
- Unified deployment experience across AWS, Azure, GCP, and on-prem
- Focus on developer experience with integrated logs, metrics, events, monitoring, alerting, and CI/CD integrations
- Real-time policy enforcement for data residency, usage quotas, rate limits, and cost control
- Framework-agnostic tracing and full agent observability
- Immutable audit logging for complete audit readiness
- Automated management of reliability and costs
- Ability to optimize for cost or latency without changing applications
- Faster experimentation and productionization (from 4-8 weeks to minutes)
- Faster feedback loop for ML models
- Significant cost savings (e.g., 60% cloud cost reduction, computing costs greater than service cost)
- High uptime (99.99%) and scalability (10B+ requests processed/month)
- Average 30% cost optimization through smart routing, batching, and budget controls
- Acquisition of Seldon AI to expand control plane for enterprise AI
Funding rounds we track
Intel Capital, Peak XV, Jump Capital
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from TrueFoundry's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.