Company profile
Patronus AI
Frontier lab building Digital World Models for AI agent training and simulation.
- Category
- AI infrastructure
- Headquarters
- San Francisco, California
- Sells to
- Developers
- Business model
- SaaS subscription, Usage-based API, Services & consulting, Freemium
- Deployment
- Cloud / SaaS, API, On-premise
- Pricing
- Tiered per user and usage · from $25/mo · free tier
- Builds own models
- Yes
- Modalities
- Text, Image, Multimodal
What Patronus AI does
Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. They are training the First Digital World Model for AI agent training and simulation, with a mission to simulate all of the world’s intelligence. They provide an end-to-end system to evaluate, monitor, and improve the performance of LLM systems, enabling developers to ship AI products safely and confidently. Their work includes developing infrastructure to power agentic products and conducting research in AI evaluation.
Products
- First Digital World ModelA foundational infrastructure for self-adaptive worlds, enabling continual learning for AI agents by predicting and simulating agent actions in digital workflows.
- PlatformA core evaluation platform providing teams with a centralized solution for experiments, logging, comparisons, and traces.
- LLM-as-a-JudgeEnables developers to score multimodal AI systems for image to text.
- GliderA powerful 3B evaluator LLM that can score any text input on user-defined criteria, designed for explainable evaluation and fine-grained rubric-based scoring. Supports multilingual reasoning and span highlighting.
- LynxA SOTA hallucination detection LLM capable of advanced reasoning, beating GPT-4 on hallucination tasks. Available in 8B and 70B versions.
- PercivalAn evaluation copilot for agentic systems built to detect 20+ failure modes in agentic traces, suggesting optimizations, and evaluating reasoning and planning errors.
- Percival Chat AssistantAn interactive AI agent that lets users unlock the power of Percival.
- Generative SimulatorsAdaptive environments that co-generate tasks, world dynamics, and reward functions.
- MemTrackA benchmark to evaluate long-term memory and state tracking in multi-platform agent environments.
- Patronus ExperimentsMeasures and automatically optimizes AI product performance against evaluation datasets.
- Patronus DatasetsOff-the-shelf, adversarial testing sets designed to break models on specific use cases.
- FinanceBenchAn industry-first benchmark for LLM performance on financial questions, with 10,000+ high-quality Q&A pairs based on publicly available financial documents.
- SimpleSafetyTestsA diagnostic test suite to identify critical safety risks in LLMs across 5 areas: suicide, child abuse, physical harm, illegal items, and scams & fraud.
- EnterprisePIIThe industry’s first LLM dataset for detecting business-sensitive information, containing 3,000 examples of annotated text excerpts from common enterprise text types.
- Patronus LogsContinuously captures evaluations, auto-generated natural language explanations, and proactively highlighted failures in production.
- Patronus ComparisonsCompares, visualizes, and benchmarks LLMs, RAG systems, and agents side by side across experiments.
- Patronus TracesAutomatically detects agent failures across 15 error modes, allows chat with traces, and auto-generates trace summaries.
- MLLM-as-a-Judge (Judge-Image)An LLM judge that supports image input to text output use cases, with a Google Gemini backbone, used to score and optimize multimodal AI systems and detect caption hallucinations.
Key capabilities
- Research-backed, real-world inspired simulations
- First Digital World Model
- Interactive Digital Worlds generated dynamically for AI agents
- 30–40% model lift on long-horizon tasks
- 1M+ world data artifacts covering diverse domains
- 85% UI/UX feature parity with real-world products
- 5k+ expert contributors across software, academia, finance
- Designing scenarios that target core model skills
- Deep research understanding and reasoning over large semantic datasets
- Multi-turn dialogue for collaborative problem solving
- Long horizon task planning and execution
- Agentic memory with context windows and other tooling
- Experimentation Framework for A/B testing and optimizing LLM system performance
- Real Time Monitoring with tracing, logging, and alerts for LLM and agent interactions
- Visualizations and Analytics for AI application performance
- Powerful Evaluation Models (Lynx, Glider) and custom evaluator definition
- Dataset Generation with proprietary algorithms for RAG, Agents, and redteaming
- Patronus API
- Patronus SDK
- Custom error taxonomy
- Prompt management (version and deploy prompts as code)
- Human-in-the-loop annotations
- On-prem / dedicated VPC security options
- Custom data retention
- SSO
- Premium Platform Features (Patronus Evaluation Runs, webhooks)
- Premium API Features (Higher rate limits, volume discounts, stability)
- AI Services (Custom eval model fine tuning, eval dataset generation)
Use cases
- Scaling the creation of high alpha simulations for frontier models
- Evaluating models on static data sets
- Improving agents on long horizon problems in real world-like settings
- Evaluating, monitoring, and improving performance of LLM systems
- Detecting hallucinations in AI applications
- Optimizing AI agents for code generation
- AI agent fleet optimization
- Accelerating complex AI agent development
- Optimizing image captioning for multimodal AI systems
- Scaling AI performance with automated evaluations and rigorous experimentation
- Evaluating and optimizing personalized message replies for Airbnb hosts
- Preventing hallucinations in AI-powered customer support chatbots
- Numerical reasoning of financial metrics
- Information retrieval from databases
- Logical reasoning for subjective financial asks
- Knowledge of accounting & finance
- Regulatory compliance testing for financial AI
- Benchmarking LLM systems against FinanceBench
- Detecting hallucinations and unexpected LLM behavior on financial questions
- Creating custom benchmarks for financial services
- Ensuring accuracy, relevance, regulatory alignment, and safety in financial AI outputs
- Scoring multimodal AI systems for image to text
- Detecting and mitigating caption hallucination from product images
- Testing whether user queries are surfacing the most relevant product screenshots
- Testing whether OCR extraction for tabular data is accurate
- Testing whether AI-generated brand images, logos, and listings are accurate
- Testing whether captions accurately describe image scenes
- Evaluating RAG applications
- Debugging agent failures
- Production LLM monitoring
- Building evaluations with AI assistance
- Adding guardrails to applications
- Configuring custom criteria for LLM-as-judges
- Generating test datasets
AI approach
Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. They are training the First Digital World Model for AI agent training and simulation, aiming to simulate all of the world’s intelligence. They develop core evaluation platforms and tools like LLM-as-a-Judge, Glider, Lynx, and Percival to evaluate, monitor, and optimize Generative AI applications and agents. They also generate high-quality custom datasets and offer redteaming algorithms.
Tech named: Digital World Models, LLM-as-a-Judge, Glider (3B evaluator LLM), Lynx (SOTA hallucination detection model, 8B and 70B versions), Percival (evaluation copilot for agentic systems), Generative Simulators, MemTrack (benchmark), FinanceBench (benchmark), SimpleSafetyTests (diagnostic test suite), EnterprisePII (LLM dataset), Patronus Evaluators, Patronus Experiments, Patronus Datasets, Patronus Logs, Patronus Comparisons, Patronus Traces, Patronus API, Google Gemini backbone (for MLLM-as-a-Judge)
Industries served
- Technology
- Information and Internet
- Software Development
- Customer Service
- Finance
- E-commerce
- Healthcare
- Legal
- Marketing & Sales
- Security
What it says sets it apart
- Frontier lab developing simulation research and infrastructure
- Training the First Digital World Model for AI agent training and simulation
- Company behind influential research in AI evaluation (FinanceBench, Lynx, SimpleSafetyTests, CopyrightCatcher)
- Research-backed, real-world inspired approach
- High model lift (30-40%) on long-horizon tasks
- Extensive world data artifacts (1M+)
- High UI/UX feature parity (85%) with real-world products
- Large network of expert contributors (5k+)
- Proprietary dataset generation algorithms
- Industry-first Multimodal LLM-as-a-Judge
- Team of AI researchers and engineers from Meta AI, Amazon AGI, Google
- Focus on engineering trust and reliability for AI agents through simulation
Funding rounds we track
Greenfield Partners, Greenfield, Lightspeed Venture Partners, Datadog, Samsung, Notable Capital, Factorial Capital, Samsung Ventures, Group 11
Notable Capital, Datadog Inc., Lightspeed Venture Partners, Datadog, Factorial Capital
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from Patronus AI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.