AI News Today.

Artificial intelligence, professionally covered

Company profile

LangWatch

AI agent testing and evaluation platform for reliable production systems.

langwatch.aiProfile compiled July 202613 source pages read
Category
MLOps
Headquarters
Amsterdam, North Holland
Sells to
Developers
Business model
Freemium, SaaS subscription
Deployment
Cloud / SaaS, Self-hosted, Hybrid
Pricing
Freemium · from $29/mo · free tier
Builds own models
No — builds on existing models
Modalities
Text, speech, Multimodal

LangWatch is an AI agent engineering platform that provides simulation-based testing and evaluation to transform unpredictable AI agents into reliable production systems. It offers tools for agent simulations, evaluations, and observability, enabling teams to test agents against realistic multi-turn scenarios, catch regressions, and monitor performance in production. The platform supports self-hosting and integrates with various LLM providers and agent frameworks, aiming to make AI agents reliable enough for teams to focus on strategy and creativity.

  • LangWatch PlatformA comprehensive platform for AI agent testing, evaluation, and observability, offering agent simulations, evaluations, and monitoring capabilities.
  • Scenario (open-source)An open-source agent testing framework that powers agent simulations, allowing users to write scenarios in Claude Code for realistic, multi-turn testing.
  • LangWatch CLIA command-line interface for documentation and platform operations related to prompts, scenarios, evaluators, datasets, monitors, traces, and analytics.
  • LangWatch SDKs (Python and TypeScript)Software Development Kits for Python and TypeScript that are OpenTelemetry-native, allowing for easy integration and instrumentation of agents.
  • AI GatewayAn optional sub-chart for self-hosted deployments that terminates LLM traffic on the user's perimeter, providing virtual keys, hierarchical budgets, multi-provider routing, guardrails, and prompt caching.
  • RedTeamAgentA drop-in replacement for UserSimulatorAgent that runs structured adversarial attacks against agents, including multi-turn escalation and refusal detection.
  • Agent Simulations
  • Evaluations
  • Pre-Production Testing
  • Spec-driven agent building
  • Replicate and fix issues from production
  • Simulate real users (text and voice)
  • Write scenarios in Claude Code
  • Local and CI integration for scenarios
  • Red teaming (adversarial simulations)
  • Judge with reasoning
  • Trace tool calls, skills, MCP
  • Whitebox and blackbox testing
  • LLM as a judge
  • Notebook or UI for evals
  • Online evaluations (production traffic)
  • Pairwise comparison of outputs
  • Multimodal evaluation (images, mixed media)
  • Observe every trace, token, and cost
  • OpenTelemetry native
  • Fast search and filtering
  • Custom views and AI search
  • Waterfall, flame graph, topology, sequence diagram views
  • Topic clustering for conversations
  • Plot any metric (cost, latency, scores)
  • Self-host option
  • Simulated users (LLM-powered user simulator)
  • Multi-turn conversation testing
  • Configurable success criteria
  • Scripted to auto-pilot simulations
  • Tool-call verification across long dialogues
  • Framework-agnostic adapters
  • Run locally or in CI/CD
  • Simulation visualizer (visual debugging)
  • Pause, evaluate & annotate mid-conversation
  • Open-source Scenario SDK (Python + TypeScript)
  • Voice agent testing (ElevenLabs, OpenAI Realtime, Twilio, Pipecat, Gemini Live)
  • Voice: latency metrics & noise/interruption injection
  • Adversarial / red-teaming (Crescendo escalation, refusal detection)
  • Offline experiments via SDK and UI
  • CI/CD integration for evals
  • Multi-modal evaluations
  • Online evaluations, Monitors
  • Evaluation by thread
  • Guardrails (code integration)
  • Built-in evals (RAGAS, hallucination, toxicity, PII, LLM-as-a-judge)
  • Create reusable evaluators org-wide
  • Custom scoring
  • Build custom evals via workflows
  • Annotations / annotation inbox
  • OpenTelemetry & REST support
  • LLMs & model gateways (OpenAI, Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, Vertex AI, Groq, Ollama, LiteLLM)
  • Agent frameworks (OpenAI Agents, LangGraph, LangChain, CrewAI, Pydantic AI, Agno, Mastra, Vercel AI SDK, Google ADK, LangFlow, Flowise)
  • Prompt optimization (DSPy)
  • Encryption (AES-256 at rest, TLS 1.2+ in transit)
  • Access control (RBAC, MFA, SSO)
  • Monitoring & response (Snyk, AWS CloudTrail + CloudWatch)
  • Backup & recovery (daily encrypted backups, geo-redundant storage)
  • Secure development (security code audits, Snyk scanning, peer review)
  • Data privacy (automatic PII detection + removal, GDPR, DPA)
  • EU data residency
  • Hybrid, self-hosted or on-prem deployment options
  • SSO, SCIM
  • Audit logs
  • Priority support
  • Testing AI agents against realistic, multi-turn scenarios
  • Catching regressions in AI agent behavior
  • Validating prompts
  • Simulating agents against realistic scenarios before deployment
  • Continuous testing and evaluation for AI agents
  • Turning requirements into agent tests automatically
  • Speeding up AI agent development
  • Replicating and fixing issues from production
  • Adversarial simulations for jailbreaks, policy breaks, and unsafe tool calls
  • Scoring single outputs to whole conversations
  • Evaluating production traffic in real time
  • Comparing two outputs side by side for model/prompt/version selection
  • Evaluating images and mixed media
  • Searching and analyzing traces, tokens, and costs
  • Building custom analytics graphs over metrics
  • Automating the creation of test plans and pull requests for AI agents
  • Benchmarking new models on production data
  • Ensuring regulatory compliance (GDPR, SOC 2)
  • Operating in air-gapped environments
  • Automating chatbot testing and quality assurance
  • Blocking harmful content in real-time using guardrails

LangWatch provides an AI agent testing platform that uses AI for agent simulations, evaluations, and red teaming. It leverages LLMs as judges and for generating realistic user scenarios. The platform also includes an AI-powered assistant, Langy, to generate test plans and fix failures.

Tech named: LLM-as-a-judge, AI-powered user simulator, AI-powered trace Ask, AI Gateway, Langy

  • Software Development
  • Facility Management
  • Cleaning Companies
  • Property Management
  • Simulation-based AI agent testing and evaluation
  • Focus on multi-turn agent failures, not just model failures
  • Open-source agent testing framework (Scenario)
  • Loop engineering approach to agent testing
  • Ability to simulate real users in text and voice
  • Red teaming for adversarial simulations
  • Comprehensive evaluation capabilities, including LLM as a judge and multimodal support
  • OpenTelemetry native observability for GenAI
  • Self-hosting, hybrid, and cloud enterprise deployment options for data sovereignty and compliance
  • Built by engineers who experienced the fragility of AI in production
  • Emphasis on actionable binary evaluations over continuous scores
  • Deep production context integration for test scenarios (e.g., with Vinny's implementation)
  • Strong security and compliance foundations (ISO 27001, GDPR, SOC 2)

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from LangWatch's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.