AI News Today.

Artificial intelligence, professionally covered

Company profile

Coval

Automated testing for AI agents, specializing in chat and voice systems.

coval.aiProfile compiled July 202615 source pages read
Category
Developer tools
Headquarters
San Francisco, San Francisco
Sells to
Enterprise
Business model
SaaS subscription, Usage-based API
Deployment
Cloud / SaaS, API
Pricing
Tiered subscription and usage based · from $100/mo
Builds own models
Yes
Modalities
speech, Text, Audio

Coval accelerates AI agent development with automated testing for chat, voice, and other objective-oriented systems. It provides an AI evaluation platform for teams to scale AI agents with confidence, offering simulation, observation, and review capabilities. Inspired by the autonomous vehicle industry, Coval uses automated simulation and evaluation techniques to boost test coverage, speed up development, and validate consistent performance. The platform is built for voice AI, handling accents, interruptions, background noise, and tool calls, and provides full lifecycle coverage from simulation to production monitoring. It incorporates human-in-the-loop judgment to sharpen the system over time. Coval supports both voice and chat agents on a single platform and is vendor-agnostic, integrating with existing observability stacks.

  • Voice AI Evaluation PlatformA comprehensive platform for simulating, observing, and reviewing voice AI agents at scale. It provides infrastructure to ship voice AI with confidence, covering every stage of agent evaluation.
  • SimulateA product within the platform to run thousands of realistic conversations before launch. It offers voice-native simulation, comprehensive coverage with diverse voices, languages, and environments, and is CI/CD ready.
  • ObserveA product within the platform for continuous evaluation and monitoring of AI agents in production. It catches failures the moment they happen, scores every production call, and provides threshold-based alerting and native integrations.
  • ReviewA product within the platform for human-in-the-loop review of AI agent evaluations. It routes critical calls to expert reviewers, allows for one-click overrides and annotations, and uses human feedback to retrain the AI judge.
  • Coval REST APIAn API that enables programmatic launching of voice and chat evaluations, managing test data, and analyzing AI agent performance. It includes endpoints for runs, agents, simulations, and test sets.
  • Automated simulation and evaluation
  • Voice-native evaluation (accents, interruptions, background noise, tool calls)
  • Full lifecycle coverage (simulation to production monitoring)
  • Human-in-the-loop review
  • Support for both voice and chat agents
  • Vendor-agnostic evaluation
  • CI/CD integration (GitHub Actions, CLI)
  • Comprehensive coverage (27 voices, 10 languages, 20 environments)
  • Custom and universal metrics (resolution, adherence, accuracy, latency, compliance)
  • Mutations and A/B testing
  • Headless operation (CLI and API first) with UI when needed
  • Continuous evaluations in production
  • Threshold-based alerting (Slack, email)
  • Native integrations with observability stacks (Langfuse, Langsmith, Arize, Datadog)
  • Smart sampling for human review queues
  • One-click overrides and annotations for human reviewers
  • Continuous quality loop (human feedback retrains AI judge)
  • Pre-built behavior tests (identity checks, escalation, hallucination, frustrated callers)
  • OpenTelemetry-native
  • Custom dashboards and visualizations
  • Anomaly detection
  • Failure-mode categorization
  • One-click monitors
  • SOC 2 Type II compliance
  • HIPAA compliance
  • GDPR compliance
  • SAML SSO
  • SCIM provisioning
  • Role-based access control (RBAC)
  • Audit logs
  • IP allowlisting
  • Data residency
  • Private / VPC deployment
  • BYO storage
  • White-label reports
  • Custom SLA
  • Accelerating AI agent development
  • Automated testing for chat, voice, and objective-oriented systems
  • Boosting test coverage for AI agents
  • Speeding up development of AI agents
  • Validating consistent performance of AI agents
  • Scaling AI agents with confidence
  • Running realistic conversations before launch (simulation)
  • Catching failures in production (continuous monitoring)
  • Sharpening the system with human-in-the-loop review
  • Evaluating agent quality across multiple customers/configurations
  • Spinning up customized test suites from customer prompts, flows, and policies
  • Catching regressions across accounts
  • Proving reliability to customers with exportable evidence
  • Testing contact center voice agents for resolution, escalation, and deflection
  • Validating agent assist suggestions
  • Testing deflection paths
  • Ensuring compliance for financial services voice agents
  • Validating verification flows, disputes, and disclosures in banking
  • Validating disclosures in lending
  • Testing card activations, freezes, and inquiries in cards & payments
  • Ensuring HIPAA-compliant voice AI for healthcare
  • Evaluating healthcare voice agents against clinical and patient-safety standards
  • Testing appointment scheduling, medical records requests, and billing/insurance verification in healthcare
  • Comparing voice AI platforms head-to-head (vendor bakeoffs)
  • Evaluating voice agents built on in-house stacks
  • Testing agent behaviors (identity checks, collecting required information, confirming before acting, transfer & escalation, handling frustrated callers, catching hallucinations, staying on topic)
  • Testing hospitality voice agents for booking flows, service requests, and loyalty policies

Coval provides an AI evaluation platform for voice and chat agents. It uses automated simulation and evaluation techniques, inspired by the autonomous vehicle industry, to test AI agents. The platform offers voice-native simulation, continuous evaluation in production, and human-in-the-loop review. It supports various voice models (basic, advanced, custom/BYO) and uses AI judges that learn from human feedback.

Tech named: OpenTelemetry, Langfuse, LangSmith, Arize, Datadog, SIP header tracing, LiveKit, Pipecat

  • Technology, Information and Internet
  • Customer Support
  • Financial Services
  • Healthcare
  • Hospitality
  • Automated simulation and evaluation techniques inspired by the autonomous vehicle industry
  • Built specifically for voice AI, not retrofitted from text
  • Full lifecycle coverage from simulation to production monitoring
  • Human-in-the-loop judgment built into every evaluation
  • Supports both voice and chat agents on a single platform
  • Vendor-agnostic design, allowing comparison of agents from multiple vendors
  • Integrates with existing observability stacks (OpenTelemetry-native)
  • Provides pre-built behavior tests based on thousands of real production agents
  • Offers audit-ready infrastructure with SOC 2 Type II, HIPAA, and GDPR compliance
  • Focuses on continuous quality loop where human feedback retrains the AI judge
  • Allows for customized test suites for each customer's prompts, flows, and policies
  • Offers a single scoreboard for vendor bakeoffs, measuring accuracy, latency, escalation handling, task completion, and cost consistently across vendors
  • Provides full traces for debugging failed runs, including transcript, audio, and turn-by-turn breakdown

Norwest, Base10 Partners, Twilio Ventures, Y Combinator, MaC Ventures, Swift Ventures

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from Coval's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.