AI News Today.

Artificial intelligence, professionally covered

Company profile

Galileo

AI observability and evaluation platform for GenAI and agentic applications.

galileo.aiProfile compiled July 202615 source pages read
Category
MLOps
Headquarters
Burlingame, California
Sells to
Mixed
Business model
Freemium, SaaS subscription
Deployment
Cloud / SaaS, On-premise, Hybrid
Pricing
Free tier, Pro tier with usage-based pricing, and Enterprise tier with custom options. · from $100/mo · free tier
Builds own models
Yes
Modalities
Text

Galileo is an AI observability and evaluation engineering platform that helps prevent AI failures by turning offline evaluations into production guardrails. It provides a comprehensive suite of products to support AI development workflows, including fine-tuning LLMs, developing, testing, monitoring, and securing AI applications. The platform is powered by research-backed evaluation metrics and is used by AI teams from startups to Fortune 50 enterprises. Galileo enables users to capture groundtruth, build datasets from synthetic, development, and live production data, and create accurate evaluations. It distills optimized evaluations into Luna models for low-cost, scalable monitoring. Key features include out-of-box evaluations for RAG, agents, safety, and security, custom evaluators, an insights engine for debugging, and the ability to create guardrail policies to block harmful responses. Galileo supports various deployment options including SaaS, Virtual Private Cloud, and On-Premises.

  • Galileo PlatformAn end-to-end platform for AI evaluation, observability, and real-time protection, designed to help ship AI applications with confidence. It supports all AI workflows, measuring AI accuracy offline and online, and blocks harmful outputs and security risks in real-time.
  • Luna modelsProprietary small language models (SLMs) that distill expensive LLM-as-judge evaluators into compact, accurate, and low-latency models for monitoring production AI applications in near real-time.
  • Galileo EvaluateA product for evaluating AI models and prompts, enabling users to test variations, build and execute golden test-sets, debug with traces, evaluate AI output with pre-built or custom metrics, and refine prompts with ML-powered insights.
  • Galileo ObserveA product for AI observability, providing real-time insights and monitoring for active AI implementations, allowing teams to manage and optimize deployments from a streamlined console.
  • Galileo ProtectA component that integrates with NVIDIA NeMo Guardrails to build safe, secure, and robust solutions, providing real-time observability and alerts for production systems.
  • Galileo AutotuneA feature for AI evaluation tuning that automates the process of optimizing evaluators.
  • Galileo SignalsA tool to find issues with AI.
  • AI observability and evaluation engineering platform
  • Offline evals become production guardrails
  • Capture groundtruth and build datasets
  • Auto-tunes metrics from live feedback
  • Distills optimized evals into Luna models for low-cost monitoring
  • 20+ out-of-box evals for RAG, agents, safety, and security
  • Custom evaluators
  • Insights engine analyzes agent behavior to identify failure modes, surface hidden patterns, and prescribe fixes
  • Unit testing and CI/CD rigor for AI development lifecycle
  • Pre-production evals become production governance
  • Eval scores automatically control agent actions, tool access, and escalation paths
  • Block harmful responses with guardrail policies
  • SaaS, Virtual Private Cloud, and On-Premises deployment options
  • AI Evaluation (integrate with code or playground UI, build and execute golden test-sets, debug with traces, evaluate AI output, refine prompts, organize prompt versions)
  • AI Observability
  • Real-time Protection
  • Prebuilt metrics (20+ out-of-the-box evaluators)
  • Custom metrics (code-based or automatically generated LLM-as-judge evaluators)
  • Auto-tune evaluators with CLHF (Continuous Learning with Human Feedback)
  • Inference server for low-latency evaluators in production
  • Proprietary Luna models for fast and accurate evaluation
  • Online and offline evaluation capabilities
  • SDKs and APIs for easy integration
  • Agent Observability Platform (Agent Development Lifecycle, Experimentation, CI/CD Testing, Real-time Monitoring, Run-time Protection, Observability and Intervention, Agent Graph, Agent Insights, Root Cause Analysis, Custom Dashboards, Guardrails, Alerting, Metrics Engine, Autogen Metrics, Luna SLMs, Inference Server, Evaluation Assets, Prompt Store, Datasets, Traces / Sessions, Guardrail Policies, Annotations)
  • Integrates with popular LLMs and agent frameworks
  • Fine-tuning LLMs
  • Developing AI applications
  • Testing AI applications
  • Monitoring AI applications
  • Securing AI applications
  • Evaluating RAG applications
  • Evaluating AI agents
  • Ensuring AI safety
  • Ensuring AI security
  • Debugging AI systems
  • Improving conversational AI precision
  • Mitigating hallucinations in AI
  • Monitoring prompts and outputs
  • Accelerating GenAI evaluations and experimentation
  • Evaluating customer sentiment from online reviews
  • Building safe and reliable AI applications
  • Experimentation in AI development
  • CI/CD testing for AI
  • Real-time monitoring of AI in production
  • Run-time protection for AI applications
  • Optimizing prompt performance
  • Detecting and correcting hallucinations, drift, and bias in data

Galileo provides an AI observability and evaluation platform that helps teams develop, test, monitor, and secure their AI applications. It offers pre-built and custom evaluators, and distills expensive LLM-as-judge evaluators into compact, low-latency Luna models for real-time production monitoring and guardrails. The platform focuses on improving AI accuracy, detecting failures like hallucinations, and ensuring reliable operations for GenAI and agentic applications.

Tech named: LLMs, RAG, Luna models, Luna-1 (3b), Luna-1 (8b), Luna-2 (3b), Luna-2 (8b), NVIDIA NeMo, NVIDIA NIM, CrewAI, Azure AI Content Safety

  • Software Development
  • Entertainment
  • Consumer Packaged Goods (CPG)
  • FinTech
  • Only Galileo distills expensive LLM-as-judge evaluators into compact Luna models that run with low-latency and low-cost.
  • Provides an insights engine that analyzes agent behavior to identify failure modes, surface hidden patterns, and prescribe fixes.
  • Brings unit testing and CI/CD rigor into the AI development lifecycle through the eval-to-guardrail lifecycle.
  • Prebuilt evaluators to get started and custom evaluators for unique applications.
  • Luna models offer high accuracy and low latency for real-time production monitoring.
  • End-to-end solution for experimentation, CI/CD, observability, and real-time monitoring.
  • Offers a free tier for developers and small teams to experiment and build.
  • Provides a consistent evaluation framework across the development lifecycle.

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from Galileo's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.