Company profile
Snorkel AI
Snorkel AI builds expert data and environments for high-performing frontier and agentic AI.
- Category
- Data platforms
- Headquarters
- Redwood City, California
- Sells to
- Mixed
- Business model
- Services & consulting
- Deployment
- Cloud / SaaS
- Pricing
- Not published
- Builds own models
- No — builds on existing models
- Modalities
- Text, Code, Multimodal
What Snorkel AI does
Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. They combine platform technology with research-driven data development to create datasets, benchmarks, evaluations, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. They led the development of Senior SWE-Bench and launched Open Benchmarks Grants with a $3 million commitment to support open-source datasets, benchmarks, and evaluation research. Snorkel's methodology involves task design, programmatic checks, calibrated expert review, and realistic evaluation environments to create measurable training signals for frontier models and agents. They focus on specialized tasks, benchmark blind spots, and failure modes at the edges of AI capabilities.
Products
- Snorkel Data SeriesCurriculum-structured datasets for task areas where frontier models are pushing hardest, with rubrics, reviewer guidance, difficulty tiers, and evaluation slices built in.
- Custom data developmentBespoke datasets, evaluations, and benchmark expansions built when off-the-shelf coverage is insufficient, targeting specific failure surfaces.
- Specialized agentsCustom agents built on specialized data and evaluated in real workflows, with pass/fail criteria tied to performance standards that drive ROI.
- Senior SWE-benchEvaluates coding agents on real-world senior engineering tasks.
- Agents' Last ExamEvaluates how well AI agents complete economically valuable real-world work.
- OSWorld 2.0Evaluates computer-use agents on long-horizon desktop workflows.
- Frontier-BenchA benchmark to measure and evolve with the frontier of agent work.
- Open Benchmarks GrantsA $3 million commitment to support open-source datasets, benchmarks, and evaluation research.
Key capabilities
- AI Data Development
- Expert Data
- Post-Training Data
- LLM Training Data
- Preference Data
- Reinforcement Learning
- RL Environments
- Reinforcement Learning from Human Feedback (RLHF)
- Supervised Fine-Tuning (SFT)
- AI Benchmarks
- AI Evaluation
- Agent Evaluation
- Agentic AI
- Frontier AI
- Research-grade datasets
- Evaluation systems
- Runnable environments
- Expert demonstrations & reasoning
- Human solution traces
- Reasoning traces
- SME Q&A rationales
- Workflow demos and decision workflows
- Tool-use demos
- Preference labels & rankings
- Patch/draft/report quality ranking
- Trajectory QA
- Risk/safety/style calibration
- Helpful/harmless ranking
- Grounding & style
- Rubrics & verifiable outcomes
- Unit tests / compile
- Deterministic graders
- Citation correctness
- Numerical consistency/scorable math/science
- Long-horizon tasks
- Standard & custom environments
- Repo + CLI tools
- Browser/GUI harness
- Multi-step/stateful workflows
- Simulated environments
- Integration with client tools, codebase, corpus, data & permissions
- Task design
- Programmatic checks
- Calibrated expert review
- Realistic evaluation environments
- Meta-evaluation of evaluators
- Evaluator development (model-based and rule-based)
- Expert correction and feedback loop
- Research-validated methodology
Use cases
- Developing specialized training data and environments for frontier models and agents
- Building research-grade datasets
- Creating evaluation systems
- Developing runnable environments
- Closing distributional gaps in specialized domains for frontier models
- Addressing benchmark blind spots for frontier models
- Solving tasks where correctness is hard to define for frontier models
- Improving coding agents
- Evaluating real-world agents
- Evaluating computer-use agents
- Benchmarking agents in insurance underwriting
- Data development for agentic systems (tool-using systems)
- Structured data and evaluation workflows for agents that need to make decisions, use tools, and complete complex tasks end to end
- Assessing LLM math reasoning skills on high school to graduate-level challenges
- Building deep, expert-crafted datasets of realistic multi-turn, multi-agent conversations for AI voice assistants
- Curating gold-standard sets of prompts, responses, and tool calls for robust agentic evaluation benchmarks
- Creating datasets to enhance models’ deep research capabilities (multi-step, multi-turn, multi-tool deep research data)
- Annotation workflows for tasks requiring domain expertise, clear review standards, and consistent judgment
- Grading LLM information retrieval and synthesis from technical documents
- Enabling FMs to understand charts (annotations of graphs, maps, and visuals for math problems)
- Data for models that need to understand codebases, generate reliable solutions, and perform across real developer workflows
- Alignment for better code generation using human feedback
- Training and evaluation data for code generation (unique prompts, verifiable solutions, unit tests)
AI approach
Snorkel AI is a 'frontier AI data lab' that helps teams develop specialized training data and environments for high-performing frontier and agentic AI models. They combine platform technology with research-driven data development to create datasets, benchmarks, evaluations, and custom solutions. Their methodology involves task design, programmatic checks, calibrated expert review, and realistic evaluation environments to generate measurable training signals. They focus on addressing failure modes in specialized tasks and benchmark blind spots. Snorkel AI also builds custom agents grounded in expert data, evaluated against task-specific rubrics and programmatic checks.
Tech named: data programming, weak supervision, RLVR (Reinforcement Learning with Value Regularization), RIFT (Rubric Failure Mode Taxonomy), LLM (Large Language Model), RLHF (Reinforcement Learning from Human Feedback), SFT (Supervised Fine-Tuning)
Industries served
- Research Services
- Tech Industry (for AI voice assistants)
- Telecom
- Insurance Underwriting
What it says sets it apart
- Founded out of the Stanford AI Lab in 2019
- Focus on frontier AI and agentic AI
- Specialization in data and environments for models that break at the edges (specialized domains, benchmark blind spots, hard-to-define correctness)
- Proprietary process built around design choices for training data
- Expert-in-the-loop methodology combining programmatic scale with human precision
- 1,000+ expert-level domains covered
- Meta-evaluation and calibration of evaluators
- Research-validated methodology with 200+ peer-reviewed papers and open benchmarks
- Ability to build custom agents grounded in expert data and evaluated against task-specific rubrics and programmatic checks
- Commitment to open-source datasets, benchmarks, and evaluation research through Open Benchmarks Grants
- Emphasis on reproducible traces and failure analysis for partner teams
Funding rounds we track
Addition, Prosperity 7 Ventures, Greylock, Lightspeed Venture Partners, BNY, QBE Ventures, Prosperity7 Ventures
Addition, BlackRock, Factory, Cooley
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from Snorkel AI's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.