AI News Today.

Artificial intelligence, professionally covered

Company profile

Reducto

Agentic document platform for AI teams, processing unstructured data at enterprise scale.

reducto.aiProfile compiled July 202612 source pages read
Category
Developer tools
Headquarters
San Francisco, CA
Sells to
Enterprise
Business model
Usage-based API
Deployment
Cloud / SaaS, On-premise, Hybrid, API
Pricing
Credit-based, pay-as-you-go, with tiered plans (Standard, Growth, Enterprise)
Builds own models
Yes
Modalities
Text, Image, Multimodal

Reducto is the complete agentic document platform for leading AI teams needing performance at enterprise scale. It provides a comprehensive toolkit for working with documents the way a human would, combining custom in-house and leading frontier models to power efficient and accurate document workflows. Reducto powers AI teams from startups to Fortune 10 companies across various industries, offering flexible deployment options from cloud to air-gapped environments, SOC II and HIPAA compliance, and zero data retention.

  • Parse APIConverts any document into structured JSON with text, tables, and figures. Layout-aware chunking optimized for LLMs.
  • Extract APIDefines a JSON schema and pulls specific fields from documents with schema-level precision.
  • Edit APIFills detected blanks, tables, and checkboxes with supplied data, dynamically identifying fillable elements regardless of document layout or format. Fills PDF forms and modifies DOCX files programmatically with natural language instructions.
  • Split APIAutomatically separates multi-document files or long forms into individually useful units. Divides documents into logical sections using natural language descriptions for targeted downstream processing.
  • Classify APIRoutes documents by type before processing. Defines categories in natural language, no training data required.
  • Reducto StudioA visual platform to build, test, and deploy document workflows. Inspects results with bounding-box-level citations. Offers unlimited seats for Growth and Enterprise tiers.
  • MCP ServerConnects AI agents to Reducto via the Model Context Protocol, working with Claude, Cursor, VS Code, and more.
  • SDKsFirst-class clients for Python, Node.js, and Go with full type safety for developers building automated pipelines.
  • CLIAllows parsing, extracting, and managing documents from the terminal, piping output into other tools or scripts.
  • PipelinesChains multiple steps (classification, parsing, extraction, editing) into reusable, single-call workflows deployed from Studio.
  • Agentic OCR reviews and corrects outputs in real-time
  • Intelligent heuristics and layout-aware splitting
  • Schema-level precision for data extraction
  • Dynamic identification of fillable elements for editing
  • Traditional computer vision for layout breakdown
  • VLMs make corrections to mistakes
  • VLMs review Reducto's outputs
  • Intelligent chunking
  • Figure summarization
  • Graph extraction
  • Automatic page rotation
  • Embedding optimization
  • 99.9%+ uptime
  • Enterprise support and SLAs
  • SOC2, HIPAA compliant
  • Flexible deployment options (cloud, VPC, on-prem, air-gapped)
  • Zero data retention policy (Growth tier and above)
  • 30+ Supported File Types (PDF, image formats, spreadsheets, presentations, text documents)
  • Multilingual OCR
  • Custom Webhooks
  • Enrichment VLM
  • Spreadsheet Table Clustering
  • Extraction Citations
  • Bounding Box Support
  • Structured Schema
  • Batch processing
  • Volume Discounts
  • Business Associate Agreement
  • Premium Rate Limits and Priority Requests
  • Priority Slack and Email Support
  • EU/AU Data Residency Endpoints
  • Custom MSA
  • Custom SLA
  • Custom Rate Limits and Throughput
  • Custom Processing Pipelines
  • Dedicated On-Call Support
  • Role-Based Access Control
  • SSO and SAML Authentication
  • API & SDKs (Python, Node.js, Go, REST)
  • Studio citation viewer
  • Encryption at rest (AES-256) and in transit (TLS 1.2+)
  • Parsing documents with complex tables
  • Separating multi-document files or long forms
  • Extracting structured data from invoices, onboarding forms, financial disclosures
  • Filling in detected blanks, tables, and checkboxes in documents
  • LLM Data Preparation
  • Building a first-party data engine for financial assets
  • Delivering document intelligence to wealth management firms
  • Parsing financial documents end to end
  • Market research & data intake (research reports, filings, news articles, slide decks)
  • Risk and holdings assessment (401(k) and brokerage holdings)
  • Financial knowledge retrieval (unstructured research into interactive knowledge base)
  • Chart and table extraction (line graphs, visual charts, multi-level tables)
  • Advisor book-of-business management (insurance statements)
  • Financial contract data extraction (credit agreements, ISDAs, term sheets, loan docs)
  • Reliable document understanding at scale for legal and professional services
  • Automating SOC 2 and ISO 27001 compliance (understanding questionnaires, validating evidence)
  • End-to-end automated workflows for midsized law firms
  • Organizing and extracting key data for registered investment advisors
  • Accelerating insurance claims review
  • Enabling anyone to create AI agents and workflows
  • Accelerating prior authorization with 99%+ accuracy (medical documents to structured data)
  • Processing documents in enterprise workflows
  • Extracting every field from handwritten property loss forms
  • Parsing loss run reports into clean claims tables
  • Semantic search and RAG across policies, claims files, ACORD forms, underwriting submissions, specialty/reinsurance slips
  • Underwriting intake (extracting risk factors, protections, exposures, limits, timelines)
  • Specialty & reinsurance processing (parsing slips for treaty terms, line sizes, exclusions, currencies, loss histories)
  • Fraud signal extraction (inconsistencies, duplicated evidence, date mismatches, conflicting statements)
  • Evidence packet structuring (photos, adjuster notes, estimates, invoices, correspondence)
  • CAT event aggregation (photos, field reports, adjuster notes from catastrophic events)
  • Automated intake pipelines that classify, split, and route documents
  • Structured data extraction from invoices, contracts, medical records
  • Document generation and editing (filling forms, modifying templates, producing new documents)
  • RAG-ready content pipelines with layout-aware chunking for LLM consumption
  • Multi-step workflows chaining classification, parsing, extraction, and editing
  • Parsing any healthcare document into clean, structured text
  • Extracting structured data from patient intake forms
  • Clinical documentation processing (physician notes, H&Ps, discharge summaries)
  • Medical records retrieval (searchable knowledge bases from patient records)
  • Claims & prior auth automation (insurance claims, EOBs, prior authorization requests)
  • Clinical trial document processing (consent forms, protocols, study documentation)
  • Revenue cycle optimization (ICD/CPT extraction from clinical documentation)
  • Care quality analytics (structuring clinical data for quality metrics, population health insights)
  • Permit & license processing (extracting data from applications, renewals, supporting documentation)
  • Benefits administration (processing eligibility forms, supporting documents, verification materials)
  • Regulatory compliance (parsing filings, reports, compliance documentation)
  • Records management (building searchable archives from legacy documents, citizen records, historical filings)
  • Defense & intelligence (multimodal understanding of military documents, maps, figures, intelligence reports)
  • Procurement processing (parsing contracts, proposals, vendor documentation)

Reducto uses a pipeline of specialized models, including custom in-house models and frontier Vision-Language Models (VLMs), with agentic multipasses that iteratively correct errors. This architecture aims for high accuracy on complex real-world documents, handling handwritten forms, rotated pages, nested tables, multi-column layouts, and degraded scans. It combines layout-aware models, Agentic OCR for real-time corrections, and VLMs to interpret regions in context.

Tech named: OCR Models, LLM Data Preparation, Machine Learning, layout-aware models, Agentic OCR, Vision-language models (VLMs), Model Context Protocol (MCP)

  • Legal
  • Finance
  • Healthcare
  • Insurance
  • Government
  • Defense
  • Professional Services
  • Wealth Management
  • Complete agentic document platform
  • Combines custom in-house and leading frontier models
  • Reads documents like a human would, capturing layout, structure, and meaning with high accuracy
  • Agentic OCR reviews and corrects outputs in real-time for near-perfect results
  • Flexible deployment options (cloud, hybrid VPC, full VPC, air-gapped on-premises)
  • SOC 2 Type II and HIPAA compliant
  • Zero data retention policy for Growth tier and above
  • Battle-tested infrastructure with 99.9%+ uptime
  • Enterprise support and SLAs, including dedicated on-call support
  • Supports 30+ file types
  • Production-ready structured JSON in minutes, not months
  • Agent-ready tooling with MCP server, CLI, and native SDKs
  • Orchestrates a pipeline of specialized models with agentic multipasses that correct errors iteratively
  • Delivers accuracy on the long tail of real-world documents (handwritten forms, rotated pages, nested tables, multi-column layouts, degraded scans)
  • Studio citation viewer for inspecting outputs against source documents at the bounding-box level
  • No data used for training purposes for Growth tier and above
  • One-day turnaround time for specific feature requests (mentioned by a customer)

Andreessen Horowitz

From the AI funding tracker — rounds as reported by the linked publications.

This profile was compiled from Reducto's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.