Company profile
The Synthetic Data Vault
Ecosystem of open-source tools for generating synthetic data.
- Category
- Data platforms
- Headquarters
- Boston, MA
- Sells to
- Enterprise
- Business model
- Freemium, SaaS subscription, Usage-based API, Licensing
- Deployment
- Self-hosted, On-premise
- Pricing
- Freemium with usage-based pricing for enterprise · from $500/mo · free tier
- Builds own models
- Yes
- Modalities
- Tabular
What The Synthetic Data Vault does
The Synthetic Data Vault (SDV) is an ecosystem of source-available software tools built to help enterprises generate synthetic data. It is currently used by more than 50 Fortune 500 companies. Synthetic data is generated by an AI model trained on real data, maintaining the same formatting, statistical properties, and patterns, but cannot be linked back to real people or data. The SDV project launched at MIT in 2018 and is maintained and commercialized by DataCebo Inc. It is the largest and most comprehensive software system available for synthetic data generation, allowing enterprises to develop generative models for tabular data stores, assess synthetic data, and benchmark public techniques.
Products
- SDV CommunityA free platform to start your synthetic data journey, offering 5 data types, 5 basic constraints, 9 models, and native support for CSV and Excel files. Support is provided via the DataCebo Forum.
- SDV Enterprise BaseA low-cost entry point for enterprises with usage-based pricing. Includes 12+ models, 10+ data types, and dedicated support. AI Connectors are available as an add-on bundle.
- SDV BundlesAdd-on bundles for SDV Enterprise to customize plans with additional features. Examples include Synthetic data accelerators, Constraint Augmented Generation (CAG), XSynthesizers, Targeted Sampling, and Differential Privacy.
- CopulasA publicly available library within the SDV ecosystem that models and generates tabular data using classic statistical methods and multivariate copulas.
- CTGANA publicly available library within the SDV ecosystem that models and generates tabular data using Deep Learning, offering CTGAN and TVAE models.
- DeepEchoA publicly available library within the SDV ecosystem that models and generates time series data with a mix of classic statistical models and Deep Learning.
- RDTA publicly available library within the SDV ecosystem that discovers properties and transforms data for data science use, then reverses the transforms to reproduce realistic data.
- Synthetic Data VaultThe core system that generates synthetic data across single table, relational, and time series data, supporting multiple models and evaluations.
- AI ConnectorsA feature for SDV Enterprise users that connects to existing databases to automatically create highly accurate metadata and robust, referentially sound training datasets, regardless of the underlying database technology.
- Constraint Augmented Generation (CAG)A feature that allows users to apply complex logic between multiple tables, ensuring generated data follows business rules and remains valid by design. It includes automatic rule detection and validity enforcement.
- Differential PrivacyA feature that transforms and creates synthetic data with provable privacy guarantees built into model training, offering privacy-utility controls and privacy risk testing.
- SDGymA publicly available benchmarking system for synthetic data generation techniques, providing continuous evaluation of synthetic data generators across multiple dimensions like quality, speed, coverage, stability, and reliability.
Key capabilities
- Generates synthetic data with same formatting, statistical properties, and patterns as real data
- On-demand synthetic data generation
- Supports single table, multi-table, and sequential/time series data
- Publicly available, source-available libraries
- Multiple generative models (e.g., Copulas, CTGAN, TVAE, DeepEcho)
- Data transformation and property discovery (RDT)
- Metadata standard with statistical types and semantic meaning
- AI Connectors for automatic metadata creation and training data extraction from databases
- Constraint Augmented Generation (CAG) for enforcing business rules and validity
- Differential Privacy for provable privacy guarantees
- On-premise and low-compute deployment for SDV Enterprise
- Transparent and customizable AI models
- Automatic Schema Discovery for database integration
- End-to-End Data Movement for importing training data and exporting synthetic database clones
- Usage-based pricing with spending caps for SDV Enterprise
- Continuous benchmarking system (SDGym) for quality and speed evaluation
Use cases
- Test Software
- Expand Access to data
- Pilot New Products
- Augment Data
- Plan Scenarios
- Software testing
- ML development
- Detecting money laundering without compromising privacy
- Improving detection of homeowner insurance fraud
- Improving IT incident forecasting
- Developing generative models for tabular data stores
- Assessing synthetic data
- Benchmarking publicly available synthetic data techniques
- Creating highly accurate metadata with minimal manual work
- Getting high-quality training data from complex databases
- Training differentially private synthesizers
- Boosting AI model performance when real data is limited or imbalanced
AI approach
DataCebo provides the Synthetic Data Vault (SDV), an ecosystem of source-available software tools for generating synthetic data using AI models. These models are trained on real data to produce synthetic data with the same formatting, statistical properties, and patterns, without being linkable to real individuals. The company offers various models for single table, multi-table, and sequential/time series data, including classic statistical methods and deep learning approaches. SDV Enterprise allows users to train their own generative AI models for their specific datasets.
Tech named: Generative AI, Machine Learning, Deep Learning, multivariate copulas, CTGAN, TVAE, DeepEcho, Differential Privacy, Constraint Augmented Generation (CAG), XSynthesizers, Targeted Sampling, GANs
Industries served
- Banking
- Insurance
- Healthcare
- Software Engineering
What it says sets it apart
- Largest and most comprehensive software system for synthetic data generation
- Ecosystem of source-available tools
- Founded at MIT, tested and verified by enterprises globally
- Synthetic data maintains statistical properties and patterns of real data without linking to individuals
- On-premise and low-compute deployment for sensitive data
- Transparent and customizable AI models, not a black box
- AI Connectors automate metadata creation and training data extraction from complex databases
- Constraint Augmented Generation (CAG) ensures data validity and preserves hidden business rules
- Differential Privacy offers provable privacy guarantees with tunable controls
- Continuous benchmarking system (SDGym) evaluates quality and speed of synthetic data generators
- Supports single table, multi-table, and time series data synthesis
- Usage-based pricing for enterprise solutions
This profile was compiled from The Synthetic Data Vault's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.