IT Brief Australia - Technology news for CIOs & IT decision-makers
Australia
DataCebo launches SDV 2.0 for synthetic enterprise data

DataCebo launches SDV 2.0 for synthetic enterprise data

Sat, 3rd Oct 2026 (Today)
Raphael Veloso
RAPHAEL VELOSO News Editor

DataCebo has released SDV 2.0, software for building generative relational models from an organisation's own databases. The Cambridge company designed the product for enterprise data held in relational systems.

The software lets companies create a model of their databases within their own computing environment rather than moving production data elsewhere. That model can then generate synthetic data and support the training and evaluation of AI agents.

The launch reflects a growing focus on keeping control of operational data as companies expand their use of generative AI. While large language models have broadened AI use for text, images and code, many businesses still rely on internal databases for the customer records, transactions and business rules that define how they operate.

DataCebo describes those internal stores as a source of company-specific intelligence that is difficult to reproduce through one-off projects. In many organisations, software teams, data engineers and machine-learning teams create separate copies or slices of production data for testing, development or model training, often rebuilding the same underlying logic.

SDV 2.0 is aimed at that problem by modelling a database as a connected system across multiple tables. The software learns relationships, structure and constraints from a representative subset of data and can usually be trained in minutes to an hour, often on a standard central processing unit, according to the company.

Automation focus

One of the main changes in the new release is greater automation in setting up models. DataCebo said early enterprise users of its commercial SDV software found that large databases often lacked complete documentation for keys, formats, relationships and business rules, leaving teams to configure much of the model-building process by hand.

The latest version is designed to automate parts of that work. It can detect schemas and relationships, including primary, foreign and composite keys; identify formats and context in values; detect and enforce business constraints; tune the model; and generate task-specific datasets, including data for rare events and edge cases.

The product connects directly to Oracle, SQL Server, BigQuery, Spanner and AlloyDB, and runs within a customer's own environment. That approach is likely to appeal to companies in tightly regulated sectors where data movement and exposure remain sensitive issues.

Users can build a model from a representative subset of data rather than a full production estate, DataCebo said. It added that the resulting system can simulate behaviours and provide enterprise context to analytics tools, models and agents while reducing reliance on production data.

A practical concern behind the launch is the cost and operational burden of production-dependent workflows. Creating masked copies for developers, staging pre-production environments and training separate models on repeated sets of records can be slow and fragmented, particularly in businesses with hundreds of tables and thousands of columns.

"Companies have spent decades building intelligence in their databases, but every team has had to reconstruct a small piece for each new use," said Kalyan Veeramachaneni, Chief Executive Officer of DataCebo.

"A generative relational model lets them capture that intelligence once and put it to work across the business," Veeramachaneni said.

Customer examples

DataCebo cited customer deployments to illustrate the software's use in synthetic data generation. ING Belgium used the company's tools to generate 10,000 synthetic payments in two minutes, delivering 100 times the test coverage in less than one-tenth of the time previously required, according to DataCebo.

Another customer, Epiconcept, created a synthetic database in 55 minutes and used it to identify optimisations that improved query performance by 105 times, the company said.

The business traces its technology to MIT's Data to AI Lab, where Veeramachaneni, also a Principal Research Scientist at MIT, and Co-Founder and Chief Product Officer Neha Patki worked on synthetic data tools for tabular and relational systems. Their work included the Synthetic Data Vault and SDMetrics, a framework used to assess synthetic data quality.

DataCebo said SDV has recorded more than 18 million downloads, been cited in more than 5,000 research papers and is used by more than 30,000 data scientists. Adoption spans financial services, healthcare, life sciences and consumer goods, sectors where governance requirements tend to shape how production data can be used.

SDV 2.0 is available with self-service pricing starting at USD $500 per month for unlimited tables.