Synthetic Test Data with AI

The article explores how AI-generated synthetic data can create realistic test environments without copying sensitive production data. It shows how preserving real-world relationships, edge cases, and business rules can make testing more effective while keeping customer data out of development environments.

Tomasz Olszowy

•

5

min.

We have a test environment and we have an application.

We even have a pretty decent set of automated tests.

There is just one thing missing. Good data.

The easiest solution would be to copy the production database. After all, it contains exactly the kind of data our application deals with every day.

Except that it also contains names, addresses, phone numbers, customer details, transaction histories and quite a few other things we definitely don't want to see on a developer's laptop.

So we could generate random data instead.

John Smith. john@example.com. Order number 12345. Amount: PLN 199.99.

Problem solved? … Not necessarily.

Random data isn't always realistic.

Let's say we're testing an e-commerce platform. We need 50,000 customers, 200,000 orders, payments, invoices, returns and a history of status changes.

A data generator will have no trouble creating 50,000 names and email addresses.

It can generate prices too:

  • 19.99

  • 128.50

  • 847.12

  • 42.00

Every one of those values looks perfectly fine.

But real data has one inconvenient characteristic. Things are connected.

  • a customer has orders,

  • every order contains products,

  • the sum of the line items has to match the order total,

  • a discount changes the price,

  • a return shouldn't happen before the purchase,

  • and the payment date probably shouldn't be three weeks before the order was created.

Then there are the less obvious relationships:

  • payment methods differ between countries,

  • not every product is available in every market,

  • new customers may behave differently from people who have been using the platform for five years.

So we can end up with a million records that look perfectly reasonable on their own, but together describe a world that could never actually exist.

And then we use that world to test our application.

Relationships are the interesting part!

Imagine a simple customer profile:

Customer: C-18472

Account age: 4 years

Orders: 37

Returns: 3

Last order: 4 days ago

Payment type: card

Country: Germany

The record itself isn't particularly interesting.

What happens around it is much more interesting.

A customer who has been buying regularly for four years will probably have a different order history from an account created five minutes ago. Some products tend to be bought together. Some customers shop mostly at weekends. Traffic in December may look completely different from traffic in February.

With a traditional generator, we have to describe these relationships ourselves.

If there are only a few rules, that's not a problem.

But what if there are hundreds?

That's where a regular data generator starts to show its limitations. And this is where Synthetic Test Data combined with AI gets really interesting.

A model can analyse the relationships present in the source data and then try to reproduce them in new records – without copying individual customers.

We don't need another random John from production.

We need someone who behaves statistically like him but has never actually existed.

Realistic doesn't have to mean real.

That's an important distinction.

Anonymising production data

Anonymising production data still means working with data created by real events. We remove or transform the elements that could be used to identify a particular person.

Synthetic data is created from scratch.

A model might learn, for example, that 18% of orders come from new customers, around 7% result in a return, and certain product categories appear together more often than others.

It can then generate a completely new dataset:

Customers: 50,000

Orders: 213,482

Returns: 14,771

Payments: 211,936

Realistic-looking numbers alone don't give us much.

If the distribution of order values, seasonality, relationships between tables or unusual cases don't resemble reality, all we've built is a very sophisticated fiction generator.

And then there is referential integrity.

If we generated a payment for order ORD-28441, that order has to exist. If the order belongs to customer C-18472, that customer also needs to exist in the correct table.

Sounds trivial.

Unfortunately, once we have dozens of tables and millions of records, it stops being trivial.

And the most valuable cases may be the ones that almost never happen.

There is another problem with copying production data.

Production mostly shows us what happens frequently.

Tests should also care about what happens rarely.

  • What about a customer with 400 orders?

  • What about an invoice containing 200 line items?

  • What about a payment accepted one second before a timeout?

  • Or an order involving a discount, a partial refund and a currency change at the same time?

Maybe production contains two such cases.

For testing, we can generate two thousand.

There is no reason to reproduce production proportions exactly. In test data, we can deliberately increase the number of rare cases.

Not because we want to create a realistic copy of an average Tuesday, but because we want to see what the application does on a very unusual Tuesday.

This can be especially useful in regression testing. A bug that occurred only once in production doesn't have to remain a one-off case. We can take its characteristics and generate an entire family of similar test scenarios.

One bug then gives us not one test, but enough material for dozens of new ones.

Different data for the developer, different data for the pipeline

We can decide both the size and the character of the dataset.

Does a developer need a small dataset for local testing? Generate 10,000 records.

Is the CI pipeline running integration tests? Increase it to 500,000.

Preparing a performance test? Generate 100 million transactions.

We don't have to wait for production to eventually accumulate that volume, and we don't have to copy a huge database just to see how one service behaves.

The dataset can also be reproducible.

A pipeline can run a test against a particular dataset, find a regression, and the developer can later recreate exactly the same scenario locally.

We can also prepare data for a specific change.

Are we modifying the returns mechanism?

Then we don't need 50 million ordinary orders. A hundred thousand cases involving returns, partial payments, discounts and adjustments will be much more useful.

With this approach, we no longer maintain one dataset for everything. We generate data for the test we actually want to run.

AI is also perfectly capable of generating very convincing nonsense.

That's why we need one more thing – quality control for the data itself.

Just because generated data looks realistic doesn't mean it's correct.

A model may preserve the age distribution of customers but lose the relationship between country and payment method. It may reproduce order values correctly while generating an unrealistic number of returns.

So the synthetic dataset needs to be tested too.

Among other things, I would compare:

  • distributions,

  • correlations,

  • cardinality,

  • boundary values,

  • relationships between tables.

I would also check the business rules.

It would be rather unfortunate to discover after a week of testing that half of our synthetic customers made a purchase before creating their accounts.

And then I would ask one more question:

Can the source data somehow be reconstructed from this dataset?

The new data needs to resemble the source closely enough for the tests to make sense. But it must not make it possible to reconstruct specific records the model learned from.

The goal isn't to build a very "clever" copy of production.

How much does bad test data cost?

That's harder to calculate than pipeline execution time.

But let's assume that a team of ten developers loses, on average, just 15 minutes a day fixing test data, manually creating missing records or recreating the required application state.

That's:

10 × 15 minutes = 150 minutes a day

Across five working days:

12.5 hours a week.

And that doesn't include QA time or situations where a test passes simply because the dataset didn't contain the case that later appeared in production.

Of course, this is an example, not a benchmark.

But it illustrates something important: test data quality is part of software development quality.

Poor data can make even very good tests check the wrong things.

The best test data doesn't have to belong to anyone

For years, we've tried to make test environments resemble production as closely as possible.

That approach still makes sense.

But it doesn't mean we have to copy the same customers, transactions and histories.

Synthetic data allows us to separate those two things. The data can behave like production data without being the data of real customers.

If we can preserve relationships, distributions, business rules and unusual cases, we can build an environment that behaves like production without copying production one-to-one.

And while we're at it, we can generate 2,000 customers who made a partial refund exactly one second before a timeout.

None of them will be real.

The problems we find thanks to them definitely will be.

© 2026 QualityMinds, All rights reserved

© 2026 QualityMinds, All rights reserved

© 2026 QualityMinds, All rights reserved