Contents
A bug shows up in production but not in test. The fix is the one everyone knows: restore a production backup into the test environment. As the day ends, nobody stops to ask how many real people's data now sits in that environment and who can reach it. This article is about the most common and least discussed personal data risk.
What Everyone Does and Nobody Discusses
Copying production data into non production environments is one of the most entrenched habits in enterprise software development. Industry surveys show that the large majority of businesses use live customer data for application testing. The reason is not negligence but practicality: real data reveals real bugs.
The same practicality is why it stays quiet: it never produces an incident. The backup is restored, the bug is found, the job is done. The environment stays up and the data inside it stays there. No alert fires and no ticket is opened.
A simple test
A customer has asked for their personal data to be erased. It was deleted from production. Was it deleted from the copies in non production environments? In most organisations the answer is no, and usually the possibility was never considered. An erasure request is the sharpest test of how invisible those copies are.
Why It Is a Problem: Three Separate Failures
This habit is not a single rule violation. It touches three separate principles of data protection law at once, and each produces its own finding.
Purpose limitation. Personal data cannot be processed in a way incompatible with the purpose it was collected for. Customer data was collected to deliver a service, not to test software. Testing is a different purpose, and the distinction is usually never drawn, because nobody thinks of copying data into a test environment as a processing activity. By definition it is one.
Data minimisation. Data processed must be limited to what the purpose requires. Restoring a backup is by definition the opposite: every table, every column, every row. Investigating one payment error does not require the identity and contact details of the entire customer base.
Unequal protection. Production is defended: access is restricted, there is encryption, there is monitoring, its backups are audited. A test environment is not required to have any of that and usually does not. The same data has been moved under much weaker protection. From an attacker's point of view the easiest target is not the best defended environment but the weakest one holding the same data.
How Many Copies Do You Have, Really?
In most organisations the answer is larger than expected, because copies do not only live in official test environments.
There is an acceptance test environment. There is a development environment. There is a third environment set aside for performance testing. There is a backup file a developer pulled onto their laptop. There is a spreadsheet exported for analysis. There is a sample data set sent to a vendor for a bug investigation. And there is usually a backup file with a date in its name, forgotten on some server's disk.
Every item on that list is a separate copy carrying the same personal data, and every one is a separate exposure surface. What they have in common is that none of them appears in the inventory. When a personal data inventory is compiled, systems get listed; copies of systems do not.
The Fair Objection and an Honest Answer
At this point the objection from development and test teams is fair and deserves to be taken seriously: synthetic data does not find real bugs.
Production data genuinely carries three things that synthetic data does not. Volume: a query that runs on ten rows may not run on ten million. Distribution: in real data one customer has thousands of transactions and another has none, and that skew changes plan selection. Dirtiness: the empty fields, inconsistent formats and unexpected characters accumulated over years exist only in real data, and most bugs come from exactly there.
So the right answer is not to stop using production data. The right answer is to keep those three properties while protecting the fields that identify a person. The two can be separated, because what surfaces the bug is the shape and volume of the data, not the customer's name.
A Three Level Answer
Organisations that try to solve this in one step usually never start. Working in levels produces benefit on day one.
| Level | What you do | What it solves and what it does not |
|---|---|---|
| 1. Visibility | An inventory of every environment and file holding production data, with an owner, a date and a stated reason for each copy | It does not reduce risk but makes it measurable. Skip it and the later levels have nowhere to be applied. |
| 2. Masking | Replacing the fields that identify a person while preserving volume, distribution and format | This is the step that closes the real risk. Deciding which column to mask is harder than the masking itself. |
| 3. Subsetting and lifetime | A consistent subset instead of everything, and an expiry date on every copy | Meets the minimisation and retention expectations. It needs care: if referential integrity is not preserved, the tests break. |
The real difficulty at level two is not technical. Whether a column should be masked cannot be decided from its name alone; not every eleven digit number is an identity number, and a column whose name implies nothing may still carry personal data. The decision requires the name and the content to be assessed together.
What Can Be Done Today
Count the copies. How many environments and how many files hold production data? Finding the number takes less than a week and it usually comes out at twice the expectation.
Simulate an erasure request. How long would it take to remove one person's data from every environment? If you cannot answer now, you will not be able to answer when a real request arrives.
Put the copying on the record. Any data leaving production, whether by restore or by export, should rest on a request, and the answers to who asked, on what stated basis and into which environment should be recorded. That alone does not prevent the copy but it makes it visible.
Set a lifetime. Give every copy an expiry date. A copy without a date is a permanent copy.
SQL Change Guard contributes directly to the second and third rows of that table: a request to extract production data goes through approval, sensitive columns are masked in the result set, and who received what, on what basis and when is recorded. It does not stop a backup being restored into a test environment; that decision stays with you. Its contribution is that data leaving production stops being invisible.
Frequently Asked Questions
The test environment is inside our own network. Is there still a risk?
Yes, because risk does not only come from outside. Test environments are usually open to a wider group: developers, testers, sometimes vendor consultants. If data that five people can see in production is visible to fifty people in test, the exposure surface has grown tenfold even though the data is identical. Test environments also tend to lack the encryption, monitoring and backup controls that production has.
Can performance testing be done on masked data?
Yes, when the masking is done properly. Performance is driven by row counts, value distribution, index selectivity and field lengths rather than the contents of the fields. Replacing a name with another string of the same length does not change the query plan. The point to watch is that masking must not destroy selectivity: giving every customer the same name changes how the index behaves and misleads the test.
We have worked this way for years without a problem. Why now?
Because the nature of this risk is that it gives no signal until it materialises. Not having had a problem shows that the control has not been tested, not that it works. Two things trigger it: a breach notification, where the number of affected people is counted across every copy, or an auditor asking for the personal data inventory and then asking about non production environments. Neither gives warning, and neither can be fixed retrospectively.
Are anonymisation and masking the same thing?
No, and the difference matters legally. Truly anonymised data is no longer personal data, but anonymity requires that the data cannot be linked back to a person by any means, and that is harder than it sounds. Column level masking usually does not clear that bar: a combination of unmasked fields such as date of birth, postcode and gender can make a person identifiable again. So while masked test data is a major security gain, it should still be treated as personal data for compliance purposes.
Does SQL Change Guard mask our test database?
No, and it is right to be clear about the scope. Tools that copy an entire database and bulk transform the personal fields inside it are a separate category, known as test data management. SQL Change Guard does not do that. What it does is govern requests to extract data from the production database: a request is opened with a stated reason, it goes through approval, sensitive columns are masked in the result set that comes back, and who received the data is recorded. Its surface is people taking data out of production, not bulk copying between environments. The two needs complement each other rather than replace each other, and conflating them in a meeting creates the wrong expectation.
Where should we start?
With the inventory. List the environments and files holding production data and put an owner and a date next to each. The length of that list starts the rest of the discussion on its own. Then pick one environment and work through masking; trying to fix every environment at once is why most of these programmes stall before they begin.
Let us look at your own columns
We will show live how the decision on which column to mask is made, and how data leaving production is put on the record.
Book a Demo →