A copy without a test is only an assumption
Green ticks in a backup console tell you one thing. The job ran and finished without an error. They do not tell you whether a working service will come back from that copy, one that a person can log in to and get their work done.
The gap between those two is wide, and it always shows up at the worst possible moment. On the day of the failure, with people waiting for the system and customers on the phone. A copy you have never restored is not protection. It is an assumption.
The reason for that gap is mundane. Backup software reports on its own work, meaning whether it managed to read the data and write it into a copy. It does not judge whether the copy holds everything needed to stand a service up from scratch, because it has no way of knowing. Only the people who run that service know, and only once they have actually tried to restore it.
Even the industry rules of thumb say so. The classic 3-2-1 rule calls for three copies of the data on two kinds of media, with one copy kept offsite. Its newer version, 3-2-1-1-0, adds one copy cut off from the network plus zero errors, where that zero means copies verified by regular restore testing. Testing is not an extra on top of backup. It is part of it.
Restoring a file and restoring a whole service
These are two completely different exercises, and they are very often called by the same name.
Restoring a file proves that the medium is alive, the backup catalogue is consistent and the data can be pulled out. It takes a few minutes and one person is enough. A useful test, but a very narrow one.
Restoring a service proves something else entirely, because it checks the whole chain. The machine, the operating system, the configuration, the database, accounts and permissions, the licence, the addressing, the certificate and the links to other systems. It takes hours and it needs several people. Only this test answers the question the board really asks, which is whether the company gets back to work.
Most companies that say their backup is tested have tested files.
What a restore test looks like
A restore test is a planned exercise with an agreed scope, not an attempt to bring everything back at once. It has four parts and none of them can be skipped.
What gets restored. You pick whatever stops the company fastest when it is missing. Usually that is one system with a database, the mail service or a file server, not the entire server room.
Where it gets restored. Into an isolated environment, meaning a separate network with no route to production. This is not a formality. A restored server does not know it is a copy. It will start collecting mail, writing to the database, replicating and sending data outside if we let it.
How the time is measured and who confirms the result. The clock runs from the decision to restore until the moment somebody is really working in the system, not from the start of the job to the end of the data copy. Completeness is confirmed by the person who uses that system every day, not by an administrator. The accountant opens a document from a specific date. The salesperson looks for their own quote. An administrator will see that the service is running, but will not see that the last two days of data are missing.
What the first test usually uncovers
A first test almost never runs smoothly, and that is where its whole value lies. The things that come up again and again:
- the service starts but waits for a domain controller, a licence server or another system that nobody put on the restore list,
- the password to a service account was set once, during the original rollout, by a person who no longer works here, so the copy exists and the password does not,
- the database copy contains no transaction logs, so the database comes back as of the last full dump rather than the moment before the failure,
- the machine backup does not cover the network configuration, so the machine runs but has no addressing, no VLAN, no firewall rules and no published service, and nobody can reach it,
- nobody knows the order in which the services have to come up, so they block each other,
- a certificate has expired or a licence is tied to the old hardware, so the service starts and immediately refuses to work.
None of these is visible in a backup console. All of them are visible in the first hour of a test.
What the first test teaches is worth more than its result, because every one of these things can be fixed in advance. A missing system gets added to the backup scope. Service account passwords move into the company password manager. Transaction log backups simply get switched on. Addressing, rules and the start order get written down once. Fixing this before a failure costs a few hours of work. The same fix during a failure costs the whole company a day.
Too slow is also a failed test
Sometimes a test succeeds and is still a failure. The data is complete, the service runs, but the restore took several times longer than anyone in the company assumed.
That is why time is a result of the test, exactly like the completeness of the data. And that is why it has to be measured before the failure, not during it. Once the real number is known, there are three ways out: change how the copies are stored, prepare a faster restore path for the most important systems, or accept that time consciously and tell people in the company about it. Any of those beats learning the truth on a Friday evening.
Measuring the time also settles the order. Since not everything comes back at once, somebody has to decide what comes back first. That is a business decision, not a technical one, and it is better made calmly.
Somebody outside the IT team should know that measured time as well. The person who answers for the company reads a sentence about a restore taking half a day very differently from a chart in a backup console. And only that person can say whether half a day is acceptable, or whether it is worth paying to make it shorter.
How often to test
A single test buys peace of mind for one day. Environments change, so a result from last year describes last year.
A sensible rhythm runs at two speeds. The small test, restoring a file or a single machine, is done often and quickly. The big test, restoring a full service with a user confirming the data, is done less often but with a planned date and real preparation.
Whatever the rhythm, some events mean the test has to be repeated:
- a new system or a new machine added to the environment,
- a version change in the operating system, the database or the backup software,
- a change to where copies are stored or how long they are kept,
- a change to addressing, the firewall or the way a service is published,
- a change to who has access to the backup console.
The rule is simple. If something changed that would affect a restore, then something changed that invalidates the previous test.
What to record in the report
A test with no record is an anecdote. The report does not have to be long, but it has to be repeatable, so that two tests a year apart can be compared.
What it should contain:
- the date of the test and the people who took part,
- what was restored and from which copy, together with the date of that copy,
- where it was restored, meaning a description of the isolated environment,
- the start time, the finish time and the measured time until the system was usable,
- who confirmed that the data was complete, and what that confirmation consisted of,
- what did not work, what was done about it and what was changed in the configuration,
- the date and scope of the next test.
This document earns its keep in three situations: during an audit, in a conversation with an insurer, and when changing IT provider, when somebody asks for the date of the last successful restore.
Its real value, though, shows up at the second and third test. That is when you can see whether restore times are growing along with the volume of data, whether the notes from last time were acted on, and whether the scope still covers the systems added during the year. One report records an event. Three reports show which way the company is heading.
One question worth asking
If you want to know what your backup really looks like, one question to the person or the company running it is enough. When did we last restore data from a copy, and what exactly did we restore?
A sentence about copies running every night without errors is not an answer to that question. The answer is a date, a scope and the name of the person who confirmed that the data was complete. If nobody can give you that, it does not yet mean the copies are bad. It only means that nobody knows.
We run restore tests for our clients, because without them keeping backups is nothing more than a promise. If you want to check your own copies and do not know where to start, call +48 662 036 615 or write to [email protected].
