Two questions that come before any purchase
A conversation about protecting a company against failure almost always starts with tools. Someone asks about backup software, someone else about a storage array or the cloud. That is the wrong order. Without two business decisions there is no way to judge whether any tool is the right one, or how much it should cost.
It is worth separating three things that usually blur into one. A backup is a copy of data and it saves you after a deleted file or a failed disk. Disaster recovery is getting back to work after losing a whole environment, for example to a fire, to flooding or to an attack that encrypts the servers. A continuity plan is wider still, because it also covers how the company works during the time the systems are not back.
Those decisions are answers to two questions. How many hours the company can run without a particular system, and how much recent work it can afford to lose. The first number is called RTO, the second RPO. The abbreviations sound technical, but the answers do not come from an administrator. They come from the person who knows what happens in the company when the orders stop coming in.
Everything else follows from those two numbers: how often to make copies, where to keep them, whether a standby environment is needed and what budget makes sense here. Doing it the other way round ends either with an expensive solution that still misses expectations, or a cheap one that everybody assumes is enough.
RTO, or how many hours the company can run without a system
RTO stands for Recovery Time Objective. The American NIST, in its contingency planning guide, describes it as the length of time a system can spend in the recovery phase before it starts to harm how the organisation operates. Put simply, it is how many hours a given system can be down and the company can still take it.
The most common mistake is looking for one number for the whole company. RTO is set separately for each system. A shop without its sales system does not work at all. An archive of documents from previous years can be unavailable for two days and nobody will notice. Those two cases cannot get the same protection, because for one of them it would be money thrown away.
The second trap is counting only the restore itself. RTO does not start when somebody clicks restore. It starts at the moment of failure and covers everything on the way: noticing the problem, deciding to activate the plan, getting to the hardware, restoring, checking that the data is correct and letting people back to work. If nobody is on duty at the weekend, the time until Monday has to be counted into the RTO as well.
It is also worth remembering that a system never works alone. An application needs its database, authentication, the network and often an exchange of data with finance or with the warehouse. RTO covers that whole chain, so the real number is the recovery time of the weakest link, not the time it takes to restore the application itself.
RPO, or how much work you can afford to lose
RPO stands for Recovery Point Objective. NIST defines it as the point in time to which data must be recovered after an outage. In practice it answers one question. How far back in time does the company go after a failure.
A clock shows it best. If a copy is made once a day at two in the morning and the server stops at four in the afternoon, fourteen hours of work disappear. This is not abstract data. It is the orders taken that day, the invoices issued, the notes from calls with customers and the changes in files the whole team worked on.
Some of that work can be recreated from memory and from paper. Some of it never comes back, because nobody remembers what exactly a customer said on the phone. That is why RPO is worth setting together with the people who enter the data, not only with the board. They are the ones who know how long it takes to type one day of work again.
RPO, like RTO, differs between systems. The order database and the folder with photos from the company picnic do not need copies at the same frequency.
A simple example that shows the cost
Take a company that receives orders through its own system. Twenty people work in it, and every order leads to a shipment and an invoice. The board decides that RPO is one hour and RTO is four hours. Look at what follows from that.
RPO drives how often copies are made and how. A copy once a day will not do, because it allows a whole day to be lost. A copy every hour means a different way of working with the database, more disk space and more load on the link if the data leaves the building. Going down to a few minutes usually means replication, that is a second copy of the system kept current all the time. Each of those three options costs differently and each takes different work to maintain.
RTO drives what the data will be started on. Four hours is not enough time to wait for a server to be delivered and install a system from scratch. So there has to be somewhere to start, meaning spare capacity on an existing hypervisor, a standby machine or an environment in the cloud prepared in advance. This is the part of the bill nobody expects, because it sounds like backup while it is plain hardware and space.
Change either number and both the solution and the price change with it. An RTO of one day allows a calm restore onto new hardware. An RTO of one hour almost always means a second environment ready to take over. So the honest question is not which solution is best, but which number the company actually needs.
The most common mistake, an RTO set without doing the maths
Asked how quickly a system has to come back, almost everybody answers immediately. That is understandable and completely useless, because immediately is the most expensive answer there is. It is better to work out one thing first. What an hour of downtime costs the company.
A few items go into that calculation:
- the number of people who cannot work during that time, and the cost of their hour
- what cannot be sold, shipped from the warehouse or invoiced while the system is down
- the work that will have to be caught up later, overtime included
- the deadlines and contractual penalties that are missed in the meantime
- the cost of the calls with customers who have been told the system is down
The result does not have to be exact. A fair approximation is enough to set it against the cost of a solution. That is the point of a business impact analysis, which NIST describes as the process of analysing operational functions and the effect a disruption might have on them.
The outcome often surprises people in the other direction. For some systems it turns out that two days of downtime would do no real harm. That is very good news, because it frees money for the one system that has to be back within an hour.
A backup is not yet a continuity plan
A backup is a copy of data. A continuity plan is the ability to work after a failure. The difference between the two is a handful of things that backup software on its own will not deliver.
At a minimum, these have to be added to the copies:
- an environment the data can actually be started on, meaning spare capacity, standby hardware or a prepared place in the cloud
- an order in which services come back, because an application restored before its database and before authentication simply will not start
- people and a procedure, meaning who declares a failure, who makes the decision and who calls whom
- connectivity and remote access, if people are supposed to work from somewhere other than usual
- a test, meaning a trial restore with the time measured, not just a report that a copy was made
This list is not invented. The NIST contingency planning guide builds a plan out of exactly these parts, from the business impact analysis, through the choice of an alternate site and the split of responsibilities, to testing and training the people who are meant to carry the plan out.
The last point matters most and is skipped most often. A copy that has never been restored is only an assumption. A trial restore is the one thing that shows how long it really takes. It is also where the gaps surface most easily, the ones no backup report will ever show, such as a dependency on another system or a password nobody remembers any more.
How to check whether the plan really exists
A few questions show the state of things quickly. They can be put to an internal IT team or to the company that looks after the infrastructure:
- when data was last restored from a copy, and how long it took
- what it was restored onto, because putting a file back on the same server proves nothing
- what RTO and RPO each system has that the company cannot work without
- whether a copy exists that cannot be reached from the company network
- who decides to activate the plan at five in the morning on a Sunday
If the answers are that copies are running and that there is monitoring for them, the company has a backup. It does not have a plan yet. That is not a complaint about anyone, it is simply information about how much work is left.
At C4PL we run real restore tests, we do not stop at deploying and maintaining the copies. The scope of support, including round the clock cover, is agreed in the contract, because it decides what RTO is realistic at night and at the weekend.
Where to start
Not with tools. First comes a list of the systems the company cannot work without. It is usually shorter than people expect. Next to each one go two numbers and one sentence about what happens when that system is down.
Filling in that table does not need an audit. One meeting with the people responsible for sales, production, the warehouse and finance is enough. The technical part starts only after that, and it is much simpler then, because it is clear what has to be met.
Once both numbers are agreed, they are worth writing down. As long as RTO and RPO live only in conversation, they are a wish. Written into the contract with the company that maintains the infrastructure they become a commitment, and only then can the response time and the scope of duty cover be matched to them.
If you would like to go through it with someone who runs these restore tests in practice, call +48 662 036 615 or write to [email protected]. A conversation about two numbers takes about fifteen minutes and usually shows straight away where the biggest gap is.
