Ask an IT manager whether they have a disaster recovery plan and the answer is almost always yes. Ask when it was last executed and the conversation changes.
This is not negligence. Testing recovery is genuinely disruptive, it needs a window nobody wants to give up, and there is a quiet incentive not to find out that it does not work. So the plan gets written, approved, filed, and reviewed annually in a meeting, and the organisation carries a risk it believes it has retired.
The distinction that matters: a plan is a document. A tested plan is a capability. Only one of those helps you at 03:00.
What a test has to produce to count
A recovery test is not a discussion and not a checklist review. To count, it produces four things:
- A system actually restored from the same backups production depends on, not a snapshot taken specially for the test.
- A measured duration, from decision-to-recover to service-verified-working, compared against the stated RTO.
- A verified data state, compared against the stated RPO: what was the most recent transaction that survived?
- A written record of what went wrong, because something always does, and what was changed as a result.
If an exercise produces none of these, it was a tabletop. Tabletops are worthwhile. They surface gaps in who decides and who calls whom, but they prove nothing about whether the technology works.
Start from the business, not the technology
The most common structural error is deriving recovery objectives from what the current infrastructure can do. That is backwards, and it guarantees the objectives are met, which is exactly why it feels comfortable.
Work the other way:
| Step | Question | Output |
|---|---|---|
| Business impact analysis | What breaks, for whom, and what does an hour of that cost? | Ranked process criticality |
| RTO per system | How long can this be unavailable before the impact is unacceptable? | Target restore time |
| RPO per system | How much data can we afford to lose, measured in time? | Backup/replication frequency |
| Architecture | What does meeting those two numbers actually require? | Design, and its cost |
Only at the last step does technology enter. And frequently the analysis reveals that a system everyone treated as critical can in fact be down for a day, while an unglamorous one (the identity provider, the certificate authority, the DNS) cannot be down for ten minutes. Dependencies are where recovery planning goes wrong, because they are invisible until you try.
How to test without taking production down
This is the objection that stops most organisations, and it is solvable.
Restore into isolation. Build an isolated network segment with no route to production, restore the system there from the real backup, and verify it. This validates the backup, the restore procedure, the runbook and the timing, without touching live service. It is the single highest-value exercise available and it can be run quarterly without a maintenance window.
What it does not validate is failover of live traffic, DNS cutover, or whether dependent systems follow. Those need a planned, announced exercise, but you should not attempt them until isolated restore has succeeded, or you will be debugging two classes of problem at once.
A sensible progression:
- Restore verification (isolated): quarterly, no outage
- Dependency exercise (isolated, several systems together): twice yearly
- Failover exercise (live, planned window): annually
- Unannounced element: once you are confident, because a rehearsed test measures the rehearsal
- Produces
- a measured duration, compared against the written objective
- Needs no outage
- the first two rungs, which is most of the value
- Order matters
- a live failover before an isolated restore debugs two problems at once
This diagram as text
- Isolated — no route to production, no outage
- 1 · Restore verification — quarterly — validates backup, procedure, runbook and timing
- 2 · Dependency exercise — twice yearly — several systems together
- Live — planned, announced, in a window
- 3 · Failover exercise — annually — DNS cutover and dependent systems
- 4 · An unannounced element — once you are confident — a rehearsed test measures the rehearsal
Relationships
- 1 · Restore verification → 2 · Dependency exercise
- 2 · Dependency exercise → 3 · Failover exercise — only after
- 3 · Failover exercise → 4 · An unannounced element
What tests always find
After enough of these, the findings repeat. Expect at least one of:
- The restore takes far longer than assumed. Backup software reports throughput under ideal conditions. Real restores contend for network and storage, and a four-hour RTO meets an eleven-hour restore.
- A dependency nobody documented. The application comes up and does not work, because it needs a licence server, an internal certificate authority or a database link that was not in scope.
- Credentials nobody holds. The person who set up the recovery account has left. The break-glass password is in a vault that requires the authentication system that is down.
- The backup was incomplete. It ran successfully every night and excluded a directory added eighteen months ago.
- The runbook is unexecutable. It was written by someone who already knew the answer, so it says “restore the database” without the steps.
- The documentation is on the system that is down. More common than it sounds.
Every one of these is cheap to fix when found in an exercise, and extremely expensive to discover during an incident.
Writing a runbook that works under pressure
The people executing recovery are stressed, possibly woken up, and possibly not the people who wrote the plan. That should shape how it is written:
- Exact commands, not descriptions of intent
- Explicit prerequisites at the top: access needed, systems that must be up first
- Decision points marked: who authorises, and what to do if they are unreachable
- Verification after every step, so failure is caught at the step that caused it
- No assumed knowledge. If it says “as usual”, it will fail
- A last-known-good date, so the reader knows whether to trust it
The test for a runbook: hand it to a competent engineer who has never seen the system, and have them execute it while the author stays silent. That is uncomfortable and it is the only honest test.
Evidence, because someone will ask
Regulators increasingly ask to see that recovery has been tested, not that a plan exists. ISO 22301 and ISO/IEC 27031 both frame this as an operating requirement rather than a documentation one.
Retain per exercise: the date and scope; who took part; what was restored and from which backup; the measured start and finish times; whether the RTO and RPO were met; every issue found; the remediation and its completion date; and the result of any re-test.
This pack is the entire point. It converts “we have a plan” into “here is a dated record of the plan working, and of what we fixed when it did not”.
A minimum viable programme
If you currently test nothing, do not start with a full failover. Start here:
- Pick your three most critical systems, chosen by business impact rather than by how interesting they are.
- Confirm their RTO and RPO in writing, agreed with the business.
- Run an isolated restore of one of them this quarter. Measure it. Write down what happened.
- Fix what it exposed.
- Repeat with the next system. Add a dependency exercise once individual restores are reliable.
Within a year you have a tested capability and an evidence trail, and you will almost certainly have found at least one thing that would have made a real incident far worse.
How STP approaches this
We design the architecture around agreed RTO and RPO figures, write runbooks intended to be executed by someone who did not author them, and run the exercises on a schedule, producing the evidence pack described above. Where a test exposes a gap, remediation and re-test are part of the engagement rather than a separate conversation.
More on disaster recovery and business continuity, or start a conversation.

