Layer 04 · Operate · Project or subscription
Disaster Recovery & Business Continuity
STP designs, documents and repeatedly tests disaster recovery: establishing recovery time and recovery point objectives per system, building the capability to meet them, and running scheduled recovery exercises that produce evidence you can show an auditor.

Disaster Recovery & Business Continuity
Almost every organisation has a disaster recovery plan. Very few have a tested one. An untested plan is a document, not a capability, and the difference only becomes visible on the worst day of the year.
We start from the business rather than the technology: which processes matter, how long each can be unavailable, and how much data loss is tolerable. Those answers become recovery time and recovery point objectives per system, which then determine the architecture: not the other way round.
Then we test, on a schedule, and write down what happened. Recovery exercises always surface something: an undocumented dependency, a credential nobody holds, a restore that takes four times longer than assumed. Finding those during an exercise is the entire point.
What is included
- Business impact analysis and process criticality mapping
- RTO and RPO definition per system, agreed with the business
- Recovery architecture design and implementation
- Backup design, immutability and offsite copies
- Replication and failover configuration
- Runbook authoring: step-by-step, executable under pressure
- Scheduled recovery testing with documented results
- Post-exercise remediation of what the test exposed
- Evidence pack suitable for auditors and regulators
- Annual plan review as the estate changes
What you get out of it
- Recovery objectives that are measured, not asserted
- Runbooks that a competent engineer can execute at 03:00
- Documented test evidence when a regulator asks for it
Questions
- What is the difference between RTO and RPO?
- Recovery Time Objective is how long a system may be unavailable before the impact is unacceptable. Recovery Point Objective is how much data you can afford to lose, measured as time: an RPO of one hour means losing at most the last hour of data. RTO drives the recovery architecture; RPO drives backup and replication frequency.
- How often should disaster recovery be tested?
- At least annually for a full exercise, and after any significant infrastructure change. Critical systems in regulated environments generally warrant more frequent partial tests. Regulators increasingly ask for evidence of testing, not just a plan document.
Disaster Recovery & Business Continuity
A live engineer calls you back within 1 hour.

