Two datacentres do not give you resilience. Two datacentres with the wrong boundary between them give you one datacentre with twice the failure surface, because an event in either site now propagates to both. The design question is not how to connect them. It is where to put the boundary so that a failure stops at it.
Each site is its own routing domain
The default we build to: every site runs its own layer-3 domain, with its own gateways and its own first-hop redundancy. Sites are joined by a routed interconnect, not by a bridge.
- Interconnect
- two paths, different physical routes
- Boundary
- routed, not bridged
- Overlay
- EVPN/VXLAN only where a layer 2 adjacency is genuinely required
- Witness
- a third location, never inside either datacentre
This diagram as text
- DC1 production
- Spine
- Spine
- Leaf
- Leaf
- Compute and storage
- DC2 production
- Spine
- Spine
- Leaf
- Leaf
- Compute and storage
- Recovery site same automation, different variables
- Reduced capacity
- Replicated data
- Promoted on declaration
Relationships
- Spine (DC1 production) → Leaf (DC1 production)
- Spine (DC1 production) → Leaf (DC1 production)
- Spine (DC1 production) → Leaf (DC1 production)
- Spine (DC1 production) → Leaf (DC1 production)
- Leaf (DC1 production) → Compute and storage (DC1 production)
- Leaf (DC1 production) → Compute and storage (DC1 production)
- Spine (DC2 production) → Leaf (DC2 production)
- Spine (DC2 production) → Leaf (DC2 production)
- Spine (DC2 production) → Leaf (DC2 production)
- Spine (DC2 production) → Leaf (DC2 production)
- Leaf (DC2 production) → Compute and storage (DC2 production)
- Leaf (DC2 production) → Compute and storage (DC2 production)
- Spine (DC1 production) ↔ Spine (DC2 production) — eBGP
- Compute and storage (DC1 production) → Replicated data — replication
Inside a site, an interior protocol — OSPF, or iBGP in a spine-and-leaf fabric — carries reachability. Between sites, eBGP, because it gives explicit control over what is advertised and accepted. That control is the point: it is what lets you fail a site out by withdrawing routes rather than by unplugging something.
Two interconnects, and they must not share a duct, a building entry or a provider. Diversity that exists only on the invoice is the most expensive kind of assumption.
Where layer 2 is genuinely required
Some clusters and some legacy applications need the same subnet in two places. Bridging the sites to achieve it is how a loop in one building takes out both.
EVPN with a VXLAN data plane is the right instrument. The underlay stays routed and stable; the overlay carries the layer-2 adjacency only where it is configured to, with MAC learning in the control plane rather than by flooding. Broadcast, unknown-unicast and multicast handling is explicit, and a storm in one tenant does not become a storm in the fabric.
The discipline that matters: extend an overlay only for the specific workloads that require it, and write down which ones and why. Every stretched segment is a shared failure domain you have chosen to accept.
Segmentation, and the firewall that enforces it
A flat network is the condition that turns one compromised workstation into an estate-wide incident. Segmentation is cheapest when it is designed before cable is pulled and most expensive when it is retrofitted into an occupied building.
What we separate as a baseline:
- Users from servers, with policy between them rather than a route.
- Operational technology and building systems — surveillance, access control, HVAC — from everything else. These devices are frequently unpatchable by design and run firmware nobody will ever update.
- Management of the infrastructure itself onto an out-of-band network, so that losing the data plane does not also lose the ability to fix it.
- Guest traffic onto its own path to the internet, with no route into the estate at all.
- Payment or regulated environments into their own zone with a documented and auditable boundary.
Inter-zone policy is enforced on a firewall and expressed as named objects and documented rules, not as a list of addresses someone added during an outage. Where the estate runs Fortinet or Sophos at the perimeter and between zones, the rule base is kept in version control alongside the rest of the infrastructure, so a change has an author and a reason.
Branches, and when SD-WAN earns its cost
A branch with one circuit and nothing latency-sensitive does not need SD-WAN. A VPN terminating on the datacentre firewall is adequate, cheaper and simpler to support.
SD-WAN earns its cost when either is true: branches have two or more links of different kinds and you want per-application steering between them, or voice and conferencing matter enough that brownouts need to be routed around rather than reported. The operational benefit people underestimate is measurement — knowing what each circuit actually delivered last week, per application, is what makes a conversation with a provider short.
Remote access is a separate question from site interconnect. For engineers and administrators we default to WireGuard with device identity and short-lived credentials, terminating into a management zone rather than into the user network.
The recovery site
The failure mode we are most often called to fix: a recovery site that was built by hand, two years ago, and has drifted from production ever since. It will not work, and nobody will know until it is needed.
Three rules make it real:
- Built from the same automation as production. Same Terraform and Ansible, different variables. If the recovery site is configured by a different method, it is a different system and your tested runbook is for the other one.
- Data replication with a measured lag. The replication lag is your recovery point objective, whatever the plan claims. It should be graphed and alerted on, because it drifts quietly.
- Promotion is rehearsed. Not a tabletop. An actual promotion of the replica, with the clock running, measured against the stated recovery time objective.
- Measured
- decision to service-verified
- Compared against
- the written recovery time objective
- Recorded
- every issue found, with a fix date
This diagram as text
- Declare — a person, at a recorded time
- Withdraw routes from the failed site
- Promote the replica
- DNS and routing converge — this interval is measured, not assumed
- Service verified — the clock stops here
- Capacity scaled up
- Fail back — a separate plan
Relationships
- Declare → Withdraw routes from the failed site
- Withdraw routes from the failed site → Promote the replica
- Promote the replica → DNS and routing converge
- DNS and routing converge → Service verified
- Service verified → Capacity scaled up
Failing back is its own exercise and its own runbook. Teams plan the failover and discover during the incident that nobody decided how to return, which is how a one-day outage becomes a two-week split estate.
What this produces for an auditor
ISO 22301 and ISO/IEC 27031 both ask for evidence that recovery has been exercised, not that a plan exists. Designed this way, the exercise produces that evidence as a by-product: the date and scope, who took part, the measured timings against the objectives, the issues found and their remediation dates.
How STP approaches this
We put the boundary at layer 3 and extend layer 2 only where an application forces it, specify two genuinely diverse interconnects, build the recovery site from the same automation as production, and run the failover with a clock before handover. Where an estate already exists, that exercise is the audit: it finds the shared duct, the stretched VLAN nobody remembers creating, and the replication lag that is twice the stated recovery point.
More on networking, disaster recovery and data centre build, or start a conversation.

