Active-active gets discussed as though it were the top of a ladder, with single-region at the bottom and everyone climbing. It is not a ladder. It is one answer to a specific business question, and the question has to be asked first.
How much does an outage cost, and is the business willing to carry that risk?
That is not an architecture question. It belongs to the people who own the revenue and the regulatory exposure, and the architecture follows from their answer rather than the other way round.
Fault-intolerant is a real category
Some businesses genuinely cannot absorb a fault. A payments platform during a settlement window. A trading venue in session. A clinical system during theatre hours. A licensed betting operator during a major event, where the outage is measured in lost revenue per minute and the licence conditions have something to say about it.
For those, the ordinary resilience answer has a gap in it. Even a well-practised failover implies a period where requests are failing, and if the business cannot accept that period, no amount of rehearsing shortens it to zero.
Active-active is one of the ways to close that gap. It removes the cutover rather than making it faster, because both sides are already serving.
It is not the only way. Depending on the failure you are defending against, the same requirement can be met by redundancy inside one region, by a read path that degrades gracefully while writes are unavailable, or by a queue that accepts work and settles it when the primary returns. The right question is always which failure you are buying protection from, and what the business does during it.
What it costs, stated plainly
Choosing it should be done with the cost in view, because the cost is mostly not infrastructure.
Two estates to patch, two sets of alerts to tune, two places a request can be during an investigation. Every future change designed and tested twice. And a data layer that now has to answer what happens when both sides accept a write for the same record, which is a design problem with no free answer: you pick conflict resolution, or you partition writes by key, or you accept a single write region and lose some of the benefit.
Those are all solvable. They are solvable at a cost, and the cost should be weighed against the cost of the outage it prevents rather than against an intuition about what a mature architecture looks like.
How Netflix reached their answer
Netflix’s published reasoning is a good worked example of the business question being answered honestly rather than a template to copy.
Their position is not that active-active is superior. It is that at their request volume a standby region may not scale fast enough when an entire region’s traffic arrives at once, that a rarely-exercised failover is a procedure nobody has recent practice in, and that any region-level event is therefore a global emergency. Given those facts, removing the cutover was worth what it cost them.
The reasoning transfers. The inputs are theirs. An organisation that runs the same reasoning on its own numbers will often reach a different answer, and that is the reasoning working correctly rather than failing.
- The left removes
- the cutover itself
- The right accepts
- a cutover, and proves how long it takes
- Chosen by
- what the business can absorb, not by scale
This diagram as text
- Active-active no cutover
- Region A — serving
- Region B — serving
- Conflict resolution — decided, not discovered
- Two alert estates — every change designed twice
- Multi-AZ with an exercised recovery
- AZ a
- AZ b
- AZ c
- One estate, one change path
- Restore exercised quarterly — timed against the written RTO
Relationships
- Region A ↔ Region B — both live
- One estate, one change path → Restore exercised quarterly
The failure nobody puts on a slide
The architecture that actually fails most often is neither of these. It is the half-built second region: provisioned eighteen months ago during a resilience push, configured by hand, drifted ever since, and never asked to serve a request.
It carries the full cost of active-active — two estates to patch, two bills, two sets of certificates to renew — and delivers none of the benefit, because the first time anyone discovers what is missing is during the incident it was bought for.
If you have one of these, you have a decision to make, and both answers are respectable. Commit to it, build it from the same automation as production, and exercise failover on a schedule. Or decommission it and put the money into proving the single-region design recovers.
What is not respectable is leaving it there because removing it would look like a step backwards.
How to decide, in order
One. Establish what an outage costs, per hour, with the business. This is the input everything else is derived from, and it is the step most often skipped.
Two. Write down the recovery time and recovery point objectives that follow from it, agreed with the business rather than derived from what the infrastructure already does. Deriving them from current capability guarantees they are met and tells you nothing.
Three. Measure what recovery actually takes today, with a clock, into an isolated environment. Most teams have never measured this and are surprised in both directions.
Four. Compare. If the measured recovery fits inside the objective, the work is to keep it that fast, which means it stays automated and it stays exercised. If it does not fit, or if the business cannot accept any cutover at all, you have a real case for active-active. Build it from the same modules as production and treat the first exercise as part of the project rather than as a follow-up that never gets scheduled.
The part that applies whatever you choose
One thing is constant across every one of these architectures, and it is the part least often copied: the organisations that get the most out of them exercise them. The design is downstream of a practice of regularly proving the thing works.
A business with a single region that restores on a schedule and times the result has more real availability than one with two regions and a plan. Active-active that nobody has ever tested by taking a region out is two estates and a hope. Whichever answer the business requirement produces, the exercise is what turns it into a capability.
How STP approaches this
We start from the two numbers and measure the rebuild before recommending an architecture, because the measurement usually settles the argument. Where a second region is justified we build it from the same automation as production and schedule the first exercise into the project. Where one already exists and has never taken traffic, we say so plainly and let the client decide whether to commit to it or remove it.
More on disaster recovery and cloud, hybrid and multi-cloud, or start a conversation.

