Chaos engineering arrives in most organisations as a story about a monkey that kills servers in production, and leaves about a minute later when somebody imagines applying it to the payment system.
That is a fair reaction to the story. It is the wrong reaction to the method, because the parts of the method that produce the value are available to estates that cannot entertain random termination in production.
What the method actually requires
Strip the tooling away and the published principles ask for four things.
A steady state defined in business terms rather than infrastructure terms. Not CPU, not pod count — orders per minute, successful logins, streams started. The measure has to be something that tells you whether users are being served.
A hypothesis that the steady state holds while something specific goes wrong. Stated in advance, and falsifiable.
A real event, not a simulation of one. An instance terminated, a dependency made slow, a disk filled, a clock skewed.
A bounded blast radius, so that being wrong is survivable.
Only the third of these has anything to do with production, and even that is a question of where you inject the event rather than whether you inject one at all.
- Measured
- the business metric, throughout
- Bounded by
- environment, scope, and a stop condition
- Finished when
- the fix is re-tested, not when the issue is filed
This diagram as text
- State the steady state — in business terms
- Write the hypothesis down — before, not after
- Bound the blast radius — one dependency · one environment · a stop condition agreed in advance
- Inject the real event — latency before death
- Watch the business metric — not the component
- Record what was learned
- Fix, then re-run to prove it — finished here, not at the ticket
Relationships
- State the steady state → Bound the blast radius
- Write the hypothesis down → Bound the blast radius
- Bound the blast radius → Inject the real event
- Bound the blast radius → Watch the business metric
- Inject the real event → Record what was learned
- Watch the business metric → Record what was learned
- Record what was learned → Fix, then re-run to prove it
Start with latency, not death
The instinct is to kill something. Resist it, because an absent dependency is the case most systems handle at least somewhat — there is a timeout, there is an error path, somebody wrote a fallback.
A slow dependency is where systems come apart. Connection pools fill with requests waiting on something that will eventually answer. Thread pools exhaust. Retries pile on top of a service that is already struggling and turn a slowdown into an outage. Health checks keep passing, because the process is alive, so nothing is taken out of rotation.
Introducing two hundred milliseconds of latency on one downstream call teaches more about a system than terminating it does, and it is far easier to get approval for.
A ladder you can actually climb
Each rung is useful on its own. Most organisations never need the top of it.
| Rung | What you do | What it finds |
|---|---|---|
| 1 | Make one dependency slow in a pre-production environment with realistic data | Timeouts that are absent or far too long, retry storms, pools that exhaust |
| 2 | Stop one instance of a redundant service, in a window, in production | Whether redundancy is real or whether one node quietly held state |
| 3 | Remove a whole availability zone in pre-production | Capacity assumptions, and quorum placement |
| 4 | Degrade a dependency during business hours, announced, with a stop condition | Whether the on-call path works when a human has to decide |
| 5 | Unannounced, bounded, in production | Whether any of the above only worked because it was expected |
The fifth rung is where the talks start. Almost all the learning is on the first three, and an organisation that does rungs one to three on a schedule is in better shape than one that has read about rung five.
The things it reliably finds
After enough of these, the same findings recur, and they are rarely the ones the team expected.
A timeout that was never set. The default is often infinite, or sixty seconds, which in a request path is the same as infinite.
Retries without backoff or a budget. Three services each retrying three times turns one slow dependency into twenty-seven times the load on it.
A health check that checks the wrong thing. Returning 200 because the process is running, while every request that needs the database is failing, keeps a broken instance in rotation.
A dependency nobody documented. A licence server, an internal certificate authority, a DNS name resolved once at startup. These are the same findings a recovery test produces, which is not a coincidence.
A fallback that was never exercised. Written two years ago, never executed, and wrong.
Making it safe enough to be allowed
The objection to chaos work is never really about the method, it is about blame. Nobody wants to be the engineer who broke the payment system in an experiment.
Four things make it approvable:
- A written stop condition, agreed before the experiment, with a named person who can call it. Experiments stop when the condition is met, regardless of what has been learned.
- A blast radius stated in advance: which environment, which service, which percentage of traffic, for how long.
- An announced window at first. Unannounced experiments test the organisation as well as the system, which is valuable and is not where you begin.
- The finding belongs to the system, not to a person. If an experiment surfaces a missing timeout, the outcome is a timeout and a test, not a conversation about who should have set it.
It is also the evidence
ISO/IEC 27031 asks for demonstrated ICT readiness, not documented intent. An experiment log with a date, a hypothesis, a measured result and a remediation date is exactly the artefact an auditor samples for.
So the practice pays twice: it finds the failures before they find you, and it produces the evidence that you looked.
How STP approaches this
We start at latency in a pre-production environment with production-shaped data, because that is where the findings are densest and the approval is easiest. The stop condition and the blast radius are written before anything is injected. Each finding gets a fix and a re-run, because an experiment that files a ticket and moves on has only told you that you have a problem.
More on disaster recovery and managed IT, or start a conversation.

