Chaos engineering arrives in most organisations as a story about a monkey that kills servers in production, and leaves about a minute later when somebody imagines applying it to the payment system.

That is a fair reaction to the story. It is the wrong reaction to the method, because the parts of the method that produce the value are available to estates that cannot entertain random termination in production.

What the method actually requires

Strip the tooling away and the published principles ask for four things.

A steady state defined in business terms rather than infrastructure terms. Not CPU, not pod count — orders per minute, successful logins, streams started. The measure has to be something that tells you whether users are being served.

A hypothesis that the steady state holds while something specific goes wrong. Stated in advance, and falsifiable.

A real event, not a simulation of one. An instance terminated, a dependency made slow, a disk filled, a clock skewed.

A bounded blast radius, so that being wrong is survivable.

Only the third of these has anything to do with production, and even that is a question of where you inject the event rather than whether you inject one at all.

One experiment, start to finish
A single chaos experiment from stated steady state to a re-tested fixState the steady state in business terms and write the hypothesis down before running anything, because an experiment nobody could be wrong about is not worth running. Bound the blast radius first: one dependency, one environment, and a stop condition agreed in advance. Inject the real event and watch the business metric rather than the component. Stop on the agreed condition whatever the result. Record what was learned, fix it, and re-run to prove the fix.State the steady statein business termsWrite the hypothesis downbefore, not afterBound the blast radiusone dependency · one environment · a stop condition agreed in advanceInject the real eventlatency before deathWatch the business metricnot the componentRecord what was learnedFix, then re-run to prove itfinished here, not at the ticket
Measured
the business metric, throughout
Bounded by
environment, scope, and a stop condition
Finished when
the fix is re-tested, not when the issue is filed
This diagram as text
  • State the steady state — in business terms
  • Write the hypothesis down — before, not after
  • Bound the blast radius — one dependency · one environment · a stop condition agreed in advance
  • Inject the real event — latency before death
  • Watch the business metric — not the component
  • Record what was learned
  • Fix, then re-run to prove it — finished here, not at the ticket

Relationships

  • State the steady state → Bound the blast radius
  • Write the hypothesis down → Bound the blast radius
  • Bound the blast radius → Inject the real event
  • Bound the blast radius → Watch the business metric
  • Inject the real event → Record what was learned
  • Watch the business metric → Record what was learned
  • Record what was learned → Fix, then re-run to prove it

Start with latency, not death

The instinct is to kill something. Resist it, because an absent dependency is the case most systems handle at least somewhat — there is a timeout, there is an error path, somebody wrote a fallback.

A slow dependency is where systems come apart. Connection pools fill with requests waiting on something that will eventually answer. Thread pools exhaust. Retries pile on top of a service that is already struggling and turn a slowdown into an outage. Health checks keep passing, because the process is alive, so nothing is taken out of rotation.

Introducing two hundred milliseconds of latency on one downstream call teaches more about a system than terminating it does, and it is far easier to get approval for.

A ladder you can actually climb

Each rung is useful on its own. Most organisations never need the top of it.

Rung What you do What it finds
1 Make one dependency slow in a pre-production environment with realistic data Timeouts that are absent or far too long, retry storms, pools that exhaust
2 Stop one instance of a redundant service, in a window, in production Whether redundancy is real or whether one node quietly held state
3 Remove a whole availability zone in pre-production Capacity assumptions, and quorum placement
4 Degrade a dependency during business hours, announced, with a stop condition Whether the on-call path works when a human has to decide
5 Unannounced, bounded, in production Whether any of the above only worked because it was expected

The fifth rung is where the talks start. Almost all the learning is on the first three, and an organisation that does rungs one to three on a schedule is in better shape than one that has read about rung five.

The things it reliably finds

After enough of these, the same findings recur, and they are rarely the ones the team expected.

A timeout that was never set. The default is often infinite, or sixty seconds, which in a request path is the same as infinite.

Retries without backoff or a budget. Three services each retrying three times turns one slow dependency into twenty-seven times the load on it.

A health check that checks the wrong thing. Returning 200 because the process is running, while every request that needs the database is failing, keeps a broken instance in rotation.

A dependency nobody documented. A licence server, an internal certificate authority, a DNS name resolved once at startup. These are the same findings a recovery test produces, which is not a coincidence.

A fallback that was never exercised. Written two years ago, never executed, and wrong.

Making it safe enough to be allowed

The objection to chaos work is never really about the method, it is about blame. Nobody wants to be the engineer who broke the payment system in an experiment.

Four things make it approvable:

  • A written stop condition, agreed before the experiment, with a named person who can call it. Experiments stop when the condition is met, regardless of what has been learned.
  • A blast radius stated in advance: which environment, which service, which percentage of traffic, for how long.
  • An announced window at first. Unannounced experiments test the organisation as well as the system, which is valuable and is not where you begin.
  • The finding belongs to the system, not to a person. If an experiment surfaces a missing timeout, the outcome is a timeout and a test, not a conversation about who should have set it.

It is also the evidence

ISO/IEC 27031 asks for demonstrated ICT readiness, not documented intent. An experiment log with a date, a hypothesis, a measured result and a remediation date is exactly the artefact an auditor samples for.

So the practice pays twice: it finds the failures before they find you, and it produces the evidence that you looked.

How STP approaches this

We start at latency in a pre-production environment with production-shaped data, because that is where the findings are densest and the approval is easiest. The stop condition and the blast radius are written before anything is injected. Each finding gets a fix and a re-run, because an experiment that files a ticket and moves on has only told you that you have a problem.

More on disaster recovery and managed IT, or start a conversation.