Blameless incident review is one of the most widely adopted practices in operations and one of the least effective, because most organisations adopt the blameless part and not the review part.

The culture work matters and it is necessary. It is not sufficient. A team can hold a calm, psychologically safe discussion about an outage, write an honest document, and have precisely nothing change.

What blameless is for

The reason to remove blame is not kindness. It is accuracy.

When naming a person is a possible outcome, people manage what they say. Details disappear. The engineer who noticed something odd an hour before does not mention it. The one who ran the command describes it in the passive voice. You end up with a document that is accurate about the timeline and silent about the causes.

“Human error” is where the investigation starts, not where it ends.

If someone ran a destructive command against the wrong environment, the useful questions are why the environments looked alike, why the command was available without confirmation, and why nothing detected it before it took effect. Each of those has a fix. “Be more careful” does not.

Why most reviews change nothing

The failure is almost always at the same point: the document ends with a list of actions and nobody owns them.

Actions written as “we should add better monitoring” have no owner, no date and no definition of done. They are a statement of regret. Three months later the same incident happens, and the review for that one notes that an earlier review had identified the cause.

Where reviews usually stop, and where they have to reach
Where an incident review usually stops and where it has to reachMost reviews produce a timeline captured during the incident, a set of contributing factors and a blameless discussion, and stop there. The step that changes anything is an action with a named owner and a date, tracked in the same backlog as planned work rather than inside the document. Done means shipped and demonstrated, and the next review opens by checking the last one.MOST REVIEWS REACH HEREAND ONLY THENin the backlogTimelinecaptured during, not afterContributing factorspluralBlameless discussionAn action with a person’s name and a datenot a team. a team is nobodyChange shippedProven by a test or an exerciseChecked at the next review
Done means
shipped and demonstrated, not filed
Owner
a person, never a team name
Checked
at the start of the next review
This diagram as text
  • Most reviews reach here
    • Timeline — captured during, not after
    • Contributing factors — plural
    • Blameless discussion
  • An action with a person’s name and a date — not a team. a team is nobody
  • And only then
    • Change shipped
    • Proven by a test or an exercise
    • Checked at the next review

Relationships

  • Timeline → An action with a person’s name and a date
  • Contributing factors → An action with a person’s name and a date
  • Blameless discussion → An action with a person’s name and a date
  • An action with a person’s name and a date → Change shipped — in the backlog
  • Change shipped → Proven by a test or an exercise
  • Proven by a test or an exercise → Checked at the next review
Open full size · after the Google SRE Book on postmortem culture

Two rules fix most of it.

Every action has a person’s name and a date. Not a team. A team is nobody. If no one will take it, that is information: the action is not actually considered worth doing, and it should be dropped rather than left to decay on a list.

Actions live where planned work lives. In the backlog, prioritised against everything else. An action that exists only inside an incident document is invisible the moment the document is filed.

Then open the next review by checking the previous one’s actions. It takes two minutes and it is the single thing that most changes whether any of this is real.

Capture the timeline during, not after

The expensive part of a review is reconstructing what happened, and it is expensive because nobody wrote it down while it was happening.

A channel for the incident, with timestamps, where people state what they are about to do and what they observe, turns an hour of archaeology into a copy and paste. It also removes the bias that creeps in afterwards, when the team already knows the answer and reconstructs a tidier path to it than the one they actually took.

What to write down

Short is better. A review nobody reads has the same effect as no review.

  • What a user experienced, and for how long. Not what a component did.
  • The timeline, from first signal to resolution, including the things that were tried and did not work. Those are often the most useful part.
  • How it was detected. If a customer told you, that is a finding in itself.
  • Contributing factors, plural. Single-cause incidents are rare, and the search for one usually stops at the most visible rather than the most fixable.
  • What made it better or worse. The runbook that was current. The dashboard nobody could find. The alert that fired at the right time.
  • Actions, each with a name and a date.

What to leave out: anything that reads as attribution, and any detail that only exists to show the investigation was thorough.

The auditor is asking for this too

ISO/IEC 27001 Annex A 5.27 asks that knowledge gained from incidents is used to reduce the likelihood or impact of future ones. A dated review with tracked actions and evidence of completion is exactly that, which means the practice that makes operations better also satisfies the control, with no separate exercise.

The common failure here is the same one: organisations retain the reviews and cannot show that anything was done as a result.

How STP approaches this

Where we run operations, the timeline is captured in the incident channel as it happens, actions carry a person and a date and go into the same backlog as everything else, and the next review opens by checking the last one. Where we are asked to review a practice that is not producing change, the finding is almost never the culture. It is that nothing on the action list has an owner.

More on managed IT and audit and compliance, or start a conversation.