Skip to content
SRE

Day 16: It's Not Who Did It. It's Why It Happened.

A postmortem that finds someone to blame ends the conversation right when it should be starting. Blameless postmortems trade the satisfaction of an answer for the harder, more useful question underneath it.

Thamunkpillai 4 min read

"Who ran that command?" is the question that ends a postmortem before it's actually started anything useful. The engineer who ran it stops talking, or starts defending, and everyone in the room learns the same lesson: if something goes wrong on your watch, the conversation afterward is about you, not about the system. That lesson gets absorbed fast, and it produces exactly the wrong incentive — hide mistakes, under-report near-misses, and definitely don't volunteer the detail that would actually explain what happened, because that detail is now evidence.

A blameless postmortem asks a different question: not "who did it," but "why did the system allow it to happen, and why did it feel like the reasonable thing to do at the time." That reframe isn't softer — it's more demanding, because it doesn't let the investigation stop at a human decision. It has to go one level deeper, to whatever set of conditions made that decision look correct in the moment.

Note

The single sentence that captures the whole practice: fix systems, not people. A person can be replaced and the exact same incident will recur, because the thing that actually caused it — a missing alert, an unclear runbook, an easy-to-misread dashboard — is still sitting there waiting for the next person to hit it.

Blameful vs. blameless, side by side

Traditional (blameful)Blameless
"Who deployed the broken change?""What in our review or testing process let it ship?"
Focus on the individualFocus on the system and process
Defensiveness, hidden detailsOpen, honest reconstruction of events
Same incident recurs with a different name attachedRoot cause addressed, recurrence actually prevented
Fear of the next postmortemCuriosity about what the next one might reveal

The recurrence row is the one that should end any argument about which approach is more rigorous. A blameful postmortem that identifies "Sarah forgot to check the dashboard" produces exactly one fix: don't be Sarah. The next person in that role, doing the exact same reasonable-seeming thing under the exact same pressure, hits the exact same failure. A blameless postmortem that identifies "the dashboard didn't surface the one metric that mattered, and nothing forced a check before deploy" produces a fix that actually holds regardless of who's on call next.

What belongs in one

Text
1. Summary       — what happened, in plain language, one paragraph.
2. Timeline      — every relevant event, with timestamps, no editorializing.
3. Impact        — who and what was affected, and for how long.
4. Root cause(s) — the actual chain of conditions, not a single scapegoat.
5. What went well — the detection, the response, whatever worked.
6. Action items  — specific, owned, dated fixes. Not "be more careful."
7. Follow-up     — did the action items actually ship, and did they help?

Step 7 is the one that separates postmortems that matter from postmortems that are theater. An action item that never gets tracked to completion is just a paragraph that made the meeting feel productive. The postmortem process itself needs the same rigor as the system it's investigating — if "add monitoring for X" sits open for six months, that's a process failure worth its own review.

Ground rules that keep it blameless in practice

Ask "why" and "what," never "who" or "how could you." Assume everyone made the most reasonable decision available to them with the information and time pressure they had — because they almost always did. Keep it constructive: the shared goal in the room is a stronger system, not a settled score.

Why blameless is the harder discipline, not the easier one

There's a version of "blameless" that's actually just "consequence-free," and that version doesn't work — it lets genuine negligence or repeated carelessness slide under a process designed for something else. Real blameless culture isn't the absence of accountability; it's accountability aimed at the system and the process that let a mistake become an incident, which is almost always the more useful target. An engineer who fat-fingered a command isn't the problem if the deploy process allowed that single command to take down production with no review, no staging step, and no automatic rollback. Fix that, and the next fat-fingered command — because there will be one — becomes a non-event instead of an outage.

Tomorrow's article closes out the SRE Foundations phase with incident management itself: the real-time discipline of detecting, triaging, and resolving an incident well while it's still in progress — which is exactly the process a good blameless postmortem exists to keep improving.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.