Post-Incident Reviews That Fix Systems, Not People

Turn incident retrospectives into real reliability gains. Learn how blameless post-incident reviews surface systemic causes and produce action items that ship.

← Back to Blog

The Wrong Question

When a service goes down, the instinctive question is "who changed something?" It is also the least useful question available. Engineers who expect to be named stop volunteering detail, and the review ends up documenting a single unlucky commit instead of the conditions that let a single commit take production down.

A blameless post-incident review inverts that instinct. It assumes competent people acted reasonably given the information they had, and asks what made the failure possible, and what made it hard to detect or recover from. The Google SRE Workbook chapter on postmortem culture makes the practical case: blame suppresses exactly the information the organisation needs to get more reliable.

Build the Timeline First

Before analysis, reconstruct what happened as a factual, timestamped narrative. Resist the urge to explain while you assemble it. A useful timeline covers six phases, and each one is a separate opportunity to improve:

  • Detection: how long until anyone knew, and was it a monitor or a customer?
  • Escalation: how long until the person who could fix it was engaged?
  • Mitigation: what stopped the bleeding, and what was tried that did not work?
  • Communication: what did stakeholders and users know, and when?
  • Recovery: when was full service actually restored, not just the alert cleared?
  • Prevention: what would stop this class of failure recurring?

Teams frequently discover their biggest win is not in the fix at all. Shaving twenty minutes off detection often delivers more user-visible reliability than eliminating the trigger, and it generalises to incidents nobody has had yet.

Contributing Factors, Not a Root Cause

Complex systems rarely fail for one reason. A deploy went out, a health check was too permissive, a dashboard was misleading, and a runbook was stale. Removing any one of those might have prevented the outage, which means all four are contributing factors and none is uniquely the root cause.

Aim for two to five systemic factors, each framed as a property of the system rather than a property of a person: "the canary stage had no automated rollback" rather than "the engineer did not watch the canary." Record what worked, too. Knowing which safeguards held is how a team learns what to invest in, and it keeps the review honest rather than purely negative.

Action Items That Actually Ship

Most reviews fail here. An action item like "improve monitoring" is not work, it is a wish. Every item needs a named owner, a due date, and a scope small enough to finish inside a normal sprint. It helps to sort them into two buckets: mitigative items close this specific gap, while preventative items address the whole class of failure. PagerDuty publishes its own postmortem process openly, including who owns the follow-up and how quickly it is scheduled.

Track them in the same backlog as feature work, because a list nobody schedules is a list nobody completes. If action items from previous incidents are consistently unfinished, that is a signal in its own right, and often a sign the team needs an explicit error budget policy to justify the time. Atlassian's guide to blameless postmortems covers practical ways to keep that follow-through visible.

Timing and Culture

Draft the review within about 72 hours while memory is fresh, and publish the final version within one to two weeks. Circulate it widely, since the value compounds when teams learn from incidents they did not experience.

Culture is set by example. When a senior engineer opens a review by describing their own misread of a dashboard, everyone else learns that candour is safe. To connect this practice to the wider response process, see the incident management guide, and browse the blog for related reliability practices.

This article was generated with the help of AI.