Illustration of an incident being written up and producing one concrete action: a monitor, a fix, or a runbook

Post-Incident Follow-Up

Back to Server Setup · Incident Handling · Root Cause Diagnosis · Resource Monitoring

After resolution, we document what happened and put in place whatever prevents it from recurring — a monitor, a config fix, a runbook. This is the step that gets skipped, because by the time it is due the crisis has passed and something else is urgent. Skipping it is also what turns one incident into a recurring one.

1. Write It Down While It Is Fresh

Within a day or two, not a fortnight. The write-up is short and factual:

2. Blameless, Because the Alternative Does Not Work

Where a review looks for who made the mistake, people stop reporting near-misses and stop being candid about what they did — and the information dries up exactly where it is most useful. The productive question is not who typed the command but why the system allowed a single typed command to cause this, why nothing caught it, and why it took so long to notice.

That is not an absence of accountability. Actions are owned, dates are real, and follow-up is tracked. It is a statement about where to look: at the conditions, because those are what can be changed.

3. Every Incident Produces at Least One Action

If the honest conclusion is “nothing to do”, the review was not thorough enough. There is always something, and it usually falls into one of these:

Each action gets one named owner and a date. Actions owned by a team are owned by nobody.

4. Track Them, or the Review Was Theatre

The most common way this fails is not a bad review but a good review whose actions were never done. So the actions go into the same tracker as everything else, with the same visibility, and they get reviewed at the next one. An incident recurring with an outstanding action from the last time is a useful and uncomfortable signal, and it should be visible rather than quietly noted.

Be realistic about prioritisation, too. Not every action justifies its cost, and it is more honest to record that a fix was considered and deliberately deferred, with the reason, than to leave it open forever. A deferred decision is a decision; a forgotten one is not.

5. Look Across Incidents, Not Just Within Them

Each incident individually may look like bad luck. Read a quarter of them together and patterns appear that no single review shows: most incidents following deployments, one component appearing repeatedly, detection consistently slow, the same undocumented knowledge being needed each time. Those patterns point at systemic fixes — better release practice, one component that needs rebuilding, a monitoring gap — which are worth far more than the individual actions.

And the numbers worth tracking over time are how long detection takes, how long resolution takes, and how many incidents recur. If detection is not improving, the monitors from the last ten reviews were not the right ones.

What You Get