Illustration of four incident stages — triage, mitigate, restore and explain — with status updates issued throughout

Direct Incident Handling

Back to Server Setup · Responsive Support · Root Cause Diagnosis · Post-Incident Follow-Up

We work the problem ourselves — triage, mitigate, restore — with status updates as we go. Not a ticket routed to a queue, not a list of things for you to try, and not a request that you gather diagnostics first. The person who answers is the person who fixes it.

1. Triage: How Bad, How Wide, How Fast

The first few minutes are about scope, not cause. Three questions, answered before anything is changed:

2. Mitigate: Stop the Bleeding

Mitigation is not the fix. It is whatever restores service fastest while the real cause is still unknown, and choosing it over a proper fix is deliberate rather than lazy. The common moves:

Every mitigation gets written down as it is applied, because temporary measures that nobody recorded become permanent by accident — and a rate limit left on for six months is its own incident later.

3. Restore: Properly, and Verified

Service restored means verified working, not “the error stopped appearing”. That means checking the actual user journey rather than the health endpoint, confirming the monitoring agrees, and looking for the second-order damage: the queue that built up while the system was down, the jobs that did not run, the retries still arriving, the data written in a degraded state. Coming back under the full weight of a backlog is a common way to fail twice, so restoration is staged where the backlog is large.

And we note explicitly which mitigations are still in place, so removing them is a task rather than an oversight.

4. Communicate While It Runs

This is the part most often done badly, and it is the part you experience. Silence during an outage is indistinguishable from nothing being done.

5. Who Does What

For anything sustained, one person holds the incident — makes decisions, keeps the timeline, handles communication — and the others work the problem. This sounds like overhead on a small team and is not: the failure mode without it is two people making conflicting changes at once, or everyone investigating and nobody telling you anything.

Everything done is logged as it happens, with timestamps. That record is what makes the diagnosis and the follow-up possible, and it cannot be reconstructed afterwards from memory.

What You Get