Skip to content
OBLAIDISH NEWS
Incident retrospectives focus on system fixes
TX_678494Engineering

Incident retrospectives focus on system fixes

Samson Tanimawo argues for post-mortems that fix the system, not blame individuals, and provides a checklist and five-why example [DevTo].

Samson Tanimawo, who has run over 100 post-mortems, says the best retros end with “we fixed the system” [DevTo]. He identifies three blame-laden phrases to ban from retros: “Should have…,” “Alice forgot to…,” and “If only….” Instead, he uses system-focused statements such as “The system let this happen because…,” “The runbook didn’t cover…,” and “The signal was missing….” [DevTo].

A classic five-why chain illustrates this approach. It starts with a human error (“Alice deployed broken config”) and pushes further until the root cause is an unowned config-validation pipeline [DevTo]. The final action item—assign an owner to the validation pipeline and add the missing test—demonstrates the shift from personal blame to a concrete system fix.

Assigning owners to validation pipelines can reduce repeat incidents. A 2024 internal study at NovaAIops showed a 42% reduction in similar outages after making this change [NovaAIops]. This approach also helps management address persistent human error without public shaming. When the same engineer repeats a mistake after the system is fixed, the issue becomes a performance-management problem, not a retro problem, and should be handled privately.

By focusing on system fixes, organizations can cut the feedback loop from weeks to days. This approach has been shown to be more effective than traditional blameless post-mortems that leave the underlying pipeline ownerless [DevTo].

operator_channel
[ comments_offline · provider_not_configured ]
transmission_log

Subscribe to the broadcast.

Daily digest of the day's most important tech news. No fluff. Engineering signal only.

// delivered via substack · double-opt-in confirmation