Incident response and postmortems
Handle outages calmly with clear roles and severities, mitigate first, and learn from every incident with blameless postmortems.
- Run an incident with defined roles, severities and regular updates
- Prioritize mitigation over diagnosis, and measure detection and recovery times
- Write a blameless postmortem and run a healthy on-call rotation
18:02, Friday. Every drone starts flying to the same rooftop. Twelve people jump into one chat channel; three of them restart different things at once; nobody tells customers anything; the CEO asks for updates every two minutes. It takes two hours to fix a bad config push that could have been rolled back in five minutes.
Incidents will happen. What separates good teams is a calm, practiced incident response:
- Detect - ideally an alert, not a customer.
- Declare - say “this is an incident”, pick a severity and open a dedicated channel.
- Mitigate - stop the bleeding: roll back, flip a flag, fail over, add capacity. Understanding why comes later.
- Resolve - fix things properly once users are safe.
- Learn - a blameless postmortem with follow-up actions.
| Role | Does |
|---|---|
| Incident commander | coordinates: decides, delegates, keeps the big picture - and doesn’t debug |
| Operations / subject experts | investigate and make the changes |
| Communications lead | updates stakeholders and the public status page on a regular schedule |
| Scribe | keeps the timeline: what was seen, decided and done, and when |
Severity levels (for example SEV1 = major customer impact, SEV3 = minor) decide who gets woken up and how often updates go out.
Measuring response
Teams track time to detect (impact start → alert), time to acknowledge (alert → a human responds), time to mitigate (impact start → users unaffected) and time to resolve. Means hide a lot with so few incidents, so look at each incident’s timeline too.
1events = {"start": "18:02", "alert": "18:15", "ack": "18:19", "mitigated": "18:43"}
2
3def minutes(clock):
4 hours, mins = map(int, clock.split(":"))
5 return hours * 60 + mins
6
7print("time to detect:", minutes(events["alert"]) - minutes(events["start"]), "min")
8print("time to mitigate:", minutes(events["mitigated"]) - minutes(events["start"]), "min")time to detect: 13 min time to mitigate: 41 min
Blameless postmortems
After the incident, write a postmortem: summary, impact, timeline, contributing factors, what went well, and action items with owners and due dates.
It must be blameless. People act sensibly given what they knew at the time; if one person could take down production with one command, the system allowed it. Ask “how did our system make this possible?”, not “who did it?”. Blame teaches people to hide mistakes; blamelessness teaches the organization. And look for several contributing factors rather than a single “root cause” - complex failures rarely have just one.
Try it
Blameless or blameful?
Sort each sentence from Byte Bakery’s postmortem drafts.
“The config push tool didn’t validate coordinates, so an empty value routed all drones to the default rooftop.”
“Sam carelessly pushed a broken config.”
“Our alert fired 13 minutes after impact began because it only watched error rates, not delivery locations.”
“On-call should have known to roll back sooner.”
“Action: add a canary stage for config pushes (owner: platform team, due Oct 30).”
Key takeaways
Detect, declare, mitigate, resolve, learn.
Clear roles - especially an incident commander who coordinates rather than debugs - and regular updates.
Measure detection, acknowledgement and mitigation times per incident.
Blameless postmortems find contributing factors and produce owned action items.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: automate DevOps chores in Python
Write the small Python tools DevOps teams really build - pipeline runners, plan checkers, metric calculators, scanners - and run them against sample inputs. They run locally in your browser; no servers or cloud accounts needed.
Build the incident timeline
Each input line is HH:MM kind description..., where kind is start (impact began), alert, ack, mitigated, resolved or note; lines may be out of order. Print timeline: and the events sorted by time as 18:02 start: drones all routed to one rooftop.
Then print time to detect, time to acknowledge (alert to ack), time to mitigate and time to resolve (both from start), formatted as 13 min or, from an hour up, 1 h 58 min. If detection took more than 10 minutes, finish with detection was slow: add an alert for this symptom.
- The rooftop incident
- Quick catch
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Schedule the on-call rotation
The first input line lists the engineers in rotation order; the second is the number of weeks. Each following line is name unavailable week,week,....
For each week, the primary is the next available engineer in rotation order, continuing after the previous week’s primary (wrapping around); the secondary is the next available engineer after the primary. The secondary doesn’t move the rotation. Print week 1: primary ana, secondary bo, with secondary none if nobody else is available, or week 3: nobody available!. Finish with primary shifts: ana 2, bo 1, ... in rotation order.
- A month of on-call
- Holiday week
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…