intermediate
Incident basics
Coordinate detection, triage, mitigation, communication, postmortem learning, and follow-up work during production failures.
Incident response coordinates detection, triage, mitigation, communication, and learning when production behavior harms users. Restore service safely first; perfect root-cause analysis can follow.
| Role | Responsibility | |------|----------------| | Incident commander | Coordinates decisions and comms | | Technical lead | Drives mitigation and investigation | | Comms lead | Stakeholder and customer updates | | Scribe | Timeline and action log |
Typical flow: acknowledge → assess severity → mitigate (rollback, scale, feature flag, traffic shift) → communicate → stabilize → postmortem with follow-up tickets.
On interviews: commander roles, severity levels, rollback vs forward fix, customer updates during uncertainty, blameless postmortems, and action-item ownership.
Common pitfalls: blaming individuals; long silence to customers; mitigating without recording timeline; postmortems without tracked follow-ups.
The trade-off is speed of mitigation versus thorough communication and learning discipline.
Checklist:
- Assign roles early.
- Mitigate before chasing perfect diagnosis.
- Communicate impact and status regularly.
- Write blameless postmortems with owned follow-ups.