intermediate

Incident basics

Coordinate detection, triage, mitigation, communication, postmortem learning, and follow-up work during production failures.

Incident response coordinates detection, triage, mitigation, communication, and learning when production behavior harms users. Restore service safely first; perfect root-cause analysis can follow.

| Role | Responsibility | |------|----------------| | Incident commander | Coordinates decisions and comms | | Technical lead | Drives mitigation and investigation | | Comms lead | Stakeholder and customer updates | | Scribe | Timeline and action log |

Typical flow: acknowledge → assess severity → mitigate (rollback, scale, feature flag, traffic shift) → communicate → stabilize → postmortem with follow-up tickets.

On interviews: commander roles, severity levels, rollback vs forward fix, customer updates during uncertainty, blameless postmortems, and action-item ownership.

Common pitfalls: blaming individuals; long silence to customers; mitigating without recording timeline; postmortems without tracked follow-ups.

The trade-off is speed of mitigation versus thorough communication and learning discipline.

Checklist:

  • Assign roles early.
  • Mitigate before chasing perfect diagnosis.
  • Communicate impact and status regularly.
  • Write blameless postmortems with owned follow-ups.