advanced

Alerting

Actionable alerts and service-level targets that detect user-impacting problems without flooding responders with noise.

Alerting interviews connect reliability engineering to on-call reality: SLOs, SLAs, error budgets, and fighting alert fatigue with symptom-based pages and runbooks.

Subtopics: SLO, SLA, error budget, alert fatigue.

On interviews: define an SLO for search latency; explain error budget policy during a launch week; when to page vs Slack ticket.

Common pitfalls: threshold alerts on CPU; duplicate pages per pod; no runbook links.

The trade-off is catching incidents early versus burning out responders.

Checklist:

  • SLIs tied to user journeys.
  • Burn-rate alerts with windows.
  • Internal SLO stricter than SLA.
  • Quarterly alert pruning.