advanced

Disaster recovery

Plan region loss, backup restore, runbooks, failover drills, data loss windows, and communication paths.

Disaster recovery (DR) prepares for region loss, datacenter failure, ransomware, or operator error destroying primary infrastructure. Plans include backup storage in separate accounts/regions, runbooks for failover, regular game days, communication templates, and defined roles during incidents.

Active-passive DR is simpler; active-active reduces RTO but adds consistency complexity. Document data loss window (RPO) and time to restore service (RTO) with realistic drill results, not aspirational slides.

On interviews: outline DR for a single-region MVP evolving to multi-region, failover steps, and what data you might lose in worst case.

Common pitfalls: DR backups in same region as primary; runbooks untested for years; DNS failover without TTL planning; no customer communication plan.

The trade-off is flexibility versus complexity—know when the simpler path is enough.

Checklist:

  • Define RTO and RPO with stakeholders.
  • Store backups in isolated region/account.
  • Run failover drills on a schedule.
  • Maintain runbooks and comms templates.