No. 006 · The Enter Key Gallery

The Typo That Took Down S3

· about 4 hours

  1. 1Engineer debugging a slow billing system runs an approved playbook command
  2. 2One input is mistyped; far more servers are removed than intended
  3. 3Two S3 subsystems in us-east-1 need a full restart
  4. 4They haven't been fully restarted in years; it takes hours
  5. 5A large part of the web goes down, including AWS's own status dashboard

Tools should refuse to remove too much capacity too fast, however it's typed.

Source: AWS service disruption summary, March 2017

AWS S3 outage, February 2017. Duration: about 4 hours. 5 hops from the first change to the failure.