5Q Quadrants·of·Risk

Risk register · entry

Q3 · Engineered

AWS US-EAST-1 outage

A routine scaling action hit a latent bug and took down half the internet.

Tightly coupled systems where one small fault cascades and takes down the whole machine.

Quadrant
Q3 Engineered
Year
2021
Impact
5-7 hrs
Sector
Cloud
Region
Global
Category
Technological

Why this quadrant

The trigger was a narrow, engineered payoff structure (one internal automation bug with a knowable, containable failure mode) but the tail behavior turned fat and correlated the instant it touched shared infrastructure, so a single-region software defect produced simultaneous, uninsured losses across unrelated industries worldwide.

The record

  • Outage began 7:30am PST / 10:30am ET, December 7, 2021certain
  • Internal DNS traffic migration completed at 9:28am PST, resolving DNS errorscertain
  • Most core services (EC2 API errors, instance launches) recovered between 1:15pm and 2:40pm PSTcertain
  • Full resolution (Amazon Connect, EventBridge) not until 4:41pm-6:40pm PST, roughly 7-9 hours total depending on servicecertain
  • Companies affected include Netflix, Disney+, Coinbase, Ring, Amazon retail/warehouse operations, Roomba, Ticketmastercertain
  • Amazon Flex and Dolphin scanner logins failed across warehouses in at least 14 named US states plus Ontario, Canadalikely
  • A second US-EAST-1 disruption followed on December 10, 2021likely
  • No official AWS-disclosed dollar cost estimate for the December 2021 event specifically (unlike the much larger, separately documented October 2025 AWS outage)uncertain

Sources

  1. Amazon Web Services (official post-event summary)
  2. The Register
  3. Vice
  4. CNBC

The newsletter

One risk story a week, taken apart the way this one was: what was known, what was ignored, and which quadrant it really belonged to.

No spam, one email a week.