On October 20, 2025, a major AWS outage disrupted hundreds of services worldwide. Beyond the technical breakdown, it revealed a deeper truth: resilience starts with people — not servers.

On 20 October 2025, a faulty update to Amazon’s DynamoDB API in the US-East-1 region (Virginia) triggered a chain reaction across AWS’s internal DNS system. EC2, S3 and Lambda went offline within minutes, disrupting hundreds of services worldwide. Full recovery took close to four hours. Amazon later confirmed there was no data loss and no security breach.
The lesson was less technical than organisational:
Resilience is as much a staffing and process question as an architecture one. Multi-cloud design, disaster-recovery drills and open post-mortems are what shorten the next outage.
A faulty update to Amazon’s DynamoDB API in the US-East-1 region (Virginia) reportedly triggered a chain reaction across AWS’s internal DNS system. Within minutes, core services including EC2, S3 and Lambda went offline, and full recovery took close to four hours.
No. Amazon confirmed afterwards that there was no data loss and no security breach. What failed was availability, not integrity, which is why the incident is best treated as a business continuity problem rather than a security one.
Because a large share of the internet runs on a small number of providers and regions, and US-East-1 is one of the most heavily used. When dependencies concentrate in one place, a local failure propagates globally. Total dependence on a single provider or region is what turns an incident into a systemic event.
It distributes workloads across regions and providers so one failure cannot take everything down. It is not free: it adds architectural complexity, operational overhead and a broader set of skills to maintain. The honest trade-off is more resilience in exchange for more cost and more expertise.
By treating outages as inevitable and running disaster-recovery exercises that test people as well as systems: response times, escalation paths and decision-making under pressure. A recovery plan nobody has rehearsed is a document, not a capability.
Site Reliability Engineers, CloudOps and DevOps engineers, cybersecurity and infrastructure specialists, and data and AI operations profiles. These are the people who design redundancy, automate recovery workflows and run the post-mortems that stop the same failure happening twice.
By subscribing to our newsletter, you agree to receive communications in accordance with our privacy policy.