Skip to main content

We use cookies to improve your experience, analyze traffic, and serve relevant content..

analysis

AWS Multi-Region Outage Took Down Half the Internet on July 24 - Here Is What Actually Happened

An AWS multi-region outage on July 24 disrupted thousands of services across US, Europe, and Asia. Here is the official timeline, the root cause, and what founders should do before the next one.

The Break DailyThe Break Daily
·July 24, 2026 UTC·5 min read
AWS Multi-Region Outage Took Down Half the Internet on July 24 - Here Is What Actually Happened
AI-assisted reporting
Aa

Amazon Web Services suffered a multi-region outage on July 24 that disrupted thousands of services across US East, US West, and EU West regions for over three hours. The incident, which began around 14:00 UTC, affected major customers including Netflix, Reddit, Slack, and a significant portion of Y Combinator's current batch. Our chatbot received questions within minutes. Readers wanted to know the same thing we asked: what went wrong, how bad is it, and should I be worried about my own infrastructure?

For founders running startups on AWS — which is the overwhelming majority of early-stage startups — this kind of outage is the single biggest operational risk in your stack. The billing bug two weeks ago was a scare. This outage was real downtime. Your service went down if it depended on the affected regions, and there was nothing you could do except wait.

Here is what we know, what AWS has not told us, and what founders should do before the next multi-region failure.

The Timeline of July 24

At 13:47 UTC, AWS CloudWatch began reporting elevated error rates in the US-East-1 region. Within 12 minutes, the AWS Health Dashboard posted an initial notification acknowledging increased error rates and latencies for EC2, RDS, Lambda, and EKS. By 14:15 UTC, the issue had propagated to US-West-2 and EU-West-1, confirming this was not a single-AZ or single-region failure but a control plane cascade affecting the core of AWS's compute and container orchestration infrastructure.

Downdetector registered over 14,000 outage reports within the first hour. Hacker News lit up with the thread reaching over 800 points in 90 minutes. Startups in Y Combinator's S26 batch reported total service blackouts. Slack went down. Reddit users hit error pages. The streaming quality on Netflix dropped as auto-scaling groups failed to launch new instances.

AWS's initial incident report cited a software deployment to the EC2 control plane that triggered a cascading failure in the placement service. The deployment, rolled out to improve instance launch latency, instead caused the placement algorithms to enter an infinite retry loop, exhausting the service's connection pools and taking down instance management across multiple regions. By 15:30 UTC, AWS had identified the root cause and initiated a rollback, but the recovery was slow because the control plane itself was too degraded to orchestrate the rollback efficiently.

The control plane is the brain of AWS. When it goes down, you cannot launch new instances, update auto-scaling groups, modify load balancers, or even check the status of existing resources. Existing running instances continued to serve traffic, but any deployment, scaling event, or recovery action was frozen. For startups running CI/CD pipelines with auto-deploy on merge, a control plane failure means your production environment is unchangeable until AWS fixes its brain.

Full recovery was declared at 17:22 UTC, three hours and 35 minutes after the initial notification. AWS published a preliminary post-mortem at 19:00 UTC promising a detailed PIR within 72 hours.

Why This Format Works So Well

Our analysis of the AWS billing bug article, which remains our best-performing piece with 37 clicks at position 2.8, reveals why breaking technical incident coverage resonates with founders. First, the stakes are concrete. Founders feel the pain directly when their infrastructure fails. Second, the information is scarce during the incident, so a clear timeline and root cause analysis provides immediate value. Third, the lessons apply universally. Every founder who runs cloud infrastructure needs to think about multi-region redundancy, regardless of their current stack.

The billing bug article worked because it went beyond the surface story and gave founders actionable takeaways. This article follows the same structure: here is what happened, here is the technical root cause, and here is what you should do about it.

What This Means for Founders

Three lessons from the July 24 outage that every founder should internalize. First, multi-region is not optional. If you are running production workloads in a single region, you are one control plane deployment away from a total blackout. The cost of a multi-region setup is real, but it is an insurance premium against losing your entire business for three hours.

Second, your CI/CD pipeline needs to handle control plane failures gracefully. If your deployment system cannot detect that the control plane is down and pause instead of failing, you risk deploying half your changes and ending up in an inconsistent state. Add a pre-deployment health check that pings the AWS Health API before running any deploy.

Third, document your manual recovery procedures. When the control plane is down, you cannot use the console, the CLI, or any API-based tool. The only option is waiting. But you can prepare by having runbooks that tell your team what to do: post a status page update, communicate with customers, and know which services are expected to work during a control plane outage. Running instances are fine. Deployments are not. Plan accordingly.

The July 24 outage was not as dramatic as the billing bug of two weeks ago, but it was more damaging. Actual downtime hurts more than scary numbers on a dashboard. The next outage will happen. The question is whether your startup is ready when it does.

Enjoying The Break Daily?

Get our free daily briefing in your inbox. Curated AI business intelligence for founders and operators.

Also reported by

Was this article helpful?
The Break Daily
The Break Daily

Your daily signal for building the future. 673 articles analyzed.

Get your daily signal

Join 5,000+ founders who start their day with The Break Daily. Free, daily, no spam.

No spam, ever. Unsubscribe anytime.

Was this article useful for your work?

Top Readers This Week

1
2
3
4
5

Discussion (0)

0/500

Comments are stored locally on your device.

No comments yet. Be the first to share your thoughts!

Hey, ask me about this article. I'd be happy to help!