EP13: AWS 28-Hour Outage - Not a Code Bug or Bad Deployment

A cooling system failed and AWS went dark for 28 hours.

EP13: AWS 28-Hour Outage - Not a Code Bug or Bad Deployment

If your cloud region is AWS us-east-1, maybe this week was painful for you too.

It was for us.

We got affected by the recent AWS incident. Not completely down, but enough to hurt operations badly.

One of our critical client-facing SFTP applications stopped working and we had to rebuild the server from backup because there was no other recovery option left.

That pushed us almost 12 hours back and the impact was significant. We had to reach out to each client individually to communicate the situation.

We manage around 1500+ servers. Many of them recovered because of autoscaling, but this one system was in bad situation.

And honestly, I never expected the reason behind all this.

The cooling system failed. That’s it.

Not bad code.

Not wrong deployment.

Not database migration.

Not hacker attack.

A cooling system inside one AWS availability zone failed and temperature became too high. Servers started shutting down to protect hardware. Then services got impacted for 28 hours.

This might be the longest single-AZ outage in AWS history.

But this incident reminded me of something important.

Sometimes infrastructure fails because the physical world fails first.

And cloud is still physical.

Your servers are still machines.
Machines are inside racks.
Racks are inside buildings.
Buildings need electricity and cooling.

I think many teams are not really prepared for this type of failure because cloud feels like magic sometimes. We think AWS will always recover automatically.

But when this happened, many people were stuck for a long time.

AWS acknowledged the failure and said it would take longer to resolve, so they asked customers to recover from backups if they needed immediate recovery. This happens rarely.

AWS outage timeline

If your systems were running only in one availability zone, that’s a problem.
If backups were old, that’s a problem too.
If you never tested restore from snapshot, that’s a big problem.

This is the kind of failure we don’t prepare for because it feels unlikely. But it happened. And it lasted longer than most software outages ever do.

The lesson is simple: Don’t assume the cloud is magic. Design for reality. Test for failure.

Thanks for reading,

Alon

PS: I wrote a detailed breakdown of the full incident—the timeline, root cause, affected services, and my own takeaways.

You can read it here: [The 28-Hour Meltdown: What Happened When AWS US-EAST-1 Overheated]