stackery.co

Slack Went Down on the First Workday of 2021, Because Everyone Came Back at Once

Slack went down for five hours on the first workday of 2021 when AWS Transit Gateways failed to scale for the post-holiday traffic surge. The culprit was the most predictable date on the corporate calendar.

Last updated 2026-09-05

Slack suffered a global outage on Monday, 4 January 2021, the first workday of the year. Packets began dropping from around 6:00 AM PT, and by mid-morning Eastern time, as offices logged on for the day, the platform was down. Slack confirmed the outage in a status update at 10:14 AM ET. Service was not fully restored until approximately 2:30-3:10 PM ET, roughly five hours later.

The root cause was straightforward: AWS Transit Gateways did not scale fast enough to handle the traffic surge. Autoscaling is reactive by design. It watches load, detects increased demand, and then provisions capacity to meet it. This works well when traffic grows gradually. It works poorly when the entire working world comes back from Christmas vacation on the same morning and logs in at once.

AWS engineers manually increased Transit Gateway capacity, and the change had rolled out by 10:40 AM PT. The fix was not complicated. The problem was that it arrived several hours after the load did, because the scaling mechanism was designed to respond to trends it had already observed rather than events it could have anticipated. The traffic pattern that broke it was not a spike from an incident or an attack. It was everybody coming back from Christmas on the same morning, which is the most predictable surge in the entire corporate calendar.

The Takeaway

Autoscaling responds to load it has already seen, which makes it weakest against exactly the surges you can predict in advance. The first Monday of January, a product launch, a campaign send: these are known dates. Pre-scaling for a known event is cheaper than discovering your ceiling during it. If your infrastructure depends on autoscaling, identify the predictable load spikes in your calendar and scale manually before they arrive. Tools like those in our guide to AWS monitoring tools can help you understand your baseline and plan accordingly, but no monitoring tool will solve a problem that requires deciding something in advance.

Sources

Get the shortlist, not the noise

One email a week. The tool we would actually buy, and why.

Join the newsletter