- Date: Monday, 24 August 2020
- Duration: About 3.5 hours
- Context: 300 million daily meeting participants
On Monday, 24 August 2020, Zoom suffered a global outage beginning just before 9:00 AM Eastern Time. Users could not visit zoom.us, start or join meetings or webinars, or sign up for paid accounts. The platform had become critical global infrastructure during the COVID-19 pandemic, carrying 300 million daily meeting participants. The service was restored at approximately 12:30 PM ET, after roughly 3.5 hours of downtime.
The company acknowledged the situation via Twitter: "If you're having trouble connecting to Zoom, we have identified the issue and are working on a fix." What they did not disclose was the technical cause. No root cause was ever made public. What is on the record is the timing and the scale: a platform that had quietly become essential to how the world conducted business became unreachable at the start of a Monday morning.
It did not fail at 3am during a quiet patch. It failed just before nine on a Monday morning, at precisely the moment the largest number of people in the world were trying to join a meeting. Zoom's stock dropped more than 2% during the outage, a minor footnote to the thousands of derailed standup meetings, client calls, and first-day-back-at-work rituals happening across every timezone where Monday morning had arrived.
The platform that had absorbed the entire weight of remote work during a pandemic chose the worst possible moment to stop working. Not because of malice or incompetence, but because peak load is exactly when load-bearing infrastructure is most likely to fail. That is what peak means.
The Takeaway
If a tool has quietly become essential to how you work, the question is not whether it will have a bad morning but what you do during one. Load-bearing infrastructure fails at peak, because peak is when it is under load. Monitoring tools can tell you when your own services are down, but they cannot conjure a backup when someone else's platform disappears. The lesson is operational, not technical: have a fallback communication method that does not depend on the same service, and make sure your team knows what it is before they need it.
More cautionary tales at Stackery Quackery.