On Wednesday the 23rd, at 14:44 UTC, Microsoft 365 started failing: Teams not loading chats, SharePoint throwing errors, OneDrive intermittent, and the admin center — the place you'd go to check what's wrong — barely loading. The cause, per Microsoft: a network configuration change on their side. If your company didn't feel it, that wasn't your architecture's merit: the impact concentrated on North American network paths, and this time geography played in your favour. Which is why this post isn't about Microsoft. It's about a question almost nobody has answered: what does your company do during the hours when the entire office lives in someone else's cloud, and someone else's cloud is down?
What happened, with data
Microsoft logged it as incident MO1437424 (with a sibling Azure incident, ZJV6-SGG). It began at 14:44 UTC and affected, to varying degrees, Teams, SharePoint Online, OneDrive, Power Automate, Copilot, Loop, Purview, Power BI and the Microsoft 365 admin center itself. On Downdetector, reports jumped from a baseline of 29 to over 2,400 in under half an hour. Mitigation was the classic one for this category of incident: revert the network change and reroute traffic through alternative paths. The bulk of the impact lasted about four hours; formal closure came that night.
Read that again from a company's point of view: for one working afternoon, chat, telephony for those who run it on Teams, files, automated workflows and the panel used to administer all of it were failing at once. Not because of an attack, not because of a burning datacenter: because of a routine configuration change on the provider's side. It's Microsoft's third major incident in nine months with that same shape, and the pattern isn't exclusive to them.
It's not lightning: it's the weather
- →October 29, 2025: an inadvertent configuration change in Azure Front Door took down Microsoft 365 services, the Azure portal, Xbox and thousands of customer websites for over eight hours.
- →November 18, 2025: Cloudflare's global outage, which left half the Internet returning errors for hours.
- →January 22, 2026: Exchange Online and Teams were degraded for about ten hours due to a load issue during maintenance. Our favourite detail: a load-balancing change applied to speed up recovery prolonged it, and the incident wasn't closed until almost 24 hours later.
- →July 16, 2026: the AWS CloudFront outage, which we wrote about last week. One week later, it was Microsoft's turn.
Three providers, five incidents and not a single attack: in every one, the failure was born inside the provider's own operations — a change or an operation of theirs that went wrong. No broken hardware, no hackers, no lightning. At this scale, the biggest threat to the cloud is the hand that administers it — and against that there is no patch you can apply and no firewall you can buy. Anyone selling you otherwise is selling you something else.
You can't disaster-recover a SaaS
When an ERP on your own infrastructure goes down, you have a runbook: a replica to promote, a backup to restore, a hypervisor to log into. That's the world of classic disaster recovery, and there is plenty to design there. When Teams goes down, none of that exists: there is no Teams replica you can boot, no server to fail over to, and your escalation consists of opening a ticket that queues behind hundreds of thousands of other organizations'. All of DR presupposes one thing: the infrastructure is yours, or at least you can stand up another one. With SaaS, neither holds.
What about the SLA? Most Microsoft 365 services carry a financially backed 99.9% SLA: if it's breached, you can claim service credits — a percentage discount on that month's bill. It's a contract about money, not about your continuity. Nobody gives you back the lost afternoon, the orders that didn't come in or the calls that didn't ring. An SLA compensates; it doesn't continue.
This changes your IT team's role (or ours, when it's us) during the outage: they are not there to fix the incident, because they can't. They are there for something else: knowing what's going on, communicating it inward with judgement, sustaining what does depend on you, and preventing nerves from creating a second incident — that one genuinely yours. That function isn't improvised at 5 p.m. It's decided beforehand.
What we wouldn't do
- ✕Standing up a second suite "just in case". A standby Google Workspace sounds great on the slide: in practice you pay two licenses, maintain two security configurations and double your attack surface so that, on day D, nobody even remembers the password. That's PowerPoint continuity, not business continuity.
- ✕Fleeing "to another cloud" as a reflex. In nine months Cloudflare, Azure, AWS and Microsoft 365 (twice) have all gone down. Switching providers doesn't remove the failure category; it just changes the logo on the status page you keep refreshing. We say this while managing Microsoft 365 tenants daily, on no vendor's commission: for the workplace, M365 still delivers more than it costs. The problem isn't your choice of provider: it's having no plan for its bad hours.
- ✕Touching your own configuration mid-outage. Reconfiguring 200 users' Outlook, changing DNS, bypassing the proxy "to see if that helps". Remember January: it was Microsoft itself that prolonged its outage with a change meant to shorten it. If that happens to the vendor on its own platform, picture your team improvising on yours. During a provider outage, rule one is not to turn their incident into yours: everything you change today you'll have to undo tomorrow.
The plan that works fits on two pages
Nothing below requires buying technology. It requires deciding calmly, writing it down and rehearsing it once. It's the part of business continuity that actually gets used:
- 1.Plan-B communication channel. The trick question: the day Teams is down, where do you announce that Teams is down? A messaging group on people's phones and a phone list that also lives outside OneDrive. It costs zero and it's the first thing that fails in almost every company.
- 2.Who watches, who tells. One person tracks status (Microsoft's status page, the admin center if it responds — on Wednesday it was struggling too, bear that in mind — and Microsoft's status account on X) and another communicates inward at fixed intervals. Everyone else works with whatever is up. Without an owner, "incident management" is forty people refreshing Downdetector.
- 3.Local work by design. Desktop apps installed rather than browser-only, Outlook in cached mode, and critical files marked to always keep on the device. A web-only organization stops dead; one with desktop clients and sync limps along, which is infinitely better. You set this up on any given Tuesday, not during the outage.
- 4.Map of non-obvious dependencies. Is your phone system Teams Phone? That day there's no phone either. Are Power Automate flows sitting in the middle of orders or alerts? That day they don't run. Do your VPN or internal apps authenticate against Entra ID? An identity outage can lock you out even of what you host at home. You don't need to redesign everything: you need to have listed it beforehand, so that day is a glance and not a discovery.
- 5.Threshold and degraded mode decided in advance. From what point does what kick in? At 30 minutes, a general notice via channel B; at the hour mark, phone for customers and the alternate mailbox for urgent orders. The exact numbers matter less than having decided them calmly, with a Tuesday-morning head rather than a Wednesday-at-five one.
- 6.And the backup nuance, for honesty's sake. A Microsoft 365 backup does not give you Teams back during an outage: there is no "restore Teams onto your server". Backup protects you from the opposite risk — deletions, ransomware, retention — which Microsoft doesn't cover for you. Two different risks, two different answers, and you need both: don't let anyone sell you one as the cure for the other.
Continuity isn't bought: it's decided
The next major provider outage isn't a hypothesis: judging by the last nine months, it's a matter of weeks before it's someone's turn again. You can't prevent it and you can't shorten it. The only thing you decide is whether those hours find your company with six written decisions or with forty people looking at each other. At everyWAN this is exactly what we work on within compliance and continuity: plans proportionate to each company's size — the kind that fit on two pages and get rehearsed once a year, not the kind gathering dust in a drawer. If yours today is "let's hope it doesn't go down", let's talk.
Sources (verified): July 23, 2026 Microsoft 365 outage, affected services, North America scope, Downdetector peak and network-change rollback — BleepingComputer; incident MO1437424 timeline (start 14:44 UTC, resolution and rollback) and full service list — Cybersecurity News; analysis of the October 29, 2025 Azure Front Door outage (configuration change, duration) — ThousandEyes; January 22, 2026 Exchange Online/Teams outage and the load-balancing change that prolonged recovery — Messageware; 99.9% SLA with service credits — Microsoft, SLA for Online Services.
Does your continuity plan depend on Teams being up?
At everyWAN we help companies decide calmly what happens when their provider has a bad day: plan-B channel, mapped dependencies, degraded mode and a plan that fits on two pages and gets rehearsed. Better to write it on a quiet Tuesday than improvise it on a Wednesday at five.
Talk to everyWAN