Back to Blog

Six Microsoft 365 services went down together. For your continuity plan they are one

An open main electrical distribution cabinet in a plant room, with rows of identical breakers all fed from a single main switch

Monday's incident ended up taking down six services you have listed separately: Exchange Online, SharePoint Online, OneDrive, Teams, Purview and Microsoft Defender XDR. They share one authentication component underneath, and that list of six names is the most useful thing to come out of Monday night. You got it for free, and it tells you which rows of your continuity plan are not independent of each other. One of them, the row fewest people will look at, is the console you monitor with.

The hours, including the ones that do not line up

Microsoft started investigating "an increase in user reports of Exchange Online problems" at 11:55 a.m. UTC on Monday 31 August. When the problem stopped being only about Exchange, the case moved to a second identifier, MO1465074, to which Microsoft assigned an official start time of 3:08 p.m. UTC.

This is where we have to stop, because the sources do not line up and a clean timeline would be more comfortable than honest. BleepingComputer gives 5:30 p.m. UTC as the moment Microsoft acknowledged incident EX1464935 on Exchange Online; Computerworld puts the first investigation message five and a half hours earlier. Both can be true —one is when the provider's clock starts, the other when most people found out— but they are not the same hour, and picking whichever suits the story would be cheating. If you need the exact time for an internal report, take it from your own tenant's admin center rather than from the press, this page included.

What is consistent is the sequence of diagnoses. First, that they had isolated "a common failure pattern across affected Exchange Online requests that is associated with authentication and protocol connectivity". Then, more precisely: "issues related to a core authentication configuration used by multiple internal services within the Exchange Online infrastructure". And the mitigation, by hand and server by server: "we're performing a manual test on the individual server level to reset to configurations to validate if this resolves the issue". By then the list already included OneDrive for Business, SharePoint Online, Teams, Purview and Defender XDR.

The return of mail flow has the same clock problem: Computerworld places it "by late Monday", while the last restoration message BleepingComputer records is at 03:43 US Eastern time, already Tuesday. What nobody disputes is that search was still broken on Tuesday.

There is a detail in those hours worth keeping for practical reasons: the official start time Microsoft assigned to the expanded incident, 3:08 p.m., is earlier than the time most coverage says it acknowledged it. The clock that counts for the provider is its own, and the place it is written down is your admin center. If at some point you need to claim, to justify a delay to a client, or simply to explain to your management what happened on Monday, export the incident history while it is still there. It is a screenshot and five minutes, and it is the one version nobody argues with later.

Nobody has confirmed it was an expired certificate

The headline that night was that Microsoft had forgotten to renew a certificate. The actual chain is this: during the outage some users hit a client-side error message quoting a specific thumbprint —19F04B8A233DD9CE916F118056D224A1751729EA— saying it had expired; a German blog published it that same night, and everything else came from there. Microsoft has never said that. It said "authentication configuration". It may turn out to be the same thing, or the certificate may have been a symptom.

We are not going to state it, because it has not been stated. And it makes little difference to what follows: a certificate that expires and a misapplied authentication configuration belong to the same family, the shared component that decides whether the others are allowed to talk to each other. When the final analysis lands, the technical write-up will change; nothing about what you should look at tomorrow will.

Six rows of your plan, one single component

Open your continuity plan, if you have one written down. Chances are email, the document intranet, chat and the security console occupy separate rows, each with its own criticality and recovery time. They sit that way because to the people doing the work they are four different things: four logins and four ways of working. Underneath they were one, and on Monday it showed.

That has an arithmetic consequence. Multiplying availabilities to conclude that "both going down at once is impossible" only works if the two are independent. A service commitment is signed per service and says nothing about correlation between them. On Monday the failure sat below the level at which those commitments are written, so none of the six protected you from the other five.

To be clear, consolidating authentication is not a Microsoft design mistake. It is what we recommend almost every time: less surface, one place to enforce policy, one thing to audit. Anyone telling you they keep it all separate probably has the problem distributed rather than solved. The price of that decision is exactly what was on display on Monday, and you pay it in one lump.

Before anyone reaches for the easy conclusion: this is not an argument against the cloud either. The on-premises version of the same failure exists and is worse. An identity federation certificate that expires on a Sunday night locks you out of email, the intranet, internal applications and —the fun part— the console you were going to fix it from, because that authenticates against the same thing. The difference is not architectural, it is identical; it is that there is no status page, nobody to call, and you are the one counting the clock at three in the morning. What changes when you move to a large provider is who pays for the on-call rota, not whether the shared component exists.

The console you turn to when something breaks was inside the thing that was breaking

Go back to the list of six. Four of those names are work tools and you miss them straight away. The other two, Microsoft Defender XDR and Microsoft Purview, are nobody's way of writing an email: they are the security console and the compliance and audit layer. They were degraded during the same window in which tens of thousands of clients were retrying authentication in a loop.

It is worth being precise about the little that is known. "Degraded" is not "off", and Microsoft did not spell out which part of each product was affected: it could be the console, an integration, or just search inside the portal. The sensors collecting telemetry on endpoints do not depend on your ability to open the panel, so the reasonable assumption is that collection continued. What stops when the console is stuttering is the human part: looking at the incident queue, running a query, and deciding.

Let us be clear about what we are not saying: there is not a shred of public evidence that anyone took advantage of that window, and we are not going to imply it to make the paragraph exciting. What we are pointing at is duller and more useful. A spike of authentication failures is exactly what an outage like this produces, and also exactly what a credential-stuffing attack produces. The tool you use to tell those two apart —and the one that afterwards lets you reconstruct what really happened— was on the same list as the fault.

The domestic version of this takes ten minutes to check and is the one we run into most often: the system that alerts you usually lives inside the system it watches. If your monitoring sends alerts through the mail of the provider that just went down, that day you find out about nothing else — and it is precisely the day the alert mattered. Same with the admin center: it is where the incident gets published, and it lives inside the same tenant. The criterion we take from this, and apply when setting up managed monitoring and response, is that neither the alerting channel nor the telemetry destination should depend on what they are watching.

On the plan B for email, briefly

It is the proposal that always shows up forty-eight hours later, and we already devoted a whole post to why standing up a second suite "just in case" rarely pays off; we will not repeat it here. What Monday adds is a question that post did not ask, and that on its own decides whether the plan B is worth anything: which identity does it authenticate against? If it is the same directory, it goes down with you and the price is irrelevant. Some do ship their own credentials precisely for this scenario; check which one you are being sold before you look at the cost.

The second day is the expensive one

Restoring mail flow is the visible part. What ate the working hours was Tuesday: search was still degraded across Exchange Online, SharePoint Online, OneDrive and Teams —four products— and Microsoft 365 Copilot prompts requiring Microsoft 365 data were failing too. People could not find the attachment from three weeks ago or the history of a conversation, and some organisations were still draining backlogged mail queues. A few weeks ago we wrote about a broken search that, for the service commitment, was not downtime, and here it is again: long partial degradation does not count on the meter and does count on the payroll.

When this does not apply to you

If you are twenty people and email can be down for three hours without breaking a single commitment to a client, do none of this. Note the date, write down what stopped working, and get on with your day. The conversation changes when email is where orders arrive, when there is a contractual obligation to reply within a deadline, or when your sector requires you to demonstrate you have a continuity plan. And it changes too, even at twenty people, if your security console turns out to depend on the same place as your email: that is worth knowing on an ordinary Tuesday rather than on the day it matters. We cannot fix a Microsoft certificate, and anyone implying otherwise is selling smoke; what a 24x7 support service does at half past seven on an August Monday evening is notice it before your users do and decide quickly whether the problem is yours. Building architecture for a two-hour-a-year event is an expensive way to feel calm.

Sources (verified 1 September 2026): identifiers EX1464935 and MO1465074, acknowledgement time (5:30 p.m. UTC), the list of six affected services and the three verbatim Microsoft statements — BleepingComputer, 31 Aug 2026; the 11:55 a.m. and 3:08 p.m. UTC timestamps, mail recovery "by late Monday", and second-day search degradation across Exchange Online, SharePoint Online, OneDrive and Teams plus Copilot prompts requiring M365 data — Computerworld, 1 Sep 2026; the client error message quoting the certificate thumbprint — Born's Tech and Windows World, 31 Aug 2026. The timestamps in the first two sources do not agree, and this post says so rather than picking one. Our own judgement, not reported fact: reading the service list as a map of shared dependencies, the observation about Defender XDR and Purview, and the rule that the alerting channel and the telemetry destination must not depend on what they watch. Microsoft has not confirmed the root cause as an expired certificate.

Could you say which parts of your plan hang off the same component?

At everyWAN we design, deploy and maintain infrastructure for companies, with 24/7 managed services, and we apply the principle that the alerting channel must not depend on what it watches. If you want someone to go through your dependency map with that list of six names in front of them, let's talk.

Talk to everyWAN

Tags:

Share:

Subscribe to our newsletter

To receive IT stories, everyWAN news and exclusive subscriber offers, sign up to our mailing list

Minorisa de Sistemas Informaticos y Gestión S.L. © 2026
everyWAN
everyWAN