Microsoft lists five impacted services in the Azure outage of 30 September. Four of them are the road to your infrastructure: ExpressRoute Gateway, VPN Gateway, Azure Firewall and Application Gateway with its WAF. The fifth is Azure VMware Solution.
If your company is in Spain and has no night shift, you probably found out in the morning. It began at 22:30 on the Wednesday, Spanish time, and was declared mitigated at 04:15 on the Thursday: five hours and forty-five minutes. The record carries the identifier 7Q30-010 and is still published in the Azure status history, which is where everything quoted here comes from.
Four doors and a destination
The first four entries on the list are the edge of your network in Azure: that is where the dedicated line lands, where the tunnels come in and where the traffic you publish to the internet goes out. The fifth, Azure VMware Solution, is a different animal: virtual machines run inside it. If your contingency plan for the VMware estate runs through AVS, on the 30th the road and the possible destination were on the same list of impacted services.
It is worth saying straight away what the record does not document, because that matters as much as what it does: it never mentions loss of machines or data. What it describes is this, verbatim: "a subset of customers using gateway services in multiple regions experienced degraded or interrupted network connectivity. Impacted customers may have observed gateways failing to load in the Azure Portal, along with failures or delays in network management operations across the following impacted services".
Note the may have observed: Microsoft leaves it conditional, and so do we. The second half of that sentence is what turns a provider incident into your incident. Gateways might fail to load in the portal, and network management operations might fail or lag. What was in doubt, then, was the ability to touch the network. If your contingency procedure says "if the main line drops, we bring up the backup tunnel from the portal", that procedure needed exactly the thing that was in doubt.
What Microsoft says happened
The explanation runs to four sentences, plus a fifth with the mitigation. Here they are in full, uncut, because the detail matters:
«Our investigation identified that a recent change to a regional gateway management service triggered a higher-than-expected load when an unrelated operating system servicing maintenance proceeded gradually through multiple regions. Normally, this gateway management service would auto scale as needed. Demand on dependent services increased due to the increased workload, preventing these regional services from scaling as expected. The operating system servicing maintenance was paused as a precaution.»
«We reverted the contributing regional gateway manager change, which reduced gateway manager load and allowed affected services to recover.»
Microsoft's attribution is clear and should not be twisted: the gateway manager change is the contributing one, and it is the one that gets reverted. The operating system maintenance is only paused as a precaution. Anyone wanting the short version already has it: somebody shipped a change, the change produced more load than expected and, on reverting it, the service came back.
What interests us is in the first sentence, and this is our own reading, flagged as such. The load appeared when an unrelated maintenance was advancing gradually across several regions. Nobody reviewing the gateway manager change had any reason to open the operating system maintenance calendar: they were two separate jobs on separate things. And advancing in phases, which is the good practice and what we always recommend, stretches out the window in which two unrelated jobs can coincide on the same infrastructure.
This re-attributes nothing: the change is still the contributing one. What it adds is a governance question that applies to any company, not just a hyperscaler. A phased rollout answers "does this change break anything?" well. It does not answer "what else is rolling through this infrastructure right now?". And that second question almost never has an owner. Microsoft itself says it is still investigating "the scaling behavior and the safeguards needed to help prevent recurrence".
The mechanism at the end of the quote deserves a line. The manager auto scaled normally; demand on the services it depends on rose because of that extra load, and that prevented the regional services from scaling as expected. There is a chain of dependencies there that saturates end to end, which is precisely what turns autoscaling into an ornament. We covered the case of redundancy that did not survive a procedure a month ago, and the problem is from the same family: an assumption nobody ever wrote down.
The clock, which is the uncomfortable part
The record publishes six timestamps. The subtractions in the last column are ours, done on those timestamps:
| UTC | What happened | Since impact began |
|---|---|---|
| 20:30 | Customer impact begins | — |
| 21:29 | They begin investigating ExpressRoute Gateway in UK South | 59 min |
| 22:27 | They identify it affects multiple regions | 1 h 57 min |
| 23:05 | They correlate with the OS servicing and pause it | 2 h 35 min |
| 01:36 | Recovery progressed; regions still being fixed | 5 h 06 min |
| 02:15 | Mitigated, after reverting the contributing change | 5 h 45 min |
Almost an hour before they began investigating, and when they did it was for one service in one region. Almost two hours before they identified that the problem spanned several. Two and a half hours before they tied it to the maintenance. From there it was closed in three hours ten. Working out what was happening cost considerably more than fixing it. This is not about singling out Microsoft: if it takes Microsoft two and a half hours to correlate two of its own jobs, with all its telemetry, your odds of correlating a change of yours with one of your provider's —blind, at night and without sight of their calendar— are rather worse.
How many regions: the figure that is not in the source
A specific figure is going round these days: 18 or 19 regions. We went looking for that number in the status history and it is not there: the official record says "multiple regions", with no count. The only list of regions it gives appears in the 01:36 entry, and it refers to those still to be recovered: "including France Central, North Europe, Southeast Asia, UK South, and UK West". Before that it names just one, UK South, which is where it started investigating at 21:29. That including also indicates that even this list is not closed. That figure may well be right; it simply cannot be verified against the primary source, which is why we do not use it as data here.
There is one more detail, written into the entry's own header, that explains a good deal about why the numbers wobble: "Incorrect Region: correction: Impacted region(s) within a previous communication has been corrected". Microsoft corrected the region list relative to an earlier communication. In practical terms: the notice you read in the heat of the moment, while deciding whether to fail over or wait, is not the notice that remains. Anyone who decided that night on the basis of "my region is not on the list" was using a figure that later changed. We wrote about this same kind of small print when reading the perimeter the provider publishes, in the piece explaining that your cloud can lose one zone, not a region.
Three things that are in your hands
You cannot stop your provider crossing two jobs. You can make the next night like that cheaper. Three concrete things, checkable this week:
- Check how you administer things when the portal will not answer. If the only way to touch the firewall or the gateway is the provider's console, a management plane incident leaves you with no hands. It is worth writing down what the alternative route is and, above all, testing it once with the portal deliberately closed.
- Measure the road, not only the destination. Monitoring that only checks whether the virtual machine responds from inside the cloud stays green throughout a gateway incident. We measure end-to-end latency and loss with Zabbix and SmokePing precisely for this: the symptom that counts is the one the user sees from where they are.
- Time the plan. The useful question is how long it took the last time you actually ran one. Our last full recovery drill came in at 14 minutes: that is an internal figure, from our own test, and we offer it as evidence rather than a promise. What counts there is having started the stopwatch.
The report that will explain it should arrive, in theory, in mid-October
The record includes a commitment: "Once that is completed, generally within 14 days, we will publish a Post Incident Review (PIR) to all impacted customers". It is worth reading carefully, because there are two assumptions chained together. The fourteen days start counting when the internal retrospective finishes, and no date has been published for that retrospective, so our mid-October estimate is exactly that, an estimate. And that to all impacted customers points at those affected: the document that would say which safeguard they put in —and therefore whether your design has to change— is addressed to them.
Hence the insistence on Azure Service Health alerts configured with a real recipient. The live notification arrives late and with the region list half corrected, but it is the channel through which the analysis arrives afterwards, and that analysis is the only thing of any use for redesigning anything.
The conflict of interest up front, as always: everyWAN makes its living designing and operating infrastructure and cloud, so when we say "measure the road" we are describing something we invoice for. What does not depend on us is the verifiable part: incident 7Q30-010 is published and linked at the end, the quotations are in English and verbatim, and anyone can redo the table's subtractions from the six timestamps.
Sources (verified on 5 October 2026): impact window, list of impacted services, symptom description, Microsoft's account of what it knows so far, reversion of the contributing change, the six timestamps, the region correction note and the PIR commitment — Azure status history, incident 7Q30-010 ("Mitigated- Multiple services experiencing connectivity issues in multiple regions", 30-09-2026). Every italicised quotation is taken verbatim from that entry. The following are ours, and are flagged as such in the text: the conversion to Spanish time (UTC+2), the durations in the table, the mid-October estimate for the PIR, and the reading about the window a phased rollout opens. The figure of 18-19 regions going round these days does not appear in the official record, which is why it is not used here. The 14-minute recovery drill figure is everyWAN internal, from our own test, and is offered as evidence rather than a contractual commitment. Cover photograph: "Cable closet bh.jpg", Wikimedia Commons, public domain; cropped and darkened by us.
How do you get into your network the day the portal will not load?
We review what each path of your infrastructure and cloud depends on, set up end-to-end monitoring and connectivity from networks and communications —we are an operator with our own network— and put a stopwatch on your disaster recovery plan. Without moving you off your cloud if there is no need.
Talk to everyWAN