There is a word buried in Uptime Institute's annual outage analysis that changes how the human-error part has to be read. Explaining what drives outages with human error behind them, this year's report says the leading cause is still "failures to follow established procedures". Established. As in: the procedure was written, approved and filed somewhere. And it was not followed.
That rules out the comfortable answer. If the problem were missing documentation, you fix it by writing documentation, and writing documentation is a task you can hand out on a Thursday afternoon. But the report does not say procedures are missing: it says the ones that exist are not executed when they should be. And a procedure nobody has ever executed is not a procedure. It is a hypothesis in the shape of a document. Flagging this now, so the headline does not sneak in through the back door: that is our reading, not the report's conclusion — it still puts power as the leading cause of impactful outages. We get to the full caveat in two minutes.
The numbers, and what they do not say
The report is Uptime Institute's Annual outage analysis 2026, published in May 2026 and authored by Douglas Donnellan, Andy Lawrence and Rose Weinschenk. These are the three figures we think matter, with their sample size alongside because you need it to read them properly:
- →87% of those who suffered an impactful outage in the past three years say it could have been prevented with better management, processes or configuration. That is
n=98, from the 2025 survey, and it is seven points more than in 2024. The figure is rising. - →92% say human error contributed, at least in a minor way, to their most recent impactful outage (
n=220). Historically Uptime puts it at between two-thirds and four-fifths of major outages. - →The leading cause of those human-error outages is failing to follow established procedures (
n=199), and the report says it remains so "as in previous years".
Now the uncomfortable part, which we would rather put before building anything on top of it. Those samples are small: ninety-eight, one hundred and ninety-nine, two hundred and twenty responses from data centre operators. They are self-reported and retrospective, and someone looking back at their own outage has every incentive in the world to see it as avoidable: they know how it ended. That 87% measures how many operators believe their outage was preventable, which is a different thing from how many were.
And there are more caveats working against the easy headline. Uptime says explicitly that it does not count human error as a root cause but as a contributing factor, and explains why: an outage from a misconfigured update may stem from a software defect, an incorrect change or testing that never happened, and pulling those three apart afterwards is close to impossible. The cause still leading genuinely impactful outages in their survey is power: UPS systems, transfer switches and generators. Not procedures.
All of that said, the thesis holds anyway, and it is our reading rather than a finding of the report: if the top item on the list of human-error reasons is skipping a procedure that already existed, then what is missing is not in the document. It is that nobody has ever done it with the clock running.
Why a procedure that exists gets skipped
It is almost never laziness. When we are called in to sort out somebody else's incident, what turns up is always a variant of the same thing: the procedure said something that was no longer true.
- ·The credential in step 4 was rotated by somebody in March and the document still has the old one.
- ·Step 7 says "log into the management console", and the management console lives inside the thing that is down.
- ·Two of the steps take half an hour each and the document lists them as if they were one click, so the timing the service was promised on was never real.
- ·Whoever wrote the document no longer works there, and whoever is on call at three in the morning is reading it for the first time.
None of those four failures is caught by reading. They are caught by executing. It is the same point we made when we wrote that your infrastructure documentation is lying to you, except there we were talking about inventory and here about instructions. A stale inventory costs you an hour. A stale procedure costs you that hour on incident day, which is the only day when time costs money.
The cost, in boardroom language
The report itself puts a price on the conversation. More than half of respondents, 57%, say their most recent major outage cost over $100,000. And for the second year running, one in five exceeds a million. Uptime attributes that less to more failures and more to the fact that ever more services depend, directly or indirectly, on a single data centre or a single availability zone.
There is a second figure that worries us more than the money: duration. Most publicly reported outages are resolved within twelve hours (55%), but the share running beyond 48 hours has grown for the second year in a row. Uptime links that partly to fibre and subsea cable cuts, which in 2025 occurred at more than double their long-term average and take longer because somebody has to physically go and repair them while coordinating several parties. It is exactly the argument behind our insistence that contractual redundancy is not physical route diversity: two carriers in the same trench are one carrier on the day the digger turns up.
And to translate it into the only format a board understands, availability arithmetic is not up for debate. 99% availability is 87.6 hours a year in the dark: more than two working weeks. 99.9% is 8.76 hours. 99.99%, 52.6 minutes. Every extra nine is not bought with luck; it is bought with architecture and with procedure. If you have not put numbers on this yet, the boring necessary part is in RTO and RPO with no fluff.
A Tuesday at ten, not a Saturday at three
The operational conclusion is uncomfortable but simple: you have to switch something off on purpose, with people watching, during office hours. Hyperscalers have been doing it for years and gave it a name — chaos engineering, Netflix's Chaos Monkey — but the SME version is not an automated beast killing machines at random. It is an appointment in the calendar, once a quarter, and a list of what breaks.
This is how we do it, and how we set it up for customers:
- Announce it. A surprise drill measures how much people panic, not how good the system is, and it burns the trust you need to run the next one. Date, time and scope in writing.
- One piece, and switch it off for real. Not "let us pretend": unplug the node, stop the service, cut the link. If the drill is a meeting about what would happen, the outcome is a meeting.
- It is run by someone who did NOT write the procedure. This is the step most people skip and the one that finds the most. Whoever wrote it fills the gaps from memory without noticing; a newcomer walks straight into them.
- Time five things. When the first symptom appears, when the first alert fires, when the first human finds out, when the decision is taken, and when the service is fully back. The gap between first symptom and first alert is your telemetry; the gap between decision and recovery is your procedure.
- Write down what was missing, not what went well. The useful output of a drill is a list of homework with an owner and a date, not a paragraph saying the team responded well.
- Next quarter, repeat with whatever went worst. If you always rehearse what already works, you are measuring your comfort zone.
Our most recent full recovery drill finished in 14 minutes. We say that with the caveat attached: it is our own internal figure, on our infrastructure and our procedure, and it is not a contractual guarantee and it is not your number. It is good for one thing, which is the thing that matters: it is a number that exists because somebody switched something off and looked at the clock. Yours, until you measure it, stays unknown, and unknown is the dangerous category: you cannot promise it, you cannot budget for it, and you do not know whether it fits inside a working day.
When you should NOT run a drill
This is where we part ways with the brochure version of this idea. Without some preconditions in place, switching something off on purpose means risking a real outage in order to write a tidy report. There are four of them, and if you are missing one, the drill is not your next step.
- ✕If you have never restored a backup. That is the drill before this one, and it is cheaper. Restore for real, stopwatch in hand — not check that the dashboard is green. It is the first thing we test when we set up managed backup. We wrote about the difference in restoring is not recovering.
- ✕If you have no one-click way back. Without rollback, switching a piece off stops being a controlled experiment and becomes an outage with witnesses. Start with the way back, as when we explained why deploying on "latest" takes away your reverse gear.
- ✕If there is no telemetry. Without metrics you measure nothing: you only learn something was wrong when somebody shouts. A drill without telemetry produces an anecdote, not a number.
- ✕If the only person who knows the system is on holiday, or it is month-end close, or peak season. A drill is a planned spend of attention. You do it when attention is plentiful, which is why Tuesday morning wins every time.
One more warning, because we have run into it: a drill does not replace change governance. Knowing how to recover from an outage does not reduce your odds of causing one next Friday with an unreviewed change. They are two separate jobs: a staging environment close to production, staged rollouts and somebody who reviews before touching anything. The classic study that put numbers on the weight of human error in large internet services is over twenty years old — Oppenheimer, Ganapathi and Patterson, USENIX 2003 — and even then it placed hardware at the bottom of the list. The Uptime report does not say the same thing with the same data, but it points the same way.
The three questions that reveal it in five minutes
You do not need an audit to know where you stand. These three questions, asked out loud in a management meeting, give you the full diagnosis from the look on people's faces:
- What happens today if any one machine in the system dies?
- When did you last restore a backup? Actually restore it, not "have" it.
- Who reviews the next change before it touches production?
If all three answers start with "well, we should" or "we have it documented", you already know which group you are in. None of this is a reproach: most of the companies we work with came to us exactly there, and getting there is normal when the business grows faster than the infrastructure. What is not normal is staying there.
Failure is inevitable; an outage is a decision
A disk breaks: that is a failure, and it will happen. That because of that disk the website stops selling, the warehouse stops shipping and the sales team stares at a frozen screen: that is an outage, and it was decided earlier, when somebody chose where the data lived. Same breakage, opposite outcome.
The Uptime report ends up, on its own and without being asked, in the same place: it says regular testing, staff training and more disciplined change management are likely to remain among the "most effective — and cost-efficient" ways to reduce outage risk. It is the most useful sentence in the whole report, and it is not about buying anything. It is about rehearsing.
We set this up as part of our compliance and continuity work: the drill calendar, who runs them, what gets timed and what gets fixed afterwards. The picture beforehand — which dependencies you actually have and which ones will not survive a test — comes out of consulting, and the getting-it-back-up part out of disaster recovery. None of the three starts by buying hardware. It starts by switching something off on a Tuesday.
Sources (verified on 3 September 2026): the 87% with better management/processes/configuration and its seven-point rise over 2024 (n=98, 2025 survey), the 92% naming human error as at least a minor contributor (n=220, Data Center Resiliency Survey 2026), "failures to follow established procedures" as the leading driver of human-error outages (n=199), the treatment of human error as a contributing factor rather than a root cause, the historical two-thirds to four-fifths range, power as the leading cause of impactful outages, the 57% above $100,000 and one in five above a million, both from the Global Data Center Survey 2025 (which the report also refers to as its 2025 annual survey), the 55% resolved within twelve hours and the second consecutive yearly rise in outages beyond 48 hours, fibre and subsea cable cuts at double their long-term average in 2025, and the sentence on regular testing and change management — Uptime Institute, "Annual outage analysis 2026" (UII Keynote Report 201, May 2026; Douglas Donnellan, Andy Lawrence and Rose Weinschenk). The weight of human error in large internet services — D. Oppenheimer, A. Ganapathi and D. Patterson, "Why Do Internet Services Fail, and What Can Be Done About It?", USENIX 2003. The nines arithmetic is direct calculation over 8,760 hours a year. The 14-minute drill is an internal everyWAN figure on our own infrastructure: a test that was run, not a service guarantee.
How long would it take you to come back? Put a number on it
We set the drill up, we run it, and we hand you the five timings and the list of what needs fixing. On a Tuesday morning.
Talk to everyWAN