On Tuesday 15 September AWS updated its service health dashboard. It was not a press release or a post-mortem: it was a line of administrative text. It said that, after a thorough assessment, it is unable to restore access to the resources and data hosted exclusively in that Region. That word — exclusively — is what separates the customers who permanently lost data from the ones who just had a very long scare.
As far as public reporting goes, it is the first time a hyperscaler has written off an entire Region because of a military attack. It is worth reading exactly what was said, slowly, because almost everything you need in order to review your own setup was already published before any of this happened.
What happened, plainly
In March 2026, during the war with Iran, drone strikes hit AWS facilities in the United Arab Emirates and physically damaged infrastructure in Bahrain. In April a second Bahrain zone went down. Six months later, on 15 September, AWS published the outcome: the Bahrain Region (me-south-1) is not coming back, and in the UAE one of the three availability zones, the one identified as mec1-az2, is not coming back either.
The technical sentence from AWS, the one that explains why, is this: "The damage to our infrastructure spanned multiple Availability Zones and exceeded what our regional and multi-AZ services are designed to withstand."
A war is not measured in availability zones, and what it leaves behind first is not data. We are not going to comment here on the conflict, or on who fired what: that is not our subject and it is not our place. What is our place is what comes right after that technical sentence, because it is literally our trade: look at an infrastructure and say what happens if a chunk of it disappears. And there is a lesson there that applies as much to a company in Sabadell with two servers as to a Gulf bank.
Because AWS also said this: most customers resumed operations elsewhere, restoring from backups or reaching data that was still accessible. The ones who lost nothing were not luckier. They had something outside.
The perimeter was published all along
Everyone in this industry knows the S3 durability number: eleven nines, 99.999999999%. It gets quoted in meetings as if it were a property of the data itself, like its size or its format. Almost nobody opens the page it comes from, and that is where the problem starts: the two sentences around it are what define what the number is actually talking about.
The first: the standard storage classes "redundantly store objects on multiple devices across a minimum of three Availability Zones in an AWS Region". A minimum of three availability zones within an AWS Region. The second, the design goal, in the singular: they are "designed to sustain data in the event of the loss of an entire Amazon S3 Availability Zone". The loss of one entire zone. One.
And there is a third fact in the same paragraph that almost nobody quotes, and it is the most revealing of the three. Zones "are physically separated by a meaningful distance, many kilometers, from any other Availability Zone, although all are within 100 km (60 miles) of each other". Separated by a meaningful distance, several kilometres, but all within 100 kilometres of one another.
A hundred kilometres is an excellent distance against a fire, a flood, a power cut or a builder with an angle grinder. It is a reasonable distance against a local earthquake. And it is a distance that, against certain kinds of event, means nothing at all, because the event is bigger than the radius. That is not an AWS failure: it is exactly what AWS wrote down that it did. The dashboard sentence closes the loop with almost uncomfortable precision: the damage exceeded what our regional and multi-AZ services are designed to withstand. The design did not fail. The event left the perimeter that the design itself publishes.
It is the same trap we have described before in different clothes: two duplicated things are only two if they do not share the point that can fail. Redundancy is not a number of copies: it is a list of things that can happen without you noticing.
The division of labour is in the contract, with its own heading
The Amazon EBS documentation opens the snapshots chapter with an Important callout that leaves no room for interpretation: "AWS does not automatically back up the data stored on your EBS volumes. For data resiliency and disaster recovery, it is your responsibility to create EBS snapshots on a regular basis, or to set up automatic snapshot creation by using Amazon Data Lifecycle Manager or AWS Backup." AWS provides the tooling to automate it; switching it on is yours.
And when it explains where that snapshot lives, it repeats the perimeter: "Snapshot data is automatically replicated across all Availability Zones in the Region." Across every zone in the Region. The snapshot you take every night, the one that lets you sleep, sits by default inside the same perimeter as the original data. It protects you from human error, from deletion, from ransomware and from losing one zone. It does not protect you from losing the Region, unless you deliberately copy it to another one.
The contract says the same thing in lawyer language. The AWS Customer Agreement has an entire section, 2.3, titled "Your Security and Backup": "You are responsible for properly configuring and using the Services and otherwise taking appropriate action to secure, protect and backup your accounts and Your Content in a manner that will provide appropriate security and protection, which might include use of encryption to protect Your Content from unauthorized access and routinely archiving Your Content."
And section 11.3, Force Majeure, lists which events excuse performance. Among acts of God, earthquakes, power outages and riots, at the end of the list, spelled out: "acts of terrorism, or war". The word was in the contract years before anybody needed it.
In case there was any doubt about what gets paid when things go wrong, the compute SLA closes it from both sides. The remedy: "Unless otherwise provided in the Agreement, this SLA sets forth your sole and exclusive remedies, and AWS' sole and exclusive obligations, for any unavailability, non-performance, or other failure by us to provide Amazon EC2." Service credits, and nothing else. And the exclusion, which in the original enumerates both SLAs and their cases: the SLAs "…do not apply to any unavailability, suspension or termination of Amazon EC2… caused by factors outside of our reasonable control, including any force majeure event…" In a force majeure event the SLA simply does not apply.
AWS has been generous beyond all of that: it waived March usage billing in the affected Regions and suspended charging while the situation lasted, a gesture Forbes estimated in April at around 150 million dollars — a press estimate, not a figure AWS has published. We mention it not as a reproach but for the contrast: a credit offsets an invoice. It does not offset data. And nobody ever claimed it would; it is written down that it does not.
This is not about AWS
The easy reading of this story is "see, the cloud". It is a bad reading, and a convenient one, because it saves whoever makes it from looking at their own setup. A hyperscaler publishes its design perimeter with names, kilometres and a durability target. Your server room publishes none of that, which does not mean it has no perimeter: it means you have not written it down.
The perimeters we most often find on site visits: the backup NAS in the same cabinet as the server it backs up. The second cluster node in the next rack, on the same electrical panel. The offsite copy that does leave, yes, but onto a disk that lives in the same building, in a drawer. And the most common case of all, the one nobody even perceives as a decision: mail and files in Microsoft 365 with no copy outside Microsoft, because "it is already in the cloud".
We repeat it often because it is the axis of everything we do: failure is inevitable, an outage is a design decision. A dying disk is a failure. Your warehouse not shipping orders because of it is an outage, and the difference between the two was decided years ago by whoever chose where to put the second copy. Redundancy is built in layers — disk, machine, room — and each layer covers the one below. What AWS has done this week is show the layer above, the one almost nobody draws because it looked theoretical: the Region.
The four questions we write on one sheet
This is not a list of best practices. It is literally what we write down, in this order, when we walk into a company to review continuity. Four questions, and all four are answered with a proper noun or a number, never with an adjective.
- If the whole ROOM disappears, what does not come back? Not the disk, not the server, not the node: the room, the building, the site. It is the question almost nobody asks out loud because it sounds alarmist, and it is the only one that orders all the others. The answer is a list of systems, not a "well, we would have the backups".
- Where is the copy that does NOT share that room? Name of the site, how many kilometres away, which provider, who holds the credentials to read it. If the answer is "it is in the cloud", the question still stands: in which Region, and whether that Region is the same one the system runs in. If both answers match, you have a copy, not two sites.
- How long does it take to be operational from there? Timed, not estimated. And dated: when was the last time somebody actually held the stopwatch. A restore nobody has performed has no RTO, it has a hope. A continuity plan nobody has rehearsed is a document, and documents do not boot servers.
- If the perimeter breaks, who is on the hook? This one is contractual, not architectural, and it is usually the unanswered one. Open your provider's agreement — whichever it is — and search for two things: backup and force majeure. You will almost certainly find that the first points at you and the second excuses them. That is not a scam: it is the division of labour, and it is better learned on signing day than on the day something breaks.
What AWS recommends to its own customers
Worth saying, because it dismantles the caricature: none of this is our discovery, or a criticism of the provider. It is in AWS's own disaster recovery whitepaper, and the sentence is explicit: "If your definition of a disaster goes beyond the disruption or loss of a physical data center to that of a Region or if you are subject to regulatory requirements that require it, then you should consider Pilot Light, Warm Standby, or Multi-Site Active/Active."
In other words: if your definition of a disaster includes losing an entire Region, then backup-and-restore inside that Region is not your strategy. And the same document warns that the most ambitious of the four, multi-site active/active, "is the most complex and costly approach to disaster recovery". That caveat matters as much as the first one: most of the companies we work with do not need active/active, and selling it to them would be selling smoke. What everyone does need is to know which box they are in and to have chosen it, rather than landing in a box through an accumulation of small decisions.
And a note from the same document, in a callout box, that we could have signed ourselves: "Your backup strategy must include testing your backups."
What we do, and what we will not promise you
On the systems we operate, recovery is timed. The last full recovery drill we ran came in at 14 minutes. That is an internal figure from one of our own tests, on a specific infrastructure with a specific procedure: it is evidence, not a promise, and it does not mean your case will take that. It means the number exists and carries a date, which is exactly what is missing from most continuity plans we come across.
What we will not tell you is "with us you are safe". Our datacentre is a place too. It is a room with a postal address, an electrical panel and a roof, just like yours and just like Amazon's. A colocation site does not stop being a dot on a map because we are the ones operating it. That is why, in the designs we put our name to, the second copy leaves our perimeter as well: another location and, when the case calls for it, another provider and another technology. We are not selling the disappearance of risk, which does not exist. We are selling risk that is written down, distributed and rehearsed.
There is a line we use a lot: an untested backup is not a backup, it is a lucky charm. The AWS customers who came back are the ones who had something outside and knew how to read it. And there is another uncomfortable lesson we saw in the case where cutting a company off took 101 minutes and reconnecting it took nine days: breaking is fast; coming back is not. Which is why the useful question is not how long it takes to lose something, but how long it takes to have it again.
Could you say today where the copy that does not share a room with your servers lives?
We work on compliance and continuity in that order: first write down the real perimeter — which systems, which room, which contract — then decide which RTO and RPO you can defend in front of a board or an auditor, and only then build the disaster recovery that holds those numbers up, with the replica outside the perimeter and a timed, dated drill. When the second location has to be physical, we put it in our colocation; and when the sensible answer is a different provider, we say so. We do not resell AWS or any other hyperscaler.
Talk to everyWANWhat we are not claiming
We are not saying AWS breached anything, or that its services are unreliable: we are saying its design perimeter is published and the event exceeded it, which is what AWS itself wrote. We do not know which specific customers lost data, how many, or of what kind; AWS has not published that and we are not going to assume it. We do not know whether those affected had copies outside the Region, beyond AWS's general statement that most customers resumed operations elsewhere. We have not run workloads in me-south-1 or in the UAE Region, so we are not reporting first-hand experience of them. Nor have we read any affected customer's actual contract: the clauses we quote are the public, general ones, and a negotiated enterprise agreement may say something different. And we have no stake in which provider you choose: we do not resell AWS, Azure, Google Cloud or anybody's licences.
Note on sources
All consulted on 19 September 2026. One, the event: the AWS service health dashboard update of 15 September 2026, reported by Reuters and the international technical press, source of the two verbatim quotes — "The damage to our infrastructure spanned multiple Availability Zones and exceeded what our regional and multi-AZ services are designed to withstand" and "we are unable to restore access to the resources and data hosted exclusively in…" — the identification of the Bahrain Region (me-south-1) and the mec1-az2 zone in the UAE, the March and April 2026 dates, and the statement that most customers resumed operations elsewhere by restoring from backups. The 150 million dollar estimate comes from Forbes, in April 2026, on the waived March billing. Two, the design perimeter: the Data protection in Amazon S3 page of the official documentation (the eleven nines, the minimum of three zones within a Region, the goal of sustaining the loss of one entire zone, and the under-100 km distance between zones) and the Amazon EBS snapshots page (the warning that AWS does not back up automatically and that responsibility is the customer's, and the replication of snapshot data across all zones in the Region). Three, the contract: the AWS Customer Agreement, sections 2.3 Your Security and Backup and 11.3 Force Majeure, and the Amazon Compute Service Level Agreement, sections Credit Request and Payment Procedures and Amazon Compute SLA Exclusions. Four, the vendor's own recommendation: the Disaster Recovery of Workloads on AWS whitepaper, chapter Disaster recovery options in the cloud. What is our opinion, flagged as such in the body: the reading that the problem is not the design but that almost nobody reads the published perimeter; the list of perimeters we find on site visits; the four questions; and the insistence that our own datacentre is also a dot on a map. The everyWAN figures — the 14-minute drill and the failure-versus-outage distinction — come from our internal resilience material.
Cover image: aerial photograph of the Ashburn (Virginia) data centre corridor, from Wikimedia Commons, public domain / CC0, cropped by us. The text and branding are ours, added on top.