Back to Blog

The backup was safe. And it had spent 47 hours inside the same outage

A warehouse order desk with the monitor switched off and pushed aside, a spiral notebook with orders written by hand, delivery notes on a spike, a tape calculator and a phone off the hook

Two sentences from the same status page, published a day apart. Sunday evening's: "there is no risk of data loss. The stored information and data are safe". Monday's: "while the incident remains active, it is temporarily not possible to access backups or to migrate the affected services to another node". Both are true. Together they describe a situation that appears in almost no recovery plan: your data is perfect and there is absolutely nothing you can do with it.

The clock, with the source in front of us

We are talking about the outage of DonWeb's Nova node, an Argentine hosting and domain registrar. The clock we use here is theirs, from their public status page, because it is the only timeline not reconstructed by third parties. Their first entry about the incident is dated Sunday 30 August at 09:23 (Argentina time, GMT-3) and already reads "Identified". The last one we read before writing this is from Tuesday 1 September at 08:43, and it still says the same: incident active, unresolved.

Between those two markers there are 47 hours and 20 minutes. That is our own subtraction, and it is deliberately conservative: it counts from the moment the provider opened the incident on its status page, and not from the moment the first customer lost service, which was earlier. Local press puts the start "on Sunday", with no hour; there is no start timestamp in a primary source, so we are not inventing one. The real number is worse than 47, and 47 is already enough for what follows.

The scope, in the provider's words, is "100% of the servers on the NOVA node", with Cloud Servers inaccessible. Local press talks about hundreds of companies: billing systems, databases, e-commerce shops. The uncomfortable figure there is the hundred per cent. You can ride out a degradation by working slower; this is the main switch going off.

Identifying the cause is not a recovery milestone

The incident is opened on Sunday at 09:23 already in the "Identified" state, with this text: "Our technical team has managed to identify the cause of the problems and is already working to restore service as quickly as possible". The 09:35 update repeats it. At 11:14 the first concrete progress appears, "the replacement of a connection component". Two days later the service was still down.

This is the thing we have most often seen misread from the other end of the phone, almost always by management rather than by engineers. "They know what it is now" gets automatically translated into "so it will be quick", and decisions get made on top of that translation —wait instead of activating plan B, do not tell customers yet, do not set up the manual process yet—. They are two different jobs, and the second can take an order of magnitude longer than the first. On shared storage infrastructure, knowing which part failed tells you very little about how long what sits on top takes to become consistent again.

The two sentences have to be read together

"There is no risk of data loss" is, technically, very good news, and credit where it is due: the provider published it early and stuck to it. It means the RPO —how much information you lose— is zero or close to it. That is the number people look at most when buying storage, and the one that appears on every datasheet.

The second sentence says the RTO —how long until you are working again— has no number. And it has none for a reason worth facing head-on: you cannot speed it up yourself. The escape route any sensible plan assumes —"if this drags on, I restore the copy somewhere else and carry on"— ran through the same panel that was down. The copy existed, it was intact, and it was every bit as unreachable as the original server. This is different from what we described a few hours ago about Microsoft 365: there the problem was six services you thought were independent hanging off one shared component; here there is only one service, and what fails is the way out.

It is the same family of mistake we have already described in a domestic version: the backup server placed inside the very domain it protects. The scale is different and the mistake is the same. A copy inherits the availability of the place you restore it from, not its own. If the only route to it goes through the provider's panel, your copy has exactly the same recovery time as the provider, and the provider decides what that is.

The missing figure: the hour at which you stop waiting

When we wrote about RTO and RPO we insisted that the long stretch of being half-up has no number in any plan. This outage adds a figure that also has none, and it is the one you genuinely miss at hour forty-seven: at what point do you stop waiting for the provider?

Without that hour written down in advance, it never gets decided. The reason is psychological and fairly human: every status update looks like the second-to-last one, and starting up elsewhere means writing off the previous hours of work. So you wait one more hour. And another. It is exactly the mechanism that keeps people in a queue: you have been there so long that leaving now feels like waste. At hour forty-seven, the decision to wait was never taken; no decision was taken at all.

The figure is different in every company and comes out of a business conversation, not a technical one: how many hours of downtime the operation can absorb before the damage stops being recoverable —customers who leave, orders that do not come in, a contractual obligation you breach—. It might be four hours or it might be three days. What matters is that it is written down, that somebody with authority to spend money signs it, and that it comes with the one thing that makes it executable: a copy that does not depend on the provider that is down, and somebody who knows how to bring it up without improvising.

What we are not going to do with this

We are not going to say what failed, because the provider has not published it and any hypothesis of ours would be invented. The only technical detail on the status page is "the replacement of a connection component", from which nothing serious can be deduced. Nor are we going to use this to imply it could not happen to us: anyone who runs hardware for long enough eventually has a Sunday like that, and whoever says otherwise either has not been at it long or does not talk about it. Publishing updates every few hours for two days, in public, is more than many do.

What is fair to criticise, and it goes for everyone —us included— is selling the backup and the restore inside the same perimeter as the service without saying out loud what that implies. Nobody is deceiving anybody; it is the default architecture of almost all regional cloud, and it is convenient for both sides. But the customer deserves to know that in the worst case that copy is not an exit, and that sentence appears in no proposal. It appears in an incident update, on a Sunday, when it is no longer any use.

Four questions to ask yourself this week

  • Where do you restore your copy from? If the answer is "from my provider's panel", your copy and your server share a fate. Having a second copy outside that perimeter is not paranoia: it is the difference between having an exit and not having one.
  • How many hours of downtime can you take before activating plan B? Write it down in digits. "Soon" and "as fast as possible" do not count.
  • Have you ever booted that copy somewhere else? If you have not, you do not know how long it takes or whether it works. One timed rehearsal a year turns an invented figure into a measured one.
  • Can you work by hand for two days? Invoice, receive goods, serve customers. If the answer is no, the recovery plan is also a paper procedure, and that one takes an afternoon to write.

The third is the one most often skipped and the only one that produces a real number. We already wrote it in connection with the wiping of the Romanian land registry, where the word that mattered was "immutable"; here the word that matters is "reachable". They are different requirements and have to be asked for separately, because a copy can meet one and fail the other without anyone noticing until the day it matters.

When this does not apply to you

If what you host is a corporate website that does not sell, two days down is an annoyance and little more: do not build architecture for that. Nor do you need to duplicate anything if you can keep invoicing on paper for a week without breaking a sweat. The conversation changes when orders go through there, when the ERP is where your inventory lives, or when you have a signed commitment to respond within a deadline. At that point what you are buying is not more availability —that is enormously expensive and almost never the answer— but an exit that does not depend on whoever went down: a copy outside that perimeter, at another provider or on your own hardware, and a recovery procedure somebody has actually run once with a stopwatch in front of them.

And an observation that keeps recurring, and is why we labour the point: the same thing happens to the big ones. In July it was a giant's turn, and we wrote then about what to learn from a CloudFront outage. What decides how it ends for you is whether your plan had a door to the outside. When the hardware is yours or sits in a data centre you choose, that door is easier to draw, though it does not appear on its own either: it still has to be designed.

Sources (verified 1 September 2026): all verbatim quotes from the provider come from DonWeb's public status page, and each one sits in the entry indicated (times GMT-3): "Our technical team has managed to identify the cause of the problems and is already working to restore service as quickly as possible" is already in the opening entry of 30 August at 09:23, marked "Identified", and is repeated at 09:35; "the replacement of a connection component was carried out" first appears at 11:14 and is repeated at 16:49; "there is no risk of data loss. The stored information and data are safe" is from 20:04 on 30 August; "while the incident remains active, it is temporarily not possible to access backups or to migrate the affected services to another node" is from 31 August at 14:09; and "100% of the servers on the NOVA node" is in the 31 August 08:48 entry. The other updates consulted are 31 August 06:27 and 1 September 08:21 and 08:43; at that last time, the most recent published, the incident was still active. The scope ("100% of the servers on the NOVA node"), the absence of an estimated resolution time and the impact on hundreds of companies — La Capital, 1 Sep 2026 and Punto Biz. Our own arithmetic: the 47 hours and 20 minutes are the difference between the first and last status entries cited, not a published figure; the incident began before the first entry and we could not pin that moment down in a primary source, so the real number is larger. Our own judgement, not reported fact: reading the two sentences together as zero RPO with an RTO that has no number, the warning about confusing "cause identified" with "nearly there", the proposal to write down the hour at which you stop waiting, and the four questions. The technical cause of the incident has not been published by the provider and this post does not infer it.

Where would you restore from if your provider went quiet for 48 hours?

At everyWAN we design backups and recovery plans that do not depend on a single perimeter, and we rehearse them with a stopwatch instead of assuming they work. If you cannot answer that question, that is exactly the conversation we have.

Talk to everyWAN

Tags:

Share:

Subscribe to our newsletter

To receive IT stories, everyWAN news and exclusive subscriber offers, sign up to our mailing list

Minorisa de Sistemas Informaticos y Gestión S.L. © 2026
everyWAN
everyWAN