Back to Blog

Retry storms: the second outage is the one you cause

Dim corridor between server cabinets, with hundreds of status lights on at the same time

On 12 June 2025 Google Cloud had a global outage, with dozens of its services down at once. 40 minutes in, the mitigation had been rolled out and regions started to recover, the smaller ones first. One did not: us-central1 was not fully resolved until 2 hours and 40 minutes after the incident began. Same bug, same people fixing it. What stretched that one region was not the failure: it was the retries.

This post is not about Google. It is about the fact that the very same mechanism is wired, right now, into the infrastructure of any company with a dozen servers and three integrations: in a cron job, in the backup agent, in every desktop mail client and in the monitoring system itself. And that nobody has written it down anywhere, because retries are not a decision: they ship switched on by default.

What exactly happened in that one region

The public incident report describes the start plainly: a quota policy change went in with blank fields, and those blank fields "exercised the code path that hit the null pointer causing the binaries to go into a crash loop". So far, a bug —and one that had been deployed since 29 May with neither proper error handling nor feature flag protection, waiting for data of the right shape to wake it up—. What came next is the part that matters: as those tasks restarted en masse, they "created a herd effect on the underlying infrastructure it depends on (i.e. that Spanner table), overloading the infrastructure". And one sentence that deserves to be framed in more machine rooms: "Service Control did not have the appropriate randomized exponential backoff implemented to avoid this".

Getting out of the hole required the opposite of what instinct demands: slowing down. They throttled task creation and routed traffic to multi-regional databases to take load off the one that was drowning. That is why the big region was the last one to close: not because the fix was harder there, but because it had to be recovered slowly so that the recovery would not knock it down again.

A retry storm is exactly this: a system that can no longer cope with the load receives more load precisely because it cannot cope. It is the traffic jam where everyone leans on the horn. The noise does not clear the road; it jams it a little harder.

Where this lives if you do not own an entire region

The normal reaction to a report like this is "sure, but that is Google scale". Scale is not what you need. You need two things: a shared resource, and several things insisting at once. The places where we find it:

  • ▸The cron job that takes longer than its own interval. It was scheduled every five minutes back when it took two. Today it takes seven because the table has grown. From that day on there is not one process: there are two, then three, each fighting over the same database lock. Nobody changed anything; only the data grew.
  • ▸The backup job that fails and comes back in. A backup window that does not close and retries overlaps with the next one. Two passes reading the same disks make the third take even longer. The symptom on the dashboard is not "the array is slow": it is "backups take a bit longer every day".
  • ▸The power comes back and everything starts at once. This is the small-company herd effect. Sixty desktops, twenty virtual machines and a couple of services all asking for DNS, domain controller and licensing in the same second. The cut lasted four minutes; getting back to work, forty.
  • ▸Monitoring, which pushes hardest exactly when it hurts. In the Nagios/Icinga family there are two intervals: the normal one and the retry one, which kicks in when a check has just gone into a bad state and is almost always configured shorter. It is only a handful of probes —the ones between SOFT and HARD state, until max_check_attempts runs out, after which it goes back to the normal interval—, so it is no storm. But it is your monitoring system pushing in the same direction as everything else, and worth knowing before you lower that interval "so we find out sooner".
  • ▸And the human retry. When something does not load, people hit F5. Forty people hitting F5 is an unauthorized load test against a server that was already asking for help.

There is nothing modern in any of these five cases: no microservices, nothing that sounds like a conference talk. There is a table, a storage array or a domain controller, and several things insisting on top of it at once. Which is exactly the us-central1 recipe, only three zeros smaller.

The multiplication nobody designed

There is a calculation in Google's SRE book worth keeping at hand when somebody suggests "just have it retry, to be safe". If the database cannot service requests, and the backend, the frontend and the browser JavaScript each issue 3 retries —4 attempts—, a single user action may create 64 attempts against the database (4³). Nobody designed 64. Each layer added a defensive 3, unaware of what the others were doing, and the threes multiplied instead of adding up.

What the chapter actually says is softer than we would like: think about the service holistically, decide whether you really need retries at a given level and, above all, "avoid amplifying retries by issuing retries at multiple levels". Our own rule is harder, and we own it as ours: retry in one place, the highest one, and let the failure travel cleanly up from the layers below. An error that reaches the user quickly is information. An error that bounces ten times on the way is load.

The four numbers you need written down

This is not a project. It is four values somebody has to decide and write down, and which today, in most places, are wherever the installer left them:

  • 1The timeout. Without an explicit timeout no retry policy matters: you get hung connections piling up until the pool is exhausted. A timeout is the promise to release the resource even if the answer never arrives.
  • 2The attempt cap. Google uses a budget of up to three attempts per request; if it has failed three times, they let the failure bubble up to the caller. The reasoning is simple: if a request has landed on overloaded tasks three times, a fourth is unlikely to help.
  • 3The growing, randomized wait. "Always use randomized exponential backoff when scheduling retries", says the book. The key word in that sentence is not "exponential": it is "randomized". If a thousand clients wait exactly one second, in one second you will have a thousand clients again. Randomness is what spreads the herd out; it is exactly the piece missing in the region that took 2 h 40.
  • 4The global budget. A per-request cap is not enough, because many capped requests still add up. That is why, at Google, each client tracks what proportion of its traffic is retries and only retries while that ratio stays below 10%. The published figure is blunt: with per-request caps alone, worst-case growth sits just below 3x; adding the 10% budget brings it down to 1.1x in the general case. The version for those without that machinery: "only allow 60 retries per minute in a process, and if the retry budget is exceeded, don't retry; just fail the request".

Four numbers. None of them costs money. What costs is the conversation to decide them, because it forces you to admit that some requests are better dropped.

When NOT to touch your retries

It would be nonsense if the message were "remove your retries". A well-placed retry is what keeps a network blip from ever reaching the user. But there are two kinds of operation that must not be retried automatically, and they are worth writing down before the four numbers: anything that cannot be repeated safely —a payment, an order dispatch, an email: if you cannot guarantee that repeating it is harmless, retries are not giving you resilience, they are giving you duplicates somebody will clean up by hand— and errors that will not change: a 401 or a 404 does not improve with insistence. Retries are for transient conditions, not for a firm "no".

And there is a third thing we would not do: build a platform to fix this. Here we part ways with the usual pitch. You do not need a service mesh, a resilience product or a redesign. You need four values somebody has actually decided, and an inventory of who retries whom. If someone sells you the former before you have written the latter, they are selling you the expensive part.

The worst moment is not the outage: it is the restart

What makes the 12 June case special is that the storm did not happen during the failure, but on the way back. Everything started at once and trampled itself. And that is precisely the part nobody rehearses, because drills usually measure "how long does a restore take" and almost never "what happens when 200 things come back at once and all of them want to authenticate, resolve names and read from the same place".

We time full recoveries on a regular basis —our last internal drill came in at 14 minutes, and we report that as evidence, not as a contractual promise— and what you learn there is not the time: it is the order. What has to be up before what, what should be staggered on start-up, and which integration gets nervous if its dependency takes thirty seconds too long. That order gets written once and holds for every later scare. On why an unrehearsed plan is just literature, we already made that case in a continuity plan nobody has rehearsed.

There is an obvious family resemblance with high availability: a cluster does not prevent downtime either, it shortens it. Retries are the same thing in reverse: they do not prevent the failure and, badly configured, they stretch it.

Five questions for your next meeting

With one rule: "yes" does not count as an answer. An answer is a number, or a place where it is written down.

  • 1.How many layers sit between the user and the database, and how many of them retry? If nobody knows, the real number is the product, not the sum.
  • 2.Which scheduled jobs now take longer than their own interval? That is a query, not an opinion: average duration is in the history.
  • 3.Do the waits between retries carry randomness, or do all clients come back in the same second?
  • 4.Which operations must never be retried automatically? If the list is empty, nobody has looked.
  • 5.When the power comes back, what starts first and what waits? If the answer is "everything at once", your next incident is already written.

The uncomfortable part

A now-classic study on why internet services fail (Oppenheimer, Ganapathi and Patterson, USENIX 2003) put operator error —and within it, in more than half of the cases, configuration errors— at the head of service failures in two of the three services it studied. With a caveat the authors themselves stress, and which we already covered when writing about why redundancy does not survive the procedure: it is not that operators fail more often than hardware, it is that their errors are masked less often. Redundancy hides the dead disk; it does not hide the person touching production. Twenty-odd years later, the June 2025 picture rhymes with it: a configuration change with blank fields, an unguarded code path and, above all, a decision nobody made —how retries work— stretching the worst case out to two hours and forty minutes.

Failure is inevitable. The outage —your company grinding to a halt, and for how long— is still a design decision. Retries are one of those decisions that make themselves if nobody makes them.

At everyWAN we look at this where it shows: in the platform that carries business applications (data and applications), in the infrastructure underneath (infrastructure and cloud) and in the monitoring that has to tell a slow service apart from a service drowning in its own insistence. We use Zabbix and SmokePing, with one rule you can see in everything we run: alerts that matter, not noise.

Sources (verified on 28 Sep 2026): public incident report for the Google Cloud outage of 12 June 2025 —crash loop, herd effect on Spanner, missing randomized exponential backoff and per-region recovery times— status.cloud.google.com; three-attempt budget, 10% retry ratio and growth from 3x to 1.1x — Google SRE Book, "Handling Overload"; the 4³ = 64 multiplication, randomized exponential backoff, the 60-retries-per-minute budget and retrying at a single level — Google SRE Book, "Addressing Cascading Failures"; causes of failure in large internet services — Oppenheimer, Ganapathi and Patterson, Why Do Internet Services Fail, and What Can Be Done About It?, USENIX 2003. The 14-minute figure is an internal everyWAN drill: a measurement, not a service commitment.

Do you know how many things are retrying against your database right now?

We look at it with you: who retries whom, which jobs overlap, and in what order everything has to come back after an outage. Without selling you a platform along the way.

Talk to everyWAN

Tags:

Share:

Subscribe to our newsletter

To receive IT stories, everyWAN news and exclusive subscriber offers, sign up to our mailing list

Minorisa de Sistemas Informaticos y Gestión S.L. © 2026
everyWAN
everyWAN