A drive dies on a Tuesday afternoon. The cluster barely notices: it keeps serving, nobody in the office finds out, an alert lands on the dashboard. That is precisely what you paid for when you paid for redundancy. But what you bought is not immunity: it is time. A margin in which to fit a new drive before the next one dies. And that margin was worked out assuming the spare arrives tomorrow.
In 2026 it no longer arrives tomorrow. And nothing in the design has changed: same drives, same replication, same array. The only thing that moved is the delivery date of a purchase order. It turns out that date was a design parameter, and we spent twenty years treating it as a footnote.
Redundancy does not buy immunity: it buys a window
The usual axis: failure is inevitable, the outage is a design decision. With your data on a single disk the two are the same thing; with your data replicated across three machines, that same dead disk never reaches the floor below.
The small print, which we tell less often, is that the design decision has an expiry date. While you are down one replica, the system has spent one of its lives. If the next failure lands inside that window and hits the wrong place, you are back to an outage. Every redundancy calculation —RAID, replica 3, erasure coding, the extra node in the cluster— is really a bet on how long you take to replace the part. The technical name is mean time to repair, and any sensible design of the last decade has assumed hours, or at most a couple of days.
This year's drives are already sold
You do not have to take an analyst's word for it: the manufacturers said it themselves on their earnings calls. Western Digital stated on 29 January, on its fiscal second-quarter call, that it was "pretty much sold out for calendar year 2026", with firm purchase orders from its seven largest customers covering the whole year; on that same call it added that it had already closed long-term agreements with two of those seven customers for 2027 and with one for 2028. Seagate said on its own call that its nearline capacity —the high-capacity drives that end up in arrays and clusters— was "fully allocated through calendar year 2026", and that it would start taking orders for the first half of 2027 in the following months.
The buyers of that output are not the industrial SME next door: they are the large cloud operators filling AI data centres. In German retail, hard drive prices rose between 20% and 50% compared with mid-2025, and SSDs of up to 2TB by around 50% since the summer of 2025. On the contract side, TrendForce forecast quarter-on-quarter increases of 13–18% for conventional DRAM and 10–15% for NAND in the third quarter of 2026. It is the same phenomenon we wrote about with server memory prices, except this time it has reached storage.
The honest caveat: those percentages are market observations, not your quote. Your price and your lead time depend on your channel, your volume and whether your specific model is on somebody's allocation list. What matters here is not the exact percentage but that lead time has stopped being a constant.
On the counter side the picture is reasonably clear. In an Omdia survey of 186 partners conducted in March 2026, more than 90% reported delayed fulfilment from their vendors, more than 70% reported price increases already applied across PCs, servers and storage, and more than 80% said their quotes now expire in under 30 days. That last number is the one that stings in an emergency purchase: the price you are quoted today is no good next month.
The same array, 66 times the risk
Let's put numbers on it — numbers are the only thing that convinces a committee. Take a group of 24 drives in which one has just died: 23 are still spinning. For the failure rate we use Backblaze's published figures, one of the few open, auditable field data sets in the industry: in its first-quarter 2026 report, across 341,263 analysed drives, the quarterly annualised failure rate was 1.24% and the lifetime figure stood at 1.39%. We use 1.39%, the more conservative of the two.
A warning before we go on, otherwise the arithmetic gets applied where it does not belong: this holds for a RAID group with no spare drive, or for a cluster with no free room left to place the missing copy. If your Ceph has free capacity, the dead disk recovers on its own within hours onto another OSD on the same host, and your distributor's lead time never touches the window — we wrote about it a few days ago, and the real free capacity is set by the fullest disk, not by the average. Lead time rules in two cases: when there is no room left to rebuild into, and when what dies is the whole node.
- →Spare in 48 hours: the probability that another of the 23 drives fails before the repair is 0.18%. One incident in 571. Paperwork, as we said.
- →Spare in 20 weeks: the same probability rises to 11.5%. One in 8.7. Same design, same drives, same money spent.
- →That is 66 times the risk, caused by a variable that appears in no architecture diagram. Even with the 20TB-plus drives, which Backblaze puts at 0.85% annualised across more than 86,000 still-young units, it only falls to 7.2%: one in fourteen.
It is worth seeing why the answer is 66 and not something else, because there is less magic in it than it looks: at probabilities this low the risk is almost linear in time, so multiplying the wait by 70 multiplies the risk by 66. The interesting part is not the factor, it is that a variable you treated as constant has moved by two orders of magnitude. (And if you divide the rounded percentages you will get 64: the 66 is calculated without rounding.)
The arithmetic is deliberately simple, and it is worth saying where it limps. It assumes failures are independent and the rate is constant, which is optimistic: the classic Schroeder and Gibson study of about 100,000 production drives found that failures do not fit an exponential distribution, that there is significant autocorrelation over windows of up to 30 weeks, and that real replacement rates ran from 2% to 4% a year — up to 13% in one system — against the 0.58–0.88% promised on the datasheets. Translated: drives from the same batch, of the same age, at the same temperature, tend to die near each other. The real number for any given rack is worse than ours, not better. And Backblaze's figures come from a vast fleet in a well-run data centre; your rack may have a different temperature, a different power quality and a different level of vibration. We do not publish a failure rate of our own, so we are not going to invent one.
Where it really hurts is not the disk: it is the node
A disk you can still keep on a shelf. A whole server, no. And that is where the shortage bites hardest, because a dead node is not replaced with a part: it is replaced by buying one, building it and commissioning it. In a three-node Ceph cluster with size=3 and min_size=2, losing a node leaves the cluster serving — yes — but with nowhere to recreate the third replica. We wrote about it a few days ago: with three nodes the cluster does not heal, it holds on. Holding on for 48 hours is a perfectly reasonable plan. Holding on for a quarter is a different conversation, and one you now have to have before drawing the design, not after.
Put in one sentence: N+1 is only N+1 while the "+1" can still be bought. If it cannot, what you have is N and a promise.
The cut nobody signs off: retention
There is a second effect, quieter and harder to spot than the first. Capacity that does not arrive leaves no visible gap in the rack: it leaves a small decision taken by whoever has to make tonight's backups fit. Retention goes from 90 days to 30. A weekly copy is dropped. A repository "hardly anyone uses" stops being backed up. Nobody takes it to a committee because it does not look like a policy change, it looks like maintenance.
And it is one. Retention is your recovery window: the number of days you can rewind when you discover something has been corrupted or encrypted for weeks. Shortening it for lack of disk is a risk decision taken for warehouse reasons. We are not saying it can never be taken — sometimes it is the only sensible way out — we are saying write it down, with a date, with who took it and with what it costs you. If the shortage is going to trim your disaster recovery plan, let it do so in plain sight.
What we ask now before signing off on a design
- The real lead time for YOUR part number, in writing and dated. Not the catalogue figure, not "it usually takes". The exact model you would need tomorrow, quoted by your distributor and kept in an email.
- Does your support contract include the part or only the labour? A "next business day" is only as good as the warehouse behind it. Ask explicitly what happens if the part is out of stock, and get the answer in writing.
- Count spares per failure domain, not in total. Four spare drives are worth little if they are the model that does not fit the array you care about most.
- Reduce the number of distinct models. One spare that fits eight machines covers more risk per euro than eight different spares. Homogeneity, boring elsewhere, is money here.
- Write down the degraded-mode policy. What gets switched off, what gets moved, what gets accepted while you wait. Deciding it on the day of the failure, with the phone ringing, is deciding it badly.
- Recalculate the whole window: procurement time plus rebuild time. That sum is what belongs in the continuity plan. The one you have written today almost certainly contains only the second half.
When this is not about you
If you are ten people with one server and your backups go somewhere you do not administer, do not buy a cupboard full of spares: tying up money in parts for a single machine makes no sense. Your question is a different and shorter one: how long does your provider take to hand you an equivalent machine, and is that written down anywhere or is it just custom? If the provider supplies the hardware — in colocation or under a managed service — the spare is their problem, but the lead time is still yours: ask for it.
Nor is this an argument for bolting for the public cloud. The buyers of the 2026 and 2027 drive output are precisely the large cloud operators: the cost of the shortage sits inside their pricing, you just see it as a per-gigabyte rate that refuses to fall rather than as a hardware invoice. It is a legitimate decision, but a cost and control decision, not a magic exit from this situation; we ran the five-year numbers here and the same side does not always win.
We say all this knowing which side we get paid on: we design and run infrastructure for other companies, so a post that ends in "review your design" suits us. Hence the counterweight: if you already have the lead time for your critical part in writing, and the spares that lead time demands, you have done the hard part and you need nobody. The one with the problem is whoever discovers the lead time on the day they call the distributor with a dead drive in their hand.
What we are not claiming
- ✗We are not saying your lead time is 20 weeks. We use twenty weeks as a middle, deliberately conservative case: enterprise drives have been reported on backorder for two years, and more than 90% of the partners Omdia surveyed report delays. Your exact lead time belongs to your distributor, not to us.
- ✗We are not saying the manufacturers are rationing. They state in their results that capacity is committed under signed contracts. That is a commercial fact, not a conspiracy.
- ✗The 11.5% is not a prediction about your array. It is the output of a simple model with stated assumptions and somebody else's failure rate. It serves to compare two scenarios against each other, not to promise anything.
Sources (verified on 29 September 2026): Western Digital's statement ("pretty much sold out for calendar year 2026", firm purchase orders with its seven largest customers) and Seagate's (nearline capacity "fully allocated through calendar year 2026" and orders opening for the first half of 2027), together with the price rises in German retail (20–50% on hard drives versus mid-2025; ~50% on SSDs of up to 2TB since summer 2025) — heise online, "WD and Seagate confirm: Hard drives for 2026 sold out" (16 Feb 2026), reporting both companies' earnings calls; Western Digital's call was held on 29 Jan 2026 (fiscal second quarter, schedule published on investor.wdc.com) and is where the long-term agreements with two of its seven largest customers for 2027 and one for 2028 were announced — Tom's Hardware; enterprise drives on backorder for up to two years — Tom's Hardware, "Hard drives on backorder for two years"; contract price forecasts for 3Q26 (conventional DRAM +13–18% QoQ, NAND +10–15%) — TrendForce (3 Jul 2026); an Omdia survey of 186 partners conducted in March 2026 (more than 90% reporting delayed fulfilment, more than 70% price increases already applied, more than 80% quotes expiring in under 30 days) — Channel Dive, "Hardware price hikes accelerate amid supply chain disruptions" (24 Mar 2026); field failure rates (341,263 drives analysed, 1.24% quarterly annualised, 1.39% lifetime, and 0.85% across the 20TB-plus drives, which Backblaze itself calls "still fairly young") — Backblaze Drive Stats Q1 2026, with the 86,000-plus unit count from the press release of 9 Jul 2026; the absence of an exponential fit, failure autocorrelation up to 30 weeks, and replacement rates of 2–4% a year (13% in one system) against the 0.58–0.88% on datasheets, across about 100,000 drives — Schroeder and Gibson, "Disk failures in the real world", USENIX FAST 2007. The probability percentages are our own calculation using the exponential survival formula over 23 drives at a 1.39% rate (and 0.85% in the variant); the assumptions are stated in the text.
How many spares are on your shelf right now?
We review your redundancy design with the real lead times in front of us, review the spares you hold per failure domain, and put the degraded-mode policy in writing before you need it.
Talk to everyWAN