The Ceph upgrade procedure on Proxmox recommends, as its fourth step, that you run ceph osd set noout: "optional, but recommended", says the wiki. Ceph's own OSD troubleshooting page says that this is "more a thought exercise" than a suggestion that anyone "in the post-Luminous world" should run it. Both sentences are live today and both are defensible. The one that decides whether your cluster recreates a lost copy tonight is the second.
We run Proxmox VE with Ceph in production, spread across several datacentres, so we have set that flag ourselves many times and we will keep setting it. This post is not about noout being wrong. It is about what it switches off, precisely, about the two warnings you will not read while it is set, and about the four commands that do the same job without disabling the automation of the entire cluster.
October is the month a lot of people will be touching their Ceph
The active releases table in Ceph's documentation says this, today: Squid, released 26 September 2024, latest 19.2.6, estimated end of life 31 October 2026. Tentacle, released 18 November 2025, latest 20.2.4, estimated end of life 1 June 2027. Losing support for a Ceph release does not mean the cluster stops: it means the backports stop arriving, security ones included. So anyone still on Squid has weeks of patches left, and that turns October into a month of maintenance windows.
That word, "estimated", is not decoration, and we have written about it: Squid's date moved from 19 September to 31 October in a commit with no announcement, and in the same move Tentacle lost 170 days of estimated useful life. What matters here is the practical consequence: a lot of people will be running the Squid-to-Tentacle upgrade script this month.
That script, on the Proxmox wiki, has a very recognisable shape: change the repository, set the flag (and it offers to do it from the UI, on the OSD tab, with a button called Manage Global Flags), apt update and apt full-upgrade, restart the monitors one at a time, the managers, the OSDs node by node, the MDS daemons if you run CephFS, set ceph osd require-osd-release tentacle and, right at the end, unset the flag. The last step on the list is the one that goes undone. Not through sloppiness: because the window does not end when the procedure ends, it ends when somebody decides it has ended, and that somebody has been staring at ceph -s for four hours at two in the morning.
What noout switches off is not rebalancing
Most people set it with one idea in mind: "I am stopping the cluster from shifting terabytes while I reboot a node". The documentation is more precise than that idea. In the health-check list, the OSDMAP_FLAGS entry describes noout as: down OSDs are not automatically being marked out after the configured interval.
That configured interval has a name and a value: mon_osd_down_out_interval, 10 minutes by default, described as "mark any OSD out that has been down for this long". And the difference between the two states is the whole conversation: down means "not answering"; out means "removed from the data distribution". It is the move to out that makes CRUSH recompute where each copy belongs, and that kicks off the creation of the missing copy on another disk. If it never reaches out, the cluster recreates nothing on its own. Marking it out by hand does work, even with the flag set; what it takes is somebody noticing.
And here is the detail that changes the conversation: the flag does not discriminate. It does not know that osd.7 is down because you just rebooted it and osd.12 is down because its electronics died. To the flag they are the same state and get the same treatment. With size 3, that means two copies where your design says three, and a PG_DEGRADED warning that defines precisely that: "data redundancy is reduced for some data". We wrote a week ago about what heals itself and what does not in a three-node Ceph; this is one step above that, because here healing is prevented by a decision of yours, taken six hours ago and written down nowhere.
Ceph's manual says, in those words, not to use it
On the OSD troubleshooting page, right after explaining how to set the cluster-wide flag, there is a warning box. It reads: "this is more a thought exercise offered for the purpose of giving the reader a sense of failure domains and CRUSH behavior than a suggestion that anyone in the post-Luminous world run ceph osd set noout". And the box goes on, worth quoting in full: "when the OSDs return to an up state, rebalancing will resume and the change introduced by the ceph osd set noout command will be reverted". Which is to say: the project discourages it because it considers it of little use. Our reason for discouraging it is a different one, and it comes next.
Luminous is from 2017. That sentence is not a forum comment: it sits in the project's official documentation, on the OSD troubleshooting page, and the next paragraph offers the alternative bluntly: in Luminous and later, it is safer to flag only the affected OSDs.
It is worth saying the other half too, or this turns into cheap point-scoring: the Proxmox procedure is not reckless. It asks for the cluster-wide flag because it is a generic procedure that has to work the same on a three-node cluster and on a thirty-node one, and because the alternative requires knowing each host's CRUSH bucket name. For a manual, that is a reasonable call. What is not reasonable is that a manual written for everybody ends up being the only thing deciding your specific cluster's redundancy posture for six hours.
The brake you already had on and did not know about
There is a parameter almost nobody looks at that already does, by default, much of what people think they are buying with noout: mon_osd_down_out_subtree_limit. Its description is "the smallest CRUSH unit type that Ceph will not automatically mark out", with an explicit example: if it is set to host and all the OSDs of a host go down, Ceph will not mark them out on its own. Its default value is rack.
Read it slowly, because it says two things at once. First: the scenario that genuinely frightens you is already braked. If a whole rack disappears — or any unit of that size or larger — Ceph will not start relocating tens of terabytes on its own. With one important caveat: if your CRUSH map is the flat default — root and hosts, with no buckets of type rack — the only level that reaches that size is the whole root, so the real protection is rather smaller than it sounds. Second, and this is the one that matters for tonight's window: a host is not braked. If you reboot a node and it takes more than ten minutes to come back, its OSDs will be taken out of the distribution. That case, and no other, is what justifies touching a flag. And before typing it is worth asking a different question: which failure domain am I about to touch? If the answer is one host, the flag belongs to that host.
And what almost nobody spells out: a degraded PG is not scrubbed
This part appears in no upgrade procedure and it is what made us write the post. The PG_NOT_SCRUBBED health check contains a sentence that, read out of context, looks like an implementation detail: "PGs are scrubbed only if they are flagged as clean… misplaced or degraded PGs will not be flagged as clean". The PG_NOT_DEEP_SCRUBBED entry says the same thing more cautiously, with a "might not".
Deep scrubbing is the mechanism that compares the copies' checksums and finds silent corruption: the OSD_SCRUB_ERRORS warning, raised by scrubs in general, exists precisely because "recent OSD scrubs have discovered inconsistencies". Now chain the pieces in the order they happen:
- You set
nooutfor the window. - A disk genuinely dies (not the one you rebooted).
- Its OSD stays
downand never reachesout: the copy is not recreated. - The PGs that lived there become
degraded, and therefore stop beingclean. - Not being
clean, they stop being scrubbed.
In other words: the only mechanism verifying that your two remaining copies are still correct stops running on exactly the data that just lost a copy. That is not two problems added together. It is one problem switching off the other one's detector.
A permanently amber cluster is a cluster with no alarm
Someone will rightly point out that Ceph does warn you that scrubbing has stopped. True, and you can work out when. The maximum shallow scrub interval, osd_scrub_max_interval, is 7 days; the warning fires once an extra fraction of that interval has passed, set by mon_warn_pg_not_scrubbed_ratio, which defaults to 0.5: another three and a half days, so 10.5 days. For deep scrubbing, osd_deep_scrub_interval is another 7 days and the ratio mon_warn_pg_not_deep_scrubbed_ratio is 0.75: 12.25 days.
Those two warnings exist and they work. Those warnings do arrive. The question is where they arrive. From the second you ran ceph osd set noout, the cluster has been in HEALTH_WARN because of OSDMAP_FLAGS. If the flag is still set twelve days later, the deep-scrub warning lands on a dashboard that has been amber for twelve days, under a line everybody has learned to ignore because "that is the upgrade one". A new amber inside an old amber is not an alert: it is yesterday's line. The flag does not only switch off a function; it switches off the channel through which Ceph would have told you.
The same OSDMAP_FLAGS entry describes the rest of the family, and it is worth putting them side by side with what people think they do:
| Flag | What people think it switches off | What the documentation says |
|---|---|---|
noout | Data movement during the reboot | A down OSD being marked out after the configured interval |
nobackfill, norecover, norebalance | The same as noout, "just in case" | "Recovery or data rebalancing is suspended" |
noscrub, nodeep_scrub | The disk noise of the scrubs | "Scrubbing is disabled" |
nodown | False positives from a hiccuping switch | "OSD failure reports are being ignored, which means that the monitors will not mark OSDs down" |
The first three rows get set together with a frequency that commands respect, usually by copying a line out of a forum thread. But the worst one to leave behind is the last. With nodown, a genuinely dead OSD still shows as up in the map, because the monitors have stopped listening to whoever reports otherwise. And that is no longer an amber you ignore: it is a green that is not true.
The four commands that do the same job without switching off the cluster
The alternative the documentation itself proposes fits in four lines. The flag stops belonging to the cluster and starts belonging to the disk or the host you are about to touch:
# one single disk you are pulling or replacingceph osd add-noout osd.12 ceph osd rm-noout osd.12 # the whole host you are about to reboot (CRUSH bucket name)ceph osd set-group noout nodo-03 ceph osd unset-group noout nodo-03
This does not leave you without a warning: there is a health check of its own, OSD_FLAGS, which fires when "one or more OSDs, CRUSH nodes, or CRUSH device classes have a flag of interest set". You still get your amber. The difference is that it is now an amber naming three specific OSDs instead of one covering the whole cluster, and that while it is set, the rest of the cluster keeps its automation: if a disk dies on another node during your window, that one does go out after ten minutes and its copy is recreated on its own, which is exactly what you wanted to happen.
The thirty-second check, and the one missing from your monitoring
Before you read any further, check whether you have one set right now. Three commands, none of which changes anything:
ceph -s # reports the flags that are set, under healthceph osd dump | grep ^flags # the cluster-wide flags, rawceph health detail # OSDMAP_FLAGS and OSD_FLAGS, with the detail
If noout, nodown, noin or noup comes back — the rest of what osd dump lists is there by default — the next question is not technical: who set it, what for, and when was it going to come off? If nobody can answer all three, the flag has been there longer than anyone remembers and your real redundancy is not the one on the design document.
And the part that is genuinely our job: we watch clusters with Zabbix, and the rule that matters here is not "tell me if the cluster is not in HEALTH_OK". That alert gets muted in the first month, for good reasons, and from then on it is worth nothing. The useful rule is a different one: tell me if a flag is set and more than N hours have passed since it was set. A maintenance flag is a manual change with an implicit expiry date; the alert has to be about the expiry, not about the state.
This is not a private obsession of ours. The classic Oppenheimer, Ganapathi and Patterson work on why internet services fail (USENIX, 2003) put operator error at the top of the failure causes in two of the three services it studied, ahead of hardware, with misconfiguration the dominant subtype within it. The paper is also precise about when it happens: while operators were "making changes to the system, e.g., deploying or upgrading software". ceph osd set noout is, literally, a manual configuration change made by an operator. It is in the winning category.
What we are not saying
We have not lost data to this, and we are not going to invent a client story to round the post off. What is here is the documentation read end to end and the arithmetic done. The cluster-wide flag is not a mistake in itself: it is a generic procedure applied to a specific cluster, and all the risk is in how long it lasts, not in the act of setting it. Half an hour of noout with somebody watching is one thing; three weeks because nobody closed the window is something else entirely.
The values we have quoted are the defaults. If somebody changed mon_osd_down_out_interval, the subtree_limit or the two warning ratios, your numbers are not these; check them with ceph config get mon mon_osd_down_out_interval and ceph config get mgr mon_warn_pg_not_deep_scrubbed_ratio before trusting the arithmetic. And the end-of-life dates in the releases table are marked as estimated by the project itself, which has already moved them more than once.
Conflict of interest, up front: we are not Proxmox resellers and we do not sell Ceph subscriptions, we take no commission on either. Operating other people's clusters and running these windows at three in the morning is work we do bill for.
Sources (verified on 2 October 2026): the active releases table with Squid 19.2.6 (initial 26-09-2024, estimated end of life 31-10-2026) and Tentacle 20.2.4 (initial 18-11-2025, estimated end of life 01-06-2027) — Ceph releases index; the literal descriptions of noout, nodown, noscrub/nodeep_scrub and nobackfill/norecover/norebalance, the OSD_FLAGS check for OSDs and buckets, the definitions of PG_DEGRADED and OSD_SCRUB_ERRORS, and the sentences stating that degraded PGs are not flagged as clean in PG_NOT_SCRUBBED and PG_NOT_DEEP_SCRUBBED — Ceph health checks; the "thought exercise" warning and the add-noout, rm-noout, set-group and unset-group commands — OSD troubleshooting; mon_osd_down_out_interval (10 minutes) and mon_osd_down_out_subtree_limit (rack) — monitor/OSD interaction; osd_scrub_max_interval and osd_deep_scrub_interval (7 days each) — OSD configuration reference; the 0.5 and 0.75 defaults of the two warning ratios — src/common/options/global.yaml.in in the Ceph source; the procedure with ceph osd set noout, the Manage Global Flags button and the final unset — Proxmox wiki, Ceph Squid to Tentacle upgrade; manual configuration changes as the leading cause of serious outages — Oppenheimer, Ganapathi and Patterson, "Why Do Internet Services Fail, and What Can Be Done About It?", USENIX 2003. The reading that the flag switches off the other problem's detector, the 10.5 and 12.25 day arithmetic and the alert-on-flag-age rule are ours, not those sources'.
How many flags does your cluster have set right now?
If the answer takes more than thirty seconds, that is already an answer. We design and operate distributed storage with Ceph and we run VMware to Proxmox migrations, boring work included: every window having an owner and a closing time, and monitoring that warns about the flag that has been set for three weeks instead of painting an amber nobody looks at any more.
Talk to everyWAN