The spreadsheet has 260 rows. In 160 of them, the column that says whether the secure value is already set answers no. That, counted, is the VMware Cloud Foundation 9.1 hardening guide.
On 5 October VMware published a post explaining its Security Configuration Guide for VCF, part of a twenty-two piece series for cybersecurity awareness month. The guide itself is not new —they have been publishing these since the vSphere 4.0 era, and the repository still holds them all— but the post is useful because it explains, column by column, what each field means. We did the obvious thing: download the CSV and count.
What is inside the spreadsheet
The guide ships as Excel and as CSV. The VCF 9.1 directory carries a VERSION file with the string 910-20260612-01, and the CSV, twenty columns per control: the identifier, the compliance framework mappings, the affected component, a prose discussion, the expected functional impact, the priority, the parameter, the factory value, the recommended value and the PowerCLI commands to check it and to change it. There are 260 controls. Everything that follows comes from there, and anyone can redo the count with the same file.
The split by component is not even, and it is worth knowing before you hand out the work: 120 controls belong to ESX, 51 to vCenter, 29 to NSX, 26 to Operations and 17 to vSAN. The remaining 17 are spread across Automation, VCF, Operations for Networks, Protection and Recovery, SDDC Manager and the installer. Almost half the controls sit in the hypervisor.
160 of 260: the short list is not that short
The blog post reassures the reader with this line: "Many controls are secure by default, so the list of real work is usually shorter than it first appears". It is sensible general advice. For this particular guide, the count runs the other way: the Is the Default? column answers NO on 160 of the 260 controls, and the Action Needed column says Modify on exactly those same 160. 61.5% of the sheet is work, not audit.
There is also a clean coincidence that saves time: every P0 and P1 control is one you have to change, and every P2 already ships correct. Not one exception across the 260 rows. So the priority column and the pending-work column say the same thing, and you can sort by either. One honest caveat: 22 of those 160 carry the tag Upon Feature Enablement, meaning they only apply if you use that feature. That leaves 138 that apply regardless.
The P0 shuts the door. The P2 leaves you the key
Priority P0 means, in the guide's own words, "where there is not a secure default, and you should do this immediately". There is one P0 control almost everyone recognises: esx-9.lockdown-mode, host Lockdown Mode. It ships as lockdownDisabled and the guide recommends lockdownNormal. With it on, nobody gets into the host through the back door: it is administered from vCenter, and vCenter's roles and audit trail can no longer be bypassed by logging straight into the machine.
The functional impact column of that same control says what happens if you get it wrong:
«Enabling Lockdown Mode blocks direct access to the host for everyone except accounts on the Exception Users list and the DCUI.Access list. Verify those lists are configured before turning on Lockdown Mode; an incomplete list combined with a vCenter outage can leave the host unreachable.»
And here is what caught our eye, offered as our reading rather than as a vendor finding. The control that keeps that rescue path open —esx-9.lockdown-dcui-access, the list of accounts that can still log in at the host's physical console if it ends up isolated from vCenter— is classified as P2. And P2 means the default is already secure and auditing it now and then is enough. Its own impact text warns that misconfiguring that list with Lockdown Mode on "can leave the host unrecoverable if vCenter is unreachable".
Put plainly: if you work the sheet in priority order, which is what the guide asks of you, you shut the door before checking you still hold a key. Priority order measures exposure risk, not operational risk. The guide's three lockdown controls —the mode, the console list and the exception users list— are read together or not at all.
A column that exists because they expect you to undo it
Among the twenty columns there is one called Installation Default Value, and the blog post explains what it is for: "What the value was before you changed it, because once in a while you may want to undo what you changed". A hardening guide that ships the previous value as standard is quietly telling you what it expects to happen.
The confirmation sits in the first FAQ. To "is the guide supported by Broadcom?" the answer starts with yes, and continues: "However, if you make a change and something is amiss, our support engineers may ask you to undo it". It is a perfectly reasonable sentence from the vendor's side, and it is the one that turns your change log into part of the control. If you hardened without writing down what you touched, the night something breaks you will be undoing blind, with the clock running.
The virtual switch: same setting, opposite default
The blog post picks the virtual switch security settings as its example of a control with a price: "Setting them to what the Guide recommends may prevent clustered applications from working correctly". We went to look at those rows, and there is something the example does not mention.
They are three classic settings: reject forged transmits, reject guest MAC address changes and reject promiscuous mode. On the standard switch, the first two ship as Accept and the guide marks them P0 as soon as you use that feature. On the distributed switch, those same two already ship as Reject and appear as P2: audit only, and at two levels, because port groups inherit the switch policy but a port group can override it. Promiscuous mode ships closed on both. The conclusion is ours: two of the three settings give you work or not depending on which kind of switch you have, a networking decision that in most installations we have seen was taken for other reasons.
And in the impact text of the MAC changes control there appears, among the workloads that depend on being able to change it, one that is not a customer application: "applications licensed by MAC address, and vCenter Reduced Downtime Upgrade". Put another way, that piece of hardening can collide with the mechanism that lets you upgrade vCenter without stopping it. The way out the guide proposes is a good one and worth noting: create a separate port group that allows it and connect only the authorised VMs there. The guide also documents an exception it creates itself: enabling vSAN File Services turns Accept on for that service's port group, because its nodes present more than one MAC, and closing it there breaks File Service traffic.
The word that never makes the headlines: recoverability
The sentence that best describes what this guide is actually for is tucked inside the explanation of the discussion column: "Many of the security controls in the Guide have secure defaults. However, many do not, and that's because they require you to make a decision, or have implications for functionality or recoverability of the system". Recoverability. We counted how many controls mention recovery in their discussion or their impact: 32 of the 260.
The Lockdown Mode control itself warns that "some operations, such as backup integrations and low-level troubleshooting, require direct host access", and recommends switching it off temporarily on the affected host for those tasks. That is the knot, and it is the axis of nearly everything we write: a failure is inevitable, an outage is a design decision. If you harden and in doing so add half an hour to the road back, you have traded a risk that may never materialise for a cost that is certainly charged on the bad day. Do it, but put a number on it first. What it costs not to measure it is something we covered in RTO and RPO without the fluff.
What the guide says it does not do
This is the part we liked most, because you rarely see it and because it saves you an awkward conversation with an auditor. The guide carries mappings to compliance frameworks —we counted NIST SP 800-53 R5 on all 260 controls, PCI DSS 4.0.1 on 255 and a DISA STIG identifier on 168, so 92 controls have none— and still answers this to the question of whether implementing it makes you compliant with PCI DSS or NIST: "Not by itself". The mappings save time; compliance, it says, depends on your scope, your processes, your people and your auditor's judgement.
It goes on to say it does not cover what runs inside the virtual machines, only the infrastructure and the machines' own configuration; and that ideas like least privilege or separation of duties are not enumerated because they are design, not parameters. There is one more answer that deserves its own paragraph: in VCF 9.1 the management functions move to a new platform, VCF Management Services, described as a Kubernetes-powered layer inside VCF itself. And since "there are no user-configurable settings inside VCFMS", there are no controls for it in the guide.
Fewer knobs means fewer ways to get it wrong, and as an engineering decision we find that defensible. But it changes what the sentence "we applied the guide" means: that plane is not hardened by you and not audited by you either. What still belongs to you is who reaches it, and that is precisely what the guide says it does not enumerate.
From 9.0 to 9.1: 34 more controls and a new column
Both guides are in the repository, and both carry the same date in their version number. Counted the same way: the VCF 9.0 one has 226 controls and 19 columns, with 135 not set; the 9.1 one has 260 and 20 columns, with 160. The column that appears in 9.1 and was not in 9.0 is the NIST SP 800-53 R5 mapping: the correspondence table with the standard is newer than the guide.
Comparing identifiers, 58 appear in 9.1 and not in 9.0, and 24 were in 9.0 and are gone. Part of that movement will be renames rather than genuinely new controls —we cannot tell from outside and we will not claim otherwise— but it is exactly what the guide warns about when asked whether an older version can be used on a newer one: "Between major versions we rename or remove parameters, change defaults, and add and retire components". A hardening done against the 9.0 sheet is not a hardening done against the 9.1 one.
The line that describes most incidents we see
Asked whether you need to recheck settings after patching or upgrading, the guide says yes and gives three reasons. The third is this: "people make changes during troubleshooting that they never put back". It is right, and it matches what the literature has said for over twenty years: in the study by Oppenheimer, Ganapathi and Patterson presented at USENIX in 2003 on why internet services fail, operator configuration changes top the list of causes, and hardware accounts for a much smaller share.
That is why periodic re-auditing earns its slot in the calendar: it is what catches the temporary change nobody put back. And it is why we push monitoring so hard —we run Zabbix and SmokePing— with one simple rule: it should alert on symptoms before a customer calls.
The work order, and the step the guide does not have
The guide proposes six steps and they are good ones: filter to what you actually run, audit before you change anything, do the P0 controls first, read the discussion and the impact before touching the parameter, write down what you decide to skip and why, and check again later. The fifth is the one most people skip and the one the guide justifies best: "Your auditors will ask about the ones you skipped, as will anyone who inherits the environment after you".
There is an assumption hidden in the fourth step, the one about trying the change in a test environment first: it presumes you have one and that it resembles production. If you do not have one, that step does not exist and what you call "testing" is deploying. Building a test environment that looks like production, a phased rollout and a one-click rollback is less glamorous than hardening, and prevents more outages.
And we add a seventh step the list does not carry, flagged as our own criterion: after applying the P0 controls, re-time the restore and the upgrade path. Controls that touch recoverability do not fail the day you apply them; they fail the day you need them, which is exactly when nobody has time to discover that direct host access was what the backup agent needed. Our last full recovery drill took 14 minutes: that is an internal figure from a test of ours, offered as evidence and not as a contractual promise.
And if you are thinking of leaving VMware
Conflict of interest up front, as always: we are not VMware resellers, nor resellers of any other platform, and we sell nobody's licences, so this pays us no commission in either direction. We have migrated companies from VMware to Proxmox and we have also recommended staying put when it made sense. This guide does not move that decision one millimetre: the spreadsheet exists just the same on the other side, with different names and the same small print. In fact, the day we published our Proxmox VE hardening checklist we deliberately included a section on what we do NOT do, for the same reason VMware includes an impact column.
If you take away one idea, make it this one: the deliverable of a compliance and continuity engagement is not a hardened cluster, it is the record. What you changed, what you decided not to change and why, and what the value was before you touched it. A hardened cluster comes undone in one bad night; the record is what lets you build it back.
Sources (verified on 6 October 2026): the column explanations, the six steps, the priority definitions and the FAQ answers quoted come from "Security Configuration Guide for VMware Cloud Foundation (VCF)", headed on the page as "VCF Security Hardening Guidance", published on 5 October 2026 on the VMware Cloud Foundation blog and signed by Bob Plankers. The VCF 9.0 and 9.1 control files are in the public vcf-security-and-compliance-guidelines repository; we used the 9.1 CSV (version 910-20260612-01) and the 9.0 one (version 902-20260612-01). Every italicised quotation is in English and verbatim. The following are OURS, produced with a script over those two CSV files so anyone can redo them: the counts of 260 and 226 controls, the 160 and 135 non-default values, the split by component, the match between priority and action needed, the 32 controls mentioning recovery, the NIST, PCI and STIG mappings, and the identifier comparison between the two versions. Also our own reading, and flagged as such in the text: the contrast between the priority of Lockdown Mode and that of the console access list, and the observation about the standard switch versus the distributed one. The 14-minute recovery drill figure is everyWAN internal, from a test of ours, and is offered as evidence rather than a contractual commitment. Study cited: Oppenheimer, Ganapathi and Patterson, "Why Do Internet Services Fail, and What Can Be Done About It?", USENIX 2003. Cover photograph: "Front of server racks at NERSC", by Derrick Coetzee, Wikimedia Commons, CC0; cropped and darkened by us.
Do you know which controls you skipped, and why?
We audit your platform against the guide that matches its version, put what was applied and what was discarded in writing through compliance and continuity, review the rest of the surface with cybersecurity and, if a control is not worth it for you, we will say so: that is what vendor-agnostic consultancy is for. And we re-time the restore afterwards.
Talk to everyWAN