Skip to content
Close Menu

    Subscribe to Updates

    Get the latest news from tastytech.

    What's Hot

    The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And How to Reclaim Focus

    July 21, 2026

    5 Free Courses to Go From AI Beginner to Practitioner

    July 21, 2026

    VCF Upgrade Field Guide: Planning the Maintenance Move

    July 21, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    tastytech.intastytech.in
    Subscribe
    • AI News & Trends
    • Tech News
    • AI Tools
    • Business & Startups
    • Guides & Tutorials
    • Tech Reviews
    • Automobiles
    • Gaming
    • movies
    tastytech.intastytech.in
    Home»AI Tools»VCF 5.2.x to 9.1 Upgrade Runbook: Exact Sequence, Dependencies, Downtime, and Validation
    VCF 5.2.x to 9.1 Upgrade Runbook: Exact Sequence, Dependencies, Downtime, and Validation
    AI Tools

    VCF 5.2.x to 9.1 Upgrade Runbook: Exact Sequence, Dependencies, Downtime, and Validation

    gvfx00@gmail.comBy gvfx00@gmail.comJuly 21, 2026No Comments30 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email


    Table of Contents

    Toggle
    • TL;DR
    • Introduction
    • Scope, Assumptions, and Support Boundary
      • Source Version Go or No-Go Gate
      • What “Exact Sequence” Means
    • The Upgrade Dependency Chain at a Glance
    • Readiness Gates Before the Change Window
      • Establish a Verified Source and Target Baseline
      • Validate Aria Operations Before It Becomes VCF Operations
      • Prepare VCF Management Services Before SDDC Manager Is Upgraded
      • Validate Core Infrastructure Compatibility
      • Prove Backup and Restore Readiness
    • The Exact VCF 5.2.x to 9.1 Upgrade Sequence
    • Component Runbook and Hold Points
      • Upgrade VCF Operations First
      • Upgrade SDDC Manager
      • Deploy VCF Management Services
      • Upgrade VCF Operations for Networks and HCX Where Present
      • Upgrade NSX Federation Global Managers Before Local Managers
      • Upgrade NSX Local Managers
      • Upgrade vCenter Server
      • Transition Baseline-Managed Clusters to vLCM Images
      • Upgrade vSAN Witness Hosts Before Data Hosts
      • Upgrade ESX Hosts
      • Upgrade NSX Edge Nodes and Finalize NSX
      • Complete Post-Core Service Transitions
    • Downtime and Maintenance Window Model
    • Validation Framework
      • Platform Control-Plane Validation
      • Infrastructure Data-Plane Validation
      • Workload and Business-Service Validation
    • Evidence Collection Package
      • Capture a Repeatable vSphere Evidence Baseline
    • Fallback Boundaries and Stop Criteria
      • Stop the Upgrade When
    • Common Blockers to Resolve Before Execution
      • Unsupported Source or Patch Level
      • Management Services IP or DNS Failure
      • VCF Operations Collection Gaps
      • vLCM Baseline Dependencies
      • ESX Hardware and Bootbank Constraints
      • Stale NSX or vCenter Inventory
      • Insufficient Maintenance Capacity
    • Operational Handoff After the Upgrade
    • Conclusion
    • External References
      • Next Post
      • Related posts:
    • US slams Russia’s ‘dangerous escalation’ in Ukraine amid new deadly strikes | Russia-Ukraine war New...
    • How the UPS cargo plane crashed in Louisville, what we know about victims | Aviation News
    • Drone captures ongoing rescue efforts after Venezuela earthquakes | Earthquakes

    TL;DR

    A VCF 5.2.x to 9.1 upgrade is not a single SDDC Manager update followed by routine infrastructure patching. It is a dependency-controlled platform transition that begins with the Operations layer, introduces mandatory VCF Management Services, and then moves through NSX, vCenter, ESX, and NSX Edge finalization.

    The first gate is source-release eligibility. Current Broadcom planning guidance provides paths for VCF 5.2.0 through 5.2.3, subject to the exact source build, bill of materials, and environment-specific planner results. VCF 5.2.4 must not be forced into the currently blocked 9.1 paths.

    Every upgrade stage needs five things:

    • A documented entry condition
    • A named technical owner
    • A realistic downtime or service-impact expectation
    • A validation and evidence package
    • A defined stop, retry, escalation, or fallback boundary

    Running virtual machines may remain available through much of the process, but management, lifecycle, observability, network control, host evacuation, storage synchronization, and north-south routing remain exposed to real operational risk.

    Introduction

    The difficult part of a VCF 5.2.x to 9.1 upgrade is not finding the upgrade workflow. The difficult part is preserving a valid dependency chain while multiple management planes, lifecycle authorities, and infrastructure layers change underneath the environment.

    Broadcom’s published sequence places the Operations components before the core VCF upgrade. That order is not an administrative preference. VCF Operations becomes part of the control path used to continue lifecycle management in VCF 9.1.

    SDDC Manager follows the Operations transition. VCF Management Services is then deployed, establishing the fleet lifecycle, instance lifecycle, software depot, identity, licensing, and runtime services required by the updated platform. Only after those layers are healthy should the core NSX, vCenter, ESX, and Edge work proceed.

    This runbook consolidates that transition into one operational pillar. It is intended for architects, VCF platform engineers, network teams, virtualization teams, storage teams, observability owners, application owners, and change managers who need one shared sequence with explicit prerequisites, validation gates, evidence requirements, downtime expectations, and fallback limits.

    This article does not replace the environment-specific output from the current VCF Upgrade Planning Tool. It provides the operational structure into which that generated plan should be inserted.

    Scope, Assumptions, and Support Boundary

    This runbook covers an SDDC Manager-managed VCF 5.2.x environment moving to VCF 9.1, with the following core components in scope:

    • Aria Operations transitioning to VCF Operations
    • Aria Suite Lifecycle dependencies where present
    • SDDC Manager
    • VCF Management Services
    • NSX Global Managers where Federation is used
    • NSX Local Managers
    • NSX Edge nodes
    • vCenter Server
    • ESX hosts
    • vSAN and stretched-cluster dependencies
    • VCF Operations for Networks where present
    • Post-upgrade identity, logging, orchestration, and lifecycle cleanup
    • End-to-end validation and evidence collection

    Conditional products such as vSphere Replication, Site Recovery Manager, Avi Load Balancer, HCX, NSX Federation, vSphere Supervisor, vSAN Data Protection, VxRail, vSAN File Service, and standalone Orchestrator must be inserted at the stage defined by the current planner.

    VCF Automation must also be treated as an environment-specific branch. Its transition depends on the installed Aria Automation, Aria Suite Lifecycle, identity, and integration topology. Do not add Automation to the sequence from memory. Generate a planner path that includes it.

    Source Version Go or No-Go Gate

    Do not begin with bundle downloads. Begin by proving that the installed source release, exact build, and component bill of materials have a supported target path.

    Source state Runbook decision Required action
    VCF 5.2.0 Potentially eligible Confirm the exact direct path and prerequisites in the current planner
    VCF 5.2.1 Potentially eligible Confirm the exact direct path and prerequisites in the current planner
    VCF 5.2.2 Eligible when the planner and interoperability checks pass Continue through readiness gates
    VCF 5.2.3 Eligible when the planner and interoperability checks pass Continue through readiness gates
    VCF 5.2.4 Current direct 9.1 paths are blocked Stop and wait for a supported target path
    Unknown or mixed bill of materials Unsupported planning condition Reconcile the live inventory before continuing

    This is a version-sensitive decision. Recheck it immediately before the maintenance window, even when the plan was approved several weeks earlier.

    A source environment is not eligible merely because SDDC Manager displays a 5.2.x version. Every managed component must align with the expected bill of materials or be explicitly supported by the planner.

    What “Exact Sequence” Means

    The sequence is exact at two levels.

    Mandatory dependency order: Operations must be handled before SDDC Manager and the core VCF stack. VCF Management Services must exist and be healthy before the remaining fleet-managed lifecycle work proceeds.

    Environment-specific inserts: Optional and adjacent products appear at specific points according to what is installed. These components do not replace the core order, but they may introduce additional validation gates and maintenance windows.

    The safe interpretation is not “run every possible product step.” It is “preserve the generated order for every component that exists in the environment.”

    The Upgrade Dependency Chain at a Glance

    The following diagram shows the governing flow. Optional products branch into the sequence without changing the order of the core platform.

    The central point is the lifecycle-control handoff. VCF Operations is not a monitoring appliance that can be upgraded after the infrastructure. It becomes part of the management path required to move the remaining environment forward.

    Readiness Gates Before the Change Window

    The upgrade should not enter an execution window until the source environment, target path, recovery state, and operational ownership have all been proven.

    The readiness phase should produce:

    • A signed component and version matrix
    • A generated environment-specific upgrade plan
    • Completed product prechecks
    • Validated backups and restore ownership
    • Reserved Management Services networking
    • Hardware and firmware compatibility evidence
    • A cluster evacuation model
    • A documented application-validation plan
    • A decision and escalation matrix
    • A shared evidence repository

    Establish a Verified Source and Target Baseline

    Create a component matrix from the live environment, not from an old architecture document.

    Component Current version and build Target version and build Upgrade owner Backup verified Precheck
    Aria or VCF Operations Record live value Planner target Operations team Yes or no Pass or fail
    Aria Suite Lifecycle Record live value Transition state Operations team Yes or no Pass or fail
    SDDC Manager Record live value Planner target VCF platform team Yes or no Pass or fail
    NSX Global Managers Record live value Planner target Network team Yes or no Pass or fail
    NSX Local Managers Record live value Planner target Network team Yes or no Pass or fail
    NSX Edge nodes Record live value Planner target Network team Yes or no Pass or fail
    vCenter Server Record live value Planner target Virtualization team Yes or no Pass or fail
    ESX hosts Record live value Planner target Virtualization team Yes or no Pass or fail
    vSAN witness hosts Record live value Planner target Platform team Yes or no Pass or fail
    Optional products Record live value Planner target Named owner Yes or no Pass or fail

    A mixed or undocumented bill of materials is a stop condition. Version drift tolerated during normal operations can become a hard lifecycle blocker during a major upgrade.

    Validate Aria Operations Before It Becomes VCF Operations

    Operations is the first production dependency, so its readiness must be proven before the core maintenance window.

    Perform the following checks:

    • Confirm the exact Aria Operations version, patch, and build.
    • Confirm the Aria Suite Lifecycle version and product inventory.
    • Run the current Aria Operations pre-upgrade readiness assessment.
    • Review discontinued or changed metrics used by dashboards, alerts, reports, and super metrics.
    • Confirm cluster health, node state, capacity, certificates, DNS, NTP, and administrative credentials.
    • Record Remote Collectors, Cloud Proxies, collector groups, adapters, outbound notifications, and authentication sources.
    • Confirm the supported path for the exact source patch level.
    • Verify the post-upgrade collector and proxy transition requirements.
    • Define how observability coverage will be maintained while Operations is unavailable.

    Patch-level boundaries matter. Broadcom currently documents a supported transition from Aria Operations 8.18.6 to VCF Operations 9.1, while the 8.18.7 scenario requires additional caution under the current guidance. Treat any unresolved patch combination as a stop condition until the source knowledge base and planner show a supported path.

    Plan enough time for the upgrade. Operating-system and database work can make progress appear static. A quiet progress bar is not proof that the process has failed. Monitor the active upgrade logs before declaring the workflow stalled.

    Prepare VCF Management Services Before SDDC Manager Is Upgraded

    VCF Management Services is mandatory for VCF 9.1. It is not an optional Day-N feature.

    Prepare the following before the change window:

    • At least 12 reserved IP addresses for the initial deployment
    • Capacity to expand the allocation toward 30 addresses for future services and scale
    • Forward and reverse DNS records for required service FQDNs
    • A and PTR records for the centralized VCF License Server
    • Management-network reachability
    • Required firewall paths for deployment, lifecycle, DNS, NTP, identity, depot, and licensing
    • Certificates matching the FQDNs used by the deployment
    • Sufficient management-domain compute, memory, and storage
    • A non-overlapping internal services range
    • Named owners for DNS, networking, licensing, certificates, and identity

    The default internal services range is 198.18.0.0/15. If that range overlaps with an existing management, laboratory, routing, or test environment, select a supported alternative before deployment.

    An overlap discovered during execution is not a minor documentation problem. It can invalidate service networking and require cleanup or redeployment.

    Validate Core Infrastructure Compatibility

    Before staging payloads, confirm:

    • Hardware compatibility for the exact server, storage controller, NIC, HBA, firmware, BIOS, and target ESX release
    • Sufficient ESX bootbank space for the target image
    • Cluster capacity to evacuate at least one host at a time
    • DRS, affinity, anti-affinity, fault-tolerance, passthrough, vGPU, local-storage, and pinned-workload constraints
    • vSAN health, resynchronization state, object compliance, witness reachability, and free capacity
    • NSX Manager cluster health, transport-node health, TEP reachability, Edge high availability, and routing stability
    • vCenter service health, certificate validity, identity configuration, backup status, and external integrations
    • Whether any cluster remains managed by vSphere Lifecycle Manager baselines
    • Whether custom VIBs, drivers, or vendor add-ons conflict with the target image
    • Whether stretched clusters have an approved witness-upgrade sequence

    VCF 9.1 lifecycle operations rely on the target image-management model. Treat baseline-managed clusters as an upgrade dependency, not as a cleanup item to address after vCenter has already moved.

    Prove Backup and Restore Readiness

    A successful backup job is not the same as a recoverable platform.

    Before the first upgrade action:

    • Verify recent file-based backups for SDDC Manager, NSX, and vCenter.
    • Confirm the backup destination does not depend on the component being upgraded.
    • Record encryption passwords, repository access, retention, and restore ownership.
    • Confirm the supported snapshot procedure for Aria or VCF Operations.
    • Ensure Operations cluster snapshots capture all required nodes consistently.
    • Define when temporary snapshots will be deleted.
    • Verify the restoration sequence for each management component.
    • Record which restore or cleanup actions require Broadcom Support.
    • Confirm that application owners understand the infrastructure recovery boundaries.
    • Preserve all backup identifiers in the change evidence repository.

    Do not call a snapshot a rollback plan unless the restore order, integration consequences, dependency state, and post-restore validation are documented.

    The Exact VCF 5.2.x to 9.1 Upgrade Sequence

    The following table provides the governing control sequence for the core VCF 5.2.x to 9.1 path. Use the current planner to remove components that do not exist and insert environment-specific instructions.

    Stage Component or activity Why it is positioned here Exit condition
    Readiness Source eligibility, inventory, backups, prechecks, evidence baseline Prevents an unsupported or unrecoverable start All go or no-go gates signed
    Operations Aria Operations to VCF Operations Establishes the updated lifecycle and observability control path Cluster healthy and collection restored
    Conditional pre-core vSphere Replication, SRM, vSAN Data Protection, Avi Keeps protection and adjacent services compatible Product-specific validation passed
    Core lifecycle SDDC Manager Moves the VCF instance control plane to 9.1 SDDC Manager healthy and synchronized
    Management layer VCF Management Services Introduces required fleet and instance services Services, DNS, identity, licensing, and depot validated
    Management add-ons VCF Operations for Networks and HCX where present Transitions adjacent management products through the supported workflow Product health and data collection validated
    Global networking NSX Federation Global Managers where present Preserves Global and Local Manager compatibility Federation healthy
    Local networking NSX Local Managers Prepares network management before compute and host lifecycle Managers, policy, transport nodes, and routing healthy
    Compute control vCenter Server Updates the management plane required for host lifecycle Services, identity, inventory, backup, and integrations healthy
    Lifecycle transition vLCM baseline-to-image migration Establishes the required target cluster lifecycle model Clusters are image-managed and compliant
    Platform services vSphere Supervisor where present Aligns Supervisor with the updated vCenter lifecycle Supervisor healthy
    Stretched clusters vSAN witness hosts Aligns witness compatibility before data hosts Witness version and connectivity validated
    Integrated HCI VxRail where applicable Follows the Dell-supported integrated lifecycle process VxRail validation passed
    Hypervisor ESX hosts Performs rolling host lifecycle after control planes are healthy All hosts connected, compliant, and healthy
    Network data plane NSX Edge nodes and NSX finalization Aligns north-south services with the updated platform Edge HA, routing, services, and traffic validated
    Post-core transition Identity, Orchestrator, Logs, ELM, vSAN File Service Completes transitional service work Replacement services validated
    Closure End-to-end validation and evidence sign-off Converts task completion into operational acceptance Acceptance criteria signed

    Component Runbook and Hold Points

    Upgrade VCF Operations First

    Entry conditions

    • The exact Aria Operations version is supported.
    • The readiness assessment is complete.
    • The Operations cluster is healthy.
    • Capacity and disk-space checks pass.
    • The supported clustered snapshot or backup procedure is approved.
    • Collector, proxy, adapter, dashboard, alert, and notification inventories are captured.
    • The Operations owner and VCF lifecycle owner are available.

    Execution focus

    Upgrade through the supported Aria Suite Lifecycle or VCF Operations workflow. Do not bypass the Operations transition because the remaining 9.1 lifecycle content depends on the updated control plane.

    Complete all required Cloud Proxy, Remote Collector, and collector-group work after the appliance transition. Metrics continuity is part of the upgrade, not a cosmetic post-task.

    Validation

    • All Operations nodes are online.
    • Cluster health is green.
    • Object and metric collection has resumed.
    • SDDC Manager, vCenter, NSX, and other adapters are collecting.
    • Alerts, dashboards, reports, notifications, and super metrics operate as expected.
    • Fleet lifecycle functions are available.
    • Authentication and administrative access succeed.
    • No unresolved database, analytics, adapter, or collector errors remain.

    Evidence

    Capture source and target builds, cluster status, node status, adapter status, collector-group state, recent metric timestamps, active alerts, upgrade logs, lifecycle inventory, and snapshot or backup identifiers.

    Fallback boundary

    Before SDDC Manager is upgraded, Operations remains the most isolated major recovery boundary. A supported clustered snapshot restore may still be possible when snapshots were taken correctly.

    Restoration is not complete until authentication, collection, VCF SSO or Identity Broker integration, and lifecycle functions have all been retested.

    Upgrade SDDC Manager

    Entry conditions

    • VCF Operations is healthy and collecting.
    • The target bundle appears through the supported lifecycle workflow.
    • The source SDDC Manager release is eligible.
    • The SDDC Manager backup is current and verified.
    • Passwords, certificates, DNS, NTP, depot access, and disk capacity pass validation.
    • No domain creation, host commissioning, password rotation, certificate replacement, or lifecycle workflow is active.

    Execution focus

    Upgrade SDDC Manager through the supported lifecycle workflow. Record the workflow ID, source build, target build, operator, and start time before execution.

    Monitor the workflow through both the user interface and the relevant service logs. Do not depend on a single progress indicator.

    Validation

    • SDDC Manager services are healthy.
    • Management-domain and workload-domain inventories are complete.
    • vCenter, NSX, clusters, hosts, and credentials are synchronized.
    • Depot and bundle configuration is available.
    • No stale or failed lifecycle tasks remain.
    • File-based backup scheduling works.
    • VCF Operations sees the updated instance state.

    Evidence

    Capture the task ID, timestamps, source and target builds, precheck report, upgrade-log summary, service health, inventory view, depot state, and backup configuration.

    Fallback boundary

    After SDDC Manager moves to 9.1, do not treat an appliance snapshot as a universal undo.

    If SDDC Manager validation fails, stop before deploying VCF Management Services. Use the documented recovery process or engage Broadcom Support. Do not continue in the hope that later components will repair an unhealthy control plane.

    Deploy VCF Management Services

    Entry conditions

    • SDDC Manager 9.1 is healthy.
    • Required IP addresses are reserved.
    • Service FQDNs have forward and reverse DNS.
    • The License Server has valid A and PTR records.
    • The internal services range does not overlap with the environment.
    • Compute, memory, storage, and network capacity is available.
    • Credentials and certificates pass validation.
    • Required firewall paths are open.

    Execution focus

    Deploy VCF Management Services from the supported 9.1 workflow.

    Treat this as a platform deployment with its own architecture and readiness requirements. Validate each service as it appears. A partially healthy Management Services deployment is not an acceptable foundation for NSX, vCenter, or host lifecycle work.

    Validation

    • All Management Services nodes and runtime components are healthy.
    • Fleet Lifecycle is available.
    • SDDC Lifecycle is available.
    • Software depot access works.
    • The License Server is reachable and registered.
    • Required FQDNs resolve forward and reverse.
    • Identity and administrative access work.
    • Capacity and IP consumption match the design.
    • VCF Operations displays the expected instance and lifecycle inventory.

    Evidence

    Capture service inventory, node health, assigned IP addresses, DNS forward and reverse tests, license registration, depot status, certificate details, deployment task ID, and support bundles where generated.

    Fallback boundary

    Rollback becomes significantly more complex after Management Services is deployed.

    A failed deployment may require a supported retry or component-specific cleanup. Preserve all task IDs and logs, back up critical data, and engage Broadcom Support before performing destructive cleanup or restoring components outside the documented workflow.

    Upgrade VCF Operations for Networks and HCX Where Present

    VCF Operations for Networks must follow the planner-defined Fleet Lifecycle path.

    Validate:

    • Platform and collector node health
    • Data-source connectivity
    • Flow collection
    • Search
    • Topology
    • Alerts
    • Retention
    • User access
    • Integration with the updated management environment

    HCX is also conditional.

    Before and after its planner-defined stage, validate:

    • Site pairing
    • Compute profiles
    • Network profiles
    • Service meshes
    • Interconnect appliances
    • Network extensions
    • Migration state
    • HCX Manager backups
    • Application traffic crossing extended networks

    Do not run planned workload migrations while the VCF control planes are being upgraded.

    Upgrade NSX Federation Global Managers Before Local Managers

    Where NSX Federation is deployed, Global Managers must lead the Local Manager transition.

    Entry conditions

    • Federation health is green.
    • Inter-site connectivity and latency are stable.
    • Global and Local Manager backups are current.
    • No active onboarding, policy redesign, Edge replacement, or transport-node remediation is running.
    • Recovery ownership is documented for every location.

    Validation

    • The Global Manager cluster is healthy.
    • Local Managers remain registered and reachable.
    • Policy synchronization succeeds.
    • Global and local configurations remain consistent.
    • No stale site, host, or Edge objects block the next stage.
    • Cross-location policy and routing behavior remains correct.

    Do not proceed to Local Managers while Federation health or synchronization remains degraded.

    Upgrade NSX Local Managers

    Entry conditions

    • VCF Management Services is healthy.
    • NSX Manager clusters are healthy.
    • Backups are current and restorable.
    • Transport nodes, TEP tunnels, host switches, and Edge clusters are healthy.
    • Stale or unsupported host inventory has been reconciled.
    • Certificates, FQDNs, DNS, and NTP are correct.
    • Unrelated firewall, segment, gateway, NAT, VPN, and routing changes are frozen.

    Execution focus

    Upgrade the NSX management plane before vCenter and ESX.

    Maintain a strict network-change freeze during this stage. The goal is to validate the upgraded management and policy planes without introducing unrelated configuration variables.

    Validation

    • All Manager nodes are healthy.
    • Management and policy clusters are stable.
    • Transport nodes remain connected.
    • TEP tunnels are healthy.
    • Distributed firewall policy is realized.
    • Tier-0 and Tier-1 gateways are healthy.
    • BGP and BFD sessions are stable.
    • NAT, VPN, DHCP, load balancing, and service insertion operate where used.
    • Edge clusters remain healthy pending their later upgrade.
    • VCF inventory reflects the expected NSX version.

    Evidence

    Capture Manager health, cluster health, transport-node status, TEP state, Edge state, route summaries, BGP and BFD state, alarms, backup state, and lifecycle task identifiers.

    Fallback boundary

    Prefer repair and retry over independent NSX rollback. Once vCenter or hosts advance, restoring only NSX can create unsupported version skew.

    Stop at the NSX validation gate when the network control plane is unhealthy.

    Upgrade vCenter Server

    Entry conditions

    • NSX Managers are healthy.
    • The vCenter file-based backup is current.
    • Root and SSO credentials are validated.
    • Certificates, FQDN, DNS, NTP, identity sources, and external integrations are documented.
    • Legacy authentication dependencies have been remediated where required.
    • No active cluster remediation, certificate replacement, host-profile operation, or inventory migration is running.

    Execution focus

    Upgrade vCenter through the VCF Operations Fleet Lifecycle workflow.

    During the appliance transition, expect management-plane unavailability. Running virtual machines normally continue, but provisioning, vMotion initiation, DRS actions requiring vCenter, automation, backup integrations, and monitoring can be constrained.

    Validation

    • All vCenter services are healthy.
    • SSO and administrative login succeed.
    • Inventory is complete.
    • Hosts and clusters are connected.
    • NSX integration is healthy.
    • Datastores and distributed switches are present.
    • Content libraries, tags, alarms, scheduled tasks, permissions, and roles remain intact.
    • Backup configuration works.
    • External integrations reconnect.
    • Enhanced Linked Mode state is correct where used.

    Evidence

    Capture the vCenter build, service health, SSO test, inventory counts, host states, alarm summary, backup result, integration tests, and lifecycle task ID.

    Fallback boundary

    Use workflow retry or repair before considering a restore.

    Once ESX hosts begin moving to the target version, an independent vCenter reversion becomes increasingly unsafe because the management and host versions may no longer form a supported combination.

    Transition Baseline-Managed Clusters to vLCM Images

    Do not wait until ESX remediation fails to discover that a cluster still depends on baselines.

    For every cluster:

    • Build or import the approved target image.
    • Include the correct vendor add-on.
    • Include required drivers and components.
    • Validate hardware compatibility.
    • Check firmware dependencies.
    • Remove conflicting custom VIBs.
    • Test image compliance.
    • Resolve remediation blockers.
    • Confirm Supervisor dependencies where present.
    • Record the approved image definition and owner.

    The exit condition is not simply that an image exists. The cluster must be image-managed, compliant, and ready for rolling host remediation.

    Upgrade vSAN Witness Hosts Before Data Hosts

    For stretched clusters, upgrade witness hosts before the data hosts.

    Validate:

    • Witness connectivity
    • Fault-domain membership
    • Witness appliance or host version
    • Object health
    • Resynchronization state
    • Network latency
    • Storage policies
    • Cluster fault tolerance
    • Recovery ownership

    Do not begin data-host remediation while witness health or object availability is degraded.

    Upgrade ESX Hosts

    Entry conditions

    • vCenter is healthy.
    • Cluster lifecycle images are valid and compliant.
    • Witness hosts are upgraded where required.
    • vSAN is healthy.
    • No uncontrolled resynchronization is running.
    • DRS and evacuation capacity are proven.
    • Workload constraints are documented.
    • NSX transport-node and Edge health is green.
    • Firmware and drivers match the approved image.
    • ESX bootbank capacity is sufficient.

    A useful host-level check is:

    df -h /bootbank
    

    Review the output against the minimum required by the target image, planner, or applicable Broadcom knowledge-base guidance.

    Execution focus

    Upgrade hosts sequentially.

    For every host:

    1. Confirm evacuation capacity.
    2. Review workload exceptions.
    3. Enter maintenance mode through the supported workflow.
    4. Verify VM migration or approved shutdown.
    5. Apply the image and reboot.
    6. Confirm management, storage, and NSX connectivity.
    7. Exit maintenance mode.
    8. Wait for vSAN stabilization.
    9. Validate the host before beginning the next one.

    Do not run the cluster as a blind batch when host hardware, firmware, networking, or workload constraints differ.

    Validation

    • The host reports the target version and build.
    • The host is connected.
    • Maintenance mode is cleared.
    • vLCM compliance is green.
    • Management, vMotion, vSAN, and TEP vmkernel interfaces are healthy.
    • Physical uplinks and distributed-switch membership are correct.
    • NSX transport-node status is healthy.
    • vSAN objects are healthy.
    • Resynchronization is within the approved threshold.
    • DRS and HA are normal.
    • Representative VMs run and migrate successfully.

    Evidence

    Capture each host’s pre-state, maintenance-mode timestamps, image result, version, build, vmkernel state, uplink state, NSX status, vSAN health, resynchronization state, failed migrations, and approved exceptions.

    Fallback boundary

    A failed host remediation is normally a host-recovery problem, not a reason to downgrade the entire platform.

    Keep the host isolated, preserve logs, restore bootability through the supported recovery method, and continue only after cluster redundancy has been restored.

    Upgrade NSX Edge Nodes and Finalize NSX

    NSX Edge nodes follow the ESX host layer in the core sequence. This aligns the north-south data plane with the updated management, networking, and hypervisor layers.

    Entry conditions

    • Applicable ESX host remediation is complete.
    • NSX Managers and transport nodes are healthy.
    • Edge cluster high availability is healthy.
    • BGP and BFD peers are stable.
    • Traffic baselines are recorded.
    • Application tests are ready.
    • A network owner is available throughout the failover period.

    Execution focus

    Upgrade Edge nodes according to their high-availability topology.

    Expect service movement, routing reconvergence, and possible brief packet loss as Edge nodes restart or services fail over. Avoid unrelated gateway, routing, NAT, VPN, DHCP, load-balancer, or DNS-forwarding changes during this stage.

    Perform NSX finalization only after Manager, host, Edge, routing, and application validation has passed.

    Validation

    • Every Edge node reports the target version.
    • Edge cluster high availability is healthy.
    • Tier-0 and Tier-1 service routers are active as designed.
    • BGP and BFD sessions recover.
    • North-south and east-west traffic passes.
    • NAT works.
    • VPN services work where present.
    • DHCP and DNS forwarding work where used.
    • Load balancing and service insertion work where used.
    • Distributed firewall enforcement and logging remain correct.
    • No stale host or Edge inventory blocks finalization.
    • NSX finalization completes successfully.

    Evidence

    Capture Edge versions, high-availability state, routing neighbors, route counts, failover timestamps, packet-loss observations, application tests, service tests, and finalization-task output.

    Fallback boundary

    Before finalization, a failed Edge upgrade may still be addressed through node-specific repair or replacement.

    After finalization, the preferred posture is forward recovery. Do not independently restore an older Manager, Edge, vCenter, or ESX state without a coordinated compatibility plan.

    Complete Post-Core Service Transitions

    After the core infrastructure is healthy, complete the planner-defined service transitions.

    These can include:

    • Migrating VMware Identity Manager dependencies to the supported Identity Broker model
    • Transitioning standalone Orchestrator
    • Deploying or enabling VCF Operations for Logs
    • Migrating required historical log data
    • Resolving Enhanced Linked Mode according to the approved topology
    • Upgrading or validating vSAN File Service
    • Completing Operations for Networks integration
    • Verifying HCX services
    • Updating backup integrations
    • Updating monitoring and automation integrations
    • Preparing Day-N workload-domain upgrade plans

    Do not decommission a legacy appliance merely because its replacement exists.

    Decommission only after functional validation passes, required data has been retained or migrated, authentication paths work, integration owners sign off, recovery requirements are met, and the replacement service has entered the support model.

    Downtime and Maintenance Window Model

    There is no responsible universal duration for the entire upgrade.

    Total time depends on:

    • Source patch levels
    • Number of optional products
    • Number of workload domains
    • Cluster and host count
    • Host evacuation time
    • vSAN resynchronization
    • Hardware performance
    • Bundle-download strategy
    • Network convergence
    • Application testing
    • Number of enforced validation holds
    • Whether remediation or retries are required

    Plan each stage according to management-plane impact and possible workload impact.

    Stage Expected management-plane impact Expected workload impact Planning guidance
    VCF Operations Operations UI, analytics, alerts, and collection are unavailable or degraded Running VMs continue, but monitoring coverage is reduced Reserve time for upgrade, stabilization, collectors, and validation
    SDDC Manager Lifecycle and domain-management workflows are unavailable Running workloads normally continue Freeze lifecycle, certificate, password, and domain changes
    VCF Management Services New services are deployed and integrated Running workloads normally continue Include time for DNS, licensing, certificates, and deployment retries
    NSX Managers Network management and policy control may be interrupted Existing forwarding should continue when the data plane is healthy Freeze policy changes and validate control-plane recovery
    vCenter Server vCenter UI, APIs, automation, DRS control, and integrations are unavailable Running VMs normally continue, but management operations are constrained Notify backup, monitoring, automation, and application teams
    ESX hosts One host enters maintenance mode and reboots No outage only when evacuation succeeds Model N+1 capacity and workload constraints
    NSX Edge nodes Edge services fail over and routing reconverges Brief north-south interruption or packet loss is possible Validate HA, peer timers, retries, and synthetic transactions
    Post-core services Identity, logging, or orchestration changes Impact depends on service dependencies Use service-specific acceptance tests

    A realistic maintenance plan separates four clocks:

    • Execution time: Time spent running lifecycle tasks.
    • Stabilization time: Time required for services, analytics, routes, collectors, and vSAN to settle.
    • Validation time: Time required for platform and application owners to prove acceptance.
    • Decision buffer: Time reserved for retry, evidence collection, escalation, or a controlled stop.

    A maintenance window that covers only execution time is incomplete.

    Validation Framework

    Validation should move from internal component health to representative business-service behavior.

    A green lifecycle task is necessary, but it is not sufficient.

    The next stage should begin only after the current component, its integrations, and the services that depend on it have passed their defined acceptance criteria.

    Platform Control-Plane Validation

    Confirm:

    • VCF Operations cluster health
    • Metrics collection
    • Adapter status
    • Alerts and dashboards
    • Lifecycle inventory
    • Administrative authentication
    • SDDC Manager services
    • Domain inventory
    • Depot connectivity
    • Credentials and certificates
    • Backup configuration
    • VCF Management Services node health
    • Fleet Lifecycle
    • SDDC Lifecycle
    • Identity
    • Licensing
    • Software depot
    • NSX Manager and policy health
    • Transport-node health
    • Edge health
    • Routing state
    • vCenter services
    • SSO
    • Host and cluster inventory
    • Datastores and distributed switches
    • ESX host connectivity and compliance
    • vSAN health
    • Witness status
    • Optional-product health

    Infrastructure Data-Plane Validation

    Test:

    • Management-network reachability
    • Forward DNS resolution
    • Reverse DNS resolution
    • NTP consistency
    • vMotion
    • vSAN I/O
    • vSAN object health
    • TEP tunnel health
    • East-west traffic across representative segments
    • North-south traffic through every critical Tier-0 path
    • BGP convergence
    • BFD convergence
    • Distributed firewall enforcement
    • NAT
    • VPN
    • Load balancing
    • DHCP
    • DNS forwarding
    • Backup connectivity
    • Monitoring connectivity
    • Administrative jump-host access

    Workload and Business-Service Validation

    Select representative services before the change window.

    For each service, document:

    • Application owner
    • Test endpoint or transaction
    • Authentication path
    • Network path
    • Database dependency
    • Storage dependency
    • Expected result
    • Maximum acceptable interruption
    • Evidence format
    • Sign-off authority

    At least one meaningful transaction should cross every critical identity, network, storage, and application boundary.

    The infrastructure upgrade is not accepted until the application owners confirm that the platform is delivering the services it was upgraded to support.

    Evidence Collection Package

    Evidence should be collected during the workflow, not reconstructed after an incident.

    Use a structure such as:

    VCF-5.2-to-9.1-Evidence/
    |-- 00-change-control/
    |   |-- approved-plan
    |   |-- contacts-and-escalation
    |   `-- decision-log
    |-- 01-source-baseline/
    |   |-- component-builds
    |   |-- health-reports
    |   |-- backup-confirmation
    |   `-- precheck-results
    |-- 02-operations/
    |-- 03-sddc-manager/
    |-- 04-management-services/
    |-- 05-nsx/
    |-- 06-vcenter/
    |-- 07-esx-and-vsan/
    |-- 08-edge-finalization/
    |-- 09-application-validation/
    |-- 10-support-bundles/
    `-- 11-final-signoff/
    

    Every stage record should include:

    • Stage name
    • Component
    • Source version
    • Target version
    • Task or workflow ID
    • Operator
    • Start time
    • End time
    • Entry-gate result
    • Backup or snapshot identifier
    • Precheck result
    • Validation result
    • Exceptions
    • Log location
    • Support-bundle location
    • Decision to proceed, retry, stop, or escalate
    • Approver

    Capture a Repeatable vSphere Evidence Baseline

    The following PowerCLI example creates a timestamped evidence folder and exports host, cluster, and datastore state.

    Change the vCenter name and evidence path before running it. Execute the script before the upgrade and after every affected workload domain.

    $VCenter = "vcsa01.example.local"
    $EvidenceRoot = "C:\VCF91-Evidence"
    $Stamp = Get-Date -Format "yyyyMMdd-HHmmss"
    $OutputPath = Join-Path $EvidenceRoot $Stamp
    
    New-Item -ItemType Directory -Path $OutputPath -Force | Out-Null
    
    Connect-VIServer -Server $VCenter
    
    Get-VMHost |
        Select-Object Name, ConnectionState, Version, Build |
        Export-Csv `
            -Path (Join-Path $OutputPath "esx-hosts.csv") `
            -NoTypeInformation
    
    Get-Cluster |
        Select-Object Name, HAEnabled, DrsEnabled |
        Export-Csv `
            -Path (Join-Path $OutputPath "clusters.csv") `
            -NoTypeInformation
    
    Get-Datastore |
        Select-Object Name, Type, CapacityGB, FreeSpaceGB |
        Export-Csv `
            -Path (Join-Path $OutputPath "datastores.csv") `
            -NoTypeInformation
    
    Disconnect-VIServer -Server $VCenter -Confirm:$false
    

    Successful execution produces timestamped CSV files containing host, cluster, and datastore state.

    This script does not replace SDDC Manager evidence, NSX health reports, VCF Operations evidence, vSAN health, application tests, or product-specific support bundles.

    Common failures include an incorrect vCenter FQDN, expired credentials, a missing PowerCLI module, certificate prompts during unattended execution, and an output path without write permission.

    Fallback Boundaries and Stop Criteria

    There is no single full-stack undo for a VCF upgrade.

    VCF is a dependency stack. Once downstream components advance, restoring one upstream appliance can create a combination that was never intended or tested.

    Boundary Practical recovery posture
    Before the Operations upgrade Abort safely, preserve the baseline, and reschedule
    Operations upgraded, SDDC Manager unchanged A supported Operations restore may remain possible, followed by complete collection and integration validation
    SDDC Manager upgraded, Management Services not deployed Stop and repair or recover SDDC Manager before adding new dependencies
    Management Services deployed Prefer supported retry and component-specific cleanup
    NSX or vCenter upgraded Prefer repair and retry because isolated reversion risks version skew
    ESX remediation started Isolate and recover failed hosts rather than improvising a platform downgrade
    NSX Edge finalization complete Treat the environment as a forward-recovery state unless vendor-supported recovery guidance says otherwise
    Legacy services decommissioned Restore only through the approved service recovery plan and retained data

    Rollback flexibility decreases as the upgrade progresses.

    The change plan should therefore become more conservative at every downstream stage.

    Stop the Upgrade When

    Stop at the current hold point when any of the following occurs:

    • The source or target path is no longer supported.
    • A mandatory precheck fails.
    • A backup cannot be verified.
    • The restore owner is unavailable.
    • VCF Operations is unhealthy.
    • Critical objects are not collecting.
    • Management Services DNS validation fails.
    • Management Services IP allocation is invalid.
    • Certificates do not match service identities.
    • The internal services range overlaps with an existing network.
    • Licensing or depot connectivity fails.
    • SDDC Manager inventory is incomplete.
    • Lifecycle services are unhealthy.
    • NSX Manager or policy health is degraded.
    • Transport nodes or TEP tunnels are unhealthy.
    • Edge high availability is degraded.
    • Routing or firewall validation fails.
    • vCenter services or SSO validation fails.
    • A cluster cannot evacuate a host safely.
    • vSAN has object unavailability or uncontrolled resynchronization.
    • More than one host is unexpectedly unavailable in the same failure domain.
    • Application interruption exceeds the approved threshold.
    • Evidence, logs, task IDs, or support bundles cannot be preserved.
    • The team reaches the decision deadline without enough time for stabilization and validation.

    A controlled stop is not a failed change. Continuing after a stop criterion has been met is what converts a recoverable issue into a larger incident.

    Common Blockers to Resolve Before Execution

    Unsupported Source or Patch Level

    Do not force VCF 5.2.4 toward a currently blocked 9.1 target.

    Do not assume every Aria Operations 8.18 patch behaves identically. Validate the exact Operations source patch and target path.

    Management Services IP or DNS Failure

    Common causes include:

    • Too few reserved IP addresses
    • Missing PTR records
    • FQDN case inconsistencies
    • Certificate names that do not match DNS
    • Firewall omissions
    • Overlap with the internal services range
    • Incorrect License Server records

    VCF Operations Collection Gaps

    The appliance can upgrade successfully while collectors, proxies, adapters, dashboards, alerts, or reports remain incomplete.

    Treat collection continuity as an acceptance criterion.

    vLCM Baseline Dependencies

    Baseline-managed clusters, conflicting vendor add-ons, unsupported drivers, and incomplete image definitions can block the host stage after the management planes have already advanced.

    ESX Hardware and Bootbank Constraints

    Older processors, unsupported firmware, limited bootbank space, custom VIBs, and unvalidated drivers can turn rolling remediation into a recovery problem.

    Stale NSX or vCenter Inventory

    Deleted hosts, stale transport nodes, mismatched certificates, unresolved Enhanced Linked Mode state, and disconnected integrations can block planning or finalization.

    Insufficient Maintenance Capacity

    A cluster may appear healthy but still be unable to evacuate a host because of reservations, affinity rules, passthrough devices, vGPU, local storage, fault tolerance, or application licensing.

    Operational Handoff After the Upgrade

    The upgrade is not finished when every component displays a 9.1 version.

    Complete the handoff by updating:

    • Fleet ownership
    • VCF instance ownership
    • Workload-domain ownership
    • Network ownership
    • Observability ownership
    • Identity ownership
    • License ownership
    • Backup and restore procedures
    • Certificate workflows
    • Password-rotation workflows
    • Software-depot procedures
    • Lifecycle procedures
    • VCF Operations dashboards and alerts
    • Collector and proxy ownership
    • NSX validation procedures
    • Edge failover procedures
    • vCenter support procedures
    • ESX image-management procedures
    • vSAN witness procedures
    • Logging and retention architecture
    • Identity Broker support procedures
    • Day-N workload-domain plans
    • Known exceptions
    • Accepted risks
    • Support-bundle locations
    • Final evidence index

    The most important post-upgrade artifact is a clean operating model.

    VCF 9.1 centralizes more lifecycle and management responsibility. Ownership must follow the new architecture rather than remain attached to the appliances and processes used in VCF 5.2.x.

    Conclusion

    A successful VCF 5.2.x to 9.1 upgrade is a controlled transfer of lifecycle authority.

    Operations comes first because the remaining upgrade depends on it. SDDC Manager follows. Mandatory VCF Management Services then establishes the lifecycle, identity, licensing, software-depot, and runtime foundation required by the updated platform.

    NSX Managers move before vCenter. vCenter moves before ESX. Witness and image-management dependencies must be resolved before host remediation. NSX Edge upgrades and finalization close the core infrastructure transition. Optional products must be inserted exactly where the current planner places them.

    The maintenance impact should be described honestly. Running virtual machines may remain available through many management-plane transitions, but observability, lifecycle control, network policy, vCenter operations, host evacuation, storage stabilization, and Edge convergence all carry real risk.

    That is why every stage needs an entry gate, validation hold, evidence package, decision deadline, and fallback boundary.

    Fallback is also stage-specific. Early in the sequence, a supported restore may remain practical. As Management Services, NSX, vCenter, ESX, and Edge states advance, isolated appliance reversion becomes increasingly dangerous. The operational strategy shifts toward repair, retry, host-level recovery, vendor-assisted cleanup, and forward recovery.

    Treat the published sequence as a dependency contract, not a suggestion. Use the planner for the exact environment, enforce the hold points in this runbook, and do not declare success until platform health, infrastructure data paths, representative applications, evidence, and operational ownership have all been validated.

    External References

    Next Post

    MCP vs A2A in 2026: Which Protocol Does Your AI Architecture Actually Need?

    TL;DR MCP and A2A solve different integration problems. Use Model Context Protocol when an AI application needs a standard way to discover and use tools, resources, prompts, APIs, or enterprise…

    Related posts:

    Missile debris injures eight in Qatar after Iran launches barrage | Israel-Iran conflict News

    Wall Street’s AI gains are here — banks plan for fewer people

    SAP outlines new approach to European AI and cloud sovereignty

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleTop 16 Artificial Intelligence Movies And TV Series To Watch In 2026
    Next Article VW Explored A Range-Extending Gas Engine For The ID. Buzz
    gvfx00@gmail.com
    • Website

    Related Posts

    AI Tools

    The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And How to Reclaim Focus

    July 21, 2026
    AI Tools

    Trump imposes 50% US tariffs on some Canadian goods, citing discrimination | International Trade News

    July 21, 2026
    AI Tools

    US public health agencies to test OpenAI and Anthropic AI models

    July 20, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Black Swans in Artificial Intelligence — Dan Rose AI

    October 2, 2025212 Views

    Every Clue That Tony Stark Was Always Doctor Doom

    October 20, 2025134 Views

    We let ChatGPT judge impossible superhero debates — here’s how it ruled

    December 31, 2025100 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram

    Subscribe to Updates

    Get the latest tech news from tastytech.

    About Us
    About Us

    TastyTech.in brings you the latest AI, tech news, cybersecurity tips, and gadget insights all in one place. Stay informed, stay secure, and stay ahead with us!

    Most Popular

    Black Swans in Artificial Intelligence — Dan Rose AI

    October 2, 2025212 Views

    Every Clue That Tony Stark Was Always Doctor Doom

    October 20, 2025134 Views

    We let ChatGPT judge impossible superhero debates — here’s how it ruled

    December 31, 2025100 Views

    Subscribe to Updates

    Get the latest news from tastytech.

    Facebook X (Twitter) Instagram Pinterest
    • Homepage
    • About Us
    • Contact Us
    • Privacy Policy
    © 2026 TastyTech. Designed by TastyTech.

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.