TL;DR
VCF 9.1 changes the economics and architecture of VMware disaster recovery by allowing virtual machines on vSAN, VMFS, and NFS datastores to replicate into a vSAN ESA target. It also supports fan-in designs where multiple source clusters use one centralized recovery site.
The important design point is that a shared recovery site is not one oversized datastore with a single recovery button. It is a collection of site-pair relationships, control-plane dependencies, replication flows, recovery plans, network mappings, capacity reservations, and operational owners that happen to converge on a common vSAN ESA destination.
A strong design sizes storage for the full protected footprint and retention model, sizes compute for the maximum credible concurrent recovery event, preserves failure-domain and control-plane resilience at the target, isolates replication traffic, and keeps protection groups and recovery plans aligned to application ownership. The same architecture can also become a controlled migration bridge from VMFS or NFS to vSAN, but only when migration sequencing, testing, and cutover responsibilities are designed deliberately.
Introduction
Many VMware environments did not arrive at their current storage architecture through one clean design decision.
They accumulated it.
One workload domain runs on vSAN. Another cluster still depends on VMFS backed by an external array. A third environment uses NFS because that was the simplest operational choice at the time. Each platform may have a different replication method, a different recovery tool, a different support team, and a different cost model.
That fragmentation becomes expensive when disaster recovery is designed around the source storage platform. Every array or datastore type can create another protection island, another target appliance, another runbook, and another operational dependency.
VCF 9.1 introduces a more useful target-state pattern. Virtual machines residing on vSAN, VMFS, or NFS can be replicated to a vSAN ESA cluster, and multiple source clusters can converge on a shared recovery site. That creates an opportunity to standardize the recovery destination even when the source estate remains heterogeneous.
The opportunity is real, but so is the architectural risk. A centralized recovery site can reduce duplicated infrastructure and simplify modernization, or it can become a concentrated failure domain that is undersized, operationally ambiguous, and impossible to recover under pressure.
This article focuses on how to design the stronger version of that architecture.
Why a Shared Recovery Site Changes the Design Conversation
The traditional recovery design often mirrors the source environment. SAN-backed clusters replicate to compatible SAN infrastructure. NFS environments use a storage-specific recovery pattern. vSAN environments use a different protection stack. Recovery capacity is purchased and operated in parallel even when the business applications share similar recovery objectives.
A VCF 9.1 fan-in model changes the center of gravity. The source environment can remain mixed, while the target becomes standardized on vSAN ESA.
| Design area | Fragmented recovery model | Shared VCF 9.1 recovery model |
|---|---|---|
| Source storage | Recovery design follows each storage platform | vSAN, VMFS, and NFS sources can converge on vSAN ESA |
| Recovery target | Multiple arrays, appliances, or storage silos | Centralized vSAN ESA recovery environment |
| Orchestration | Separate recovery methods and runbooks | Consistent protection groups and recovery plans by site pair |
| Capacity | Duplicated by platform or source cluster | Shared pool sized for defined concurrent recovery scenarios |
| Operations | Storage-team-centric recovery ownership | Cross-functional platform, network, application, and security ownership |
| Modernization | Recovery and storage migration are separate projects | Replication can become a migration bridge to vSAN |
| Primary risk | Tool and platform fragmentation | Concentrated target-site and control-plane dependency |
The consolidation benefit is not simply fewer storage products. The larger value is that recovery architecture, testing, operational ownership, and migration planning can use a common destination model.
The tradeoff is concentration. The shared site now matters to several source environments at the same time. Its vCenter, Protection and Recovery services, WAN connectivity, vSAN capacity, recovery networks, DNS, identity services, and application dependencies become part of a broader resilience boundary.
The Fan-In Recovery Topology at a Glance
The diagram below shows the key relationship that should shape the design. Each source site retains its own management and protection context. The shared recovery site centralizes the target infrastructure, but it must still preserve separate site-pair relationships.
The most important observation is that the shared site is common infrastructure, not a common administrative object with no internal boundaries.
Each protected site requires its own vCenter Server and Protection and Recovery appliance. The shared recovery site requires its own vCenter Server and a Protection and Recovery deployment that can represent the separate pairings. Broadcom documents multiple Protection and Recovery server instances or add-on appliances for shared protected-site and shared recovery-site topologies.
That means the design should treat each source-to-target relationship as an explicit recovery boundary with its own pairing, permissions, inventory, protection groups, network mappings, test schedule, and operational owner.
Support and Scope Guardrails
Before sizing anything, establish what VCF 9.1 is actually doing in this pattern.
Local Protection and Remote Protection Are Different Capabilities
Local snapshot-based protection remains tied to vSAN. A virtual machine on a VMFS or NFS datastore does not gain local vSAN snapshot protection merely because the environment also contains a Protection and Recovery appliance.
Remote replication is broader. In VCF 9.1, virtual machines residing on vSAN, VMFS, and NFS datastores can be replicated to a vSAN ESA target.
That distinction matters because the operational model is different:
- Local protection addresses fast rollback and operational recovery for workloads already running on vSAN.
- Remote replication moves protected state to a different site and failure domain.
- Recovery orchestration sequences failover, test, planned migration, reprotection, and failback activities.
- Cyber recovery adds isolated validation, security inspection, and clean-room workflow requirements that are not automatically satisfied by ordinary disaster recovery.
A shared recovery site can support several of these outcomes, but the terms should not be collapsed into one generic backup service.
The Target Is vSAN ESA
The shared destination is not an arbitrary datastore. The multi-source replication capability depends on a vSAN ESA target.
That has design implications for hardware compatibility, networking, storage policies, host count, capacity overhead, failure domains, lifecycle management, and operational skill requirements. A recovery site built around an aging array cannot simply be relabeled as the VCF 9.1 target.
Shared Recovery Does Not Mean Shared Failure Assumptions
The architecture supports several source clusters converging on one destination. It does not prove that all protected sites should share one destination.
A shared site is a good fit when the source environments have compatible risk assumptions, recovery ownership, regulatory boundaries, and concurrency expectations. It is a weaker fit when business units require hard administrative separation, when multiple source sites can fail from the same regional event, or when the cyber-recovery boundary must remain isolated from conventional disaster recovery.
Designing the Recovery-Site Control Plane
Storage receives most of the attention in recovery designs because it holds the replicas. The control plane is often the more fragile dependency.
A replicated VM is not operationally useful if the recovery team cannot authenticate, resolve names, access vCenter, run the recovery plan, connect recovery networks, or reach application dependencies.
Recovery vCenter Placement
The recovery vCenter should be treated as a first-class resilience component.
It should not depend on the protected site for the services required to manage recovery. At minimum, validate the target site’s access to:
- DNS and reverse DNS
- NTP
- certificate trust and revocation dependencies
- identity providers or break-glass accounts
- network management
- storage management
- logging and monitoring
- backup and restore procedures for the recovery vCenter itself
The recovery vCenter also becomes a shared administrative boundary. Role design should prevent one source-site team from unintentionally modifying another site’s recovery objects, networks, folders, or plans.
Protection and Recovery Appliance Placement
Each source site needs a local appliance associated with its vCenter. At the recovery site, the Protection and Recovery design must support the required number of site-pair contexts.
Do not assume that one base appliance automatically represents every protected site with no scale or topology implications. Shared-site designs can require additional Protection and Recovery server instances or add-on appliances.
The design record should document:
| Item | Required design decision |
|---|---|
| Source appliance placement | Cluster, datastore, network, backup, and lifecycle owner |
| Recovery appliance placement | Cluster, anti-affinity, management network, and dependency mapping |
| Site-pair ownership | Named team accountable for each protected-to-recovery relationship |
| Add-on instance requirement | Count based on fan-in topology, scale, and current product limits |
| Administrative roles | Least-privilege access for platform, recovery, application, and audit teams |
| Monitoring | Appliance health, replication health, RPO state, certificate state, and service availability |
| Recovery of the recovery platform | Documented restoration path for vCenter and Protection and Recovery components |
The recovery appliance should not be placed on infrastructure that disappears during the same event it is intended to manage.
Sizing the Recovery Cluster Around Credible Failure Scenarios
A shared recovery cluster is rarely sized correctly by adding every source site’s production totals together. That can be unnecessarily expensive. It is also rarely sized correctly by purchasing the smallest cluster that can store the replicas. That can produce a target that holds data but cannot run the recovered services.
The right method separates storage, compute, network, and concurrency.
Start with a Recovery Scenario Matrix
Define the events the target must support.
| Scenario | Workloads active at recovery site | Storage implication | Compute implication | Network implication |
|---|---|---|---|---|
| Single-site disaster | One source site’s recovery plans | All replicas retained | Capacity for one site’s recovery tiers | Failover traffic plus user and dependency access |
| Correlated multi-site event | Selected plans from two or more sites | No major change in stored footprint | Capacity for maximum concurrent recovery set | Highest WAN and north-south demand |
| Planned migration | Migration wave plus steady-state replicas | Temporary overlap and changed-block catch-up | Capacity for migration wave and validation | Controlled cutover and synchronization traffic |
| Recovery-plan testing | Isolated test copies | Additional temporary capacity and snapshot activity | Test bubble capacity | Isolated test networks and admin access |
| Cyber validation | Selected recovery set in isolated environment | Additional retention and staging capacity | Validation compute plus security tooling | Strictly controlled inspection and egress paths |
The business must choose which scenarios are contractual requirements and which are best-effort. Architecture cannot infer that decision from VM counts.
Size Storage for Protection, Retention, and Operations
Target storage should include more than the current used capacity of the protected VMs.
A practical sizing model includes:
Replica footprint = protected used capacity after validated data-reduction assumptions
Retention footprint = additional snapshots or recovery points retained at the target
Change reserve = expected changed blocks waiting to synchronize during interruption or maintenance
Seed staging = temporary space required during manual seeding or migration waves
Recovery overhead = swap, logs, application growth, test copies, and powered-on recovery activity
Policy overhead = vSAN resilience overhead based on the target topology and storage policy
Operational reserve = space preserved for vSAN internal operations, maintenance, rebuild, and host failure
Growth reserve = forecasted protected-capacity growth across the design horizon
Do not size the target from optimistic effective-capacity claims alone. Use measured source consumption, observed change rates, actual retention policies, and the storage policy that the target cluster will enforce.
VCF 9.1 Auto-RAID can simplify policy selection on new vSAN ESA clusters, but it does not remove the need to understand capacity overhead. The resulting resilience model still consumes physical capacity, and the overhead changes with cluster topology.
Size Compute for the Maximum Concurrent Recovery Set
Recovery compute should be based on the largest recovery event the organization has committed to support, not on the total number of protected VMs.
A practical model is:
Required recovery compute = resources for the maximum concurrent recovery set + vSphere HA reserve + management overhead + test or validation concurrency + performance headroom
The maximum concurrent set should be defined by recovery plans and business tiers.
For example, a recovery site protecting three source clusters might not need enough compute to run every VM from all three sites at once. It may need enough capacity to run:
- all Tier 0 and Tier 1 services from one failed site,
- shared infrastructure dependencies,
- selected Tier 2 services,
- the recovery management stack,
- and a controlled test bubble for another site.
That is a business continuity requirement, not a storage calculation.
Preserve Failure-Domain Resilience at the Target
A recovery site is not a place to accept weaker infrastructure standards without documenting the consequence.
Ask what happens when the recovery cluster is already running failed-over workloads and then loses:
- one host,
- one rack,
- one top-of-rack switch,
- one storage device class,
- one power feed,
- or one management service.
If a rack failure drops the cluster below the storage policy’s tolerance or removes too much recovery compute, the shared site has converted multiple source-site risks into one target-site risk.
Where rack-level independence is required, design vSAN fault domains deliberately and confirm that host count, network topology, and capacity remain valid after the modeled failure.
Separating Recovery Compute and Storage Decisions
Even when recovery compute and storage use the same vSAN ESA hosts, they should be modeled as separate demand curves.
Storage is usually steady-state. Replicas, retention, and growth consume capacity every day.
Compute is usually event-driven. Most recovery CPU and memory demand appears during a test, migration, disaster, or cyber-validation event.
This difference creates three common design options.
Fully Hyperconverged Recovery Cluster
The same hosts provide vSAN ESA capacity and recovery compute.
Best fit: Moderate environments where storage and compute growth remain reasonably aligned.
Strengths: Simpler lifecycle, fewer cluster types, direct operational ownership, and straightforward recovery placement.
Tradeoffs: Adding storage may require buying unnecessary CPU and memory. Adding compute may require buying unnecessary storage devices. Maintenance consumes both compute and storage fault-domain capacity.
Storage-Heavy Recovery Cluster with Reserved Compute
The cluster is optimized for replica capacity but still contains enough CPU and memory for defined recovery tiers.
Best fit: Environments with a large protected footprint and a lower expected concurrent power-on set.
Strengths: Better economics for data-heavy protection designs.
Tradeoffs: A correlated event can exceed compute even when storage is healthy. Recovery-plan priorities must be enforced operationally.
Disaggregated or Shared vSAN ESA Pattern
Storage and recovery compute are separated where the supported VCF and vSAN design permits it.
Best fit: Larger environments where storage and compute scale at different rates.
Strengths: Independent scaling and potentially better resource utilization.
Tradeoffs: More network and lifecycle complexity, additional compatibility validation, and a greater need for precise ownership across compute and storage clusters.
The design should not select one of these patterns based on hardware preference alone. It should select the pattern that matches the protected-capacity curve, recovery-concurrency curve, failure-domain requirements, and operating model.
Planning Bandwidth, Latency, and Replication Isolation
Replication bandwidth is driven primarily by changed data within the recovery point objective, not by the provisioned size of the VM.
A 10 TB VM with a low change rate may be easier to protect than a 2 TB database that rewrites hundreds of gigabytes every hour.
Use Change Rate as the Primary Input
A simple payload estimate is:
Required payload Mbps = changed GiB per RPO interval x 8192 / RPO interval in seconds
If a workload set changes 300 GiB during a four-hour RPO window, the payload requirement is approximately 171 Mbps before protocol overhead, retransmissions, contention, and catch-up margin.
The production design should then add headroom for:
- peak rather than average change rate,
- multiple source sites replicating concurrently,
- retransmissions and packet loss,
- maintenance or WAN outages that create a backlog,
- initial synchronization,
- manual-seed catch-up,
- recovery tests,
- and other traffic sharing the link.
A design based only on average daily change rate can meet the monthly bandwidth budget and still miss the RPO every afternoon.
Latency Is an Operational Variable, Not a Checkbox
There is no useful universal statement that a link is acceptable simply because it is reachable.
Round-trip latency, packet loss, jitter, congestion, and traffic shaping affect synchronization behavior and catch-up time. The design should validate the actual WAN path under representative load, then measure whether protected workloads remain within RPO during peak periods.
The important metric is not theoretical link speed. It is sustained replication throughput and observed RPO compliance under the worst expected operating condition.
Isolate Replication Traffic
VCF Protection and Recovery documentation includes separate VMkernel configuration guidance for source and target replication traffic. Use that capability to prevent replication from competing unpredictably with management, vMotion, storage, or tenant traffic.
The target site should have distinct network paths for:
Replication, recovery, and testing are different traffic classes. Treating them as one recovery VLAN makes troubleshooting and capacity control harder.
Building Protection Groups Around Applications, Not Datastores
The source datastore type determines how the data is protected, but it should not determine the entire recovery organization model.
Protection groups should align to things the business can reason about:
- application or service boundary,
- recovery tier,
- RPO and RTO class,
- dependency order,
- data ownership,
- regulatory boundary,
- recovery network mapping,
- and operational approver.
A datastore-based protection group can be convenient when one datastore cleanly maps to one service. In many environments, it does not.
A better naming model might look like:
PG-SITEA-ERP-TIER0PG-SITEA-IDENTITY-TIER0PG-SITEB-FILE-TIER1PG-SITEC-APPDEV-TIER2
VCF 9.1 also supports vSphere tags for protection-group membership. That can improve scale, but only when tag governance is strong. Dynamic membership is useful when new workloads inherit the correct protection automatically. It is dangerous when application teams can assign recovery tags without validation, or when stale tags silently keep decommissioned workloads in scope.
Treat protection tags as policy objects with ownership, review, and change control.
Using Recovery Plans as the Execution Boundary
Protection groups define what is protected. Recovery plans define how the organization intends to recover it.
A useful recovery plan should encode:
- startup groups,
- explicit dependencies,
- recovery priorities,
- network mappings,
- IP customization where required,
- pre-power-on and post-power-on actions,
- application validation checkpoints,
- operator approvals,
- and stop conditions.
Do not create one enormous recovery plan for every workload from a source site. Large plans are harder to test, harder to troubleshoot, and harder to govern.
Break plans at application or service boundaries that can be tested independently, then coordinate them through an incident runbook.
In a fan-in design, keep the site-pair boundary visible. A multi-site incident may require several recovery plans across several pair contexts. The orchestration platform can execute each plan, but the enterprise still needs an incident-level runbook that decides which sites, applications, and recovery waves take precedence when the shared compute pool is constrained.
Manual Replica Seeding for Constrained WAN Links
Initial synchronization can overwhelm a WAN long before steady-state replication becomes a problem.
VCF 9.1 introduces manual replica seeding so that a large initial dataset can be transferred to the recovery site through another method, including physical transport, then synchronized differentially.
This is useful when:
- the initial protected footprint is measured in tens or hundreds of terabytes,
- the WAN can sustain changed-block replication but not the first full copy,
- migration deadlines are shorter than the full-sync duration,
- or the same link must continue supporting production traffic.
A disciplined seeding workflow should include:
- Freeze the source inventory and record the VM, disk, datastore, and intended target path.
- Create and secure the seed copy using an approved transfer method.
- Protect the media or transfer path with encryption and chain-of-custody controls.
- Stage the seed data in a uniquely named target location.
- Configure replication to use the intended seed.
- Allow differential synchronization to catch up.
- Validate replication health, recovery-point creation, and test recovery before declaring the workload protected.
Manual seeding does not fix an undersized steady-state WAN. It only reduces the cost of the initial copy.
VCF 9.1 release notes also call out naming considerations in fan-in environments. When different source clusters contain VMs with identical names, use unique names during replication setup or specify unique seed paths. A shared recovery site should adopt an enterprise naming standard before the first migration wave, not after collisions appear.
Using the Shared Site as a Migration Bridge to vSAN
The architecture becomes more valuable when recovery and modernization are planned together.
A VM on VMFS or NFS can be replicated into the vSAN ESA recovery site, tested there, and later moved through a planned migration workflow. That creates a controlled bridge from legacy storage to vSAN without requiring the recovery target to duplicate the original array platform.
The migration pattern is:
This pattern can reduce duplicate migration tooling, but it is not an in-place datastore conversion. The workload is changing site, recovery context, network attachment, and operational ownership.
The migration plan must address:
- source shutdown authority,
- final synchronization,
- DNS and load-balancer changes,
- firewall policy,
- identity and certificate dependencies,
- application quiescence,
- validation ownership,
- rollback criteria,
- and the future protection direction after cutover.
The strongest use case is a planned modernization program where the recovery site is intended to become a production landing zone. The weakest use case is treating the recovery cluster as permanent production without revisiting its compute, support, network, and operating model.
A Phased Implementation Strategy
A fan-in architecture should be built one site pair at a time. The shared target may be common, but the risks should be introduced incrementally.
Establish the Support and Version Baseline
Before deployment:
- confirm the exact VCF, vCenter, ESX, Protection and Recovery, and vSAN versions,
- review the current Protection and Recovery patch release,
- validate interoperability and supported source versions,
- confirm the required appliance and add-on instance count,
- document current known issues,
- and resolve naming collisions in the protected inventory.
Exit criterion: The architecture has a signed compatibility baseline and no unresolved support assumption.
Build the Recovery-Site Foundation
Deploy and validate:
- recovery vCenter,
- vSAN ESA target cluster,
- failure domains and storage policy,
- management services,
- identity and break-glass access,
- DNS and NTP,
- replication networks,
- recovery networks,
- test isolation networks,
- logging and monitoring,
- and Protection and Recovery components.
Exit criterion: The target site can be administered independently of any protected site.
Pilot One Source Site
Choose a limited application set with representative characteristics:
- one low-change VM,
- one high-change VM,
- one multi-VM application,
- one workload requiring network remapping,
- and one workload from the relevant source datastore type.
Test initial sync, steady-state RPO, isolated recovery, cleanup, planned migration, reprotection, and failback where applicable.
Exit criterion: The site pair completes a documented recovery test with measured RTO and RPO evidence.
Add Source Sites Incrementally
For each additional site:
- deploy and pair the required appliance context,
- validate permissions,
- confirm unique names and paths,
- add WAN capacity forecasts,
- onboard protection groups,
- create recovery plans,
- and run an isolated test before accepting production scope.
Exit criterion: Each pair is independently supportable and does not reduce existing pairs below their service objectives.
Test Concurrent Recovery Conditions
The architecture is not proven until the shared target is tested under the concurrency it was sized to support.
Run exercises that combine:
- multiple active replication streams,
- a recovery-plan test,
- a maintenance event,
- a simulated WAN interruption and catch-up,
- and the loss of a target host or failure domain.
Exit criterion: The shared site continues meeting defined recovery objectives under the maximum credible concurrent condition.
Introduce Migration Waves
Only after disaster-recovery operations are stable should the same platform be used as a migration bridge for VMFS or NFS workloads.
Exit criterion: Each migration wave has application acceptance, rollback criteria, future protection configuration, and legacy-storage decommission approval.
Automation should improve consistency without hiding recovery risk.
Build a Protection Inventory
The inventory should capture at least:
| Field | Why it matters |
|---|---|
| Source site and vCenter | Defines the pairing and ownership boundary |
| Cluster and datastore type | Confirms source support and placement |
| VM name and unique identifier | Prevents fan-in naming collisions |
| Used capacity and growth | Drives target storage sizing |
| Observed write rate | Drives WAN and RPO sizing |
| Business service and owner | Defines recovery accountability |
| Recovery tier | Controls sequencing and compute priority |
| Network dependencies | Drives mappings, firewalling, and validation |
| Protection-group tag | Enables governed dynamic membership |
| Recovery plan | Defines execution and testing scope |
PowerCLI can collect much of the vCenter inventory, tag state, datastore placement, and performance history. The point is not to automate every recovery decision. The point is to make the protected estate measurable and reviewable.
Use Tags Carefully
Tag-driven protection can reduce manual drift, but it should include:
- a controlled tag namespace,
- named tag owners,
- automated checks for unprotected production VMs,
- automated checks for stale or conflicting tags,
- change logging,
- and periodic reconciliation between application inventory and protection groups.
Monitor the Recovery Service, Not Just the Appliances
A green appliance status does not prove recoverability.
Operational dashboards should include:
- replication health,
- current and historical RPO violations,
- backlog after WAN interruptions,
- target capacity and reserve consumption,
- Protection and Recovery service health,
- certificate expiration,
- recovery-plan test age,
- failed cleanup operations,
- and the percentage of protected workloads with a current successful test.
Keep Human Approval at the Recovery Boundary
Automate inventory, validation, drift detection, report generation, and test preparation.
Do not make production failover a silent infrastructure automation action. Recovery plans can execute the technical sequence, but incident command should retain authority over business priority, communications, source shutdown, data-consistency decisions, and failback timing.
Operational Ownership for a Shared Recovery Site
A centralized platform requires explicit decision rights.
| Capability | Accountable owner | Responsible teams |
|---|---|---|
| Recovery-site architecture | VCF or infrastructure architecture | Platform, storage, network, security |
| vSAN ESA capacity and health | Storage or VCF platform owner | vSAN operations |
| Recovery vCenter and appliances | VCF platform owner | Virtualization operations |
| WAN and replication networks | Network owner | Network operations |
| Protection groups and tags | Application service owner | Platform automation and application teams |
| Recovery plans | Business service owner | DR team, platform team, application team |
| Recovery testing | Business continuity owner | All technical owners |
| Cyber validation | Security owner | Incident response, EDR, platform, application teams |
| Migration cutover | Program owner | Application, platform, network, storage, change management |
| Capacity funding | Infrastructure sponsor | Finance, platform, business continuity |
The shared site fails organizationally when everyone owns the infrastructure and no one owns the recovery decision.
Risks, Caveats, and Common Design Failures
Treating the Target as Cheap Capacity
A recovery site that stores replicas but lacks compute, network, dependencies, or operational access is not a recovery site. It is a remote copy.
Ignoring the Shared vCenter Blast Radius
The target vCenter is a common control-plane dependency. Its failure can affect several protected sites at once. Protect, back up, monitor, and test it accordingly.
Underestimating Add-On Appliance Requirements
Fan-in and fan-out topologies can require multiple Protection and Recovery server instances or add-on appliances. Validate the current product limits and deployment model before finalizing IP addresses, DNS records, certificates, and management capacity.
Allowing Duplicate VM Names
Duplicate names from different source sites can create replication and seeding complications. Enforce site-aware naming or unique target paths before onboarding.
Sizing WAN from Provisioned Capacity
Provisioned terabytes do not determine steady-state bandwidth. Changed blocks, RPO, peak periods, catch-up behavior, and concurrent sources do.
Assuming Seeding Solves Bandwidth
Manual seeding accelerates the initial copy. It does not reduce daily change rate or eliminate the need for steady-state RPO capacity.
Creating Oversized Protection Groups and Recovery Plans
Large groups create large fault domains. Smaller application-aligned units are easier to test, sequence, troubleshoot, and prioritize.
Testing Each Site Independently but Never Together
A shared recovery architecture must be tested as shared infrastructure. Independent success does not prove concurrent capacity, control-plane scale, network throughput, or operational coordination.
Mixing Disaster Recovery and Cyber Recovery Without Isolation
A conventional DR site may share identity, management, and network trust with production. Cyber recovery can require stronger isolation, separate identity, controlled egress, security inspection, and clean-room procedures. Do not assume that a fan-in DR site is automatically a cyber vault.
Using the Recovery Site as Production Without Redesign
A migration bridge can become a production landing zone, but the target must then be reassessed for production performance, support, observability, backup, lifecycle, security, and availability. Recovery sizing is not automatically production sizing.
When Not to Use the Shared Fan-In Pattern
A shared recovery site is not the default answer for every environment.
Use separate recovery domains when:
- source sites share a correlated regional risk that can exceed target capacity,
- regulatory or contractual requirements prohibit shared administrative boundaries,
- business units require independent change control and recovery authority,
- the target vCenter blast radius is unacceptable,
- cyber-recovery isolation cannot coexist with ordinary DR operations,
- WAN paths cannot sustain the required RPOs,
- or the recovery site cannot be funded and operated as critical infrastructure.
Centralization is valuable only when it reduces complexity without hiding risk.
Practical Design Checklist
Before approving the architecture, confirm that the design can answer each question directly.
- Which source vCenters and datastore types are in scope?
- Which target vSAN ESA cluster and storage policy will receive the replicas?
- How many Protection and Recovery server instances or add-on appliances are required?
- What is the maximum concurrent recovery scenario?
- How much compute is reserved for each recovery tier?
- How much raw and effective storage is required after resilience, retention, reserve, and growth?
- Which rack, host, network, and control-plane failures must the target tolerate?
- What are the measured change rates and peak RPO bandwidth requirements?
- How are replication, management, recovery, and test networks separated?
- How are duplicate VM names and seed paths controlled?
- Which protection groups map to which applications and owners?
- Which recovery plans have current successful test evidence?
- Who can authorize planned migration, disaster recovery, reprotection, and failback?
- How will multiple source sites be prioritized during a shared event?
- Which workloads will use the design as a migration bridge to vSAN?
- What conditions require a separate recovery site instead?
A design that cannot answer these questions is not ready for production, regardless of whether replication can be configured successfully.
Conclusion
VCF 9.1 makes a centralized vSAN ESA recovery site a realistic target for heterogeneous VMware environments. vSAN, VMFS, and NFS workloads no longer have to preserve their original storage platform at the recovery site simply to achieve remote protection.
That is the technical capability. The architecture still has to make it operationally safe.
The shared site should be designed as a collection of explicit site-pair relationships supported by a resilient recovery vCenter, the required Protection and Recovery server instances, isolated replication paths, application-aligned protection groups, tested recovery plans, and capacity models based on credible concurrent events.
Storage should be sized for replicas, retention, policy overhead, operations, and growth. Compute should be sized for the maximum recovery set the business has committed to run. Network planning should be based on measured change rate and observed RPO performance. Recovery ownership should be assigned before an incident, not discovered during one.
When those disciplines are in place, the shared recovery site becomes more than a cost-consolidation exercise. It becomes a common resilience platform and a practical migration bridge from VMFS or NFS to vSAN ESA. Without them, it becomes a concentrated dependency that several environments may discover they share only after the failure begins.
