VMware vSAN Architecture in Production
vSAN architecture is ultimately about where data lives when something goes wrong. Capacity and benchmark numbers matter, but production design has to answer harder questions: which failures the cluster can tolerate, how quickly objects can rebuild, how maintenance changes risk, whether remote datastores create new dependencies, and how recovery works when the problem is larger than one disk or one host.
The current VCF 9.1 platform broadens those choices. vSAN ESA continues to be the preferred modern architecture for new performance-oriented deployments, while 9.1 adds more flexibility for mixed ESA and OSA environments, remote datastore use, recovery topologies, and storage efficiency. Those features are useful only when they serve an intentional target architecture. The design should make migration and resilience simpler, not turn temporary interoperability into permanent complexity.
Raw usable terabytes do not describe the service a vSAN cluster can provide. Architects must define failures to tolerate, maintenance expectations, rebuild capacity, and the performance required while degraded. A policy that protects against a host failure still needs enough surviving resources to serve I/O and rebuild data. Capacity planning should therefore reserve both space and performance headroom for recovery, not just steady-state growth.
Model simultaneous operational conditions. A host may be in maintenance when a device fails. A rebuild may coincide with backup or replication. A site or rack failure may remove several components that were assumed to be independent. Fault domains should reflect the physical infrastructure—racks, power feeds, or other shared dependencies—so policy intent corresponds to real failure isolation.
vSAN Express Storage Architecture separates itself from the older OSA model through a storage stack designed for modern NVMe devices and different data-placement behavior. Current guidance and performance work are increasingly centered on ESA. That makes ESA the natural target for many new deployments, but existing estates may continue to operate OSA clusters for years due to hardware life cycles and migration risk.
VCF 9.1 makes mixed-mode environments more flexible, including remote datastore scenarios across ESA and OSA. Treat that flexibility as a transition tool or a conscious tiering strategy. Record which architecture is preferred for new capacity, which workloads may remain on OSA, and what triggers migration. Without that target state, teams can accumulate storage paths and operational exceptions that become harder to reason about during incidents.
A vSAN storage cluster can provide centralized shared storage to compute-focused clusters, changing the traditional assumption that every HCI cluster must contribute both compute and local storage. This can improve hardware utilization and allow older compute clusters to consume newer storage. It also creates a new dependency: the consuming cluster now depends on the availability, network path, capacity, and operational schedule of the storage cluster.
Design the relationship as a service boundary. Define performance expectations, failure behavior, maintenance coordination, and ownership. Verify what happens when the storage cluster is degraded or disconnected and how many consuming clusters are affected. Centralization can reduce duplicated capacity, but it can also concentrate blast radius. The architecture should justify that trade with measurable efficiency or operational benefit.
Storage policies express availability and placement intent. The correct policy depends on the tolerated failure, cluster topology, performance profile, and capacity economics. vSAN ESA has reduced some of the old performance trade-offs between mirroring and erasure coding, and VCF 9.1 introduces more system-managed behavior such as Auto-RAID. That can simplify policy management, but architects still need to understand the service objective behind the policy.
Do not choose a policy simply because it is the default or because it offers the highest nominal resilience. A policy that consumes excessive capacity may reduce the headroom needed for rebuilds and growth. A policy optimized for efficiency may be inappropriate for a latency-sensitive workload or a small failure domain. Validate the result under failure, not only when the cluster is healthy.
Disks and hosts will fail, and components will enter maintenance. The architecture should make repair behavior predictable. Estimate how long a rebuild can take at realistic utilization, how much network bandwidth it consumes, and how application latency changes while repair traffic is active. Ensure there is enough free space in the correct fault domains to complete recovery.
Operational monitoring should expose object health, resynchronization backlog, capacity pressure, device health, and performance impact. A cluster that is technically compliant but spends long periods in a vulnerable degraded state does not meet a strong resilience objective. Set escalation thresholds for prolonged repair, repeated device faults, or insufficient rebuild headroom so teams act before the next failure compounds the problem.
VCF 9.1 expands vSAN efficiency with updated compression and global deduplication capabilities. These can improve usable capacity and reduce storage cost, but the benefit depends on data characteristics and the current architecture. Compression may deliver meaningful savings for some datasets and very little for others. Deduplication effectiveness varies with duplication patterns, application behavior, and the amount of common data.
Measure actual reduction and performance impact rather than relying on a generic ratio. Include CPU overhead, write behavior, recovery operations, and capacity forecasting. Efficiency should not be used to postpone necessary hardware growth if the cluster is already operating too close to capacity limits. Space savings are valuable only when the system retains enough operational reserve to repair and maintain itself safely.
vSAN depends on network performance and consistency. Latency, packet loss, congestion, MTU mismatch, and asymmetric paths can appear as storage symptoms because the data path crosses hosts. Remote datastore designs make this even more important because storage traffic may traverse more of the network. Define bandwidth and redundancy for normal traffic, rebuilds, remote access, and maintenance concurrently.
Troubleshooting should correlate network and storage evidence. If one host shows elevated latency, compare link errors, congestion, path behavior, and resynchronization activity before replacing storage hardware. The storage team and network team should share a common diagram of the data path and failure domains. The architecture is incomplete if each side only understands its own component.
Local redundancy protects against component failures; it does not replace backup, replication, or disaster recovery. VCF 9.1 expands vSAN Protection and Recovery with more flexible replication and centralized recovery patterns. Architects should define what local policy protects, what snapshots or backup protect, what remote replication protects, and which business event each mechanism is intended to address.
Recovery design should specify copy independence, retention, recovery sequence, and the site or cluster capacity needed to run restored workloads. A centralized recovery vSAN cluster can simplify operations, but only if network, authentication, catalog, and compute dependencies are also recoverable. Test restores and failovers with representative workloads; a protection policy that has never been exercised is only an assumption.
Firmware changes, host upgrades, device replacements, and cluster expansion expose whether the storage architecture has sufficient headroom and clear ownership. Maintenance mode should be part of capacity modeling. If routine maintenance pushes a cluster into risk, the design is too tight. Upgrade plans should account for version compatibility across mixed ESA/OSA or remote-datastore relationships.
Track exceptions created during change. Temporary policy changes, delayed repairs, forced object placement, or capacity workarounds should have owners and expiration dates. Otherwise, the post-maintenance environment may differ materially from the architecture document. Successful lifecycle management is not simply finishing the upgrade; it is returning the cluster to a known, supportable state with resilience verified.
The strongest vSAN architecture makes data placement, failure tolerance, capacity reserve, recovery, and maintenance easy to explain. ESA, storage clusters, mixed-mode support, Auto-RAID, compression, deduplication, and protection features provide more design options in VCF 9.1, but options are useful only when they map to service objectives and operational evidence.
Review the architecture as the estate changes. Hardware generations, workload mix, recovery targets, and capacity economics will move over time. Preserve a clear target state, measure whether current policies still satisfy it, and test degraded operation regularly. Production storage earns trust when failure and recovery behave as deliberately as normal reads and writes.
Capacity expansion deserves the same rigor as initial sizing. Adding hosts or storage can change failure-domain balance, rebuild behavior, and the amount of data moved while the cluster rebalances. Plan expansion before the cluster reaches emergency thresholds so new hardware can be introduced and validated without competing with urgent recovery activity. Growth should restore operating margin, not merely postpone the next capacity alarm.
Finally, preserve a recovery evidence pack: policy assignments, object-health state, recent resynchronization history, hardware alerts, network counters, and the last successful restore or replication test. During a complex storage incident, this evidence shortens the time spent guessing which layer failed. It also lets post-incident reviews distinguish a design weakness from a component fault and turn the result into a concrete architecture change.
Hardware compatibility and firmware discipline also belong in storage architecture. Certified devices and supported combinations reduce risk, but operations must still track firmware, controller behavior, drive health, and upgrade sequencing. Storage reliability depends on the full hardware-software stack, so changes should be validated against both support guidance and observed cluster health before broad rollout. Keep those assumptions versioned with the cluster design.
