TL;DR
An NVIDIA NVL72 AI factory is not a collection of GPU servers connected to a fast Ethernet network. Each rack is a tightly coupled, 72-GPU NVLink scale-up domain with rack-level firmware, power, cooling, management, and service dependencies. Additional racks are joined through a separate scale-out fabric, commonly a dual-plane, rail-optimized Spectrum-X design for current GB300 reference architectures, or a validated InfiniBand design in applicable GB200 deployments.
The practical design rule is simple: treat the rack as the atomic compute system, treat the fabric as a separate distributed system, and validate both through an evidence chain that starts with facilities and firmware and ends with NCCL, storage, inference, resilience, and operational handover. NVIDIA certification can narrow procurement risk, but it does not replace site-specific architecture review or cluster acceptance testing.
Introduction
The hardest part of an NVIDIA GB200 or GB300 NVL72 deployment is not selecting a GPU. It is making a rack-scale computer behave like a dependable enterprise platform inside a real data center.
That distinction matters because an NVL72 rack compresses several normally separate design problems into one operational unit. Compute trays, NVLink switch trays, DPUs, network adapters, rack power shelves, liquid cooling, firmware, Fabric Manager, telemetry, orchestration, storage, and the scale-out network must all agree. A weakness in any one layer can prevent the cluster from reaching the performance, availability, or maintainability expected from the hardware investment.
A successful design therefore begins below Kubernetes and below the leaf-spine fabric. It begins with ownership, facility readiness, supported component recipes, and a clear understanding of where the scale-up domain ends and the scale-out fabric begins.
This deep dive provides a practical architecture and validation framework for teams designing GB200 NVL72 or GB300 NVL72 AI factories. It does not prescribe one universal bill of materials. OEM implementations, validated designs, software releases, and site constraints still govern the final solution.
Design the AI Factory as a System, Not a Rack Purchase
The most important architecture decision is the unit of design.
Traditional server clusters often treat the node as the smallest independently deployable and serviceable component. An NVL72 system changes that model. The rack contains 18 compute trays and nine NVLink switch trays that form one 72-GPU NVLink domain. The rack also depends on coordinated power, cooling, management, and software services. A compute tray is replaceable, but it is not an isolated four-GPU island from the perspective of the complete platform.
The design should use three nested units:
- The tray is a hardware service unit.
- The rack is the scale-up compute and NVLink failure domain.
- The cluster is the scale-out scheduling, storage, and network domain.
This framing prevents a common design mistake: treating all high-speed links as one undifferentiated fabric.
| Design dimension | Scale-up domain | Scale-out domain |
|---|---|---|
| Primary boundary | Inside one NVL72 rack | Between NVL72 racks and shared services |
| Main technology | NVLink and NVSwitch | Spectrum-X Ethernet or a validated InfiniBand fabric |
| Typical unit | 72 GPUs | Multiple 72-GPU racks |
| Primary purpose | Fine-grained GPU-to-GPU communication | Distributed training, inference, storage, and cluster communication |
| Failure concern | Rack-wide NVLink health and topology | Rail, plane, leaf, spine, optics, congestion, and routing behavior |
| Core validation | NVLink topology, Fabric Manager, DCGM, intra-rack collectives | Per-rail RDMA, NCCL across racks, storage, congestion, and resilience |
| Change model | Rack-wide supported software recipe | Fabric-wide topology and capacity plan |
The architecture is not complete until each unit has an owner, an acceptance gate, and a maintenance procedure.
Establish Rack-Scale Ownership Boundaries
NVL72 projects often fail organizationally before they fail technically. The infrastructure crosses facilities, networking, platform engineering, storage, security, AI operations, and vendor support. When responsibility is vague, design gaps surface during commissioning, when changes are expensive and schedule pressure is high.
A practical ownership model should assign accountability before hardware arrives.
| Capability | Accountable owner | Required responsibilities |
|---|---|---|
| Facility power | Data-center facilities | Utility capacity, distribution, protection, grounding, busbar integration, emergency procedures |
| Liquid cooling | Facilities and cooling integrator | CDU design, secondary loop, water quality, temperatures, flow, leak detection, isolation, service procedures |
| Rack integration | OEM or system integrator | Exact BOM, rack assembly, cabling, firmware baseline, factory evidence, shipping and installation controls |
| Scale-out compute fabric | Network engineering | Plane design, rail mapping, addressing, optics, routing, congestion control, telemetry, fault isolation |
| Management network | Platform and network operations | Out-of-band access, management addressing, DNS, NTP, identity, jump hosts, logging |
| Storage fabric and platform | Storage engineering | Data paths, metadata paths, RDMA support, capacity, performance, protection, backup, GDS validation |
| Rack software | AI infrastructure team | Supported release recipe, Fabric Manager dependencies, driver and CUDA alignment, health monitoring |
| Kubernetes or Slurm | Platform engineering | Cluster lifecycle, scheduling, operators, admission control, quotas, tenancy, upgrades |
| Workload acceptance | AI platform and model teams | NCCL baseline, training or inference tests, workload SLOs, reproducibility |
| Security and governance | Security engineering | Segmentation, privileged access, supply-chain evidence, logging, vulnerability and exception processes |
| Service lifecycle | Joint operations team | Drain, maintenance, spares, rollback, vendor escalation, revalidation after changes |
The ownership boundary between the OEM and the customer deserves special attention. A rack can arrive factory-integrated and still require customer-owned design decisions for fabric topology, IP addressing, storage, control-plane placement, tenant isolation, operational tooling, and workload validation.
Understand the NVL72 Scale-Up Domain
An NVL72 rack contains 18 compute trays. Each tray contributes four Blackwell-family GPUs and two Grace CPUs. Nine NVLink switch trays provide the rack-scale switching layer that joins all 72 GPUs into one NVLink domain.
That topology is the defining property of the system. It allows workloads to use the rack as a tightly coupled accelerator complex, but it also creates dependencies that do not exist in loosely coupled PCIe GPU servers. Fabric Manager, NVSwitch health, rack-level firmware compatibility, power sequencing, and topology discovery become first-class platform requirements.
The following diagram shows the architecture boundary that should drive the rest of the design.
The reader should notice two separate communication systems. NVLink and NVSwitch create the scale-up domain inside the rack. Spectrum-X or InfiniBand creates the scale-out domain between racks. They solve different problems and have different owners, telemetry, failure modes, and acceptance tests.
GB200 and GB300 Are Related, Not Interchangeable
The design method is common, but the component details are not identical.
| Area | GB200 NVL72 | GB300 NVL72 | Design implication |
|---|---|---|---|
| GPU generation | Blackwell | Blackwell Ultra | Workload, memory, power, firmware, and software baselines differ |
| Compute trays | 18, four GPUs per tray | 18, four GPUs per tray | Both form a 72-GPU rack domain |
| NVLink switch trays | Nine | Nine | Both require rack-scale NVLink management |
| Scale-out adapters | Current NVIDIA rack guidance commonly shows ConnectX-7 and BlueField-3 | Current reference architecture shows ConnectX-8 and BlueField-3 | Cabling, optics, port speeds, rail mapping, and software matrices differ |
| Rack power envelope | Approximately 120 kW in current NVIDIA rack guidance | Up to approximately 142 kW in the current NVIDIA reference architecture | Facility design must use the exact OEM and release-specific values |
| Common scale-out pattern | Validated InfiniBand or Ethernet design, depending on platform and solution | Spectrum-X dual-plane reference designs are prominent in current guidance | Do not copy a GB300 network BOM into a GB200 design without validation |
Power figures are planning anchors, not construction values. The final facility design must use the selected OEM rack, exact power shelves, configured components, redundancy mode, regional electrical requirements, and measured commissioning data.
Separate Scale-Up from Scale-Out Networking
Scale-up and scale-out are often discussed as if they are only bandwidth tiers. The more useful distinction is the communication behavior they support.
NVLink is optimized for tightly coupled GPU communication inside the rack. It allows the 72 GPUs to participate in a common high-bandwidth domain through NVSwitch. The scale-out network connects separate NVLink domains and external services. It carries inter-rack collectives, checkpoint traffic, storage access, orchestration traffic, telemetry, and service communication according to the detailed design.
This separation changes several architecture decisions:
- A healthy Ethernet or InfiniBand fabric does not prove that the NVLink domain is healthy.
- A healthy NVLink domain does not prove that multi-rack NCCL will perform correctly.
- A per-node network benchmark cannot prove that rail mapping is correct across the cluster.
- A successful synthetic collective does not prove that storage or inference SLOs are met.
- Maintenance inside one rack can affect a 72-GPU scheduling unit even when the scale-out network remains available.
The acceptance plan must test each boundary independently and then test the complete path as one system.
Build a Dual-Plane, Rail-Optimized Spectrum-X Fabric
Current NVIDIA GB300 reference material uses an RDMA-based, full nonblocking leaf-spine compute fabric with rail-optimized GPU connectivity. The reference design separates the scale-out network into two independent planes.
A rail is the repeated network path associated with a consistent GPU position across compute trays. Instead of connecting ports arbitrarily to any leaf, the design maps equivalent GPU-facing ports to the same logical rail. This gives collective communication software a predictable topology and reduces the chance that traffic concentration or asymmetric paths undermine performance.
A dual-plane design gives each GPU-facing path access to two independently routed fabrics. It is not simply two VLANs on one physical network. The planes should have independent links and switching paths, with the degree of physical separation defined by the resilience objective.
Fabric Design Criteria
The detailed design should explicitly document:
- Nonblocking target: Define whether the compute fabric is truly 1:1 at the intended cluster scale, including any super-spine tier.
- Plane independence: Identify shared power, shared racks, shared optics paths, and shared management dependencies that could defeat the apparent redundancy.
- Rail mapping: Maintain a port-to-GPU-to-leaf map that survives installation, replacement, and expansion.
- RDMA configuration: Standardize RoCE, priority flow control where required, ECN, queue design, MTU, DSCP, and congestion-control behavior according to the validated recipe.
- Adaptive routing and load balancing: Confirm the selected switch and NIC software features are supported together and measure their behavior under collective workloads.
- Telemetry: Collect interface counters, congestion signals, buffer events, FEC, optical health, RDMA counters, retransmissions, and flow-level evidence.
- Cabling discipline: Use a machine-readable cable plan, unique labels, optical budgets, inspection procedures, and post-install link validation.
- Expansion headroom: Reserve leaf, spine, super-spine, optics, IP, and power capacity for the next planned rack unit.
The network team should be able to prove the mapping from a logical NCCL rail to a physical GPU port, cable, switch port, plane, and spine path. A diagram without this operational mapping is not enough.
Keep Compute, Storage, Customer, and Management Fabrics Distinct
A practical NVL72 design usually contains several networks even when some functions share physical switches. The architecture should describe the functions separately before deciding which can be converged.
| Fabric or network | Primary traffic | Design priority | Common failure concern |
|---|---|---|---|
| Scale-out compute | Inter-rack NCCL, distributed training and inference | Predictable latency, RDMA, nonblocking behavior, rail symmetry | Congestion, bad rail mapping, plane failure, optical degradation |
| North-south or converged data | Storage, customer services, checkpointing, data ingest | Throughput, service reachability, segmentation, storage compatibility | Shared congestion, security policy, storage hotspots |
| Storage backend | Parallel file system, object, metadata, checkpoint traffic | Data durability, metadata scaling, GDS or RDMA path validation | Metadata bottleneck, data-path fallback, recovery time |
| Out-of-band management | BMCs, switches, power, cooling, service access | Isolation, dependable reachability, privileged access | Management lockout, DNS or NTP failure, credential sprawl |
| Cluster control plane | Kubernetes or Slurm control, registries, APIs, automation | Availability, consistency, security, backup | Control-plane quorum, certificate expiry, dependency outage |
| Telemetry and logging | Metrics, traces, logs, events, fabric telemetry | Loss resistance, retention, time synchronization | Blind spots, clock skew, overload during incidents |
Convergence can be valid, but it must be a decision rather than an assumption. Sharing a physical fabric between compute, storage, and customer traffic changes congestion control, security ownership, blast radius, and acceptance testing. The design should state what is shared, what is isolated, and what happens when one traffic class becomes abusive.
Treat Power, Cooling, and Rack Weight as Architecture Inputs
A high-density NVL72 rack is a facility project and a computing project at the same time. The cluster design is not ready when the logical topology is complete. It is ready when the data center can deliver and sustain the required electrical, thermal, structural, and service conditions.
Facility Readiness Checklist
| Area | Design questions | Required evidence |
|---|---|---|
| Electrical capacity | Can the site supply the rack under normal, redundant, and maintenance states? | One-line diagrams, load study, breaker and busway schedule, commissioning results |
| Rack power shelves | Does the selected rack use the required feeds and redundancy mode? | OEM power design, feed mapping, phase balance, failover test |
| Cooling capacity | Can the secondary loop remove the rack heat at the required temperatures and flow? | CDU sizing, facility-water design, flow and temperature calculations, measured test |
| Residual air cooling | Which components still depend on room airflow? | Airflow and containment review, rack inlet limits, fan and sensor validation |
| Water quality | Are chemistry, filtration, materials, pressure, and maintenance compatible? | Water specification, sampling plan, flushing and service procedure |
| Leak detection | Can leaks be detected, localized, contained, and escalated? | Sensor map, BMS or DCIM integration, alarm test, incident runbook |
| Floor loading | Can the route, slab, raised floor, and final position support the configured rack? | OEM shipping and operating weight, point-load analysis, structural approval |
| Delivery and rigging | Can the rack pass through docks, corridors, doors, elevators, and turns? | Route survey, dimensions, packaging plan, rigging method |
| Service clearances | Can trays, power shelves, manifolds, cables, and switches be serviced safely? | Floor layout, front and rear clearance validation, maintenance simulation |
| Emergency operations | What happens during power, cooling, water, fire, or facility-control events? | Shutdown sequence, escalation matrix, drills, recovery and restart procedure |
Rack weight must be taken from the exact selected system, not a generic NVL72 label. Integrated rack weight changes with the OEM frame, power shelves, manifolds, switches, cabling, shipping configuration, and optional components. The floor-load study should include the delivery route and point loads, not only the final rack footprint.
The power design should also distinguish nameplate capacity from measured operating demand. Capacity planning needs enough headroom for workload peaks, component aging, redundancy, and future software behavior. It should not assume that an average training run represents the worst electrical state.
Make Fabric Manager and Rack Software Dependencies Explicit
The NVLink domain depends on a coordinated software stack. Fabric Manager is not an optional monitoring add-on. It participates in discovering, configuring, routing, and monitoring the NVLink and NVSwitch fabric, while coordinating with GPU drivers and the broader rack software environment.
The design should document at least these dependencies:
- NVSwitch firmware and switch operating software
- Fabric Manager and NVLink Subnet Manager functions
- GPU driver and CUDA compatibility
- NVIDIA Collective Communications Library
- Inter-Memory Exchange services where used
- Data Center GPU Manager health and telemetry
- Network operating systems, NIC and DPU firmware, and RDMA software
- Mission Control or the selected rack-management and lifecycle platform
- Baseboard management, rack power, cooling, and leak telemetry
The names are easy to confuse. Fabric Manager manages the NVLink fabric. Ethernet fabric operations use the selected Spectrum-X management and telemetry tooling. InfiniBand deployments use the appropriate InfiniBand fabric-management stack. These systems can be integrated operationally, but they are not interchangeable.
Power-On and Restart Sequencing
Rack power sequencing is an architecture dependency because the NVLink switch layer must be available for the compute nodes to discover and join the intended domain. The operational runbook should specify the supported order for power-on, shutdown, and recovery. It should also define when compute nodes must be restarted after a switch-layer or management-layer event.
A generic data-center instruction such as “power on the rack” is insufficient. The procedure should include:
- Facility and cooling readiness confirmation.
- Rack power and management verification.
- NVLink switch-layer initialization.
- Fabric Manager and topology confirmation.
- Compute-tray initialization.
- GPU, NVLink, NIC, DPU, and storage health checks.
- Scheduler admission only after the rack passes its gate.
Treat the Release Recipe as One Configuration Item
The supported state is a recipe, not a collection of independently latest packages. GPU drivers, CUDA, Fabric Manager, NCCL, DCGM, switch firmware, NIC firmware, DPU software, DGX OS or the selected host OS, and cluster operators must be tested as a combination.
At the July 25, 2026 research cutoff, NVIDIA had posted GB200 rack release 1.3.10 and GB300 rack release 1.0.10, both dated July 15, 2026. The release notes also identify minimum rack-wide release levels for NVLink Recovery compatibility. Those values are useful examples of why lifecycle management must be rack-wide, but they should be rechecked immediately before publication, procurement, or upgrade execution.
Layer Kubernetes and Cluster Software Above a Stable Rack Baseline
Kubernetes does not remove the need for rack-level validation. It adds another control plane above the hardware and low-level software stack. The cluster should not admit the NVL72 rack until the rack has already passed hardware, firmware, NVLink, network, and storage gates.
The reader should notice that Kubernetes is near the top of the dependency chain. It can schedule a resource only after lower layers expose a truthful and stable resource state.
Separate Administration from Workload Tenancy
For larger AI factories, use a dedicated management or administration cluster for lifecycle services, deployment automation, images, telemetry, and platform control functions. Avoid placing critical rack-management services exclusively inside the same user cluster they must recover.
The user cluster should enforce:
- homogeneous node pools for the relevant GB200 or GB300 configuration
- explicit GPU and topology labels
- topology-aware scheduling and placement
- taints, tolerations, and admission policies
- controlled access to RDMA devices and host capabilities
- version-pinned GPU and Network Operators
- namespace, quota, and priority controls
- maintenance drain and return-to-service gates
- centralized logs, metrics, fabric telemetry, and audit trails
The current Network Operator and GPU Operator support matrices should be treated as release inputs. Kubernetes, operating-system, driver, OFED, NIC, and operator compatibility changes over time. Copying an operator version from a lab build into production without checking the complete matrix is a preventable failure.
Design Failure Domains Before You Schedule Jobs
A cluster is not resilient because it has redundant links. It is resilient when the scheduler, network, storage, facility systems, and operators agree on what can fail together and what action follows.
| Failure domain | Potential impact | Design response | Acceptance evidence |
|---|---|---|---|
| Single GPU | Job failure or degraded workload, depending on software | Detect, isolate, drain affected resource, preserve fault evidence | DCGM event, scheduler action, workload result |
| Compute tray | Four GPUs and two Grace CPUs unavailable | Treat tray as a service unit, drain and replace under rack procedure | Inventory delta, topology check, post-repair test |
| NVLink switch tray | NVLink topology or routing impact may extend beyond one tray | Treat as a rack-level event until the supported degraded behavior is proven | Fabric Manager state, NVLink diagnostics, vendor procedure |
| Entire NVL72 rack | Loss of one 72-GPU scheduling unit | Multi-rack capacity plan, job checkpointing, queue policy | Rack isolation test, scheduler behavior, recovery timing |
| One scale-out link | Reduced path capacity or loss of one GPU path | Dual-plane connectivity, fast detection, controlled remediation | Link-down test, counters, NCCL comparison |
| One leaf or one plane | Possible reduced bandwidth with continued reachability through remaining paths | Independent planes, capacity headroom, topology-aware routing | Plane isolation test, workload continuity and performance delta |
| Spine or super-spine path | Multi-rack path loss or congestion shift | Redundant spine paths, tested convergence, telemetry | Failure injection, convergence timing, collective results |
| Storage path or metadata service | Checkpoint, load, or data-pipeline interruption | Redundant paths, failure-aware clients, recovery procedure | GDS or storage failover test, integrity check |
| Management plane | Reduced visibility or inability to operate safely | Separate management HA, break-glass access, backup and restore | Control-plane failover and restoration evidence |
| Power or cooling zone | Rack or multi-rack shutdown | Facility zoning, controlled shutdown, capacity and restart plan | Facility simulation, alarm and shutdown records |
The expected behavior should be written as a testable hypothesis. For example, “Plane B preserves reachability after Plane A isolation” is a design objective, not a fact, until the exact installed topology and software recipe prove it under representative load.
Use Cluster Acceptance Testing as a Gated Evidence Chain
Acceptance testing should prevent an unhealthy layer from being hidden by a higher-level benchmark. A large language model can sometimes run on a partially degraded platform, but that does not mean the platform is ready for production.
A useful evidence chain moves from static configuration to dynamic workload behavior.
| Gate | What to validate | Minimum evidence |
|---|---|---|
| Design and BOM | Certified or supported components, topology, ownership, firmware and software recipe | Approved design, BOM, certification status, support matrix, exceptions |
| Facility | Power, cooling, water, floor, access, alarms, emergency procedures | Commissioning reports, readings, alarm tests, approvals |
| Physical installation | Correct trays, switches, power shelves, cables, optics, labels, leak sensors | Inventory, photos, cable scan, optical and link results |
| Firmware and supply chain | Version alignment, signed images, SBOM, known exceptions | Machine-readable inventory, hashes, SBOM, vulnerability review |
| Tray health | CPUs, GPUs, memory, local storage, NICs, DPUs, sensors | BMC and DCGM results, burn-in records, error logs |
| NVLink domain | 72-GPU topology, NVSwitch state, Fabric Manager health, error-free operation | Topology capture, FM state, NVLink counters, intra-rack tests |
| Per-rail network | Port-to-GPU mapping, RDMA, bandwidth, latency, plane independence | Cable map, RDMA tests, counters, plane-specific results |
| Storage | Data and metadata performance, GDS or RDMA path, integrity, failover | Storage benchmarks, path verification, integrity evidence |
| Kubernetes or Slurm | Node identity, resource discovery, scheduling, drain, operators, recovery | Cluster-state export, scheduling tests, failover records |
| NCCL | Correctness and collective performance at tray, rack, multi-rack, and cluster scope | NCCL logs, bus bandwidth, algorithm bandwidth, error and variance analysis |
| Workload | Representative training or inference behavior and SLOs | Model, dataset, configuration, throughput, latency, accuracy or quality evidence |
| Resilience | Link, plane, node, service, and management failures | Failure-injection records, observed impact, recovery timing |
| Soak and handover | Stability over time and operational readiness | Soak report, incident log, runbooks, owner sign-off, support contacts |
Use an Acceptance Manifest
A versioned acceptance manifest prevents the test team from validating one configuration while operations inherits another. The following YAML is a framework, not a vendor-supplied schema.
acceptance_baseline: platform: gb300-nvl72 rack_count: 4 hardware_bom: "" software_recipe: " " fabric: type: spectrum-x planes: 2 oversubscription: "1:1" rail_map: " " required_tests: - facility_commissioning - firmware_inventory - dcgm_health - nvlink_domain - per_rail_rdma - nccl_collectives - gpudirect_storage - scheduler_drain_and_return - inference_slo - failure_injection - soak_test evidence: repository: " " retention_days: 365 required_approvers: - facilities - network - storage - platform - ai_operations - security
The fields to change are the platform, rack count, approved BOM, release recipe, fabric type, rail-map artifact, test list, retention period, and approvers. Success means every test result can be traced to the installed hardware and software baseline. A common failure is retaining benchmark output without the topology, versions, workload parameters, or environmental state needed to interpret it.
Validate NCCL and Inference at the Topology Boundaries
NCCL testing is necessary because it exposes the interaction between GPUs, NVLink, RDMA, topology discovery, routing, and collective algorithms. It is not sufficient because production workloads add model, framework, storage, serving, and concurrency behavior.
Build a Layered NCCL Matrix
Run collective tests at multiple scopes:
| Scope | Purpose | Representative operations | What to compare |
|---|---|---|---|
| One compute tray | Establish a local hardware baseline | All-reduce, all-gather, reduce-scatter | Correctness, consistency, local topology |
| One NVL72 rack | Validate the complete 72-GPU NVLink domain | All-reduce, all-gather, reduce-scatter, all-to-all | Fabric Manager state, bus bandwidth, algorithm bandwidth, errors |
| Two racks | Validate first scale-out boundary | All-reduce, reduce-scatter, send-receive | Rail mapping, plane symmetry, RDMA path, variance |
| Multiple racks | Validate leaf-spine behavior under distribution | All-reduce, all-to-all, broadcast | Congestion, adaptive routing, hot spots, scaling efficiency |
| Full cluster | Establish production baseline | Workload-relevant collectives and message sizes | Repeatability, tail behavior, failure sensitivity, capacity effect |
Record more than the headline bandwidth number. Retain message sizes, process placement, GPU mapping, NCCL environment, topology files, algorithm selection, transport selection, CPU and NUMA affinity, rail counters, switch counters, optical health, retransmissions, and run-to-run variance.
Pass thresholds should come from the validated design, vendor or integrator baseline, and the exact installed bill of materials. An arbitrary value copied from a different cluster can hide topology differences or create a false failure.
Validate Inference as a Service, Not a Kernel
Inference acceptance should use representative models and serving behavior. At minimum, measure:
- model load and warm-up time
- time to first token
- inter-token latency
- request throughput
- tokens per second per user and across the service
- P50, P95, and P99 latency
- prefill and decode behavior
- concurrency and queueing
- HBM and KV-cache utilization
- storage and model-loading paths
- GPU, network, CPU, memory, and power utilization
- behavior during a failed node, link, plane, or service component
- recovery, rescheduling, and return to full capacity
The test should state the model, precision, framework, serving stack, tensor and pipeline parallelism, sequence lengths, input and output distributions, concurrency, batching, scheduler policy, and SLO. Without those details, an inference number is not reusable evidence.
For training workloads, apply the same principle. Record samples or tokens per second, scaling efficiency, checkpoint behavior, restart time, convergence or quality controls, and the infrastructure state during the run.
Expand Capacity in Scalable Units, Not Random Racks
Capacity expansion should preserve the original architecture rather than append hardware wherever space remains.
Before the first rack is installed, the design should reserve:
- rail-aligned leaf ports
- spine or super-spine capacity
- optics, patching, and cable pathways
- IP address, VLAN, routing, and management capacity
- rack power and upstream electrical capacity
- CDU, secondary-loop, and facility-water headroom
- storage throughput and metadata capacity
- management-cluster and telemetry capacity
- scheduler, quota, and queue capacity
- service access, spares, and maintenance space
Add racks symmetrically across the two network planes. Preserve the GPU-to-rail mapping. Update the source-of-truth inventory and cable map before installation. Then rerun the acceptance tests affected by the change, including per-rail RDMA, multi-rack NCCL, storage, workload, and failure tests.
Mixing GB200 and GB300 in one administrative estate may be operationally reasonable, but it should not imply one homogeneous scheduling pool. Different GPUs, NIC generations, power envelopes, firmware, drivers, operators, and performance profiles can complicate placement and support. Use separate node pools or clusters unless the exact mixed design and workload behavior have been validated.
Plan Maintenance and Lifecycle Around the NVLink Domain
A rack-scale system changes maintenance economics. A component replacement may be physically local but operationally rack-wide. The maintenance model should assume that some interventions require draining the entire NVL72 domain, preserving fault evidence, applying a coordinated recipe, and rerunning acceptance gates.
Define Maintenance Classes
| Maintenance class | Example | Likely scope | Required revalidation |
|---|---|---|---|
| Observation only | Log collection, telemetry review | Component or rack | No workload disruption, evidence preserved |
| Hot-serviceable component | Approved cable, optic, fan, or management component | Link, tray, or rack, depending on procedure | Link health, topology, counters, affected workload test |
| Compute-tray service | Tray replacement or firmware intervention | Tray and potentially full rack drain | Inventory, DCGM, NVLink, NCCL, scheduler return |
| NVLink switch service | Switch tray, FM, NVOS, or topology change | Treat as rack-wide | Complete NVLink-domain and rack acceptance |
| Scale-out fabric change | Leaf, spine, optics, routing, congestion policy | Plane, rail, or cluster | Per-rail RDMA, plane isolation, multi-rack NCCL |
| Rack software upgrade | Driver, CUDA, FM, NCCL, DCGM, OS, operators | Rack or cluster rollout wave | Compatibility, functional, performance, resilience, rollback |
| Facility maintenance | Power feed, CDU, water loop, alarms | Rack, row, or facility zone | Facility commissioning subset, controlled restart, rack acceptance |
Use Canary and Wave-Based Upgrades
Do not update all racks at once. Maintain an approved production recipe, a test or canary environment that is representative enough to expose compatibility problems, and a wave plan with explicit stop conditions.
A sound lifecycle workflow is:
- Validate the candidate recipe in a nonproduction or canary rack.
- Run the complete affected acceptance suite.
- Drain one production rack.
- Apply the change with captured before-and-after inventory.
- Revalidate the rack and its fabric paths.
- Return it to service under observation.
- Continue only when the stop conditions remain clear.
- Retain the previous recipe, rollback method, and evidence.
Firmware and software “latest” is not a lifecycle strategy. Supported compatibility and reproducible evidence are the strategy.
Follow a Practical Design and Validation Sequence
The work should progress through decision gates rather than a single installation project.
Discover and Scope
Define workloads, model sizes, training and inference patterns, data volumes, RTO and RPO, security boundaries, tenancy, growth, and operational ownership. Decide whether GB200 or GB300 is the correct platform and whether the scale-out design will use Spectrum-X Ethernet or a supported InfiniBand architecture.
Exit criterion: Approved requirements, assumptions, workload profiles, and ownership model.
Qualify the Facility
Complete electrical, cooling, structural, delivery-route, service-clearance, water, leak-detection, and emergency-procedure studies against the exact OEM rack.
Exit criterion: Signed facility design and capacity reservation for the initial and planned expansion phases.
Complete the Detailed Architecture
Produce logical and physical network designs, rail maps, plane boundaries, storage architecture, management and security zones, IP plans, control-plane placement, software recipe, failure domains, and observability design.
Exit criterion: A buildable, reviewable design with no unresolved ownership boundary.
Factory and Integration Validation
Confirm certified components, exact BOM, firmware inventory, cable plan, SBOM, factory test evidence, shipment controls, and on-site prerequisites.
Exit criterion: The delivered configuration matches the approved design and support baseline.
Rack Bring-Up
Commission cooling and power, initialize management and switch layers, establish Fabric Manager health, bring up compute trays, validate inventory, and confirm the 72-GPU NVLink topology.
Exit criterion: Each rack passes its standalone health and NVLink gate.
Fabric and Storage Acceptance
Validate rail mapping, both planes, RDMA, leaf-spine behavior, congestion controls, storage paths, GDS where required, and management reachability.
Exit criterion: Every rack has symmetric, supportable paths to the required cluster services.
Cluster and Workload Acceptance
Validate Kubernetes or Slurm, operators, scheduling, NCCL at increasing scale, representative training or inference, failure behavior, and soak stability.
Exit criterion: The platform meets documented technical and service SLOs with reproducible evidence.
Operational Handover
Deliver runbooks, dashboards, alerts, spares, support paths, escalation, maintenance classes, upgrade procedure, rollback, evidence retention, and capacity thresholds.
Exit criterion: Operations can detect, isolate, recover, maintain, and expand the platform without relying on undocumented project knowledge.
Conclusion
A GB200 or GB300 NVL72 AI factory should be designed from the facility upward and validated from the tray outward. The rack is a 72-GPU NVLink scale-up domain, not a loose collection of servers. The scale-out network is a separate distributed system that must preserve rail symmetry, plane independence, RDMA behavior, and predictable performance across racks.
The strongest design does not stop at a reference diagram. It assigns ownership, fixes a supported release recipe, makes failure domains explicit, preserves expansion headroom, and defines an acceptance evidence chain. Facilities, Fabric Manager, Spectrum-X or InfiniBand, storage, Kubernetes, NCCL, and the workload are all part of the same production outcome.
NVIDIA certification and reference architectures are valuable inputs. They reduce uncertainty about supported components and known patterns. They do not prove that a specific site, fabric, storage system, software recipe, or workload has been implemented correctly. That proof comes from disciplined commissioning, topology-aware testing, failure injection, workload validation, and repeatable lifecycle operations.
The practical goal is not merely to power on 72 GPUs. It is to build an AI factory that can be measured, maintained, expanded, and trusted under production conditions.
