Hardware architecture

HGX, NVL and NVL72: understanding GPU connectivity

Why GPU count and total memory do not establish communication speed: NVLink domains, CPU RAM and networking between nodes.

7 min readUpdated 6 October 2026
Editorial illustration

Two servers containing eight GPUs can have different internal architectures. That affects distribution of one model, synchronisation latency and useful concurrency. Comparing offers requires a connectivity map as well as a GPU list. Begin with the boundary of the fast domain: which GPUs communicate through a common NVLink fabric, and where traffic takes another path.

1. Separate GPU count from domain size

A communication domain is a group of accelerators joined by a particular high-speed interconnect. Its boundary need not match a chassis. HGX connects SXM modules on a common baseboard through NVLink and NVSwitch. H200 NVL connects separate PCIe cards using bridges. In GB200 and GB300 NVL72, one NVLink domain extends across the compute trays of an entire rack.

Eight H200 NVL cards equipped with four-way bridges form two groups of four GPUs. Paired arrangements are also possible; confirm the bridge inventory in the quotation. Communication between those groups does not become an eight-GPU NVSwitch fabric. Eight H200 GPUs in both machines are insufficient grounds to treat NVL and HGX configurations as equivalent.

2. Map the name to a physical platform

NVL and NVL72 sound similar but describe systems at different scales. One PCIe card, an HGX baseboard and a complete rack are different units of supply. An accelerator image can explain architecture without establishing the chassis contents. Match part numbers, physical layouts and component lists, especially when a seller uses a generic family name.

PlatformFast domainRequired specification detail
2× H200 NVLA pair linked by NVLink bridgePCIe cards and a compatible two-way bridge
8× H200 NVL2×4 with NVL4, or paired groupsBridge count and type; paths between groups
HGX H200 / B200 / B3008 GPUs through NVSwitchExact HGX baseboard and server model
GB200 / GB300 NVL7272 GPUs through rack NVLink36 Grace CPUs, compute and switch trays

3. Do not combine different memory types into one VRAM figure

HBM belongs to GPUs; system RAM serves CPUs. For Grace Blackwell, vendors also publish combined fast memory that includes both categories. A large fast-memory figure is not an equal amount of HBM. Shared addressing or coherence does not make every memory access equally fast: the physical placement of weights and KV state still matters.

Even within an NVLink domain, the runtime distributes the model and its state. Check available memory on each GPU, layer placement, potential replication and runtime headroom. CPU memory can participate in offloading, but that is a separate execution mode with measurable transfer costs. Request GPU HBM and CPU RAM as separate entries in a model budget.

4. Match the distribution strategy to the links

Tensor parallelism partitions computation inside layers and needs frequent communication. Pipeline parallelism places groups of layers on different devices. Independent model replicas can serve separate requests with different synchronisation requirements. A larger domain is particularly useful when a workload benefits from it; the number of connected GPUs alone does not determine service speed.

Check networking between nodes separately: NIC count, GPU and CPU affinity, switches, cables, RDMA support and measured bandwidth. Do not compare a 400 Gb/s port numerically with 900 GB/s NVLink: letter case changes the unit by a factor of eight, and published figures may describe different directions or aggregate links.

5. Request evidence for the actual OEM configuration

Obtain the block diagram covering GPUs, CPUs, PCIe switches and network cards. Establish which bridges are installed, which ports share lanes and whether transfers cross another CPU. Compare that diagram with the running server’s connectivity matrix and communication measurements. An accelerator specification does not describe every route inside an OEM system.

Validate the complete model path: placement, startup, a long request and target concurrency. Compare configurations using identical software versions and inputs. If a supplier substitutes the chassis or network adapter, update the map and repeat relevant measurements. This makes topology a verifiable purchasing condition rather than a product label.

One large model depends on domain boundaries, communication paths and memory placement. Many independent workloads may benefit from a different layout.

Checklist

What to request about topology

Sources and documentation

These guides help you prepare requirements. The exact configuration and terms are set out in the quotation.