Architecture explanation · reviewed October 11, 2026
Large AI workloads can split a model or its data across many GPUs. Those GPUs must exchange intermediate results and synchronize their work, so slow communication can leave computing capacity waiting. The network’s bandwidth, latency, and topology therefore affect how efficiently a job uses a cluster. Adding GPUs does not guarantee a proportional increase in useful performance.
How to read this finding
NVIDIA distinguishes scale-up communication within a GB200 NVL72 rack, using NVLink, from scale-out connections between racks, using Ethernet or InfiniBand. Storage networking also matters because the GPUs need a steady supply of data. These are different parts of the system, not interchangeable specifications. The index counts hardware; it does not benchmark each cluster’s interconnect, storage throughput, or performance on the same training job. Compare those details before treating two equal-sized GPU clusters as equivalent.
Citation estimates use inputs through September 1, 2026; fleet counts have an October 1, 2026 cutoff. The research-topic sample covers January 1 to June 1, 2025. See each answer for its period and limitations.
Benaich, Nathan. “Why does networking matter in a GPU cluster?” State of AI Report Compute Index. Web page updated 2026-10-11; data periods as specified above.