All topicsSTATE OF AI REPORT COMPUTE INDEX.

Why does networking matter in a GPU cluster?

Architecture explanation · reviewed October 11, 2026

Large AI workloads can split a model or its data across many GPUs. Those GPUs must exchange intermediate results and synchronize their work, so slow communication can leave computing capacity waiting. The network’s bandwidth, latency, and topology therefore affect how efficiently a job uses a cluster. Adding GPUs does not guarantee a proportional increase in useful performance.

How to read this finding

NVIDIA distinguishes scale-up communication within a GB200 NVL72 rack, using NVLink, from scale-out connections between racks, using Ethernet or InfiniBand. Storage networking also matters because the GPUs need a steady supply of data. These are different parts of the system, not interchangeable specifications. The index counts hardware; it does not benchmark each cluster’s interconnect, storage throughput, or performance on the same training job. Compare those details before treating two equal-sized GPU clusters as equivalent.

Charts and sources

  1. NVIDIA data-center network architecture
  2. Cluster inventory
  3. Inventory sources

Citation estimates use inputs through September 1, 2026; fleet counts have an October 1, 2026 cutoff. The research-topic sample covers January 1 to June 1, 2025. See each answer for its period and limitations.

All data sources

Cite this page

Benaich, Nathan. “Why does networking matter in a GPU cluster?” State of AI Report Compute Index. Web page updated 2026-10-11; data periods as specified above.

Related questions

Who has the largest GPU cluster in this index?How many GPUs do OpenAI and Anthropic have?How many GPUs does Europe’s JUPITER supercomputer have?What is a GPU cluster, and how is it different from a fleet?How much power does a GPU cluster need?