A graphics processing unit cluster combines accelerator servers with high-speed networking, storage and software that schedules work across devices. Artificial intelligence training and inference use clusters when one device lacks sufficient memory or throughput.
Cluster performance depends on communication topology, utilization, reliability and software efficiency. A managed cluster adds operational services such as provisioning, monitoring and repair so customers do not need to operate the hardware directly.
Acronyms and aliases
GPU cluster acronymaccelerator cluster synonym
General terms
Specialised terms
Frequently asked questions
Why do artificial intelligence models use graphics processing unit clusters?
Large models and workloads exceed the memory or processing capability of one device, so computation and data are distributed across many accelerators.
Does adding more graphics processing units always make a cluster faster?
No. Communication overhead, workload partitioning, software scaling and idle devices can limit the benefit of additional hardware.
Videos explaining graphics processing unit cluster