A practical walkthrough of building a lossless Ethernet fabric for AI training and inference, covering RoCEv2, PFC, and ECN configuration across the 7060X, 7260X, and 7800R platform families.
Talk to a SpecialistGenerative AI training workloads operate on all-to-all communication patterns between GPU accelerators. During collective operations such as AllReduce, every GPU exchanges gradient data with every other GPU in the training cluster simultaneously. When any single packet in that exchange is dropped, the entire collective stalls while the sending GPU retransmits, and every other GPU in the operation idles waiting for the missing data to arrive. At the cost of current GPU accelerators, idle time during retransmission represents a measurable and preventable cost overhead. The solution is a network fabric that prevents packet drops under load rather than recovering from them after the fact: a lossless Ethernet infrastructure built on Priority Flow Control (PFC), Explicit Congestion Notification (ECN), and RoCEv2 transport, tuned to eliminate the drop events that cause stalls.
Arista's data centre switching portfolio provides the hardware capabilities required at each tier of the AI fabric: the 7800R4 modular spine and 7700R4 Distributed Etherlink provide the high-radix, non-blocking interconnect at the spine tier; 7060X5 fixed spine and 7260X3 fixed leaf-spine provide 400G and 100G density with the deep buffers required to absorb in-fabric congestion without dropping; and the 7010X series provides power-efficient 1/10/25G access at under 0.3W per Gbps for scale-out server connectivity. All platforms run EOS, which implements PFC and ECN tuning consistently across tiers and provides LANZ (Latency Analyzer) streaming telemetry at nanosecond precision for detecting micro-burst congestion events before they propagate into retransmissions. Zero Touch Provisioning and In-Service Software Updates eliminate the operational overhead that would otherwise slow large-cluster deployments and ongoing maintenance.
RDMA over Converged Ethernet v2 (RoCEv2) moves data between GPU memory directly over the Ethernet fabric without CPU involvement, reducing latency and CPU overhead compared to TCP-based data movement. PFC creates a lossless transport layer by signalling upstream devices to pause sending before buffer occupancy reaches the drop threshold, while ECN marks packets early to signal congestion before PFC is triggered. The combination (ECN for early notification, PFC as a backstop) is tuned differently per deployment based on fabric RTT and buffer sizes; Arista EOS exposes the specific queue thresholds, pause thresholds, and ECN marking rates required to achieve lossless operation at production traffic levels.
The 7260X3 provides up to 12.8 Tbps in 2RU with 64 MB of dynamic shared buffer, a significantly larger buffer than fixed-function shallow-buffer switches that rely on PFC alone to prevent drops. The large shared buffer absorbs the traffic bursts that occur when many GPU-to-GPU flows complete a collective operation simultaneously and simultaneously begin the next one, a pattern that generates very short, very high-rate bursts at the spine tier. Absorbing those bursts in the switch buffer rather than triggering PFC pause frames keeps the fabric flowing without the stall behaviour that PFC can introduce when triggered frequently. The 7800R4 modular spine extends this with 800G port density for the highest-scale AI fabric tiers where 400G leaf-spine is insufficient.
LANZ (Latency Analyzer) instruments every switch buffer in the fabric at nanosecond precision, recording the peak buffer occupancy and latency of each queue during each sampling interval and streaming that data to CloudVision in real time. For an AI fabric operator trying to determine whether occasional training job slowdowns are caused by in-fabric congestion, LANZ provides the buffer-level evidence required to make that determination: a peak queue depth that correlates with the slowdown timestamp confirms in-fabric congestion, whereas flat LANZ data during a slowdown directs investigation to the compute tier rather than the network. Without LANZ, diagnosing transient congestion in a 400G or 800G fabric requires inference from job completion times rather than direct measurement of where the congestion occurred.
The 7010X series achieves under 0.3W per Gbps for 1/10/25G server access connectivity, a power efficiency figure that becomes significant at the scale of GPU clusters where thousands of server-to-fabric connections each consume measurable power at the switch port level. At 7W per 100G port, the 7060X5 provides similar efficiency at the 400G leaf tier. For a facility where power is the constraining resource limiting GPU cluster expansion, choosing switching hardware at the power-efficient end of the port density/watt spectrum directly expands the number of GPUs that can be powered and cooled within the facility's available power envelope. Platinum-rated PSUs (above 93% efficiency) further reduce the facility power draw for the same switch output power.
Full specifications for Bridge the AI Infrastructure Gap
Detailed specifications are coming soon — see the datasheet in the Documentation tab for full details in the meantime.
Download product documentation and resources
Arista switches validated for lossless RoCEv2 AI networking.
Modular 800G spine for the largest AI fabric deployments — high-radix non-blocking interconnect with deep buffers and EOS-native RoCEv2 tuning.
400G fixed-configuration workhorse leaf and spine — 450 ns cut-through latency, Platinum-rated PSUs, deep shared buffer, full PFC/ECN support.
12.8 Tbps in 2RU with 64 MB dynamic shared buffer — high-capacity leaf-spine for burst-heavy AI collective operations without triggering PFC.
Power-efficient 1/10/25G access switching at under 0.3W per Gbps — scale-out server connectivity for dense GPU compute rows.
Our team of experts is ready to help you find the perfect solution for your business needs. Get personalized advice and competitive quotes.
We're here to help with any questions