How Arista's jointly validated 400G and 800G AI fabric, built with Broadcom, scales cleanly across two-tier topologies to support GPU clusters numbering in the tens of thousands.
Deploying an AI data centre fabric at 400G or 800G port density involves selecting, integrating, and validating switching hardware, optics, cabling, and EOS configuration profiles that together produce a lossless fabric capable of sustaining the traffic patterns of GPU training and inference workloads. Arista and Broadcom have jointly developed and validated a pre-tested 400G and 800G leaf-spine reference architecture that covers hardware selection, cabling topology, EOS PFC and ECN tuning profiles, and AI Job Observability configuration, reducing the validation effort that would otherwise fall on the customer's network engineering team before a production cluster can come online. The pre-validation covers 2-tier leaf-spine topologies designed to cluster tens of thousands of accelerators, with both fixed-configuration and modular-chassis options at the spine tier to match the scale of the deployment.
The 7800R4 AI Spine is the anchor platform at the spine tier for the largest deployments, providing 800G port density in a modular chassis configuration with the non-blocking, high-radix interconnect required to prevent bottlenecks at the spine layer as the cluster scales. The 7060X6 Series provides 800G port density at the leaf tier in a fixed 1RU configuration, pairing with the 7800R4 spine in a non-blocking 2-tier design. For deployments where the AI fabric is part of a mixed 400G/800G environment (GPU training on 800G, inference and general compute on 400G), the 7060X5 at 400G and 7060X4 for mixed workloads provide the leaf options. Per-flow and per-job visibility through LANZ and AI Job Observability allows operators to trace the source of tail latency and throughput variance to specific flows, jobs, or fabric tiers, the diagnostic capability that distinguishes a production-ready AI fabric from a research deployment.
Joint validation between Arista and Broadcom covers the full stack of hardware, EOS software, and tuning profiles required for a production AI fabric: switch ASICs, optics compatibility, PFC/ECN parameter sets, and RoCEv2 transport configuration profiles. Pre-validation means the customer organisation does not need to design and execute its own validation test suite before deploying the fabric into production, which is the work that extends typical AI cluster deployment timelines by weeks to months. The validated reference also defines the upgrade path when new GPU generations require higher per-port bandwidth, since the reference architecture is designed to accommodate interface speed upgrades without requiring a complete topology redesign.
The 2-tier leaf-spine topology used in the validated fabric provides equal-cost, non-blocking connectivity between every accelerator in the cluster: every GPU can reach every other GPU through the same number of hops with the same bandwidth and the same latency. This uniformity is the network property that allows AI framework schedulers to place training jobs across any available set of accelerators without regard to their physical location in the rack or row, because the network does not introduce location-dependent latency or bandwidth asymmetry that would need to be factored into scheduling decisions. Equal-cost multi-path (ECMP) load-balancing across multiple parallel spine paths prevents any single spine switch from being the bottleneck in the collective communication pattern.
AI Job Observability provides per-job and per-flow network telemetry that correlates training job identifiers with the specific network flows carrying that job's traffic, allowing operators to answer the question of whether a specific training job is running slower than expected because of a network bottleneck or a compute bottleneck. When a training job stalls, job observability data shows whether the stall correlates with congestion on specific fabric links, whether the flow distribution across ECMP paths is uneven, or whether a specific GPU's flows are contributing disproportionate traffic. Without job-level observability, diagnosing training slowdowns requires inference from aggregate switch statistics that cannot be attributed to specific jobs running simultaneously on the same fabric.
The validated fabric uses open-standards Ethernet, with no proprietary transport protocols, no vendor-specific fabric adapters, and no custom optics that restrict sourcing to a single vendor. Standard RoCEv2 over Ethernet is supported by all major AI accelerator vendors, meaning the same Arista fabric layer supports NVIDIA, AMD, Intel, and custom AI ASIC accelerators without requiring separate fabrics or protocol bridges. ECMP load-balancing, BGP routing between fabric tiers, and EVPN for multi-tenancy are all implemented using IETF-standard protocols and configurations that are compatible with standard network automation tooling, removing the operational overhead of proprietary fabric management software that would otherwise be required alongside the standard EOS management tools.
See the hardware pages below for detailed technical specifications of each platform in this validated fabric.
This solution brief does not have standalone specifications; see the individual hardware pages in the Featured Products section below for full technical details of each platform.
Download product documentation and resources
Validated Arista hardware for AI cluster leaf-spine deployments.
800G modular spine chassis, providing non-blocking high-radix interconnect for large AI cluster tiers.
800G fixed leaf/spine with maximum port density for GPU-rack connectivity at the leaf tier.
400G leaf, the 400G workhorse for AI inference tiers and mixed-workload fabric deployments.
Disaggregated 800G chassis for cross-row spine interconnect at scale beyond a single chassis.
Explore other Arista solution areas that complement this fabric deployment.
Recommended
A practical walkthrough of building a lossless Ethernet fabric for AI training and inference, covering RoCEv2, PFC, and ECN configuration across the 7060X, 7260X, and 7800R platform families.
Recommended
Pairing Arista's AI networking fabric with the VAST Data platform to deliver a single stack that scales from initial training runs through to production inference.
Recommended
What deterministic AI-ready performance actually looks like on the 7010X, 7060X, and 7260X, including port-speed flexibility and how EOS handles self-healing during fabric events.
Our team of experts is ready to help you find the perfect solution for your business needs. Get personalised advice and competitive quotes.
We're here to help with any questions