4RU modular AI/ML switch at 12.8 Tbps with hot-swappable interface modules and 900 ns latency supporting 100G, 200G, and 400G port densities.
The Arista 7358X4 Series is a purpose-built modular AI fabric switch platform addressing the bandwidth, congestion management, and lossless Ethernet requirements of large-scale GPU training clusters. Available in 4-slot (7358X4-4) and 8-slot (7358X4-8) chassis configurations, the series reaches 460.8 Tbps of aggregate switching capacity when fully populated with 36x QSFP-DD 400G line cards (7358X4-LC-36DD), or equivalent capacity using 800G line cards as they become available. The platform supports leaf attachment at 400G and spine interconnection at 800G in the same chassis, enabling a two-tier fabric without changing hardware at the spine.
Advanced congestion control capabilities including hardware-based ECN marking, per-destination output queue management, and deep ingress buffering address the collective communication traffic patterns (AllReduce, AllGather, ReduceScatter) that produce incast and congestion events in AI training networks. Full lossless Ethernet operation via PFC (Priority Flow Control) and DCQCN (Data Center Quantized Congestion Notification) eliminates packet loss that would degrade GPU utilisation during training iterations, while RDMA over Converged Ethernet (RoCE) support provides the low-latency GPU-to-GPU communication path that modern AI frameworks require.
Up to 460.8 Tbps of switching capacity in the 8-slot chassis provides enough aggregate bandwidth to support tens of thousands of connected GPUs in a two-tier fabric. The modular design means capacity scales linearly as line cards are added, allowing initial deployments to start at 4 slots and expand without replacing any chassis hardware or reconfiguring fabric topology.
Hardware Priority Flow Control (PFC) and DCQCN (Data Center Quantized Congestion Notification) enforce lossless Ethernet delivery across the fabric, eliminating the retransmission events that cause GPU stalls during all-reduce operations. Unlike software-based congestion management, the hardware implementation responds within microseconds, keeping GPU utilisation high even during worst-case incast scenarios.
Native RoCE (RDMA over Converged Ethernet) support provides the RDMA transport that GPU interconnect frameworks such as NCCL require to bypass the CPU and transfer tensor data directly between GPU memory spaces across the network fabric. Combined with hardware ECN marking and deep per-port buffers, RoCE operates at consistent low latency under full training-cluster load.
The 7358X4 supports line-card mixing between 400G leaf-facing cards and 800G spine-facing cards within the same chassis, enabling a two-tier AI fabric where the leaf-to-spine uplinks operate at 800G without requiring a separate spine-only chassis. This halves the hardware inventory and rack footprint needed to build a non-blocking fabric connecting 400G-attached GPU nodes.
Full specifications for Arista 7358X4 Series
Download product documentation and resources
Explore other configurations and models that might suit your needs.
Recommended
#DCS-7060PX4-32
Recommended
#DCS-7060DX5-64S
Recommended
#DCS-7060X6-64PE
Our team of experts is ready to help you find the perfect solution for your business needs. Get personalized advice and competitive quotes.
We're here to help with any questions