
The call objective is to demystify the "Backend Network" and explore the networking design requirements for building distributed GPU clusters. Participants will move beyond traditional Ethernet concepts to understand why AI workloads demand lossless fabrics, specialized topologies, and a new transport evolution through the Ultra Ethernet Consortium (UEC). By the end of the call, you will be able to speak the basic language of AI infrastructure—from NCCL collectives to packet spraying and "Scalable Units" (SUs).
The AI Infrastructure Stack: Understanding the relationship between programming frameworks (PyTorch/TensorFlow) and the underlying network.
Distributed Training Patterns: Dive into Data, Tensor, and Pipeline parallelism and their unique traffic signatures.
The Lossless Requirement: How RDMA over RoCEv2 uses PFC and ECN to prevent the "tail latency" that stalls GPU compute.
Modular Cluster Design: Building with Scalable Units (SU), Rail-optimized topologies, and non-blocking fat-trees.
Storage Tiers for AI: High-performance NVMe fabrics and the impact of checkpointing on training efficiency.
The Future: Ultra Ethernet (UEC): Why we are moving from connection-oriented legacy (InfiniBand/RoCE) to connectionless, packet-spraying transport.
UEC Architectural Innovations: Exploring Packet Trimming, Link-Level Retry (LLR).
A foundational understanding of IP networking (clos-topologies, BGP, ECMP).
No prior experience with AI/ML is required;
Feel free to ask your specific questions relevant to your current or future projects.
Network Architects and Engineers tasked with designing or operating GPU backends.
Data Center Operators looking to transition from traditional "Front-end" clouds to "Back-end" AI fabrics.
Systems Engineers curious about the convergence of HPC (High-Performance Computing) and Ethernet.