hpe cray slingshot network

**HPE Slingshot and Dragonfly+ HPC Interconnect** is the **high-performance network fabric deployed in the Frontier exascale supercomputer that combines Ethernet protocol compatibility with low-latency RDMA semantics over a dragonfly+ topology — achieving 200 Gbps per port bandwidth with adaptive routing that dynamically avoids congested links, enabling the all-to-all communication patterns of MPI collective operations at scale across 74,000 compute nodes**. **Slingshot Architecture** HPE Cray Slingshot is a purpose-built HPC interconnect: - **Physical layer**: 200 Gbps per port (400 Gbps planned), standard Ethernet electrical (but custom protocol extensions). - **Protocol**: Rosetta ASIC (switch chip) + Cassini NIC (host adapter), compatible with standard Ethernet frames but adding RDMA (via libfabric CXI provider) and enhanced QoS. - **Fabric topology**: dragonfly+ (see below). - **Congestion control**: hardware adaptive routing + injection throttling (no PFC needed — avoids head-of-line blocking without lossless Ethernet). - **Multitenancy**: traffic classes (bulk data, latency-sensitive, system management) with QoS isolation. **Dragonfly+ Topology** - **Groups**: each group is a fat-tree within a rack (local switches fully connected within group). - **Global links**: each group has global links to all other groups (1 or few links per group pair). - **Bisection bandwidth**: O(N) links for N groups → O(1) bandwidth per node (vs fat-tree which scales O(N log N) cost for full bisection). - **Path diversity**: between any two nodes, multiple paths exist (local routing within group + different global links). - **Diameter**: 3 hops (source group → inter-group → destination group) for any all-to-all communication. **Adaptive Routing** Static routing (fixed path per source-destination pair) suffers from hot spots when many flows share the same global link. Adaptive routing: - Each Rosetta switch monitors queue depths on output ports. - For each packet: choose output port with lowest congestion (not just shortest path). - Minimal vs non-minimal adaptive: UGAL (Universal Globally Adaptive Load-balancing) allows longer paths if they are less congested. - Result: uniform traffic spreading across all global links, near-bisection bandwidth for all-to-all MPI. **Frontier Deployment** - 74,000 compute nodes (AMD EPYC + MI250X). - 90 dragonfly+ groups × 64 ports per group = 5760 inter-group links. - MPI allreduce performance: near-linear scaling to 74K nodes for bandwidth-bound collectives. - Slingshot vs InfiniBand: Ethernet compatibility (standard switches usable for storage/management), vs IB's lower latency and native RDMA. **Software Integration** - libfabric CXI provider: RDMA semantics over Slingshot, used by OpenMPI, MPICH, SHMEM. - PMI (Process Management Interface): job launch and rank-to-node mapping. - NUMA-aware allocation: HPE PBS/SLURM integration for Slingshot topology-aware job placement. HPE Slingshot is **the network fabric that enables exascale computation by combining the cost and compatibility benefits of Ethernet with the performance and congestion management of purpose-built HPC interconnects — proving that a dragonfly+ topology with adaptive routing can deliver near-theoretical bisection bandwidth to tens of thousands of GPU-accelerated nodes**.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account