AI & Data Center · Deployment Guide · 14 May 2026

400G Spine-Leaf Switch Design for Enterprise SONiC Fabrics: A Deployment Playbook

Engineering guidance on 400G Spine-Leaf Switch Design for Enterprise SONiC Fabrics for Australian infrastructure teams, covering form factor tradeoffs, capacity.

an engineer testing enterprise open-networking switches for “400G Spine-Leaf Switch Design for Enterprise SONiC Fabrics: A Deployment Playbook”
SONiCopen networkingdata centerAI fabricEthernet

In brief

Engineering guidance on 400G Spine-Leaf Switch Design for Enterprise SONiC Fabrics for Australian infrastructure teams, covering form factor tradeoffs, capacity.

Key takeaways

  • Engineering guidance on 400G Spine-Leaf Switch Design for Enterprise SONiC Fabrics for Australian infrastructure teams, covering form factor tradeoffs, capacity.

Why 400G Spine-Leaf Matters for Enterprise SONiC Fabrics

Use 400G spine-leaf when the fabric needs predictable two-hop east-west paths and when 100G aggregation links are no longer enough for AI, storage, or dense virtualization racks. A 64-port 400GbE spine is commonly quoted as 51.2 Tb/s full-duplex switching capacity, but the real design question is more practical: how many leaf uplinks are required to keep each rack inside the target oversubscription ratio under production traffic?

Spine-leaf (also called Clos fabric) architecture eliminates these bottlenecks by giving every leaf switch a direct uplink path to every spine switch. The result: predictable hop count, deterministic latency, and horizontal scalability. When combined with Enterprise SONiC as the network operating system, buyers gain hardware-software decoupling, container-based modularity, and a Linux-native operational model that network teams already understand.

SONiC (Software for Open Networking in the Cloud) is a Linux Foundation project that runs on switches from multiple vendors and ASICs. The SONiC architecture uses Docker-based services, Redis-backed state databases, and SAI as the hardware abstraction layer between the NOS and switch silicon. That is useful for open networking, but it does not remove the need for platform validation. The ASIC SDK, SAI implementation, optics qualification, buffer behavior, and SONiC image all need to be tested together.

This playbook walks through the engineering decisions required to design, size, and deploy a 400G spine-leaf fabric on Enterprise SONiC, with specific attention to Australian market considerations such as optics lead times, local support availability, and rack power density. Treat it as an architecture review checklist before the bill of materials is locked.

Spine-Leaf Architecture Fundamentals at 400G

A 400G spine-leaf fabric replaces the traditional core-aggregation-access hierarchy with a two-tier leaf-spine topology. Every leaf switch connects to every spine switch with one or more 400GbE uplinks. Server-facing ports on leaf switches operate at 10G, 25G, 50G, or 100G depending on the workload.

The key design parameters are:

Oversubscription ratio: This is the ratio of total server-facing bandwidth to total uplink bandwidth. Common targets are 3:1 for general-purpose workloads, 2:1 for storage-intensive environments, and 1:1 for AI/ML training clusters running collective communication operations (AllReduce, AllGather) that generate heavy east-west traffic.

Spine count: The number of spine switches determines the maximum number of leaf switches in the fabric. A leaf switch with 32x 400GbE uplinks can connect to up to 32 spine switches. Each spine switch in turn needs sufficient 400GbE port density to accept one uplink from every leaf.

Leaf count: Driven by the number of server racks. Each leaf switch typically serves one rack (Top-of-Rack design). A 1U leaf switch with 48x 25GbE or 48x 100GbE server-facing ports plus 8-16x 400GbE uplinks serves a standard rack.

Fabric scale example: A fabric with 16 spines (each with 64x 400GbE ports) and 48 leaf switches (each with 8x 400GbE uplinks) provides 48 racks of connectivity with a 6:1 oversubscription ratio (assuming 48x 25GbE server ports per leaf = 1.2Tb/s server-facing vs 3.2Tb/s uplink per leaf). Increasing uplinks to 16 per leaf reduces oversubscription to 3:1.

Latency characteristics: Do not accept a single idle port-to-port latency number as proof that the fabric is suitable for RoCE or storage. Validate loaded latency, queue depth, ECMP distribution, PFC pause counters, and retransmission/error counters while traffic crosses the exact leaf-spine-leaf path that will exist in production.

The 400G speed class is the current sweet spot for enterprise spine-leaf: it delivers enough bandwidth to avoid oversubscription problems at scale while the ecosystem of optics, cables, and compatible switches is mature enough to avoid early-adopter risk.

Platform Selection: Decision Criteria for 400G Spine and Leaf Switches

Selecting the right 400G switch platform is the most consequential decision in the design process. The following criteria should guide your evaluation:

1. ASIC generation and throughput: The switch ASIC determines forwarding capacity, buffer architecture, feature support, and power efficiency. Ask for the forwarding ASIC, SAI version, SONiC image version, supported optics list, and known feature exceptions. A 400G port count alone is not enough evidence.

2. SONiC compatibility and support maturity: Not all 400G switches run SONiC equally well. Check the SONiC Foundation supported devices list for validated platforms. Look for platforms where the ASIC SAI (Switch Abstraction Interface) layer is mature, with all required features (BGP unnumbered, EVPN-VXLAN, RoCE v2, PFC, ECN) production-ready, not just lab-tested.

3. Port configuration flexibility: Evaluate the port breakout options. A 400GbE QSFP-DD or OSFP port can often break out to 4x 100GbE, 2x 200GbE, or operate as a single 400GbE link. Flexibility matters when you need to mix 100G server uplinks with 400G spine interconnects on the same platform.

4. Buffer and traffic management: For RoCE v2 workloads, PFC, ECN marking, ETS behavior, and queue telemetry are more important than raw packet forwarding claims. Evaluate shared buffer behavior, per-port buffer allocation, egress queue depth, pause duration, CNP/ECN response, and whether those counters are exposed cleanly through SONiC telemetry.

5. Power, cooling, and form factor: A 400G switch typically draws 400-800W depending on ASIC, port count, and optics. In Australian colocation facilities, power costs are a significant OpEx factor. Compare power-per-port and throughput-per-watt metrics.

6. Bare-metal vs branded options: Bare-metal switches running SONiC can lower hardware lock-in, but the engineering burden shifts to image qualification, optics qualification, and operational runbooks. Branded or commercially supported SONiC switches add pre-validated images, support contracts, and sometimes proprietary management features. The choice depends on your team’s Linux, automation, and packet-level troubleshooting capability.

7. Acceptance evidence: Require a lab report that includes the SONiC image, ASIC/SDK version, optics models, cabling type, traffic generator profile, ECMP hash distribution, PFC/ECN counters, packet loss, loaded latency, reboot behavior, and rollback procedure. Without those artifacts, the platform is not production-proven for your use case.

Decision table summary (to be populated with xSONiC-specific platforms once datasheets are approved):

CriterionSpine SwitchLeaf Switch
Minimum ASIC throughput12.8 Tb/s or higher6.4 Tb/s or higher
400GbE port count (spine)32-648-16 (uplinks)
Server-facing ports (leaf)N/A48x 25/50/100GbE
SONiC SAI maturityProduction-readyProduction-ready
RoCE v2 / PFC / ECNRequiredRequired
Buffer depthDeep (shared)Moderate
Form factor1U or 2U1U
Optics connectorOSFP or QSFP-DDQSFP-DD or QSFP28 (breakout)

Oversubscription and Port Mapping Design

Oversubscription design is where most spine-leaf fabric mistakes happen. The goal is to match fabric bandwidth to workload requirements without over-provisioning (wasting capital) or under-provisioning (causing congestion).

Step 1: Profile workload east-west traffic

  • AI/ML training (collective operations): 1:1 oversubscription target
  • Distributed storage (Ceph, vSAN, NVMe-oF): 2:1 acceptable
  • General compute (VMs, containers, microservices): 3:1 to 4:1 acceptable
  • Mixed environments: design for the most demanding workload

Step 2: Calculate per-rack bandwidth

  • Count server NIC ports and speeds per rack
  • Example: 20 servers x 2x 100GbE NICs = 4Tb/s per rack
  • Leaf switch must have at least 40x 100GbE server ports or use breakout cables

Step 3: Size uplinks

  • For 1:1 oversubscription: 4Tb/s uplink bandwidth = 10x 400GbE uplinks per leaf
  • For 2:1: 5x 400GbE uplinks per leaf
  • For 3:1: 3-4x 400GbE uplinks per leaf (round to available port count)

Step 4: Size spine count

  • Each spine switch needs one 400GbE port per leaf
  • 48 leaves with 8 uplinks each = 48 ports per spine x 8 spines = 384 total 400GbE spine ports
  • If spine switches have 64x 400GbE ports: ceil(48/64) = 1 spine switch needed per uplink pair, but you need 8 spines for 8 uplinks per leaf
  • Adjust spine count to match leaf uplink count

Step 5: Validate non-blocking fabric

  • Total spine-to-leaf bandwidth should equal or exceed total server-facing bandwidth for 1:1 designs
  • Use ECMP (Equal-Cost Multi-Path) across all spine uplinks for load distribution

Port mapping example for a 32-rack fabric:

  • 32 leaf switches, each with 48x 100GbE + 8x 400GbE
  • 8 spine switches, each with 32x 400GbE (one per leaf)
  • Per-rack server bandwidth: 48 x 100GbE = 4.8Tb/s
  • Per-rack uplink bandwidth: 8 x 400GbE = 3.2Tb/s
  • Oversubscription: 4.8 / 3.2 = 1.5:1
  • Total fabric bandwidth: 32 x 3.2Tb/s (up) + 32 x 4.8Tb/s (server) = 102.4Tb/s + 153.6Tb/s

Breakout cable planning: If server NICs are 25GbE but leaf ports are 100GbE, use 4x 25GbE breakout cables from QSFP28 to SFP28. Similarly, 400GbE OSFP or QSFP-DD ports can break out to 4x 100GbE for leaf-to-spine uplinks if needed. Factor breakout cable costs and MPO/MTP fiber complexity into your cabling plan.

Engineering FAQ

How should NVMe form factor selection be made? Start with workload profile, usable capacity, serviceability, thermal envelope, write endurance, PCIe generation, slot layout, and replacement process. U.2, E1.S, M.2, and AIC devices solve different mechanical and operational problems.

What matters more than peak sequential speed? Sustained performance, thermal throttling behaviour, write endurance, latency under load, firmware stability, power-loss protection, and fleet manageability usually matter more than a single benchmark number.

How should storage be validated for AI or cloud workloads? Test the selected form factor in the real chassis with expected airflow, queue depth, write mix, temperature range, and monitoring stack. Validation should include steady-state and recovery behaviour, not only fresh-drive performance.

Sources Reviewed

Product fit

Where xSONiC fits

xSONiC can help validate the switch, optics, software image, telemetry, and support assumptions against the actual deployment before a production order is released.

Continue reading

Related articles