AI & Data Center · Validation Checklist · 29 December 2025

The RFP Questions That Matter When Procuring SONiC-Based AI Fabric Switches in Australian Data Centers

A practical RFP question set for Australian teams procuring SONiC-based 400G and 800G AI fabric switches.

an engineer commissioning a high-speed Ethernet AI fabric for “The RFP Questions That Matter When Procuring SONiC-Based AI Fabric Switches in Australian...
SONiCopen networkingdata centerAI fabricEthernetautomation

In brief

A practical RFP question set for Australian teams procuring SONiC-based 400G and 800G AI fabric switches.

Key takeaways

  • A practical RFP question set for Australian teams procuring SONiC-based 400G and 800G AI fabric switches.

Why RFP Structure Matters for AI Fabric Switching in Australia

The Australian data center market is in a period of sustained investment, driven by AI workload demand, cloud region expansion, and sovereign data requirements. In a recent Open Compute Project podcast episode, Macquarie Data Centres CEO David Hirst described how AI workloads are shifting data center design from a real estate model to a chip-out model, with liquid cooling, megawatt-per-rack density, and early collaboration between hyperscalers and supply chain partners becoming essential.

This context matters for enterprise and MSP procurement teams issuing RFPs for data center switching. When the underlying workload is AI training or inference — with bursty, unpredictable traffic patterns and strict latency requirements — the RFP evaluation criteria must go beyond traditional port density and price-per-port metrics. Questions about RDMA over Converged Ethernet (RoCE v2) support, congestion notification mechanisms, data center bridging configuration, and network telemetry become foundational rather than optional.

For Australian buyers evaluating open networking alternatives alongside incumbent switch vendors, the RFP question set is the primary lever that determines whether SONiC-based switches get a fair hearing. A vague or proprietary-locked RFP — one that specifies a single-vendor NOS by name, for example — can inadvertently exclude competitive bids from vendors offering SONiC-based solutions. Conversely, a well-structured RFP that specifies capability requirements rather than brand names opens the evaluation to the best-value solution.

The RFP Process Framework: Five Steps Applied to Open Networking Procurement

Industry best practices for RFP execution follow a structured five-step process: (1) identify the need and scope, (2) draft the RFP document, (3) distribute and engage vendors, (4) evaluate proposals systematically, and (5) select and onboard the chosen vendor. For data center AI fabric switching, each step carries specific considerations.

Step 1 — Scope definition is where Australian buyers should decide whether they are procuring a proprietary stack or open to disaggregated hardware-software models. If the scope assumes a single-vendor NOS, the RFP will naturally exclude SONiC-based alternatives. If the scope defines required capabilities — BGP/EVPN, RoCE v2, PFC, ECN, INT telemetry, NETCONF/YANG programmability — then open networking vendors can compete on merit.

Step 2 — RFP drafting should include evaluation criteria that weight technical capability, operational model, and total cost of ownership separately from hardware unit price. Industry procurement guidance emphasizes that RFPs are designed to illuminate new ideas and evaluate overall value, not just secure the lowest bid. For AI fabric switching, this means testing whether each vendor can demonstrate: sub-microsecond latency for RDMA traffic, lossless fabric behavior under congestion, automated telemetry and visibility, and integration with existing or planned AI cluster orchestration.

Step 3 — Vendor engagement benefits from pre-proposal conferences where all interested vendors — including open networking suppliers — can ask clarifying questions. Distributing Q and A addenda to all vendors ensures transparency.

Step 4 — Systematic evaluation should use a weighted scoring matrix that includes NOS flexibility, community and commercial support ecosystem, upgrade and lifecycle management, and ASIC portability — factors that differentiate SONiC-based offerings from proprietary alternatives.

Step 5 — Onboarding and contract management should include provisions for proof-of-concept testing, which is particularly important for buyers evaluating a new NOS platform for the first time.

RFP Question Set: 15 Questions for Evaluating SONiC-Based 400G/800G AI Fabric Switches

The following question set is designed for Australian MSPs, integrators, and enterprise procurement teams to include in RFPs for data center AI fabric switching. These questions are vendor-neutral and can be answered by any qualified supplier, whether offering a proprietary or open networking stack.

NOS and Architecture

  1. What network operating system does the switch run, and is it based on an open-source or proprietary codebase? If open source, which upstream project and version?
  2. Describe the NOS architecture. Is it monolithic or containerized? What is the upgrade and rollback mechanism for individual components?
  3. What is the supported ecosystem of management, automation, and telemetry tools? List NETCONF/YANG, gNMI, gRPC, SNMP, and API support.

AI Fabric and RoCE Capabilities 4. Does the switch support RoCE v2 with Priority Flow Control (PFC) and Explicit Congestion Notification (ECN)? Provide configuration reference or documentation. 5. Describe the DCBX (Data Center Bridging Capability Exchange) implementation and any Fast CNP (Congestion Notification Packet) capabilities for reducing RDMA tail latency. 6. What in-band network telemetry (INT) or equivalent per-hop visibility is supported for RDMA and AI training traffic flows? 7. What is the worst-case port-to-port latency at 400G and 800G, and under what traffic conditions is this measured? 8. Describe EVPN-VXLAN support, including multi-tenancy, symmetric and asymmetric routing, and integration with BGP unnumbered underlay.

Hardware and Optics 9. What switching ASIC(s) are used, and what is the path for ASIC portability if the buyer wants to change silicon vendors in the future? 10. List available 400G and 800G port configurations, including support for QSFP-DD and OSFP form factors. 11. What optical transceiver compatibility is guaranteed? List supported transceiver types, reach, and whether third-party optics are supported without voiding support agreements.

This question set is not exhaustive. Buyers should adapt it to their specific workload profiles, scale requirements, and operational maturity.

What the Australian Market Context Adds to the Evaluation

Several factors make the Australian data center market distinct from other regions, and these should influence how procurement teams weight their RFP evaluation criteria.

Sovereign data and compliance. Australian government and critical infrastructure operators face data sovereignty requirements that can limit cloud region choices and increase demand for on-premises or locally hosted AI infrastructure. This increases the value of network visibility and telemetry — the ability to audit and monitor traffic flows within sovereign boundaries. SONiC’s containerized architecture and open telemetry stack can be an advantage here, as buyers retain full visibility into the NOS internals rather than relying on vendor-proprietary monitoring.

Power and cooling constraints. As the Macquarie Data Centres podcast discussion highlighted, Australian data centers in dense urban environments face power availability challenges. AI workloads that require liquid cooling and megawatt-per-rack designs push the boundaries of existing facility infrastructure. For network procurement, this means evaluating switch power consumption per port at 400G and 800G speeds, and whether the switch platform supports the thermal envelope of the planned facility.

Supply chain and support geography. Australia’s geographic distance from major network equipment manufacturing and support hubs means that RFP questions about local spare parts inventory, RMA turnaround times, and Australian-based technical support are not optional — they are critical. This applies equally to proprietary and open networking vendors.

AI workload growth. The OCP podcast episode featuring David Hirst described AI workloads as behaving differently from traditional cloud workloads, with bursty traffic patterns that stress network designs built for predictable east-west flows. RFP evaluation criteria should test whether each vendor can demonstrate lossless fabric behavior under realistic AI training traffic, not just theoretical line-rate throughput.

Why Open Networking Deserves a Place in the Evaluation

The SONiC project, hosted under the SONiC Foundation as a Linux Foundation project, describes itself as an open-source network operating system based on Linux that runs on switches from multiple vendors and ASICs. It offers a full suite of network functionality — including BGP and RDMA — that has been production-hardened in large-scale cloud service provider data centers. The Open Compute Project’s Networking project lists SONiC as a sub-project alongside ONIE and SAI, positioning it as a core element of the disaggregated networking stack.

For Australian enterprise buyers and MSPs, the practical implication is that SONiC-based switches offer a credible alternative to proprietary NOS platforms for data center AI fabric deployments. The key advantages — hardware-software disaggregation, multi-vendor ASIC support, containerized modularity, and community-driven development — translate into procurement outcomes like reduced vendor lock-in, competitive pricing through multi-source hardware, and operational flexibility.

However, open networking also carries procurement risks that RFP questions should surface. These include the maturity of commercial support for SONiC in the Australian market, the availability of local engineering expertise, and the buyer’s own operational readiness to manage a disaggregated stack. A well-structured RFP allows buyers to evaluate these risks transparently rather than defaulting to a proprietary vendor out of familiarity.

xSONiC’s Enterprise SONiC data center switch portfolio is positioned in this open networking category, offering switches built for spine-leaf fabrics, AI/ML clusters, and RoCE workloads at 100G, 400G, and 800G speeds. For Australian buyers issuing RFPs, xSONiC is one of several vendors that can answer the question set outlined above. Whether xSONiC is the right fit depends on how its responses compare to alternatives across the weighted evaluation criteria.

RFP Scoring Matrix

Evaluation categorySuggested weightingEvidence vendors should provide
AI fabric capability30%RoCE v2, PFC, ECN, DCBX, ECMP hashing, failure behaviour and telemetry proof for the target 400G or 800G topology
Open networking control20%SONiC image lineage, SAI version, ASIC support matrix, configuration model and upgrade or rollback workflow
Operations and support20%Australian-hours escalation, spare strategy, observability integration, automation interfaces and defect ownership
Commercial resilience15%3-to-5-year hardware, NOS, optics, licensing and support cost with no hidden per-feature dependency
Proof-of-concept quality15%Lab report using representative optics, NICs, switch SKUs, workload profile and at least 48 hours of stability evidence

The most useful RFP response is not the longest one. It is the one that names the tested switch, NOS image, ASIC, optics, NIC firmware, topology and failure cases clearly enough that an engineering team can reproduce the claim.

Engineering FAQ

What should be proven before adopting EVPN-VXLAN on SONiC for an AI fabric? Prove underlay routing, BGP sessions, VTEP behaviour, MAC/IP learning, route scale, multi-homing design, failure convergence, and observability across the target 400G or 800G ports. The overlay should be accepted as a system, not a feature checkbox.

Why does the underlay design still matter in an overlay network? EVPN-VXLAN depends on a stable routed underlay. MTU, ECMP, addressing, route policy, 9000 byte jumbo-frame policy, link failure behaviour, and telemetry determine whether the overlay remains predictable under load and during faults.

What should be included in an EVPN-VXLAN operations runbook? Include naming, IP plan, BGP policy, VNI mapping, change process, rollback commands, failure checks, telemetry fields, backup and restore steps, and escalation ownership for the selected SONiC image, then rehearse the runbook during a 48 hours proof of concept.

Sources Reviewed

Product fit

Where xSONiC fits

xSONiC can help validate the switch, optics, software image, telemetry, and support assumptions against the actual deployment before a production order is released.

Continue reading

Related articles