- Private AI infrastructure
- MI355X and MI300X platform options
- Workload-based sizing review
AI Infrastructure
Private AI inference platform for local LLM, RAG, coding, document intelligence, and multimodal services, with xSONiC deployment support across compute, network, storage, and validation.
Model Coverage
One strongest deployable representative per model family, including parameter scale, active-parameter profile, and the workload each target helps prove before token/s validation.
Use as the GLM family top-end target when validating 1M-context workflows, tool use, and local engineering agents.
Use for the heaviest DeepSeek deployment proof with tensor-parallel serving and selected low-precision weights.
Use as the Qwen flagship target for Chinese-English enterprise copilots, document agents, and private RAG.
Use when validating European open-model deployments, document intelligence, multilingual chat, and tool workflows.
Use for open-weight assistant, multimodal RAG, and enterprise copilot comparisons against Qwen, GLM, and DeepSeek.
Use for lower-footprint VLM testing, document review, image-text Q&A, and edge-to-server comparison baselines.
Use for efficient local assistants, math and science reasoning with images, UI understanding, and latency baselines.
Use as the strongest retrieval-layer target before sizing vector search, reranking, and private RAG ingestion.
Deployment Context
Review positioning, capability notes, and deployment guidance for this xSONiC platform.
xSONiC AI Inference Server is for teams that want to run AI services inside infrastructure they control, not only through a public API. It combines AMD Instinct MI355X and MI300X platform options with xSONiC help for sizing, deployment, and handover.
Token/s should be treated as a validation result for the chosen model and serving stack. xSONiC can size the platform after the workload profile is known.
| Question | Required before publishing a number |
|---|---|
| How many tokens per second? | Target model, precision, context length, batch policy, serving framework, latency target |
| How many users? | Concurrent sessions, request pattern, p95 latency target, prompt length, output length |
| Can it run this LLM? | Model size, memory footprint, framework support, quantization plan, deployment configuration |
xSONiC positions the AI Inference Server as the compute layer of a wider private AI infrastructure stack.
For network planning, compare the published XS-DC-32X400-SP-G2 32x400G SONiC spine switch and XS-DC-64X800-AI-G1 64x800G AI fabric switch against the actual host-I/O, oversubscription, optics, storage, and failure-test requirements. These are shortlist candidates, not a claim that either switch is compatible with every server configuration.
Platform-level specifications remain aligned with the linked xSONiC datasheet.
| Area | MI355X Platform Option | MI300X Platform Option |
|---|---|---|
| GPU configuration | 8 AMD Instinct MI355X OAM GPUs on UBB 2.0 module | 8 AMD Instinct MI300X OAM GPUs on UBB 2.0 module |
| GPU memory | 288 GB HBM3E per GPU, approx. 2.304 TB total | 192 GB HBM3 per GPU, 1.5 TB total |
| Memory bandwidth | Up to 8 TB/s per GPU | Up to 5.3 TB/s per GPU |
| GPU-to-GPU fabric | 7 bidirectional AMD Infinity Fabric links per GPU at 153.6 GB/s | 7 bidirectional AMD Infinity Fabric links per GPU at 128 GB/s |
| Host I/O | 8 PCIe Gen5 x16 connections to host CPU | 8 PCIe Gen5 x16 connections with 128 GB/s per GPU scale-out network bandwidth |
| Precision support | FP16, BF16, FP8, MXFP6, MXFP4 | FP32/FP64 for HPC plus FP16/BF16/FP8/INT8 for AI |
| Power and cooling | Direct liquid-cooled option with up to 1400 W module TBP | 750 W maximum TBP per GPU in the platform specification |
The accelerator-level fields above can be cross-checked against AMD’s official Instinct MI355X product specification and Instinct MI300X platform data sheet. The linked xSONiC AI Inference Server datasheet records the current xSONiC evaluation scope.
Those sources establish accelerator and platform reference fields; they do not guarantee token throughput, model compatibility, latency, power, cooling, or availability for a quoted system. Confirm the exact server revision and serving stack, then measure the target model, precision, context length, concurrency, and p95 latency before procurement. Use the private AI network and storage planning guide and AI fabric solution to define the separate network evidence package.
Datasheet and procurement documentation: Download Datasheet or request a configuration review.
Specification Overview
Use this table as the fast path for platform sizing, port planning, and software compatibility checks.
| Category | AI Infrastructure |
|---|---|
| Rack Units | 8U |
| Ports | Platform-dependent PCIe Gen5 host I/O / High-speed network integration |
| Switching Capacity | Integrated through xSONiC AI fabric and switching design |
| Forwarding Rate | Platform dependent |
| OS Version | xSONiC validated platform software with ROCm ecosystem support |
| Protocols | PCIe Gen5, AMD Infinity Fabric, Ethernet, RoCE, Kubernetes-ready service integration |
| Management | BMC, CLI/API, Telemetry, Deployment and lifecycle service options |
Buying FAQ
Short answers for buyers comparing fit, support, and quote requirements before contacting xSONiC.
xSONiC AI Inference Server is a quote-only xSONiC product in the AI Infrastructure family. It is positioned for Private AI inference and accelerated compute deployments sized to the published system configuration.
Start with Category: AI Infrastructure; Rack Units: 8U; Ports: Platform-dependent PCIe Gen5 host I/O / High-speed network integration; Switching Capacity: Integrated through xSONiC AI fabric and switching design. Buyers should also confirm the deployment role, operating software profile, optics or cabling requirements, lead time, support scope, and any environment-specific constraints before purchase.
No. xSONiC AI Inference Server is handled through a quote flow so xSONiC can confirm current pricing, availability, lead time, configuration, and deployment requirements for the buyer's environment.
xSONiC is operated by XGY Pty Ltd in Australia. Support scope, configuration assistance, documentation, and handover requirements are confirmed during quotation and delivered by the xSONiC team in Australian time zones.
Workload Mapping
Map common private AI workloads to the sizing signal that should drive validation after the hardware specification is confirmed.
| Workload | Typical use | Sizing focus |
|---|---|---|
| Private assistant | Internal chat and document Q&A | Users, model family, context length |
| Enterprise RAG | Knowledge retrieval and controlled answers | Documents, embeddings, reranker, vector store |
| Coding assistant | Code Q&A and engineering knowledge search | Repository size, context strategy, concurrency |
| Multimodal workflow | Image-text review, extraction, classification | Input size, model type, GPU memory |
| Internal inference API | Department or product-facing AI endpoint | TPS, p95 latency, batching, monitoring |
Deployment Readiness
Bring your target model, expected user count, context length, data profile, and latency goals. xSONiC will help validate the infrastructure fit before you commit to a deployment.
Use this step to confirm model compatibility, GPU memory fit, serving framework, token throughput target, p95 latency, concurrency, storage, network, and handover requirements.