AI Infrastructure

xSONiC AI Inference Server

Private AI infrastructure for local LLM, RAG, multimodal, and enterprise model services.

Back to AI Infrastructure

Private AI inference platform for local LLM, RAG, coding, document intelligence, and multimodal services, with xSONiC deployment support across compute, network, storage, and validation.

  • Private AI infrastructure
  • MI355X and MI300X platform options
  • Workload-based sizing review
  • Local deployment support
8U Private AI Inference xSONiC AI Inference Server front chassis view
8 OAM GPUs 2.304 TB HBM3E option PCIe Gen5 Host I/O

Model Coverage

Deployable Model Matrix.

One strongest deployable representative per model family, including parameter scale, active-parameter profile, and the workload each target helps prove before token/s validation.

Model deployment proof Strongest deployable representative per model family. Each card shows one top-end target only. Final support still depends on selected precision, context length, serving framework, GPU memory footprint, and licensing path.
Z.AI GLMAgentic engineering
Strongest deployable target GLM-5.2 Latest GLM flagship for long-horizon private agent, coding, and knowledge-work validation.
744B total40B active1M context

Use as the GLM family top-end target when validating 1M-context workflows, tool use, and local engineering agents.

DeepSeekReasoning / code
Strongest deployable target DeepSeek-V4-Pro Open-weight MoE target for high-end reasoning, coding, STEM, and agentic workload validation.
1.6T total49B active1M context

Use for the heaviest DeepSeek deployment proof with tensor-parallel serving and selected low-precision weights.

Alibaba QwenMultilingual / coding
Strongest deployable target Qwen3.5-397B-A17B Top Qwen deployable representative for multilingual agents, coding, vision-language, and long-context tests.
397B total17B active262K context

Use as the Qwen flagship target for Chinese-English enterprise copilots, document agents, and private RAG.

Mistral AIEnterprise LLM
Strongest deployable target Mistral Large 3 Mistral flagship open-weight MoE for multilingual, multimodal, and enterprise customization trials.
675B total41B activeApache 2.0

Use when validating European open-model deployments, document intelligence, multilingual chat, and tool workflows.

Meta LlamaAssistant / RAG
Strongest deployable target Llama-4-Maverick-17B-128E-Instruct Strongest released Llama 4 deployable representative, using MoE scale with efficient active parameters.
400B total17B activeOpen weights

Use for open-weight assistant, multimodal RAG, and enterprise copilot comparisons against Qwen, GLM, and DeepSeek.

Google GemmaVLM benchmark
Strongest deployable target Gemma 4 31B Gemma 4 high-end dense target for compact open multimodal reasoning and image-text validation.
31B denseMultimodalOpen model

Use for lower-footprint VLM testing, document review, image-text Q&A, and edge-to-server comparison baselines.

Microsoft PhiSmall VLM
Strongest deployable target Phi-4-reasoning-vision-15B Phi family top compact reasoning-vision target for fast multimodal validation.
15B paramsVision reasoningOpen-weight

Use for efficient local assistants, math and science reasoning with images, UI understanding, and latency baselines.

Jina EmbeddingsRAG retrieval
Strongest deployable target jina-embeddings-v4 Jina flagship multimodal embedding model for visually rich documents, text, image, and mixed-media retrieval.
3.8B params32K contextText + image

Use as the strongest retrieval-layer target before sizing vector search, reranking, and private RAG ingestion.

Deployment Context

Where this platform fits.

Review positioning, capability notes, and deployment guidance for this xSONiC platform.

Overview

xSONiC AI Inference Server is for teams that want to run AI services inside infrastructure they control, not only through a public API. It combines AMD Instinct MI355X and MI300X platform options with xSONiC help for sizing, deployment, and handover.

MI300X UBB 2.0 accelerator platform option for xSONiC AI Inference Server
MI300X UBB 2.0 accelerator platform option for physical product reference.
8AMD Instinct OAM GPUsMI355X or MI300X platform options
2.304 TBHBM3E GPU memoryMI355X option, 288 GB per GPU
8 TB/sMemory bandwidth per GPUMI355X option, workload dependent
PCIe Gen5Host I/O path8 x16 connections to host CPU
Run locallyPrivate LLM and RAGKeep prompts, documents, embeddings, outputs, and logs inside controlled infrastructure.
Size correctlyModel-led validationReview model size, precision, context length, concurrency, and latency before final configuration.
Deploy as a stackServer, network, storageConnect GPU compute with switching, optics, NVMe storage, visibility, and site readiness.
Operate with evidenceToken/s by workloadPublish throughput after target-model testing, not as a single generic GPU number.

Token Throughput Planning

Token/s should be treated as a validation result for the chosen model and serving stack. xSONiC can size the platform after the workload profile is known.

Output token/sMeasured per target modelChanges with model size, precision, batch size, context length, and serving framework.
Prefill token/sMeasured with prompt profileLong context, RAG prompts, and document workloads can shift the bottleneck.
Concurrent usersSized from latency targetReviewed with p50/p95 response time, queueing behavior, and request mix.
Cost per tokenEstimated after validationRequires utilization, power, cooling, operations, and lifecycle assumptions.
QuestionRequired before publishing a number
How many tokens per second?Target model, precision, context length, batch policy, serving framework, latency target
How many users?Concurrent sessions, request pattern, p95 latency target, prompt length, output length
Can it run this LLM?Model size, memory footprint, framework support, quantization plan, deployment configuration

Infrastructure Stack

xSONiC positions the AI Inference Server as the compute layer of a wider private AI infrastructure stack.

xSONiC AI Inference Server angled 8U chassis product view
Angled 8U chassis view for private AI inference deployment planning.
01AI Inference ServerGPU compute for serving, evaluation, and controlled internal AI services.
02Data Center SwitchingInference node, storage, service ingress, and management connectivity.
03Optical InterconnectHigh-speed links between server, switch, storage, and aggregation layers.
04NVMe StorageModel files, vector data, document content, indexes, and workload datasets.
05Packet VisibilityTraffic patterns, service paths, and operational troubleshooting support.
06Deployment SupportHardware selection, site readiness, software bring-up, validation, and handover.

Validation Package

Inputs to prepare
  • Target model or model shortlist
  • Context length and precision target
  • Expected users, sessions, or request rate
  • Prompt and output length profile
  • RAG, embedding, reranking, or multimodal needs
  • Rack, power, cooling, switching, and storage constraints
xSONiC validates
  • GPU memory fit and serving framework compatibility
  • Token throughput, p50/p95 latency, and concurrency behavior
  • Storage, network, and system-level bottlenecks
  • Thermal, power, rack, cabling, and handover readiness
1DiscoverModel, users, data, latency
2SizeGPU memory, storage, network
3IntegrateServer, switch, optics, SSD
4ValidateToken/s, p95, concurrency
5HandoverDocs, monitoring, support path

Platform Options

Platform-level specifications remain aligned with the linked xSONiC datasheet.

8USystem classHigh-density GPU inference server platform
7 linksInfinity Fabric per GPU153.6 GB/s link bandwidth on MI355X option
FP8 / MXFP6 / MXFP4AI precision supportMI355X option for modern inference formats
1.5 TBHBM3 on MI300X option192 GB per GPU, 8-GPU platform
AreaMI355X Platform OptionMI300X Platform Option
GPU configuration8 AMD Instinct MI355X OAM GPUs on UBB 2.0 module8 AMD Instinct MI300X OAM GPUs on UBB 2.0 module
GPU memory288 GB HBM3E per GPU, approx. 2.304 TB total192 GB HBM3 per GPU, 1.5 TB total
Memory bandwidthUp to 8 TB/s per GPUUp to 5.3 TB/s per GPU
GPU-to-GPU fabric7 bidirectional AMD Infinity Fabric links per GPU at 153.6 GB/s7 bidirectional AMD Infinity Fabric links per GPU at 128 GB/s
Host I/O8 PCIe Gen5 x16 connections to host CPU8 PCIe Gen5 x16 connections with 128 GB/s per GPU scale-out network bandwidth
Precision supportFP16, BF16, FP8, MXFP6, MXFP4FP32/FP64 for HPC plus FP16/BF16/FP8/INT8 for AI
Power and coolingDirect liquid-cooled option with up to 1400 W module TBP750 W maximum TBP per GPU in the platform specification

Specification Overview

Technical specifications.

Use this table as the fast path for platform sizing, port planning, and software compatibility checks.

Category AI Infrastructure
Rack Units 8U
Ports Platform-dependent PCIe Gen5 host I/O / High-speed network integration
Switching Capacity Integrated through xSONiC AI fabric and switching design
Forwarding Rate Platform dependent
OS Version xSONiC validated platform software with ROCm ecosystem support
Protocols PCIe Gen5, AMD Infinity Fabric, Ethernet, RoCE, Kubernetes-ready service integration
Management BMC, CLI/API, Telemetry, Deployment and lifecycle service options

Buying FAQ

Procurement questions.

Short answers for buyers comparing fit, support, and quote requirements before contacting xSONiC.

What is xSONiC AI Inference Server used for?

xSONiC AI Inference Server is a quote-only xSONiC product in the AI Infrastructure family. It is positioned for private LLM inference.

What specifications should buyers check before quoting xSONiC AI Inference Server?

Start with Category: AI Infrastructure; Rack Units: 8U; Ports: Platform-dependent PCIe Gen5 host I/O / High-speed network integration; Switching Capacity: Integrated through xSONiC AI fabric and switching design. Buyers should also confirm the deployment role, operating software profile, optics or cabling requirements, lead time, support scope, and any environment-specific constraints before purchase.

Is xSONiC AI Inference Server sold with public pricing?

No. xSONiC AI Inference Server is handled through a quote flow so xSONiC can confirm current pricing, availability, lead time, configuration, and deployment requirements for the buyer's environment.

How is xSONiC AI Inference Server supported in Australia?

xSONiC is operated by XGY Pty Ltd in Australia. Support scope, configuration assistance, documentation, and handover requirements are confirmed during quotation and delivered by the xSONiC team in Australian time zones.

Workload Mapping

Workload Fit.

Map common private AI workloads to the sizing signal that should drive validation after the hardware specification is confirmed.

Workload Typical use Sizing focus
Private assistant Internal chat and document Q&A Users, model family, context length
Enterprise RAG Knowledge retrieval and controlled answers Documents, embeddings, reranker, vector store
Coding assistant Code Q&A and engineering knowledge search Repository size, context strategy, concurrency
Multimodal workflow Image-text review, extraction, classification Input size, model type, GPU memory
Internal inference API Department or product-facing AI endpoint TPS, p95 latency, batching, monitoring

Deployment Readiness

Deployment Readiness Review.

Bring your target model, expected user count, context length, data profile, and latency goals. xSONiC will help validate the infrastructure fit before you commit to a deployment.

Final evaluation package Validate the infrastructure fit before deployment commitment.

Use this step to confirm model compatibility, GPU memory fit, serving framework, token throughput target, p95 latency, concurrency, storage, network, and handover requirements.