Private AI inference platform for local LLM, RAG, coding, document intelligence, and multimodal services, with xSONiC deployment support across compute, network, storage, and validation.
One strongest deployable representative per model family, including parameter scale, active-parameter profile, and the workload each target helps prove before token/s validation.
Model deployment proofStrongest deployable representative per model family. Each card shows one top-end target only. Final support still depends on selected precision, context length, serving framework, GPU memory footprint, and licensing path.
Z.AI GLMAgentic engineering
Strongest deployable targetGLM-5.2 Latest GLM flagship for long-horizon private agent, coding, and knowledge-work validation.
744B total40B active1M context
Use as the GLM family top-end target when validating 1M-context workflows, tool use, and local engineering agents.
DeepSeekReasoning / code
Strongest deployable targetDeepSeek-V4-Pro Open-weight MoE target for high-end reasoning, coding, STEM, and agentic workload validation.
1.6T total49B active1M context
Use for the heaviest DeepSeek deployment proof with tensor-parallel serving and selected low-precision weights.
Alibaba QwenMultilingual / coding
Strongest deployable targetQwen3.5-397B-A17B Top Qwen deployable representative for multilingual agents, coding, vision-language, and long-context tests.
397B total17B active262K context
Use as the Qwen flagship target for Chinese-English enterprise copilots, document agents, and private RAG.
Mistral AIEnterprise LLM
Strongest deployable targetMistral Large 3 Mistral flagship open-weight MoE for multilingual, multimodal, and enterprise customization trials.
675B total41B activeApache 2.0
Use when validating European open-model deployments, document intelligence, multilingual chat, and tool workflows.
Meta LlamaAssistant / RAG
Strongest deployable targetLlama-4-Maverick-17B-128E-Instruct Strongest released Llama 4 deployable representative, using MoE scale with efficient active parameters.
400B total17B activeOpen weights
Use for open-weight assistant, multimodal RAG, and enterprise copilot comparisons against Qwen, GLM, and DeepSeek.
Google GemmaVLM benchmark
Strongest deployable targetGemma 4 31B Gemma 4 high-end dense target for compact open multimodal reasoning and image-text validation.
31B denseMultimodalOpen model
Use for lower-footprint VLM testing, document review, image-text Q&A, and edge-to-server comparison baselines.
Microsoft PhiSmall VLM
Strongest deployable targetPhi-4-reasoning-vision-15B Phi family top compact reasoning-vision target for fast multimodal validation.
15B paramsVision reasoningOpen-weight
Use for efficient local assistants, math and science reasoning with images, UI understanding, and latency baselines.
Jina EmbeddingsRAG retrieval
Strongest deployable targetjina-embeddings-v4 Jina flagship multimodal embedding model for visually rich documents, text, image, and mixed-media retrieval.
3.8B params32K contextText + image
Use as the strongest retrieval-layer target before sizing vector search, reranking, and private RAG ingestion.
Deployment Context
Where this platform fits.
Review positioning, capability notes, and deployment guidance for this xSONiC platform.
Overview
xSONiC AI Inference Server is for teams that want to run AI services inside infrastructure they control, not only through a public API. It combines AMD Instinct MI355X and MI300X platform options with xSONiC help for sizing, deployment, and handover.
MI300X UBB 2.0 accelerator platform option for physical product reference.
8AMD Instinct OAM GPUsMI355X or MI300X platform options
2.304 TBHBM3E GPU memoryMI355X option, 288 GB per GPU
8 TB/sMemory bandwidth per GPUMI355X option, workload dependent
PCIe Gen5Host I/O path8 x16 connections to host CPU
Run locallyPrivate LLM and RAGKeep prompts, documents, embeddings, outputs, and logs inside controlled infrastructure.
Size correctlyModel-led validationReview model size, precision, context length, concurrency, and latency before final configuration.
Deploy as a stackServer, network, storageConnect GPU compute with switching, optics, NVMe storage, visibility, and site readiness.
Operate with evidenceToken/s by workloadPublish throughput after target-model testing, not as a single generic GPU number.
Token Throughput Planning
Token/s should be treated as a validation result for the chosen model and serving stack. xSONiC can size the platform after the workload profile is known.
Output token/sMeasured per target modelChanges with model size, precision, batch size, context length, and serving framework.
Prefill token/sMeasured with prompt profileLong context, RAG prompts, and document workloads can shift the bottleneck.
Concurrent usersSized from latency targetReviewed with p50/p95 response time, queueing behavior, and request mix.
Cost per tokenEstimated after validationRequires utilization, power, cooling, operations, and lifecycle assumptions.
Integrated through xSONiC AI fabric and switching design
Forwarding Rate
Platform dependent
OS Version
xSONiC validated platform software with ROCm ecosystem support
Protocols
PCIe Gen5, AMD Infinity Fabric, Ethernet, RoCE, Kubernetes-ready service integration
Management
BMC, CLI/API, Telemetry, Deployment and lifecycle service options
Buying FAQ
Procurement questions.
Short answers for buyers comparing fit, support, and quote requirements before contacting xSONiC.
What is xSONiC AI Inference Server used for?
xSONiC AI Inference Server is a quote-only xSONiC product in the AI Infrastructure family. It is positioned for private LLM inference.
What specifications should buyers check before quoting xSONiC AI Inference Server?
Start with Category: AI Infrastructure; Rack Units: 8U; Ports: Platform-dependent PCIe Gen5 host I/O / High-speed network integration; Switching Capacity: Integrated through xSONiC AI fabric and switching design. Buyers should also confirm the deployment role, operating software profile, optics or cabling requirements, lead time, support scope, and any environment-specific constraints before purchase.
Is xSONiC AI Inference Server sold with public pricing?
No. xSONiC AI Inference Server is handled through a quote flow so xSONiC can confirm current pricing, availability, lead time, configuration, and deployment requirements for the buyer's environment.
How is xSONiC AI Inference Server supported in Australia?
xSONiC is operated by XGY Pty Ltd in Australia. Support scope, configuration assistance, documentation, and handover requirements are confirmed during quotation and delivered by the xSONiC team in Australian time zones.
Workload Mapping
Workload Fit.
Map common private AI workloads to the sizing signal that should drive validation after the hardware specification is confirmed.
Workload
Typical use
Sizing focus
Private assistant
Internal chat and document Q&A
Users, model family, context length
Enterprise RAG
Knowledge retrieval and controlled answers
Documents, embeddings, reranker, vector store
Coding assistant
Code Q&A and engineering knowledge search
Repository size, context strategy, concurrency
Multimodal workflow
Image-text review, extraction, classification
Input size, model type, GPU memory
Internal inference API
Department or product-facing AI endpoint
TPS, p95 latency, batching, monitoring
Deployment Readiness
Deployment Readiness Review.
Bring your target model, expected user count, context length, data profile, and latency goals. xSONiC will help validate the infrastructure fit before you commit to a deployment.
Final evaluation packageValidate the infrastructure fit before deployment commitment.
Use this step to confirm model compatibility, GPU memory fit, serving framework, token throughput target, p95 latency, concurrency, storage, network, and handover requirements.
Cookie preferences. xSONiC uses essential cookies to run this site.
With your permission, optional cookies help us understand page usage and improve navigation.
Read our privacy policy