Skip to main content

Rafay outlines AI factory stack with Aviz ONES and Spectrum-X

Companies mentioned

GPU infrastructure is the focus of a vendor blog about Rafay Platform, Aviz ONES, and NVIDIA Spectrum-X, which are presented together as a fabric-to-cloud stack for AI factories and neoclouds. The post says the combination is meant to improve multi-tenant service delivery, governance, observability, and GPU utilization for enterprise AI workloads. ## Research Overview The blog frames AI factories as a newer operating model for enterprises and neocloud providers, but says raw GPU capacity alone does not address service delivery needs. It positions the three products as separate layers of one architecture: Spectrum-X for Ethernet fabric, Aviz ONES for fabric orchestration, and Rafay for workload scheduling and governance. The post says the gap is between GPU infrastructure and enterprise-grade, multi-tenant delivery. It presents the partnership as a way to connect networking, orchestration, and workload control across private, hybrid, and sovereign environments. ## Key Findings Rafay Platform is described as the scheduler, governance layer, and scaling platform for AI factories and neoclouds. The blog says it manages deployment, scaling, upgrades, and policy controls while supporting tenant isolation and cost tracking. Aviz ONES and Rafay are described as handling GPU-aware orchestration, self-service access, telemetry, and billing support. NVIDIA Spectrum-X is presented as the lossless Ethernet foundation that supports the rest of the stack. ## Technical Breakdown The blog lists lifecycle orchestration, secure multi-tenancy, policy and cost governance, GPU-aware allocation, workload orchestration, and ecosystem integration as core capabilities. It says the platform works with tools including NVIDIA NIM, Run:AI, Kubeflow, Ray, and Jupyter. A table in the post says Aviz ONES and Rafay together provide predictable GPU performance, network and Kubernetes isolation, unified telemetry, and self-service GPU consumption. It also says Rafay can orchestrate governed compute across bare metal, Kubernetes, and virtual machines. ### How the stack is described Spectrum-X is described as a lossless, congestion-free Ethernet fabric with SuperNICs. Aviz ONES handles GPU-aware fabric orchestration and tenant segmentation, while Rafay manages networked GPUs as governed compute. The blog says the architecture is intended to support AI/ML training, inference, enterprise AI factories, and GPU platform-as-a-service use cases. It also says usage metrics can support billing and chargebacks across tenants. ## Operational Impact The blog says the combined stack is designed to let platform teams move GPU infrastructure from fixed cost management to a service model that can be billed and attributed. It says self-service catalogs let tenants request GPU clusters and AI workbenches without manual intervention. The post also says the stack is meant to improve utilization and delivery speed across the AI infrastructure layer. It presents observability as a single view of link health, ECMP balance, workload telemetry, and GPU utilization. ## Conclusion The blog presents Rafay, Aviz ONES, and NVIDIA Spectrum-X as a combined architecture for governing GPU infrastructure across AI factories and neoclouds. For enterprise leaders, the post centers on multi-tenancy, cost attribution, orchestration, and observability. This Blog Signals brief is a fact-based summary of the vendor blog.

Graph Connections

3 This is Nvidia's 88th mention on Decision Insights this quarter, following coverage of its Amazon Web Services and NVIDIA Expand Collaboration to Add GPUs and Vera CPUs in August.