Skip to main content

Rafay Platform details workload orchestration for AI factories

Companies mentioned

Rafay Platform is presented as a multi-tenant workload scheduler and governance layer that connects NVIDIA Spectrum-X lossless Ethernet fabric and Aviz ONES GPU-aware orchestration to deliver governed GPU services for AI factories and neoclouds.

Research Overview

The post frames AI factories as requiring more than GPU capacity, citing a gap between available GPU resources and enterprise-ready multi-tenant service delivery. It positions the Rafay Platform, Aviz ONES, and NVIDIA Spectrum-X partnership as an end-to-end fabric-to-cloud approach intended to cover network, GPU orchestration, and application delivery governance.

The article describes the objective as providing a governed, secure, and observable path from fabric through to workloads, with emphasis on lifecycle management, isolation, and workload orchestration.

Key Findings

The Rafay Platform is described as integrating and orchestrating cloud-based AI fabrics with NVIDIA Spectrum-X Ethernet using lossless technology, alongside Aviz ONES fabric orchestration. The combined approach is described as providing continuum fabric-to-cloud and aiming to improve utilization and acceleration of AI workloads for platform teams.

The post also links the integration to cost governance and self-service GPU consumption, describing usage metrics exposure for billing and chargebacks and catalogs for on-demand access to GPU clusters and AI workbenches on a per-tenant basis.

Technical Breakdown

The article assigns roles across the stack: NVIDIA Spectrum-X provides a lossless, congestion-free Ethernet fabric foundation with SuperNICs, while Aviz ONES provides GPU-aware fabric orchestration, tenant segmentation, and lifecycle automation. Rafay Platform is described as orchestrating networked GPUs into governed compute across bare metal and Kubernetes or via VMs as a Service.

Rafay Platform capabilities are described across lifecycle orchestration for private, hybrid, and sovereign environments; secure multi-tenancy using RBAC, policy controls, and audit trails; policy and cost governance with usage metrics; and GPU-aware allocation through dynamic binding for deterministic performance and utilization.

Operational Impact

The post states that telemetry and observability come from Aviz ONES, describing a single view for fabric and workload telemetry including link health, ECMP balance, and GPU utilization. It also describes self-service GPU consumption as on-demand access through governed, single-click catalogs intended to reduce manual platform-admin intervention.

For workload types, the article lists AI/ML training and inference, enterprise AI factory deployments, and GPU Platform-as-a-Service use cases, while describing integration across the AI pipeline with NVIDIA NIM, Run:AI, Kubeflow, Ray, and Jupyter.

This blog content describes a fabric-to-cloud stack in which NVIDIA Spectrum-X, Aviz ONES, and Rafay Platform work together to support governed multi-tenant orchestration, telemetry, and on-demand GPU service delivery for AI factories and neoclouds. Blog Signals brief is a fact-based summary of the vendor blog.

Source: aviznetworks.com, by Ram Mohan Hariprasad.

Graph Connections

4 This is Nvidia's 93rd mention on Decision Insights this quarter, following coverage of its AVIZ details AI fabric simulation in NVIDIA DSX Air in August.