Skip to main content

ONES Fabric Manager outlines hardware-enforced isolation for NVIDIA GB200 and GB300 NVL72

NVIDIA GB200 and GB300 NVL72 systems use NVLink to cluster 72 Blackwell GPUs, but multi-tenant sharing requires isolating both NVLink and cross-rack InfiniBand traffic. ONES Fabric Manager centralizes GPU partitioning and InfiniBand partition coordination to support workloads that span one or multiple racks.

Research Overview

The vendor describes rack-scale “AI factories” built around large GPU clusters connected by NVLink, with GB200 and GB300 NVL72 systems grouping 72 Blackwell GPUs in a single rack-scale environment. In shared infrastructure, the core operational question is how to separate tenants so that traffic from different customers does not mix.

The post states that VLAN techniques used for Ethernet do not address isolation for NVLink, because NVLink is not treated as a traditional Ethernet network. It also highlights an additional complexity when a tenant workload extends beyond a single rack, which requires coordination between NVLink scale-up fabric and InfiniBand scale-out fabric.

Key Findings

ONES Fabric Manager is presented as a unified control plane for both tiers of connectivity, coordinating GPU partitioning within each rack and cross-rack isolation when tenants span racks. The system identifies the involved topology from the requested tenant placement and configures the needed partitions accordingly.

The post also describes operational controls for provisioning and lifecycle management, including atomic provisioning and rollback, live inventory monitoring, and automated coordination with NVIDIA UFM. It further states that the platform supports allocations down to per-GPU and per-tray levels.

Technical Breakdown

According to the post, a rack-scale GB200 or GB300 cluster involves two fabrics: NVLink within the rack for the scale-up tier and InfiniBand between racks for the scale-out tier. Dual-plane fabrics are described as providing separate rails to trays within a rack for bandwidth and reliability.

The write-up says NVIDIA provides an agent for controlling the fabric layer for each tier, while ONES Fabric Manager acts as a shared management plane across both tiers using a unified API. It further states that the system partitions NVLink for intra-rack isolation and, when needed, uses UFM to enforce a corresponding InfiniBand partition at rack boundaries.

Operational Impact

The post frames tenant provisioning as a single control-plane workflow where the caller specifies the target machines and GPUs for a tenant, rather than tenant location details like rack boundaries. It states the platform determines which racks are implied, partitions NVLink within those racks, and constrains cross-rack communication when separate racks are involved.

For isolation, the post describes hardware-enforced GPU partitions that provide each tenant full NVLink bandwidth without interconnectivity between partitions. For multi-rack tenants, it adds a private cross-rack partition that the post says is not accessible or joinable by other tenants, with UFM remaining a black-box enforcement layer from the perspective of the tenant and operator.

Leadership Perspective

The vendor content positions the approach as consolidating topology handling, isolation orchestration, and lifecycle operations into a single control plane and API across both NVLink and InfiniBand tiers. It describes the outcome as a shift away from script- and tool-based coordination toward hardware enforced isolation coupled with automated partition management.

The post lists operational behaviors aimed at fault containment and predictable failure handling, including self-healing of failed controller connections with back-off, fail-fast request behavior, and all-or-nothing provisioning that rolls back entirely on errors. It also states that each rack is self-contained so controller failure in one rack does not stop or affect other racks.

Blog Signals brief is a fact-based summary of the vendor blog describing how ONES Fabric Manager coordinates NVLink GPU partitioning and UFM-driven InfiniBand partitioning for NVIDIA GB200 and GB300 NVL72 systems, including atomic provisioning and multi-rack tenant isolation through a unified control plane.

Source: aviznetworks.com, by Nikhil Morey.

Graph Connections

2This is Nvidia's 95th mention on Decision Insights this quarter, following coverage of its Aviz Networks outlines AI fabric lifecycle management and RDMA visibility at NVIDIA GTC Berlin 2026 in August.