Skip to main content

AVIZ ONES in NVIDIA DSX Air outlines AI fabric lifecycle simulation

AVIZ ONES running inside NVIDIA DSX Air is presented as a simulated environment for validating and operating an AI fabric across design, deployment, multi-tenant provisioning, and observability. The approach targets network teams managing GPU east-west traffic, north-south connectivity, and tenant isolation under availability demands.

Research Overview

The post frames AI networking as more than connectivity between GPU servers, arguing that the network must be designed, deployed, segmented, and operated with automation aligned to compute and application layers. It describes increasing enterprise requirements as AI clusters expand, including high-bandwidth GPU communication, storage connectivity, multi-tenancy, and stricter availability needs.

To support earlier lifecycle validation, the vendor describes an AI fabric simulation built with AVIZ ONES inside NVIDIA DSX Air, intended to let infrastructure teams validate the operational workflow before production hardware is deployed. The lab is described as a compact NVIDIA Spectrum-X-based setup with logically separate east-west and north-south network fabrics.

Technical Breakdown

The simulated east-west fabric uses leaf and spine switching to connect simulated GPU compute nodes, representing the path for GPU-to-GPU communication used in distributed training and inference. The simulated north-south fabric connects the same compute nodes to frontend, storage, firewall, and external services via a set of compute leaf, frontend spine, border leaf, and storage leaf switches.

A dedicated management network is described for out-of-band access, while AVIZ ONES is positioned as the centralized design, automation, and operations platform. Although the environment is simulated, the post states that the operational workflow reflects how teams deploy and manage a physical AI fabric.

Key Findings: Day 0, Day 1, and Day 2

The lab is described as guiding users through three operational stages for AI infrastructure: Day 0, Day 1, and Day 2. Day 0 focuses on validated fabric design-to-deploy, using the NVIDIA Spectrum-X reference architecture as a base for network design.

For Day 0, the post says users define the fabric through ONES graphical tooling, which generates validated switch configurations and compute-facing connection parameters. It also states that this enables predeployment checks of cabling model, addressing plan, routing relationships, and infrastructure dependencies before physical rollout.

Day 1 is described as multi-tenant GPU infrastructure, where users create a tenant and allocate GPU resources from shared AI infrastructure. The post states that ONES coordinates the network configuration needed for tenant connectivity and isolation, and that the model links compute allocation and network segmentation within the same workflow.

Day 2 is described as observability, rules, and alerts, covering ongoing visibility into device health, interface status, topology, utilization, and network behavior. The lab is described as collecting agentless telemetry via interfaces such as gNMI and NVIDIA NVUE API, with centralized viewing of inventory, time-series metrics, topology, and operational status.

The post also describes configurable monitoring rules and alerts, intended to move beyond static dashboards by defining conditions requiring operational attention. It states that network degradation may not always produce an outage and can show up as increased training time, reduced GPU utilization, application latency, storage delays, or inconsistent workload performance.

Operational Impact and Integration

The post argues that simulation can answer deployment questions earlier because AI infrastructure is expensive and deployment windows can be compressed. It lists example questions around matching the proposed topology to the reference design, consistent provisioning, GPU-to-tenant assignment, tenant isolation implementation, available telemetry, alert thresholds, and API integration with external orchestration platforms.

It further describes an integration model based on ONES REST APIs for tenant lifecycle operations, including creating, retrieving, updating, and deleting tenant resources. The post states that requests can originate from external platforms and that ONES translates tenant or workload intent into required network configuration while continuously monitoring the infrastructure results.

The overall emphasis is on experiencing the end-to-end operational lifecycle of an AI fabric in a single simulated environment, including design, deployment, GPU resource allocation, tenant segmentation, device onboarding for monitoring, telemetry review, alert configuration, and API-driven operations. This “Blog Signals brief” is a fact-based summary of the vendor blog.

Source: aviznetworks.com, by Kasinath Rajendran.

Graph Connections

2This is Nvidia's 97th mention on Decision Insights this quarter, following coverage of its NVIDIA ONES 4.3.1 details UFM and NMX-C integrations plus ONES Forecast in August.