AVIZ details AI fabric simulation in NVIDIA DSX Air
AVIZ describes a simulation environment in NVIDIA DSX Air that lets teams model an AI fabric before physical hardware is installed. The post focuses on design, tenant allocation, telemetry, and alerting, with an emphasis on operational readiness for enterprise network teams.
Research Overview
The blog says an AI factory requires more than connecting GPU servers to high-speed switches. It frames the network as a programmable system that must support GPU traffic, storage, front-end services, multi-tenancy, and availability requirements.
To illustrate that model, AVIZ uses ONES inside NVIDIA DSX Air to simulate a Spectrum-X-based AI infrastructure. The simulated lab includes separate east-west and north-south fabrics, along with a management network and a centralized platform for design and operations.
Technical Breakdown
The east-west fabric in the lab links simulated GPU nodes through leaf and spine switches for GPU-to-GPU traffic. The north-south fabric connects those nodes to front-end, storage, firewall, and external services through compute leaf, front-end spine, border leaf, and storage leaf switches.
AVIZ says Day 0 of the workflow covers validated design and deployment, including fabric definition, configuration generation, cabling, addressing, routing, and infrastructure dependencies. Day 1 covers tenant creation and GPU allocation, with ONES coordinating connectivity and isolation.
Operational Impact
The Day 2 phase focuses on observability, rules, and alerts. ONES collects agentless telemetry through gNMI and NVIDIA NVUE API, and operators can view inventory, metrics, topology, and operational status in one interface.
The blog says configurable rules and alerts are intended to help teams detect network degradation that may show up as longer training times, lower GPU utilization, application latency, or storage delays. AVIZ also says simulation can help teams test workflows, train staff, and assess API-based integration before production deployment.
Leadership Perspective
The post argues that simulation can reduce risk in AI infrastructure programs by allowing teams to validate topology, provisioning, tenant isolation, telemetry, and alert thresholds before equipment arrives. It also says the approach supports coordinated automation across GPU schedulers, Kubernetes platforms, cloud management systems, service portals, and IT operations tools.
AVIZ says the broader objective is to let infrastructure teams experience the full lifecycle of an AI fabric in one environment, from design through operations. This Blog Signals brief is a fact-based summary of the vendor blog.