Skip to main content

NVIDIA ONES 3.1 details real-time telemetry across GPUs, NICs, and storage

Companies mentioned

ONES 3.1 adds real-time observability for GPU-accelerated AI and HPC workloads by collecting and correlating telemetry across NICs, compute, memory, PCIe, and storage to help operators identify latency, loss, and throttling causes.

Research Overview

The vendor describes ONES 3.1 as a unified approach to monitoring that correlates host, accelerator, and network telemetry along the data path from PCIe and memory through link layers.

The stated goal is to provide a single operational view intended to support locating bottlenecks and monitoring system health across multi-vendor environments.

Key Findings

The release focuses on interface-level NIC visibility, GPU and host monitoring, and storage and platform health metrics, with correlation intended to isolate latency, loss, and throttling sources.

It also emphasizes tracking network and compute counters over time windows to surface hot spots and imbalance across nodes, devices, and time periods.

Technical Breakdown

For network links, ONES 3.1 describes visibility into administrative and operational status, MTU, port speed, and auto-negotiation, along with tracking FEC modes and LLDP counters such as TX, RX, and discards.

For GPU observability, the update states that it uses NVIDIA SMI to capture GPU temperature, utilization, power draw, memory allocation, bus ID, and serial number, and then correlate power and thermal spikes with workload phases.

For compute health, it describes continuous monitoring of CPU utilization, memory pressure, temperatures, platform metadata, and uptime with threshold and trend views for early warning related to thermal throttling and resource starvation.

For storage, ONES 3.1 describes disk health and utilization metrics including used percentage, absolute usage in MB, temperature, and health state, alongside chassis and platform indicators for node readiness.

Operational Impact

The vendor states ONES 3.1 supports collection through standard Linux interfaces and provides multi-NIC-vendor monitoring, with the example of Intel and Mellanox/NVIDIA.

The update also describes centralized monitoring across servers hosting GPUs and NICs as designed to avoid overloading resources and to avoid locking monitoring to a single hardware stack.

Blog Signals brief is a fact-based summary of the vendor blog, highlighting how ONES 3.1 correlates real-time telemetry across NICs, compute, GPUs, and storage for performance and health monitoring in multi-vendor AI and HPC clusters.

Blog post, originally published at aviznetworks.com.

Graph Connections

3 companies named across 8 categories, one of 518 sources referencing Nvidia. Previous coverage: Aviz Networks Details OPB and ASN for Outage Visibility (July).

  • Think launches Think Grid™, a new alternative to hyperscaler AI compute

    Decision Insights Signals · August 31, 2026

    Think Grid is a hosted, dedicated AI compute service in Riyadh that provides bare-metal NVIDIA PRO 6000 Blackwell accelerators with ILM orchestration. The announcement states an all-inclusive monthly subscription with 27% lower pricing versus comparable hyperscaler dedicated configurations, including storage, support, and no egress charges.

  • AVIZ ONES in NVIDIA DSX Air outlines AI fabric lifecycle simulation

    Decision Insights Coverage · August 27, 2026

    AVIZ ONES in NVIDIA DSX Air simulates AI fabric lifecycle: design, deploy, multi-tenant GPU allocation, telemetry, rules, alerts.

  • Aviz Networks Details ONES Forecast in ONES 4.3

    Decision Insights Coverage · August 27, 2026

    ONES Forecast in ONES 4.3 projects short-term metric trends from existing monitoring, displaying dashed forecast lines on ONES dashboards.