Breaking
SecurityDeveloping Story

Nvidia Bug Leaves GPU Metrics Exposed

Researchers found thousands of GPU servers leaking monitoring data online, with some vulnerable to a high-severity flaw.

··4 hours ago·4 min read
Close-up of server cooling fans in a vibrant data center
Photo by Winston Chen on Unsplash

GPU telemetry was never meant to be public. Yet researchers scanning the internet found thousands of servers broadcasting detailed hardware and performance data in the clear, and a subset of them running a monitoring service with a high-severity flaw that could let unauthenticated attackers disrupt AI workloads.

What DCGM Exporter Actually Leaks

DCGM Exporters read telemetry from the GPUs on a host: hardware details, utilization, memory usage, power consumption, and error events. Each GPU carries a unique ID, or UUID, and all of these metrics are exposed in plaintext over HTTP. That exposure hands would-be attackers a detailed reconnaissance picture, including the ability to map GPU infrastructure, identify potentially vulnerable systems, and monitor workload activity.

Michael Katchinskiy and the Lava Discovery

Michael Katchinskiy, a researcher at datacenter security startup Lava, found and reported the bug in the GPU health and performance monitoring service. In September, Nvidia released a fix for the flaw, tracked as CVE-2026-47483, and gave it an 8.2 CVSS high-severity rating.

"Once we realized how much these endpoints revealed, the next question was: How many of them are exposed to the internet?"

— Michael Katchinskiy, researcher at Lava

Scanning the Open Internet for GPUs

Katchinskiy said the researchers started scanning the internet for exposed DCGM Exporters, and the scale of exposure proved "especially significant." Over the course of four scans between March and May, the threat hunters found about 2,100 GPU servers exposing DCGM Exporter metrics to the open internet. These included 12,000 GPU UUIDs, and none of them required authentication.

Where the Exposed Hardware Sits

The hosts belonged to about 300 organizations, according to Katchinskiy, and nearly half — 5,274 of the exposed GPUs, or 44 percent of the total — were located in the US. These GPUs represented about $100 million in hardware, and included Nvidia Blackwell Ultra B300 GPUs, H200s, and H100s used to run large-scale AI workloads, plus consumer RTX 5090 and 4090 systems.

The Profiler and the Crash Path

While investigating the exposed systems, the Lava team found that about 25 percent of the exposed DCGM hosts also revealed data from Go's /debug/pprof/ built-in profiling tool. That profiler collects and exposes runtime performance data such as CPU and memory usage for running Go applications, including CPU usage, memory allocations, goroutine states, and blocking events.

That combination creates a denial-of-service path. Katchinskiy described the mechanics plainly:

"With enough concurrent unauthenticated requests, the exporter could run out of memory and crash, cutting off visibility into GPU health and activity."

The CPU and memory pressure could also affect AI training or inference workloads. Nvidia fixed the issue in version 4.8.2, and operators should upgrade to that version or later.

Node Exporter Adds More Exposure

Beyond GPU telemetry, Lava looked into Prometheus Node Exporter, which monitors server hardware and operating systems and also exposes metrics over HTTP. The team found 12,096 public Node Exporter hosts exposing data on server models, operating systems, firmware versions, hostnames, storage paths, and networking hardware commonly used in GPU clusters. That information reveals how environments are built and configured, which could also be used by attackers for reconnaissance, matching systems to known vulnerabilities.

Providers Named and Notified

The publicly exposed monitoring services affected customer infrastructure across neocloud and GPU cloud providers including Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean. Lava reported all of this to the affected providers, and Katchinskiy says those providers worked with customers to address the exposures.

For operators, the security shop says Nvidia DCGM Exporter, Node Exporter, and Prometheus services should not be directly reachable from the public internet, and recommends restricting them to authorized monitoring infrastructure.

The Numbers Behind the Exposure

  • 2,100 GPU servers found exposing DCGM Exporter metrics to the open internet across four scans between March and May
  • 12,000 GPU UUIDs included in those exposed metrics, none requiring authentication
  • 300 organizations the hosts belonged to, per Katchinskiy
  • 5,274 exposed GPUs, or 44 percent of the total, located in the US
  • $100 million in hardware represented by the exposed GPUs
  • 25 percent of exposed DCGM hosts also revealed data from Go's /debug/pprof/ profiler
  • 12,096 public Node Exporter hosts leaking server configuration data
  • 8.2 CVSS rating assigned to CVE-2026-47483

Why It Matters for AI Infrastructure

The exposure puts a spotlight on a gap between how much organizations spend on GPUs and how much attention they pay to the monitoring software attached to them. Katchinskiy framed the problem directly:

"The findings highlight a growing security gap in AI infrastructure: companies are spending millions on GPUs while leaving critical systems exposed. Those exposures can reveal how AI environments are built and, in some cases, allow attackers to disrupt them."

For businesses running AI workloads, the practical takeaway is that patching CVE-2026-47483 alone doesn't close the door. The metrics themselves remain readable in plaintext over HTTP, and the reconnaissance value of that data persists regardless of whether the exporter is updated. Operators who treat monitoring endpoints as internal plumbing may be surprised to learn they are reachable from anywhere — and that the information they emit describes hardware worth millions of dollars.

The affected providers have addressed the exposures Lava flagged, according to Katchinskiy, but the underlying pattern is not specific to any one vendor. Any organization that deploys monitoring agents for convenience and forgets to restrict them faces the same problem: the tooling meant to keep systems healthy becomes a map of the environment for anyone who finds it.

#nvidia#dcgm#gpu-monitoring#vulnerability#ai-infrastructure

Sources

Iliyas

Founder & Editor, Xploitwire

This article was written and reviewed against the sources listed above before publication, under editorial policies set by Iliyas. Read our Editorial Policy →

← Back to all stories