Nvidia GPU monitoring endpoints exposed details of 12,000 AI accelerators
Key takeaways
- 2,100 GPU servers exposed Nvidia DCGM metrics without authentication
- 44 percent of the exposed GPUs were located in the US, representing roughly $100 million in hardware
- A quarter of exposed hosts also leaked Go profiling data, allowing attackers to crash monitoring
- Operators must upgrade to DCGM Exporter 4.8.2 or later and firewall monitoring ports
Nvidia GPU monitoring endpoints exposed details of 12,000 AI accelerators
More than 2,100 GPU servers were found leaking Nvidia telemetry to the open internet, exposing the unique identifiers of roughly 12,000 accelerators and a high-severity flaw that lets unauthenticated attackers crash the monitoring service. The finding, published by datacentre security startup Lava, turns a routine observability tool into a reconnaissance goldmine for anyone mapping the world's AI compute.
The flaw itself, tracked as CVE-2026-47483, carries a CVSS rating of 8.2. Nvidia released a fix in DCGM Exporter version 4.8.2 in September. What makes the disclosure noteworthy is not the bug alone, but the scale of the exposure that surrounds it. Lava's scans between March and May found that none of the 2,100 hosts required authentication to read GPU metrics, and about 44 percent of the affected GPUs sat in the United States.
The headline numbers
The data points below come from Lava's four internet scans, its analysis of the exposed telemetry, and Nvidia's own advisory. They describe an infrastructure category that most organisations treat as internal plumbing.
| Metric | Figure |
|---|---|
| Exposed DCGM Exporter servers | ~2,100 |
| Unique GPU UUIDs visible | 12,000 |
| Organisations affected | ~300 |
| GPUs located in the US | 5,274 (44 percent) |
| Estimated hardware value exposed | ~$100 million |
| Hosts also leaking Go pprof data | ~25 percent |
| Exposed Prometheus Node Exporter hosts | 12,096 |
| CVSS severity of CVE-2026-47483 | 8.2 (high) |
| Fixed version | DCGM Exporter 4.8.2 |
The hardware mix is striking. Lava saw Nvidia Blackwell Ultra B300 GPUs, H200s and H100s, the accelerators that underpin large-scale training and inference, alongside consumer RTX 5090 and 4090 cards. A single H100 list price runs into the tens of thousands of dollars, which helps explain the $100 million estimate. This is not a story about hobbyist rigs. It is a story about production AI factories with their diagnostics ports wide open.
What the data actually shows
DCGM Exporter is a Prometheus-compatible agent that reads telemetry from Nvidia's Data Center GPU Manager: hardware health, utilisation, memory usage, power draw and error events. Each GPU carries a UUID, and the metrics are served in plaintext over HTTP. That combination is useful for operators and equally useful for anyone probing from outside.
Three patterns emerge from the Lava data. First, the exposures cluster around GPU cloud providers and neoclouds, including Nebius, Voltage Park, Lambda, Northern Data and DigitalOcean. These are companies whose entire business model is renting out accelerators, and Lava reported the findings to each, with the providers working to close the gaps. Second, the Go pprof exposure on roughly a quarter of hosts compounds the risk. The pprof profiler exposes runtime performance data such as CPU usage, memory allocations, goroutine states and blocking events. As Lava researcher Michael Katchinskiy noted, enough concurrent unauthenticated requests can exhaust memory and crash the exporter. That cuts off GPU visibility and can spill over into training or inference workloads sharing the host.
Third, and easily overlooked, are the 12,096 public Prometheus Node Exporter hosts. Node Exporter reports server-level data: models, operating systems, firmware versions, hostnames, storage paths and networking hardware. On its own, that is inventory detail. Combined with GPU UUIDs and cluster topology, it becomes a blueprint. An attacker can match a specific firmware version to a known vulnerability, or infer which workloads run where based on GPU class and node configuration. This is the reconnaissance phase of an intrusion, handed over for free.
The mapping of AI infrastructure is the quiet danger. A GPU UUID is a persistent identifier. If it appears in one scan and again six months later, an observer can track how a provider's fleet grows, which customers are expanding, and when capacity shifts between regions. That is competitive intelligence at minimum and targeting data at worst.
What the data does not tell us
The 2,100 and 12,000 figures are a floor, not a ceiling. Lava ran four scans between March and May 2026, so any host that came online after May, or was firewalled during the scan window, is invisible to the count. Internet-wide scanning also depends on address space coverage and how quickly services respond. Hosts behind load balancers or non-standard ports may have been missed entirely.
There is no data on exploitation. The presence of an exposed endpoint does not prove anyone read the metrics, and neither Lava nor Nvidia has reported in-the-wild attacks tied to CVE-2026-47483. The crash vector is real but so far theoretical in public reporting.
Attribution of the ~300 organisations is also imprecise. Lava identified providers whose infrastructure hosted exposed customers, but the customers themselves own the misconfiguration. A neocloud that gives tenants raw IP access and no default firewall is a different problem from a tenant that deliberately published metrics. The data does not cleanly separate the two.
Finally, the $100 million hardware estimate is a rough proxy based on list prices and GPU classes. It says nothing about the value of the data those GPUs process, which for a model training run could exceed the hardware cost many times over.
How this compares
Exposed monitoring endpoints are a familiar genre. Prometheus Node Exporter has been a recurring finding in internet scans for years, and the same class of misconfiguration has hit Elasticsearch, Redis, MongoDB and Docker APIs. The Centre for Internet Security and Shodan-based research have documented millions of exposed databases over the past decade.
What is new is the target. Historically, exposed telemetry belonged to web infrastructure, where the worst case was leaking traffic patterns. Here it belongs to AI compute, a category with acute supply constraints, enormous capital costs and growing geopolitical sensitivity. Earlier this year, a Californian was accused of shipping $300 million worth of Nvidia chips to China, and the US has tightened export controls repeatedly. In that context, a public map of where B300 and H200 GPUs sit is not merely an operational lapse.
The comparison with the wider memory and accelerator crunch matters too. With DRAM, HBM and NAND all short and South Korea selling $60 billion of chips in a single month, every deployed accelerator is a scarce asset. Exposing its identity and status to the internet raises the stakes on both theft-of-service and physical targeting scenarios.
Nvidia's response, a version bump and an advisory, follows the industry norm. But the DCGM issue sits alongside a broader pattern of AI infrastructure security catching up with AI infrastructure spending. Nvidia's own OpenShell and Sentry agent safety platform, which runs an AI agent kill switch on a separate chip, shows the company thinking hard about agent security. Monitoring planes have received less attention.
So what?
For operators, the action list is short. Upgrade DCGM Exporter to 4.8.2 or later. Confirm that port 9400, the default DCGM metrics port, and the Node Exporter port 9100 are not reachable from the public internet. Restrict them to authorised monitoring infrastructure, ideally over a private network or a VPN. Audit Prometheus scrape configurations for public targets, because a scrape config that points at an external address will pull data across the open internet.
For providers, the lesson is about defaults. If a tenant can expose a monitoring port without a warning, some proportion will. Sensible platforms ship with inbound rules that block metrics endpoints unless explicitly opened.
For the security industry, the finding argues for treating observability as attack surface rather than overhead. The same telemetry that helps an SRE debug a stalled training job helps an adversary decide which cluster to hit and how. Lava's number, 12,000 GPU UUIDs and roughly $100 million in hardware, is the kind of figure that should appear in board risk registers, not just security blogs.
The most important takeaway is this: the AI buildout has outpaced its operational hygiene, and the exposed endpoints are a measurable symptom. Fixing CVE-2026-47483 is a patch. Closing the port is a policy. Getting the second one right is what stops the next scan from finding 20,000 GPUs instead of 12,000.