Skip to content

GPU Process Exporter

GPU Process Exporter is a Kubex Prometheus exporter for Kubernetes nodes with NVIDIA GPUs. It reads per-process GPU data through NVIDIA Management Library, maps host PIDs back to Kubernetes pods and containers, and exposes container-labelled metrics.

The project is owned by Evenkeel Inc. d/b/a Kubex and licensed under Apache License 2.0. The first supported open-source release is v1.1.0. Earlier image tags are pre-open-source builds and are supported by Kubex but do not have their source open.

When to use it

Use this exporter when node-level or device-level GPU metrics are not enough and you need per-container attribution. It is designed for Linux Kubernetes nodes running NVIDIA GPUs and a container runtime that exposes host process and container metadata to privileged DaemonSet pods.

DCGM Exporter is still useful for device health and broad GPU telemetry. This exporter focuses on the process-to-container mapping path.

Assumptions

  • NVIDIA GPU nodes with NVML available on the host.
  • Linux nodes.
  • Kubernetes pods and containers visible through the host /proc tree.
  • A deployment model that can mount host paths and run with the permissions needed to inspect host processes.
  • Container IDs that can be resolved from cgroup data under the mounted host /proc tree.

Helm

Helm is the easiest way to use this exporter. Get the chart here.

Configuration

VariableDefaultRequiredDescription
NODE_NAMEnoneyesKubernetes node name to use in metric labels and pod lookups.
HOST_PROC_MOUNT_POINT/host/procnoHost /proc mount inside the exporter container.
DRIVER_POLL_INTERVAL1snoHow often the exporter polls NVML.
SCRAPE_INTERVAL60snoExporter's expected matching interval for Prometheus scrapes. It does not configure Prometheus and must be a multiple of DRIVER_POLL_INTERVAL.
EXPORTER_ENDPOINT/metricsnoHTTP path for Prometheus metrics.
EXPORTER_PORT9494noMetrics port. Must be between 1024 and 65535.
NVML_SEARCH_PATHauto-detectnoHost path containing libnvidia-ml.so if auto-detection does not find it.

Metrics

All metrics use Kubernetes container labels. Per-device metrics also include GPU UUID and model.

Container labels:

  • node
  • namespace
  • pod
  • container
  • container_id
  • gpu_allocation_type

Per-device labels add:

  • gpu_uuid
  • gpu_model

Exported metrics:

MetricTypeDescription
kubex_gpu_container_requestsgaugeRequested GPUs or GPU fractions.
kubex_gpu_container_limitsgaugeGPU limits or GPU fractions.
kubex_gpu_container_memory_bytesgaugeGPU memory used by the container on a GPU.
kubex_gpu_container_memory_total_bytesgaugeTotal memory for the GPU.
kubex_gpu_container_memory_footprint_percentgaugeContainer memory use as a percent of GPU memory.
kubex_gpu_container_sm_utilization_percent_seconds_totalcounterAccumulated SM utilization.
kubex_gpu_container_memory_utilization_percent_seconds_totalcounterAccumulated memory r/w activity utilization.
kubex_gpu_container_enc_utilization_percent_seconds_totalcounterAccumulated encoder utilization.
kubex_gpu_container_dec_utilization_percent_seconds_totalcounterAccumulated decoder utilization.
kubex_gpu_container_ofa_utilization_percent_seconds_totalcounterAccumulated OFA utilization.
kubex_gpu_container_jpg_utilization_percent_seconds_totalcounterAccumulated JPG utilization.
kubex_gpu_container_protected_memory_bytesgaugeProtected memory reported by NVML.
kubex_gpu_container_accounting_gpu_percentgaugeNVML accounting GPU utilization.
kubex_gpu_container_accounting_memory_percentgaugeNVML accounting memory utilization.
kubex_gpu_container_accounting_max_memory_bytesgaugeMaximum memory from NVML accounting.
kubex_gpu_container_accounting_time_usgaugeAccounting runtime in microseconds.
kubex_gpu_container_accounting_start_time_usgaugeAccounting start time in microseconds since epoch.
kubex_gpu_container_accounting_is_runninggauge1 if the process is still running, 0 otherwise.

Grafana dashboard

Import grafana/gpu-process-exporter-dashboard.json into Grafana 10 or newer:

  1. Open Dashboards and select New > Import.
  2. Upload the JSON file, or paste its contents.
  3. Select the Prometheus datasource used to scrape GPU Process Exporter when Grafana asks for Prometheus.
  4. Select Import.

The dashboard starts with the last hour of data and refreshes every minute. Its workload filters are populated from kubex_gpu_container_memory_total_bytes and support node, namespace, pod, container, and GPU UUID selections.

The utilization panels use Grafana's $__rate_interval. Configure Prometheus separately with the scrape interval used for the exporter, then set SCRAPE_INTERVAL to that same value. A mismatch can make rate queries return gaps or no data. SCRAPE_INTERVAL only sets the exporter's expected matching interval and defaults to 60 seconds.

Build and test

go test ./...
go test -race ./...
go build ./...

Build local binaries for amd64 and arm64:

./build.sh

Build a container image without pushing:

DOCKERHUB_TAG=1.1.0 ./build-docker-image.sh

Push requires an explicit confirmation:

PUSH_IMAGE=true CONFIRM_PUSH=yes DOCKERHUB_TAG=1.1.0 ./build-docker-image.sh

Container runtime requirements

A Kubernetes deployment must mount at least:

  • Host /proc at /host/proc, or set HOST_PROC_MOUNT_POINT to the mount path you use.
  • Host root at /host/root, so entrypoint.sh can locate NVML libraries.

The pod needs permissions to read host process metadata and NVIDIA driver libraries. In practice this usually means a privileged DaemonSet or an equivalent security context with the required host mounts.

Support

Support is best effort. Kubex does not provide an SLA or a backport promise for this exporter unless otherwise agreed.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages