BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage News Kubernetes Teams Get a Safer Way to See Their Own GPU Metrics

Kubernetes Teams Get a Safer Way to See Their Own GPU Metrics

Listen to this article -  0:00

Adobe engineers have described an open-source approach for giving teams self-service access to their own Prometheus metrics in multi-tenant Kubernetes environments, without exposing the shared metrics of other teams. The design is particularly aimed at GPU-heavy environments, where teams may need visibility into utilisation and power consumption to understand whether expensive accelerator capacity is actually being used.

The problem is straightforward but difficult to solve safely. A central Prometheus instance may contain metrics from thousands of namespaces, making it unsuitable for unrestricted tenant access. Giving every team query access could expose another team's data, while allowing large numbers of users to query the shared store can also create a performance and noisy-neighbour problem. Adobe's solution places a tenant-aware proxy between users and the central Prometheus and optionally gives each tenant its own smaller Prometheus instance.

The architecture uses existing Kubernetes and CNCF technologies rather than introducing another metrics platform. Requests pass through NGINX and kube-rbac-proxy for authentication and authorisation before reaching a multi-tenant Prometheus proxy. The proxy identifies the tenant, discovers available Prometheus backends, and restricts queries to that tenant's namespace.

A key component is prom-label-proxy, which modifies incoming PromQL queries to enforce a namespace constraint. This is important because simply asking developers to include the correct namespace in their queries would not provide a sufficient security boundary. The restriction is applied by the proxy before the query reaches Prometheus, preventing a tenant from deliberately constructing a query that accesses another namespace.

The platform also introduces a Kubernetes custom resource called MetricAccess, allowing teams to declare which metrics they require. Metric definitions can use exact metric names, regular expressions, or PromQL selectors. This provides a self-service mechanism while leaving the underlying collection and access controls with the platform team.

For teams requiring their own dashboards and alerts, the design can periodically remote-write a curated set of metrics into a tenant-specific Prometheus instance. With metricIsolation enabled, only the tenant's own series are collected, reducing the amount of data stored in the tenant's Prometheus as well as limiting what can be queried.

Adobe reports that one example configuration reduced a tenant's stored series from more than 10,000 to roughly 300. Apart from reducing storage and query requirements, this creates another isolation boundary: data that is never collected into the tenant's store cannot subsequently be exposed through that store.

This separation also changes the operational model. The central Prometheus continues to provide infrastructure-wide collection, while tenant-specific Prometheus instances handle the dashboards and queries that individual teams need. The shared system therefore does not have to serve every developer's routine observability queries.

The immediate use case is GPU visibility. Accelerator capacity can represent a significant infrastructure cost, but allocation alone does not indicate whether those GPUs are doing useful work. The authors describe finding a GPU that had remained at zero utilisation for 11 consecutive days despite being allocated and powered on.

With access to metrics such as GPU utilisation, framebuffer memory, power consumption and request rates, teams can build queries to identify idle GPUs, GPUs consuming power without corresponding application traffic, or workloads receiving traffic while their GPU capacity remains underutilised.

Although GPUs provide the most obvious example, the architecture is not GPU-specific. The underlying pattern is a tenant-aware observability layer built around authentication, query isolation, curated metric access, and optional per-tenant storage.

This is increasingly relevant as Kubernetes clusters become shared platforms for application teams, data workloads, and AI workloads. Platform teams need to provide sufficient observability for developers without turning a central telemetry system into either a security risk or a performance bottleneck.

There are several established approaches to the same underlying multi-tenancy problem. Grafana Mimir, for example, has multi-tenant isolation built into its architecture, using tenant identifiers to scope metric queries and requiring an authentication layer to establish the appropriate tenant context. Mimir also provides query-front-end capabilities for managing and scaling the read path. Cortex, the project on which several Prometheus-compatible managed services are based, uses a similar X-Scope-OrgID model to isolate metric data between tenants. The Adobe approach is somewhat different in that it keeps a shared Prometheus environment while adding Kubernetes-aware access controls and namespace-based filtering around it.

About the Author

Rate this Article

Adoption
Style

BT