In Prometheus, memory usage is dominated by the number of time series currently stored in the head block (plus indexes). Your goal is to find which metrics/labels/targets create the most series and which ones churn (get created/removed frequently).


Use the built-in TSDB status endpoint

Prometheus exposes a helper endpoint that lists the top offenders:

http://<prometheus>:9090/api/v1/status/tsdb?limit=50

Useful sections in the JSON:

  • seriesCountByMetricName – metric names with the most series
  • labelValueCountByLabelName – label names with the most distinct values
  • seriesCountByLabelPair – specific label=value pairs that explode series
  • (Recent versions) memoryInBytesByLabelName – memory cost per label name

You can filter further with match[], for example:

/api/v1/status/tsdb?limit=50&match[]={job="kubelet"}&match[]={job="node-exporter"}

Find noisy scrape targets (who is adding series)

Two scrape-time metrics help pinpoint offenders:

# Which targets recently added many series (churn)?
topk(20, increase(scrape_series_added[1h]))

# Which targets export lots of samples after relabeling?
topk(20, max_over_time(scrape_samples_post_metric_relabeling[5m]))

Targets at the top are strong candidates for metric filtering or exporter reconfiguration (e.g., kube-state-metrics, cAdvisor, app exporters).


Check head size (sanity check)

prometheus_tsdb_head_series
prometheus_tsdb_head_chunks

These show whether the live series count is abnormally high now. Large values usually mean high memory use in the head block.


What to do with the findings

  • Drop or relabel high-cardinality labels/series you do not need (e.g., request IDs, pod UIDs). Metric/target relabeling is the fastest win.
  • Tame noisy exporters (tighten kube-state-metrics selectors, reduce cAdvisor detail, trim app metrics).
  • Reduce churn (avoid ephemeral label values; aggregate at the app; use a longer scrape interval for volatile metrics).
  • Shard or federate if a single Prometheus is handling too much. These are standard strategies for memory relief.

Helpful PromQL: where series come from

# Per target (group by job/instance/metrics_path/endpoint)
topk(30, count by (job, instance, metrics_path, endpoint) ({job=~".+"}))

# Per metric within kube-state-metrics
topk(30, count by (__name__) ({job="kube-state-metrics"}))

# Per metric within kubelet/cAdvisor
topk(30, count by (__name__) ({job="kubelet", metrics_path="/metrics/cadvisor"}))

Helpful PromQL: which labels drive cardinality

count by (__name__) ({name!=""})
count by (__name__) ({id!=""})
count by (__name__) ({uid!=""})

Replace name/id/uid with the label you want to inspect.


Bonus: find unused metrics with Grafana mimirtool

mimirtool collects metric names used in Grafana dashboards and Prometheus rules, then diffs them against what Prometheus actually has to list unused metrics (with counts). It works with plain Prometheus + Grafana—no need to run Mimir.

Great article: https://0xdc.me/blog/how-to-find-unused-prometheus-metrics-using-mimirtool/


Conclusion

Memory pressure in Prometheus almost always traces back to too many active series or excessive churn. Start with the TSDB status endpoint to identify hot metrics and labels, use the scrape-time metrics to find noisy targets, confirm with head size gauges, and then reduce cardinality via relabeling or exporter tuning. If a single server still struggles, shard or federate. Repeat this process regularly to keep memory stable as your environment evolves.