grafana/skills

dpm-finder

Find the Prometheus metrics that drive your Grafana Cloud bill.

Zobacz źródło
Oryginalny dokument Skill

Treść z repozytorium z zachowaniem nagłówków, przykładów, kodu, tabel, linków i obrazów.

dpm-finder

Grafana PS tool ranking Prometheus metrics by DPM with per-series breakdown. Source: https://github.com/grafana-ps/dpm-finder

Prerequisites

  • Python 3.9+
  • A Grafana Cloud Prometheus endpoint URL + numeric stack ID + API key (glc_…, metrics:read scope)

Common Workflows

One-shot analysis (most common)

bash
# 1. Clone + venv + install
git clone https://github.com/grafana-ps/dpm-finder.git
cd dpm-finder
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt

# 2. Configure creds — copy .env_example → .env and fill in:
#    PROMETHEUS_ENDPOINT  https://prometheus-<cluster_slug>.grafana.net  (NOTHING after .net)
#    PROMETHEUS_USERNAME  <numeric stack id>
#    PROMETHEUS_API_KEY   glc_…

# 3. Verify creds before scanning — should return >0 series count
curl -s -u "$PROMETHEUS_USERNAME:$PROMETHEUS_API_KEY" \
  "$PROMETHEUS_ENDPOINT/api/v1/label/__name__/values" | jq '.data | length'

# 4. Run the scan (10-min lookback, 2.0 DPM minimum, top output)
./dpm-finder.py -f json -m 2.0 -t 8 --timeout 120 -l 10

# 5. Read the result — top 10 metrics by DPM
jq -r '.metrics | sort_by(-.dpm) | .[:10][] | "\(.dpm)\t\(.series_count)\t\(.metric_name)"' metric_rates.json

If step 5 is empty, lower -m or confirm the endpoint URL has no trailing path after .net.

Discover stack details with gcx

If gcx is installed it can derive the endpoint + username:

bash
gcx config check          # active stack context
gcx config list-contexts  # all configured stacks
gcx config view           # full config with endpoints

The Prometheus endpoint pattern is https://prometheus-{cluster_slug}.grafana.net. Username is the numeric stack ID.

Without gcx: look up in the Grafana Cloud portal, or query grafanacloud_instance_info{name=~"STACK_NAME.*"} on the usage datasource.

Multi-stack runs

Limit to max 3 concurrent runs to avoid GCloud rate limits. Batch the stacks and wait for each batch before the next.

Interpreting results

  • DPM = max data points per minute across that metric's series
  • series_count = active time-series count for that metric
  • series_detail[] (JSON / text only) = per-label-combination DPM breakdown — use this to spot the offending label
  • Sort by DPM descending → noisiest metrics; combine with --cost-per-1000-series to prioritize by spend

Troubleshooting

  • 401 / 403 — API key invalid or missing metrics:read; confirm PROMETHEUS_USERNAME is the numeric stack ID
  • Timeouts — bump --timeout to 120+ for stacks with thousands of metrics
  • HTTP 422 — metric has aggregation rules; tool warns + skips automatically
  • Empty results — lower -m; verify endpoint has no trailing path
  • Connection errors — exponential backoff retries up to 10 times; persistent failure usually = network/firewall

References

  • `references/cli.md` — full flag reference, output-format details, exporter mode, Docker invocation, auto-exclusion rules, retry behavior

Resources

z tego samego repozytorium

Więcej Skills

Wszystkie Skills
grafana
Społeczność

alerting-irm

Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook), notification policies with hierarchical matchers, silences, mute timings, on-call schedules and escalation chains, incident-management integrations, and SLOs with multi-window burn-rate alerts. Use when configuring alerts, debugging notification routing, setting up on-call rotations, declaring or managing incidents, defining SLOs, provisioning alerting via YAML or API, picking matchers for a notification policy, building a PagerDuty/Slack webhook receiver, or troubleshooting why an alert isn't firing — even when the user says "page me on errors", "alert me when X happens", "route this to the platform team", or "set up an SLO" without naming Alerting or IRM.

instalacje
5
GitHub Stars
246
Aktualizacja
8 wrz
grafana
Społeczność

alloy

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo / Pyroscope. Covers the Alloy config language (blocks, sys.env, component refs), prometheus.scrape → remotewrite, loki.source.file + loki.process → loki.write, otelcol.receiver.otlp → otelcol.exporter.otlp, pyroscope.scrape, K8s / Docker / EC2 discovery, relabeling, modules (import.file/git/http), clustering, Fleet Management remotecfg, the Alloy UI at :12345, and alloy fmt / alloy validate. Use when writing a config.alloy, replacing Grafana Agent / OTel Collector, scraping K8s pods, parsing logs, ingesting OTLP, or debugging "Alloy isn't sending anything" — even when the user says "set up the agent", "write me a scrape config", "drop these logs before sending", or "OTel collector config" without naming Alloy.

instalacje
5
GitHub Stars
246
Aktualizacja
8 wrz
grafana
Społeczność

beyla

Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart. Covers requirements (Linux 5.8+ with BTF, CAPSYSADMIN, host PID), language matrix (Go / Java / Python / Ruby / Node / .NET / Rust / C++ / PHP), Docker + Helm + DaemonSet install, port- / process- / Kubernetes-metadata discovery, OTLP traces + Prometheus metrics export, routes decorator (cardinality control), trace sampling, and Grafana Cloud via Alloy. Use when adding observability to a service you can't recompile, instrumenting a closed-source binary, getting RED metrics + spans onto Tempo/Mimir without touching the app, or rolling Beyla as a cluster-wide DaemonSet — even when the user says "zero-code APM", "instrument legacy app", "trace this binary", "eBPF observability", or "no SDK" without naming Beyla.

instalacje
5
GitHub Stars
246
Aktualizacja
8 wrz
grafana
Społeczność

dashboarding

Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units, thresholds, template + datasource + chained variables, transformations (organize / calculateField / filterByValue), panel + dashboard links with ${field.labels.x} / ${from}, and Loki/Prometheus annotations. Use when scripting dashboard creation, writing the dashboard JSON for a new service, adding a $job dropdown variable, computing an "Error %" column with a transformation, overlaying deploys as annotations, or pushing a dashboard via POST /api/dashboards/db — even when the user says "create a dashboard for this metric", "add a service dropdown", "show errors as percentage", "overlay our deploys", or "export the dashboard JSON" without naming the API or schema. After every API push, verify with the returned version plus a GET on the dashboard UID.

instalacje
5
GitHub Stars
246
Aktualizacja
8 wrz