grafana/skills

ml-ai

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), K…

Vedi sorgente
Documento Skill originale

Contenuto dal repository con titoli, esempi, codice, tabelle, link e immagini preservati.

Grafana Cloud AI & ML

Docs: https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/

ML alerting + automated RCA + LLM-powered Assistant in one Grafana Cloud stack.

Prerequisites

  • Grafana Cloud stack (Pro / Advanced — most features GA, some in preview)
  • API token with plugins:write for ML / Sift / LLM-plugin endpoints
  • For Dynamic Alerting: at least 14 days (ideally 90d) of history for the metric you want to forecast

Common Workflows

1. Forecasting alert with Dynamic Alerting

bash
# 1. Create forecast job (Prophet — learns daily/weekly seasonality)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/forecast \
  -H "Authorization: Bearer <token>" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "cpu-forecast",
    "metric": "avg(rate(node_cpu_seconds_total{mode=\"user\"}[5m]))",
    "datasourceId": 1,
    "interval": 300,
    "trainingWindow": "90d",
    "forecastWindow": "7d",
    "algorithm": { "name": "prophet", "config": {} }
  }'

# 2. Verify job is producing the predicted-value metric (may take a few minutes).
#    <datasourceId> must match the datasourceId used above (find it via
#    GET /api/datasources), or run the query from Explore instead.
curl -s -H "Authorization: Bearer <token>" \
  'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_forecast_upper{job="cpu-forecast"}' \
  | jq '.data.result | length'
# Expect > 0

# 3. Add an alert that fires when actual exceeds the upper bound
# expr:  avg(rate(node_cpu_seconds_total{mode="user"}[5m]))
#         > ml_forecast_upper{job="cpu-forecast"} * 1.1

2. Outlier alert — one service deviates from peers

bash
# 1. Create outlier job (DBSCAN — groups peers, flags the odd one)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/outlier \
  -H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
  -d '{
    "name": "service-error-outliers",
    "metric": "sum(rate(http_requests_total{status=~\"5..\"}[5m])) by (service)",
    "datasourceId": 1,
    "interval": 300,
    "algorithm": { "name": "dbscan", "sensitivity": 0.5, "config": { "epsilon": 0.5 } }
  }'

# 2. Verify the score metric exists (<datasourceId> must match the
#    datasourceId used above, or run the query from Explore instead)
curl -s -H "Authorization: Bearer <token>" \
  'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_outlier_score{job="service-error-outliers"}' \
  | jq '.data.result | length'

# 3. Alert when ml_outlier_score{job="service-error-outliers"} > 0.8 for 5m

3. Run a Sift investigation

bash
# 1. Trigger from API (or from Explore / Incident / OnCall)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-sift-app/resources/sift/v1/investigations \
  -H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
  -d '{ "name":"checkout-spike","start":"2024-02-01T10:00:00Z","end":"2024-02-01T10:30:00Z",
        "filters":{"service":"checkout","namespace":"production"} }'

# 2. The response includes an investigation ID — open it in the UI:
#    https://<stack>.grafana.net/a/grafana-sift-app/investigations/<id>
# 3. Verify analyses ran — each of the 8 checks shows ✔ or ✖ with linked evidence.

See `references/sift.md` for the full 8-analysis table.

4. Wire up the LLM Plugin

yaml
# 1. Provision (provisioning/plugins/llm.yaml — see references/llm-and-graph.md)
apiVersion: 1
apps:
  - type: grafana-llm-app
    jsonData: { openAIUrl: https://api.openai.com, openAIModel: gpt-4o }
    secureJsonData: { openAIKey: sk-... }
bash
# 2. Restart Grafana, then verify the health endpoint reports the configured provider
curl -s -H "Authorization: Bearer <token>" \
  https://<stack>.grafana.net/api/plugins/grafana-llm-app/health | jq
# Expect: {"status":"ok", ...}

# 3. Verify in a panel — open any panel, click the Assistant icon, ask "what does this query do?"

See `references/llm-and-graph.md` for Assistant capabilities, Knowledge Graph search syntax, and Adaptive Metrics recommendations.

Resources

dallo stesso repository

Altri Skills

Tutti gli Skills
grafana
Community

alerting-irm

Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook), notification policies with hierarchical matchers, silences, mute timings, on-call schedules and escalation chains, incident-management integrations, and SLOs with multi-window burn-rate alerts. Use when configuring alerts, debugging notification routing, setting up on-call rotations, declaring or managing incidents, defining SLOs, provisioning alerting via YAML or API, picking matchers for a notification policy, building a PagerDuty/Slack webhook receiver, or troubleshooting why an alert isn't firing — even when the user says "page me on errors", "alert me when X happens", "route this to the platform team", or "set up an SLO" without naming Alerting or IRM.

installazioni
5
GitHub Stars
246
Aggiornato
8 set
grafana
Community

alloy

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo / Pyroscope. Covers the Alloy config language (blocks, sys.env, component refs), prometheus.scrape → remotewrite, loki.source.file + loki.process → loki.write, otelcol.receiver.otlp → otelcol.exporter.otlp, pyroscope.scrape, K8s / Docker / EC2 discovery, relabeling, modules (import.file/git/http), clustering, Fleet Management remotecfg, the Alloy UI at :12345, and alloy fmt / alloy validate. Use when writing a config.alloy, replacing Grafana Agent / OTel Collector, scraping K8s pods, parsing logs, ingesting OTLP, or debugging "Alloy isn't sending anything" — even when the user says "set up the agent", "write me a scrape config", "drop these logs before sending", or "OTel collector config" without naming Alloy.

installazioni
5
GitHub Stars
246
Aggiornato
8 set
grafana
Community

beyla

Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart. Covers requirements (Linux 5.8+ with BTF, CAPSYSADMIN, host PID), language matrix (Go / Java / Python / Ruby / Node / .NET / Rust / C++ / PHP), Docker + Helm + DaemonSet install, port- / process- / Kubernetes-metadata discovery, OTLP traces + Prometheus metrics export, routes decorator (cardinality control), trace sampling, and Grafana Cloud via Alloy. Use when adding observability to a service you can't recompile, instrumenting a closed-source binary, getting RED metrics + spans onto Tempo/Mimir without touching the app, or rolling Beyla as a cluster-wide DaemonSet — even when the user says "zero-code APM", "instrument legacy app", "trace this binary", "eBPF observability", or "no SDK" without naming Beyla.

installazioni
5
GitHub Stars
246
Aggiornato
8 set
grafana
Community

dashboarding

Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units, thresholds, template + datasource + chained variables, transformations (organize / calculateField / filterByValue), panel + dashboard links with ${field.labels.x} / ${from}, and Loki/Prometheus annotations. Use when scripting dashboard creation, writing the dashboard JSON for a new service, adding a $job dropdown variable, computing an "Error %" column with a transformation, overlaying deploys as annotations, or pushing a dashboard via POST /api/dashboards/db — even when the user says "create a dashboard for this metric", "add a service dropdown", "show errors as percentage", "overlay our deploys", or "export the dashboard JSON" without naming the API or schema. After every API push, verify with the returned version plus a GET on the dashboard UID.

installazioni
5
GitHub Stars
246
Aggiornato
8 set