grafana/skills

oncall-irm

Route alerts, run on-call rotations, and drive incidents in Grafana IRM / OnCall — integrations (Alertmanager / Grafana Alerting / generic webhook / PagerDuty), Jinja2 routing + grouping templates, escalation chains (wait → notify schedule → notify team → w…

ソースを見る
リポジトリの原文

見出し、例、コード、表、リンク、参照画像を含む原文を表示しています。

Grafana OnCall & IRM

OnCall docs: https://grafana.com/docs/oncall/latest/ IRM docs: https://grafana.com/docs/grafana-cloud/alerting-and-irm/
Grafana OnCall OSS is in maintenance mode (archived March 2026). Cloud users → IRM. Concepts (chains, schedules, integrations) are identical.

Prerequisites

  • Grafana Cloud stack with IRM/OnCall enabled
  • API token (Authorization: <token>)
  • Slack workspace + admin to install the OnCall app (for ChatOps)

Core concepts

ConceptDescription
IntegrationWebhook URL accepting alerts; one per source
RouteJinja2 condition that maps to an escalation chain (first True wins)
Escalation ChainWait / notify schedule / notify team / webhook / auto-resolve steps
ScheduleCalendar-based rotation (web / iCal / Terraform)
Alert GroupRelated alerts collapsed by a Grouping ID template
Notification PolicyPer-user channels — Slack, mobile push, SMS, phone, email

Flow: alert → integration → routing template → escalation chain → notifications → ack / resolve.

Common Workflows

1. Wire Alertmanager → IRM and verify routing

yaml
# 1. In IRM: New integration → Alertmanager. Copy the webhook URL.
# 2. alertmanager.yml (see references/integrations.md for full block):
receivers:
  - name: grafana-oncall
    webhook_configs:
      - url: https://<stack>.grafana.net/integrations/v1/alertmanager/<id>/
        send_resolved: true
        max_alerts: 100
bash
# 3. Test the routing template BEFORE going live (UI: Integration → Route → "Preview")
#    Paste a sample payload (use a real Alertmanager test webhook). Expect:
#      Routing result: True
#      Escalation chain selected: <expected chain>
#    If False — your Jinja expression is wrong; fix and re-preview.

# 4. End-to-end test — fire amtool (or any test webhook) at the URL
amtool alert add foo severity=critical team=platform --alertmanager.url http://localhost:9093

# 5. Verify in IRM → Alert Groups (should appear within ~5s) — confirm:
#    - Correct route fired
#    - Correct schedule/user notified
#    - Slack message appeared in the configured channel

Routing template syntax + Jinja helpers: `references/templates-schedules.md`. Other integrations (Grafana Alerting, generic webhook, Slack): `references/integrations.md`.

2. Build an escalation chain + verify

1. Notify "Primary On-Call" (Important Notifications)
2. Wait 5 min
3. Notify "Primary On-Call" (Default Notifications)
4. Wait 10 min
5. Notify team "Platform"
6. Trigger outgoing webhook (PagerDuty / ticket)
bash
# Verify: create a test alert (UI → Integration → "Send demo alert"),
# then watch the alert-group timeline tick through steps 1 → 6.
# Use IRM → Escalation Chains → "Test" if available, otherwise the demo alert is the canonical check.

3. Create a schedule from iCal + verify "who is on-call right now"

bash
# 1. Create the schedule
curl -X POST https://<stack>.grafana.net/api/v1/schedules/ \
  -H "Authorization: <token>" -H "Content-Type: application/json" \
  -d '{"name":"Platform On-Call",
       "ical_url_primary":"https://calendar.example.com/platform.ics",
       "slack":{"channel_id":"C123456ABC","user_group_id":"S123456ABC"}}'

# 2. Verify the schedule was created
curl -s https://<stack>.grafana.net/api/v1/schedules/ \
  -H "Authorization: <token>" | jq '.results[] | select(.name=="Platform On-Call")'

# 3. Verify who is on-call right now
curl -s https://<stack>.grafana.net/api/v1/schedules/<schedule_id>/next_shifts/ \
  -H "Authorization: <token>" | jq '.results[0]'
# Expect a shift starting now (or recently) with the right user_id.

Terraform variant + shift block: `references/templates-schedules.md`.

Best practices

  • Keep chains ≤4 levels with a definitive final step (webhook to PagerDuty or auto-resolve)
  • Always set send_resolved: true in Alertmanager so OnCall auto-resolves
  • Use max_alerts: 100 in Alertmanager webhook config
  • Combine Slack + mobile push for delivery reliability
  • Assign integrations/schedules to teams for RBAC

Resources

同じリポジトリから

関連する Skills

すべての Skills
grafana
コミュニティ

alerting-irm

Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook), notification policies with hierarchical matchers, silences, mute timings, on-call schedules and escalation chains, incident-management integrations, and SLOs with multi-window burn-rate alerts. Use when configuring alerts, debugging notification routing, setting up on-call rotations, declaring or managing incidents, defining SLOs, provisioning alerting via YAML or API, picking matchers for a notification policy, building a PagerDuty/Slack webhook receiver, or troubleshooting why an alert isn't firing — even when the user says "page me on errors", "alert me when X happens", "route this to the platform team", or "set up an SLO" without naming Alerting or IRM.

導入数
5
GitHub Stars
246
更新日
9月8日
grafana
コミュニティ

alloy

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo / Pyroscope. Covers the Alloy config language (blocks, sys.env, component refs), prometheus.scrape → remotewrite, loki.source.file + loki.process → loki.write, otelcol.receiver.otlp → otelcol.exporter.otlp, pyroscope.scrape, K8s / Docker / EC2 discovery, relabeling, modules (import.file/git/http), clustering, Fleet Management remotecfg, the Alloy UI at :12345, and alloy fmt / alloy validate. Use when writing a config.alloy, replacing Grafana Agent / OTel Collector, scraping K8s pods, parsing logs, ingesting OTLP, or debugging "Alloy isn't sending anything" — even when the user says "set up the agent", "write me a scrape config", "drop these logs before sending", or "OTel collector config" without naming Alloy.

導入数
5
GitHub Stars
246
更新日
9月8日
grafana
コミュニティ

beyla

Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart. Covers requirements (Linux 5.8+ with BTF, CAPSYSADMIN, host PID), language matrix (Go / Java / Python / Ruby / Node / .NET / Rust / C++ / PHP), Docker + Helm + DaemonSet install, port- / process- / Kubernetes-metadata discovery, OTLP traces + Prometheus metrics export, routes decorator (cardinality control), trace sampling, and Grafana Cloud via Alloy. Use when adding observability to a service you can't recompile, instrumenting a closed-source binary, getting RED metrics + spans onto Tempo/Mimir without touching the app, or rolling Beyla as a cluster-wide DaemonSet — even when the user says "zero-code APM", "instrument legacy app", "trace this binary", "eBPF observability", or "no SDK" without naming Beyla.

導入数
5
GitHub Stars
246
更新日
9月8日
grafana
コミュニティ

dashboarding

Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units, thresholds, template + datasource + chained variables, transformations (organize / calculateField / filterByValue), panel + dashboard links with ${field.labels.x} / ${from}, and Loki/Prometheus annotations. Use when scripting dashboard creation, writing the dashboard JSON for a new service, adding a $job dropdown variable, computing an "Error %" column with a transformation, overlaying deploys as annotations, or pushing a dashboard via POST /api/dashboards/db — even when the user says "create a dashboard for this metric", "add a service dropdown", "show errors as percentage", "overlay our deploys", or "export the dashboard JSON" without naming the API or schema. After every API push, verify with the returned version plus a GET on the dashboard UID.

導入数
5
GitHub Stars
246
更新日
9月8日