posthog/ai-plugin

investigating-metric-anomalies

Investigates server/infrastructure metric anomalies in PostHog Metrics — from "this metric is rising/dropping/spiking" or a fired alert to a probable cause with evidence.

Zobacz źródło
Oryginalny dokument Skill

Treść z repozytorium z zachowaniem nagłówków, przykładów, kodu, tabel, linków i obrazów.

Investigating metric anomalies

The job: go from a metric symptom ("ingestion lag is rising") to a probable cause with evidence, fast. The metric tells you what and when; logs and traces tell you why. Follow the loop below — it front-loads the cheap, high-information calls and only fans out when the blast radius is unclear.

The loop

1. Pin down the metric

If you have the exact metric name, skip ahead. Otherwise call metric-names-list with a substring from the symptom (lag, error, latency, queue). The returned metric_type decides the lens: counters (sum) are only meaningful as rate/increase, gauges as avg, histograms as histogram_quantile.

2. Characterize first — one call, three answers

Call characterize-metric-anomaly with the metric name and anomalyFrom (the alert fire time, or when the user says it started looking wrong; subtract some margin if unsure). It compares against the preceding window by default and answers:

  • How bad: direction, change_ratio, anomaly_peak vs baseline_mean. If direction is flat, your window or metric is wrong — widen the window, or compare against the same window yesterday via baselineFrom/baselineTo (daily-pattern metrics often look "anomalous" against the immediately-preceding hours).
  • When: onset_time — treat this timestamp as the pivot for everything that follows.
  • Where: top_movers — label values whose behavior changed. One mover (a single pod, shard, or endpoint) means a localized culprit; everything moving together means a shared cause (an upstream dependency, a deploy, infra).

3. Sharpen with targeted metric queries

Use query-metrics to test the hypotheses the report raises:

  • Drill a mover: re-query with filters pinning the suspicious label value, grouped by a second key, to localize further (pod → container, endpoint → status code).
  • Normalize: a rising error count means nothing if traffic doubled — use clauses + formula (errors / requests) to separate rate changes from volume changes.
  • Check the neighbors: query the obvious companion metrics over the same window (for lag: throughput and error counters of the same service; for latency: request rate and saturation gauges). Use the same interval so the grids align visually.

4. Correlate across signals at the onset

Pivot into logs and traces using the same service and a window bracketing `onset_time` (a few buckets before, through the peak):

  • Logs: use query-logs (follow its own discover-first workflow) filtered to the implicated service.name and window, severity error first, then warn. Restarts, crash loops, connection errors, and deploy markers right before onset are the classic causes. Widen to other services in the request path if the service's own logs are clean.
  • Traces: use the APM span tools (query-apm-spans etc.) for the same service/window — slow or erroring spans show which dependency degraded, and a trace_id from an exemplar or log line links a concrete request across all three signals.

5. Conclude with evidence, not vibes

State: the symptom (metric, magnitude, onset), the probable cause (what you found in logs/traces and how its timing aligns with the onset), the blast radius (which services/labels are affected, from the movers and grouped queries), and the confidence level. If the cause is still ambiguous, say which hypothesis the evidence favors and what would disambiguate (e.g. "the lag began draining at 20:12 — consistent with a consumer restart; check who restarted it").

Worked example: "ingestion lag is rising — why?"

  1. metric-names-list with value: "lag"logs_rate_limiter_message_lag_seconds (histogram) and friends.
  2. characterize-metric-anomaly on it with anomalyFrom = alert time → direction up, change ratio 40x, onset_time 20:10, top mover service_name = logs-ingestion (the other services' lag stayed flat) — so the logs consumer specifically is behind, not the whole pipeline.
  3. query-metrics: rate of the consumer's throughput counter over the same window → throughput was zero during the gap and spiked after onset: the consumer wasn't slow, it was down, and the "rising lag" is it draining the backlog.
  4. query-logs for service.name = logs-ingestion (and its neighbors) around 20:00–20:15 → process exit + restart lines at the gap boundaries.
  5. Conclusion: consumer outage 20:01–20:10 (restart visible in logs); lag spike is backlog drain, self-recovering; affected signal: logs freshness only — no data loss (topic retained messages). Evidence: zero throughput during the window, message-age spike equal to outage duration, restart log lines.

Pitfalls

  • Counters reset on process restart — rate/increase already handle this; never eyeball raw cumulative counter values.
  • A metric that stops reporting is not "zero" — a gap in series with top_movers showing a vanished label value means the emitter died; pivot to logs immediately.
  • Don't trust a single aggregation: a flat avg can hide a screaming p95. For latency-like gauges and histograms, characterize the tail too.
  • Scraped Prometheus metrics arrive ~15–60s behind real time; don't read the last bucket of a series as "current".
z tego samego repozytorium

Więcej Skills

Wszystkie Skills
posthog
Oficjalne

assessing-heatmaps

Assesses what a page's heatmap is telling you and recommends concrete changes. Pulls click / rageclick / scroll-depth data for a URL, names the hot elements by cross-referencing autocapture events on the same page, and can create a saved heatmap the user opens in PostHog, then summarizes the behavior and proposes improvements.\nTRIGGER when: user asks what a heatmap shows, why people aren't clicking something, where users rage-click, how far they scroll, what to change on a page based on heatmap/click data, or to 'analyze/assess/review the heatmap' for a URL.\nDO NOT TRIGGER when: the user only wants to create a saved heatmap screenshot with no analysis (use heatmaps-saved-create directly), or is asking about session replay in general (use investigating-replay).

instalacje
1
GitHub Stars
80
Aktualizacja
4 wrz
posthog
Oficjalne

auditing-endpoints

Audit every endpoint in a PostHog project for staleness, failed materialisations, and unused materialised versions. Use when the user asks "what endpoints can I clean up?", "are any of my endpoints broken?", "which materialised versions are still being called?", or wants a one-shot cleanup pass over the Endpoints product. Produces a prioritised report grouped by issue type, with recommended actions but does not modify anything without explicit confirmation.

instalacje
1
GitHub Stars
80
Aktualizacja
4 wrz
posthog
Oficjalne

auditing-experiments-flags

Audit PostHog experiments and feature flags for configuration issues, staleness, and best-practice violations. Read when the user asks to audit, health-check, or review experiments or feature flags, check flag hygiene, or verify experiment setup.

instalacje
1
GitHub Stars
80
Aktualizacja
4 wrz
posthog
Oficjalne

authoring-data-quality-checks

Adds and runs data quality checks (dbt-test style assertions) on a project's warehouse tables and saved-query views: not-null, uniqueness, accepted values, referential integrity, row-count bounds, freshness, and custom HogQL. Use when asked to test a model, validate a view, check for nulls or duplicates, add data quality checks, find out why a number looks wrong, or judge whether a warehouse table is trustworthy before using it in an analysis. To describe what data means (metrics, certifications, joins), see setting-up-data-catalog instead. Trigger terms: data quality, data test, dbt test, not null check, uniqueness check, freshness check, referential integrity, row count check, validate model, is this table trustworthy.

instalacje
1
GitHub Stars
80
Aktualizacja
4 wrz