posthog/ai-plugin

improving-mcp-tools

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded fix, and keeps it only if before/after scores improve.

Ver código-fonte
Documento original do Skill

Renderizado do repositório de origem, preservando títulos, exemplos, código, tabelas, links e imagens.

Improving MCP tools

An MCP server gets better only in ways you can measure. This skill is the campaign procedure: score the current agent experience, fix the biggest problem, re-score, and only ship changes the numbers justify. It is the operating manual for the "improve my MCP" loop — one iteration per pass, journaled so a later iteration (or a different agent) can resume without repeating work.

The objective function

services/mcp/evals/ is the harness. benchmark/tasks.yaml is a fixed set of agent tasks with expected_tools and success_criteria; scores are only comparable across runs of the same benchmark version.

  • Probe mode (deterministic, no LLM):

LIVE_MCP_URL=... LIVE_MCP_TOKEN=... pnpm exec tsx evals/runner/probe.ts --out score.json from services/mcp/. Reports tool-presence misses (discoverability), probe failures, and latency p50/p95. Non-zero exit = regression.

  • Agent mode (LLM replay + judge): scores task success and tool-selection

accuracy. Use it for description/discoverability changes — probes cannot detect that an agent picks the wrong tool.

Run the harness against a seeded local or devbox stack, never against a customer project. Local recipe: NODE_ENV=development PORT=9876 POSTHOG_API_BASE_URL=http://localhost:8000 pnpm dev:hono, personal API key as LIVE_MCP_TOKEN.

One iteration

  1. Measure. Run the harness for a baseline. Pull production evidence with

the MCP analytics tools (query-mcp-tool-stats, query-mcp-tool-failures, query-mcp-tool-descriptions, query-mcp-tool-sample-intents) and the lenses in the signals scout cookbook (products/signals/skills/signals-scout-mcp-tool-calls/references/queries.md): failure leaderboard, retry/struggle, latency, intents that matched no tool.

  1. Pick one issue. Rank by reach × severity. Skip anything the journal

shows with two failed attempts. One issue per iteration — a PR that fixes three things can't be attributed to any of them when scores move.

  1. Fix, bounded. Only files inside the allowlist (below). Typical fixes:

sharpen a tool description so the right intent finds it, tighten an input schema that agents keep getting wrong, fix an annotation, update a skill.

  1. Validate. Re-run the affected benchmark slice plus a no-regression

sample. Keep the change only if the target metric improves and nothing else degrades. A discarded change is a normal outcome — journal it and move on.

  1. Ship. One PR per iteration with before/after scores in the body (format

in references/campaign-journal.md). Keep it stampable: ≤400 changed lines, only files inside the allowlist below, apply the stamphog label. Autonomy level comes from the campaign config — default is draft PR for human review; only arm auto-merge when the operator has explicitly enabled the self-driving experiment (see guardrails).

  1. Journal. Append the iteration record before ending the pass.

Hard guardrails

These are not suggestions; violating any of them ends the campaign pass.

  • Allowlist — a campaign PR may only touch: products/*/mcp/tools.yaml,

products/*/skills/**, services/mcp/evals/**, the codegen outputs of pnpm generate-tools / scaffold-yaml (services/mcp/src/tools/generated/** and services/mcp/schema/generated-tool-definitions.json), and docs. Anything else (handler code, package manifests, workflows, migrations, auth paths) → stop and hand the finding to a human as a draft PR or report instead.

  • Read-only against data. The harness and all production queries are

read-only. Never create, mutate, or delete customer-visible objects while measuring.

  • Evidence or it didn't happen. No PR without a baseline score, an after

score, and the exact harness commands used.

  • Benchmark integrity. Never edit benchmark/tasks.yaml in the same PR as

a fix it validates — changing the exam and the answer together proves nothing. Benchmark changes are their own PR and bump version.

  • Budgets. Respect the operator's iteration/token/PR caps (default: stop

after 3 open unmerged campaign PRs). Two failed attempts on an issue parks it permanently.

  • Kill switch. If the campaign config, its feature flag, or the operator

says stop — stop mid-iteration, journal state, end cleanly.

Failure modes to expect

  • A description change that helps one intent can steal traffic from the right

tool for another — that's why the no-regression sample is mandatory. The intent-cluster snapshot's tool_overlaps (see `exploring-mcp-intent-clusters`) lists exactly which pairs compete for which intents: snapshot it before a description rewrite and recompute after, and treat a capture shift in an overlapping pair as the regression signal.

  • Probe latency varies with stack warmth; compare medians across ≥3 runs

before attributing a latency change to your fix.

  • Tool-presence misses can be feature-flag gating, not catalog absence —

check getToolsForFeatures gating before "fixing" discoverability.

do mesmo repositório

Mais Skills

Todos os Skills
posthog
Oficial

assessing-heatmaps

Assesses what a page's heatmap is telling you and recommends concrete changes. Pulls click / rageclick / scroll-depth data for a URL, names the hot elements by cross-referencing autocapture events on the same page, and can create a saved heatmap the user opens in PostHog, then summarizes the behavior and proposes improvements.\nTRIGGER when: user asks what a heatmap shows, why people aren't clicking something, where users rage-click, how far they scroll, what to change on a page based on heatmap/click data, or to 'analyze/assess/review the heatmap' for a URL.\nDO NOT TRIGGER when: the user only wants to create a saved heatmap screenshot with no analysis (use heatmaps-saved-create directly), or is asking about session replay in general (use investigating-replay).

instalações
1
GitHub Stars
80
Atualizado
4 de set.
posthog
Oficial

auditing-endpoints

Audit every endpoint in a PostHog project for staleness, failed materialisations, and unused materialised versions. Use when the user asks "what endpoints can I clean up?", "are any of my endpoints broken?", "which materialised versions are still being called?", or wants a one-shot cleanup pass over the Endpoints product. Produces a prioritised report grouped by issue type, with recommended actions but does not modify anything without explicit confirmation.

instalações
1
GitHub Stars
80
Atualizado
4 de set.
posthog
Oficial

auditing-experiments-flags

Audit PostHog experiments and feature flags for configuration issues, staleness, and best-practice violations. Read when the user asks to audit, health-check, or review experiments or feature flags, check flag hygiene, or verify experiment setup.

instalações
1
GitHub Stars
80
Atualizado
4 de set.
posthog
Oficial

authoring-data-quality-checks

Adds and runs data quality checks (dbt-test style assertions) on a project's warehouse tables and saved-query views: not-null, uniqueness, accepted values, referential integrity, row-count bounds, freshness, and custom HogQL. Use when asked to test a model, validate a view, check for nulls or duplicates, add data quality checks, find out why a number looks wrong, or judge whether a warehouse table is trustworthy before using it in an analysis. To describe what data means (metrics, certifications, joins), see setting-up-data-catalog instead. Trigger terms: data quality, data test, dbt test, not null check, uniqueness check, freshness check, referential integrity, row count check, validate model, is this table trustworthy.

instalações
1
GitHub Stars
80
Atualizado
4 de set.