dotnet/skills

create-skill-test

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository.

Ver código-fonte
Documento original do Skill

Renderizado do repositório de origem, preservando títulos, exemplos, código, tabelas, links e imagens.

Create Skill Test

Scaffold an evaluation spec (eval.yaml) for a skill or agent so it conforms to the Vally schema, passes skill-validator check and check_eval_quality.py, is powerful enough to return a verdict, and does not overfit to the skill's own wording.

When to Use

  • Creating a new eval.yaml for a skill or agent
  • Adding stimuli to an existing eval
  • Sizing an eval so the pass gate can actually be reached
  • Setting up or repairing fixture files alongside an eval
  • Reviewing whether rubric items and graders risk overfitting

When Not to Use

  • Diagnosing a failing or regressed eval — use improve-skill-quality
  • Modifying the skill-validator or the evaluation workflows
  • Creating or editing SKILL.md files — use create-skill

Inputs

InputRequiredDescription
Skill or agent nameYesMust exist under plugins/<plugin>/skills/ or plugins/<plugin>/agents/
Plugin nameYese.g. dotnet-msbuild
Skill contentYesRead it — you cannot write non-overfitted rubric items without it
Failure modes to discriminateRecommendedEach becomes one stimulus

Workflow

Step 1: Locate the target and the test directory

text
tests/<plugin>/<skill-name>/eval.yaml          # skills
tests/<plugin>/agent.<agent-name>/eval.yaml    # agents (the agent. prefix disambiguates)

Verify the target exists at plugins/<plugin>/skills/<skill-name>/SKILL.md or plugins/<plugin>/agents/<agent-name>.agent.md, and read it.

Agent evals sit outside the verdict flow. The canonical experiment declares evals: tests/*/!(agent.*)/eval.yaml, so agent.* specs are excluded: no verdict is ever computed for them, the stimulus floor does not apply, and ./eng/run-skill-evals.sh drops them even when you name one explicitly (its --eval-filter is intersected with that glob). The distinct-stimulus floor therefore applies to skill evals only. Author agent evals for the scenario coverage and the deterministic graders, and run them as described in Step 10.

Be careful with a skill that sets `disable-model-invocation: true`. The model cannot invoke it, so the skill is absent from the model-facing skilled arm and any direct eval compares two identical arms. Answer-content graders do not create a difference between those arms. The honest coverage for such skills is dependency-level — through the outcome evals of the skills that load them, and through the plugin arm. For example, filter-syntax is covered by the filtered-command scenarios in tests/dotnet-test/run-tests/eval.yaml.

Step 2: Write the spec skeleton

The spec is Vally format. Every eval in this repo uses stimuli: and graders:; scenarios: and assertions: are a pre-Vally format that no longer loads.

yaml
name: <skill-name>
description: Evaluates the <plugin>/<skill-name> skill
type: capability
defaults:
  timeout: 5m
  runs: 1
stimuli:
  - name: <what the agent must accomplish>
    prompt: <natural developer request>
    environment:
      files:
        - src: fixtures/<case>/Project.csproj
          dest: Project.csproj
    graders:
      - type: output-matches
        config:
          pattern: (root cause|underlying issue)
      - type: exit-success
      - type: prompt
    rubric:
      - <outcome the agent should have reached>
`defaults:` replaces `config:` — it does not join it. config is a deprecated alias for the same block and vally throws on a spec declaring both. Some existing evals still open with config:; when you change settings, replace it with one defaults: block. The failure is invisible otherwise: the job exits 0 with no verdicts and the PR comment blames "transient infrastructure".

Step 3: Size the eval for power before writing content

The gate gives each distinct stimulus one vote. Repeated runs for one stimulus collapse to one majority-direction vote and remain available as reliability evidence.

  1. Distinct stimuli ≥ 5, else the verdict is underpowered — never a pass, never a regression.
  2. p ≤ 0.05 on an exact one-sided sign test over discordant (non-tie) stimulus votes.* Ties are not

discarded; they hold the discordant count down.

discordant stimulus votesrecords that passp
≤ 4none≥ 0.0625
5–7zero losses only (5W/0L)0.031
8one loss survivable (7W/1L)0.035

At exactly 5 stimuli, one tie is fatal because it leaves 4 discordant votes. At 6 stimuli one tie is survivable; at 7, up to two are. A loss is not. Five is an eligibility floor, not adequate power. For example, 80% power needs 8 discordant votes only for a true 90% conditional win rate; it needs 18 at 80%, 37 at 70%, and 158 at 60%. Size for the effect and tie rate you need to detect.

Use runs for reliability, not task breadth. Vally recommends 3 runs in CI and 5–10 nightly for pass rate, pass@k, pass^k, and flakiness. Extra runs never clear the five-stimulus floor.

Do not set runs in dotnet-skills.experiment.yaml; experiment overrides overwrite every eval's own value rather than defaulting it.

Step 4: Write stimuli

  • Name describes what is tested, not how.
  • Prompt is a natural developer request. Never mention the skill, the agent, or its vocabulary —

cued prompts inflate the overfit score and bias the baseline.

  • Each stimulus should discriminate a different property of the skill. Five stimuli covering one

property give arithmetic, not evidence.

  • Give every stimulus a stable, unique name. Vally pairs comparison trajectories by

(stimulus name, trial index); duplicate names make slot identity ambiguous.

  • Include a boundary / no-op stimulus for any skill that migrates or rewrites code, proving it

leaves already-correct input alone.

Step 5: Configure the environment

yaml
environment:
  files:
    - src: fixtures/broken-build/App.csproj      # path relative to eval.yaml
      dest: App.csproj                           # path in the agent's working directory
    - src: fixtures/broken-build                 # a directory
      dest: .
  commands:
    - dotnet build -bl || exit 0                 # guard intentional failures

Do not set `environment.skills` in a skill eval. The experiment declares vary: /environment/skills and supplies the value itself — [] for the baseline arm and plugins/<plugin>/skills/<skill> for the skilled arm — so anything the eval declares is replaced, in every arm. It cannot add a skill to one arm only. environment.skills is meaningful only in an agent.* eval, which the experiment does not vary; there it is the set of skills the agent may invoke. Copy the shape from an existing agent eval such as tests/dotnet-test/agent.test-quality-auditor/eval.yaml rather than reproducing a remembered form — the specs in this repo are not consistent about how they spell those entries.

Fixture rules — each one has already cost a real result:

  • Every referenced fixture must be tracked by git. .gitignore (e.g. coverage*.xml) has

silently swallowed a committed fixture: the eval passed locally and failed at setup in CI. Verify with git ls-files, not by looking at the working tree.

  • Every fixture must behave as its stimulus assumes. A fixture meant to be healthy must build; a

fixture meant to be broken must fail for the exact reason the stimulus is about, and no other. Judges penalize agents for unrelated "pre-existing build issues" that the fixture author introduced.

  • Every fixture must reproduce the bug its stimulus is named for. If it does not, the baseline

scores well and the skill has nothing to add.

  • Coverage fixtures must be internally consistent. A Cobertura report whose declared

line-rate, summary totals (lines-covered/lines-valid), and <line> elements disagree lets the two arms read different truths, and the loss is the fixture's fault. Update any rubric item or prompt that quotes a figure in the same change.

  • Do not wire duplicate fixtures to raise n; rename leftovers add trials without evidence.
  • A setup command that is expected to fail while still producing its artifact must be guarded

(|| exit 0), or vally drops the trial.

  • A cleanup command that strips sources must skip directories containing SKILL.md — the staged

skill lives there, and deleting it aborts only the skilled arm.

Step 6: Write graders

Graders are hard pass/fail checks evaluated on every arm.

TypeRequired configPurpose
output-matches / output-not-matchespatternRegex over agent output
output-contains / output-not-containssubstringLiteral text in output
file-exists / file-not-existspathGlob against the work directory
file-contains / file-not-containspath, valueContent of a produced file
run-commandcommand (plus optional expected_exit_code, timeout, stdout_matches)Verify produced code actually builds/runs
exit-successAgent produced non-empty output
promptRuns the LLM judge against the rubric

Rules:

  • A grader whose config is absent or missing its required key parses fine and enforces nothing.

The usual cause is an indentation slip during an edit; check_eval_quality.py blocks it.

  • Prefer broad patterns that several valid approaches satisfy:

(root cause|primary error|underlying issue).

  • If the skill mandates an output shape, assert on it. A skill required to emit a decisive

Recommendation: line can silently stop doing so while the eval still passes.

  • Use file-not-contains / file-not-exists to prove the agent avoided an incorrect action.

Step 7: Write rubric items

Rubric items are judged pairwise (baseline vs. skilled). The overfitting judge classifies each item:

ClassificationDescriptionGoal
outcomeWhether the agent reached a correct result — WHAT, not HOWTarget this
techniqueWhether the agent used a skill-specific procedureMinimize
vocabularyWhether the agent used the skill's terminologyAvoid
  1. Test outcomes, not methods: "Identified the root cause of the build failure", not "Replayed the

binlog using dotnet build /flp".

  1. Accept any valid approach.
  2. Never reference the skill by name, and never reuse SKILL.md phrasing.
  3. Never reward using the skill — the harness reports activation separately, so a rubric item that

does this measures nothing and inflates the overfit score.

  1. Do not test knowledge the model already has; it adds no delta.
  2. Keep each item independently evaluable.
  3. Do not reward raw volume (test count, report length); judges will compare it when both arms act.

Good:

yaml
rubric:
  - Correctly identified the missing NuGet package as the root cause of the build failure
  - Recognized that downstream failures cascaded from that root cause
  - Suggested a concrete fix that resolves it

Overfitted:

yaml
rubric:
  - Replayed the binary log using 'dotnet build /flp:v=diag'   # technique
  - Measured cold, warm, and no-op build scenarios             # vocabulary
  - Used the template-comparison skill                         # rewards activation

Step 8: Add constraints sparingly

yaml
constraints:
  expect_tools: [bash]
  reject_tools: [edit, create]
  reject_skills: [some-skill]
  • expect_tools: [bash] on an advisory question forces a restore or build and converts an

answer into a timeout with no quality benefit. Only require tools when the task genuinely needs them.

  • reject_tools is the right way to keep a read-only stimulus read-only.

Step 9: Add dormancy guards

A dormancy guard proves the skill stays dormant on an off-target request that superficially matches it. Add one per real "when not to use" boundary: wrong input format, out-of-scope request, incompatible project type, wrong framework version, prerequisite absent.

yaml
  - name: Decline dump analysis request
    prompt: |
      I already have a .dmp crash dump from my .NET app. Can you help me
      analyze it to find the root cause of the crash?
    expect_activation: false
    graders:
      - type: output-matches
        config:
          pattern: (out of scope|not cover|does not|cannot|only.*collect)
      - type: prompt
    rubric:
      - Stated that dump analysis is out of scope
      - Did not open or analyze the dump file
      - Did not install analysis tools such as dotnet-dump analyze, lldb, or windbg
      - Suggested the correct alternative
Never combine `expect_activation: false` with `constraints.reject_skills`. That forces the skilled arm to run skill-free, so the harness cannot observe whether the target skill hijacks the request. The comparison remains visible as report-only evidence but does not vote in preference; unexpected isolated activation blocks a pass. expect_activation: false alone is the repo convention.

Guard rubrics verify three things: recognition (why it does not apply), restraint (no workflow, no file changes, no installs), redirection (the correct next step).

Step 10: Validate

bash
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/<plugin>
python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill-name>

For an agent eval, the third command is a no-op: agent.* is outside the experiment's evals: glob. Exercise one by pointing the runner at an experiment file whose glob includes it:

bash
# copy dotnet-skills.experiment.yaml, widen its evals: glob to tests/*/agent.*/eval.yaml
EXPERIMENT_FILE=my-agent.experiment.yaml ./eng/run-skill-evals.sh <plugin>

Read the trajectories rather than the verdict — there is no sign-test result for an agent eval.

check_eval_quality.py blocks eleven structural defect classes that can corrupt a result: missing or untracked fixtures, self-contradicting coverage fixtures, empty grader configs, dormancy guards with reject_skills, sub-floor stimulus counts, duplicate YAML keys or stimulus names, and config:/defaults: collisions. Do not add a new eval to eng/eval-quality/underpowered-allowlist.txt — the gate rejects allowlist entries that are new relative to the base branch.

For the official run, submit a PR review containing /evaluate so it binds to the reviewed commit.

Validation Checklist

  • [ ] Directory is tests/<plugin>/<skill-name>/ or tests/<plugin>/agent.<agent-name>/
  • [ ] Spec uses stimuli: / graders:, and exactly one of defaults: or config:
  • [ ] For a skill eval, at least 5 preference-eligible distinct stimuli exist; dormancy contracts do not count toward this floor (agent evals are exempt)
  • [ ] Each stimulus discriminates a different property and has a stable, unique name
  • [ ] Prompts never name the skill, the agent, or its vocabulary
  • [ ] Every referenced fixture exists and is tracked by git ls-files
  • [ ] Every fixture behaves as its stimulus assumes — healthy ones build, deliberately broken ones fail only for the stated reason
  • [ ] Every grader has its required config key
  • [ ] Any output shape the skill mandates has a grader
  • [ ] Rubric items are outcome-shaped and never reward using the skill
  • [ ] Dormancy guards use expect_activation: false alone
  • [ ] skill-validator check and check_eval_quality.py pass

Common Pitfalls

PitfallSolution
Writing scenarios: / assertions:That format no longer loads; use stimuli: / graders:
Adding defaults: runs: beside an existing config:Merge into one defaults: block
Landing an eval at exactly 5 stimuliA single tie makes a pass unreachable; size for the effect and tie rate
Raising runs to clear the floorRepeats measure reliability for one task; add stimuli
Prompt mentions the skill or agent by nameRewrite as a natural developer request
Rubric rewards using the skillDrop the item — the harness reports activation separately; rubrics measure outcomes
Fixture present but ignored by gitVerify with git ls-files; CI setup will fail otherwise
Fixture that does not build, or breaks for the wrong reasonFix the fixture before blaming the skill
Dormancy guard with reject_skillsUse expect_activation: false alone
expect_tools: [bash] on an advisory questionDrop it; it causes timeouts, not quality
Timeout too short for code generationUse ~360s; empty output fails every grader
Duplicate YAML key left behind by an editIt overwrites the next stimulus field by field — delete the stray block
Duplicate stimulus namesVally uses names as comparison identity — give every stimulus a stable, unique name
Direct eval for a disable-model-invocation: true skillRemove it and cover the reference through consumer outcomes
Agent eval sized for the stimulus flooragent.* evals get no verdict; size them for scenario coverage instead
Agent eval "run" with ./eng/run-skill-evals.shThe glob drops it — use a widened EXPERIMENT_FILE
Agent eval missing environment.skillsDeclare the skills the agent routes to, or it cannot invoke them
environment.skills set in a skill evalThe experiment varies that key and replaces it in every arm; the declaration does nothing
do mesmo repositório

Mais Skills

Todos os Skills
dotnet
Comunidade

collect-user-input

Build forms, validate data, and react to user input in Blazor. USE FOR adding forms, search boxes, filter panels, inline editing, data-entry UI, file uploads, validation (annotations or custom), handling form submissions, and binding input controls. Covers EditForm, built-in input components, DataAnnotationsValidator, custom validation, SSR form patterns (SupplyParameterFromForm, FormName, AntiforgeryToken, Enhance), and @bind for simple interactive controls. DO NOT USE for project scaffolding (see create-blazor-project) or prerendering issues (see support-prerendering).

instalações
3
GitHub Stars
5,5 mil
Atualizado
23 de set.
dotnet
Comunidade

dotnet-webapi

Guides creation and modification of ASP.NET Core Web API endpoints with correct HTTP semantics, OpenAPI metadata, and error handling. USE FOR: adding new API endpoints (controllers or minimal APIs), wiring up OpenAPI/Swagger, creating .http test files, setting up global error handling middleware. DO NOT USE FOR: general C coding style, EF Core data access or query optimization (use optimizing-ef-core-queries), frontend/Blazor work, gRPC services, or SignalR hubs.

instalações
3
GitHub Stars
5,5 mil
Atualizado
23 de set.
dotnet
Comunidade

test-anti-patterns

Audit a test file or suite; produce a severity-ranked diagnostic report. ALWAYS USE for tests that verify nothing, missing/tautological assertions, swallowed/broad exceptions, flaky/order-dependent tests, duplication, or magic values. Polyglot. DO NOT USE for direct edits: writing-mstest-tests owns supplied MSTest assertions/attributes/lifecycle; code-testing-agent owns new tests. Exclude running tests, migration, assertion metrics (assertion-quality), raw .NET coverage collection (run-tests), non-.NET coverage collection/analysis (native tooling), project-wide .NET coverage/CRAP (coverage-analysis), named-target .NET CRAP (crap-score), behavioral/pseudo-mutation gaps (test-gap-analysis), test-mix/ happy-vs-error classification and trait distributions (test-tagging), or the testsmells.org catalog (test-smell-detection).

instalações
4
GitHub Stars
5,5 mil
Atualizado
23 de set.
dotnet
Comunidade

binlog-failure-analysis

Analyze MSBuild binary logs to diagnose build failures. USE FOR: build errors that are unclear from console output, diagnosing cascading failures across multi-project builds, tracing MSBuild target execution order, and generally any MSBuild build issues. Requires an existing .binlog file. DO NOT USE FOR: generating binlogs (use binlog-generation), non-MSBuild build systems.

instalações
1
GitHub Stars
5,5 mil
Atualizado
22 de set.