Renderizado do repositório de origem, preservando títulos, exemplos, código, tabelas, links e imagens.
Create Skill Test
Scaffold an evaluation spec (eval.yaml) for a skill or agent so it conforms to the Vally schema, passes skill-validator check and check_eval_quality.py, is powerful enough to return a verdict, and does not overfit to the skill's own wording.
When to Use
- Creating a new
eval.yamlfor a skill or agent - Adding stimuli to an existing eval
- Sizing an eval so the pass gate can actually be reached
- Setting up or repairing fixture files alongside an eval
- Reviewing whether rubric items and graders risk overfitting
When Not to Use
- Diagnosing a failing or regressed eval — use
improve-skill-quality - Modifying the skill-validator or the evaluation workflows
- Creating or editing
SKILL.mdfiles — usecreate-skill
Inputs
| Input | Required | Description |
|---|---|---|
| Skill or agent name | Yes | Must exist under plugins/<plugin>/skills/ or plugins/<plugin>/agents/ |
| Plugin name | Yes | e.g. dotnet-msbuild |
| Skill content | Yes | Read it — you cannot write non-overfitted rubric items without it |
| Failure modes to discriminate | Recommended | Each becomes one stimulus |
Workflow
Step 1: Locate the target and the test directory
tests/<plugin>/<skill-name>/eval.yaml # skills
tests/<plugin>/agent.<agent-name>/eval.yaml # agents (the agent. prefix disambiguates)Verify the target exists at plugins/<plugin>/skills/<skill-name>/SKILL.md or plugins/<plugin>/agents/<agent-name>.agent.md, and read it.
Agent evals sit outside the verdict flow. The canonical experiment declares evals: tests/*/!(agent.*)/eval.yaml, so agent.* specs are excluded: no verdict is ever computed for them, the stimulus floor does not apply, and ./eng/run-skill-evals.sh drops them even when you name one explicitly (its --eval-filter is intersected with that glob). The distinct-stimulus floor therefore applies to skill evals only. Author agent evals for the scenario coverage and the deterministic graders, and run them as described in Step 10.
Be careful with a skill that sets `disable-model-invocation: true`. The model cannot invoke it, so the skill is absent from the model-facing skilled arm and any direct eval compares two identical arms. Answer-content graders do not create a difference between those arms. The honest coverage for such skills is dependency-level — through the outcome evals of the skills that load them, and through the plugin arm. For example, filter-syntax is covered by the filtered-command scenarios in tests/dotnet-test/run-tests/eval.yaml.
Step 2: Write the spec skeleton
The spec is Vally format. Every eval in this repo uses stimuli: and graders:; scenarios: and assertions: are a pre-Vally format that no longer loads.
name: <skill-name>
description: Evaluates the <plugin>/<skill-name> skill
type: capability
defaults:
timeout: 5m
runs: 1
stimuli:
- name: <what the agent must accomplish>
prompt: <natural developer request>
environment:
files:
- src: fixtures/<case>/Project.csproj
dest: Project.csproj
graders:
- type: output-matches
config:
pattern: (root cause|underlying issue)
- type: exit-success
- type: prompt
rubric:
- <outcome the agent should have reached>`defaults:` replaces `config:` — it does not join it.configis a deprecated alias for the same block and vally throws on a spec declaring both. Some existing evals still open withconfig:; when you change settings, replace it with onedefaults:block. The failure is invisible otherwise: the job exits 0 with no verdicts and the PR comment blames "transient infrastructure".
Step 3: Size the eval for power before writing content
The gate gives each distinct stimulus one vote. Repeated runs for one stimulus collapse to one majority-direction vote and remain available as reliability evidence.
- Distinct stimuli ≥ 5, else the verdict is
underpowered— never a pass, never a regression. - p ≤ 0.05 on an exact one-sided sign test over discordant (non-tie) stimulus votes.* Ties are not
discarded; they hold the discordant count down.
| discordant stimulus votes | records that pass | p |
|---|---|---|
| ≤ 4 | none | ≥ 0.0625 |
| 5–7 | zero losses only (5W/0L) | 0.031 |
| 8 | one loss survivable (7W/1L) | 0.035 |
At exactly 5 stimuli, one tie is fatal because it leaves 4 discordant votes. At 6 stimuli one tie is survivable; at 7, up to two are. A loss is not. Five is an eligibility floor, not adequate power. For example, 80% power needs 8 discordant votes only for a true 90% conditional win rate; it needs 18 at 80%, 37 at 70%, and 158 at 60%. Size for the effect and tie rate you need to detect.
Use runs for reliability, not task breadth. Vally recommends 3 runs in CI and 5–10 nightly for pass rate, pass@k, pass^k, and flakiness. Extra runs never clear the five-stimulus floor.
Do not set runs in dotnet-skills.experiment.yaml; experiment overrides overwrite every eval's own value rather than defaulting it.
Step 4: Write stimuli
- Name describes what is tested, not how.
- Prompt is a natural developer request. Never mention the skill, the agent, or its vocabulary —
cued prompts inflate the overfit score and bias the baseline.
- Each stimulus should discriminate a different property of the skill. Five stimuli covering one
property give arithmetic, not evidence.
- Give every stimulus a stable, unique
name. Vally pairs comparison trajectories by
(stimulus name, trial index); duplicate names make slot identity ambiguous.
- Include a boundary / no-op stimulus for any skill that migrates or rewrites code, proving it
leaves already-correct input alone.
Step 5: Configure the environment
environment:
files:
- src: fixtures/broken-build/App.csproj # path relative to eval.yaml
dest: App.csproj # path in the agent's working directory
- src: fixtures/broken-build # a directory
dest: .
commands:
- dotnet build -bl || exit 0 # guard intentional failuresDo not set `environment.skills` in a skill eval. The experiment declares vary: /environment/skills and supplies the value itself — [] for the baseline arm and plugins/<plugin>/skills/<skill> for the skilled arm — so anything the eval declares is replaced, in every arm. It cannot add a skill to one arm only. environment.skills is meaningful only in an agent.* eval, which the experiment does not vary; there it is the set of skills the agent may invoke. Copy the shape from an existing agent eval such as tests/dotnet-test/agent.test-quality-auditor/eval.yaml rather than reproducing a remembered form — the specs in this repo are not consistent about how they spell those entries.
Fixture rules — each one has already cost a real result:
- Every referenced fixture must be tracked by git.
.gitignore(e.g.coverage*.xml) has
silently swallowed a committed fixture: the eval passed locally and failed at setup in CI. Verify with git ls-files, not by looking at the working tree.
- Every fixture must behave as its stimulus assumes. A fixture meant to be healthy must build; a
fixture meant to be broken must fail for the exact reason the stimulus is about, and no other. Judges penalize agents for unrelated "pre-existing build issues" that the fixture author introduced.
- Every fixture must reproduce the bug its stimulus is named for. If it does not, the baseline
scores well and the skill has nothing to add.
- Coverage fixtures must be internally consistent. A Cobertura report whose declared
line-rate, summary totals (lines-covered/lines-valid), and <line> elements disagree lets the two arms read different truths, and the loss is the fixture's fault. Update any rubric item or prompt that quotes a figure in the same change.
- Do not wire duplicate fixtures to raise
n; rename leftovers add trials without evidence. - A setup command that is expected to fail while still producing its artifact must be guarded
(|| exit 0), or vally drops the trial.
- A cleanup command that strips sources must skip directories containing
SKILL.md— the staged
skill lives there, and deleting it aborts only the skilled arm.
Step 6: Write graders
Graders are hard pass/fail checks evaluated on every arm.
| Type | Required config | Purpose |
|---|---|---|
output-matches / output-not-matches | pattern | Regex over agent output |
output-contains / output-not-contains | substring | Literal text in output |
file-exists / file-not-exists | path | Glob against the work directory |
file-contains / file-not-contains | path, value | Content of a produced file |
run-command | command (plus optional expected_exit_code, timeout, stdout_matches) | Verify produced code actually builds/runs |
exit-success | — | Agent produced non-empty output |
prompt | — | Runs the LLM judge against the rubric |
Rules:
- A grader whose
configis absent or missing its required key parses fine and enforces nothing.
The usual cause is an indentation slip during an edit; check_eval_quality.py blocks it.
- Prefer broad patterns that several valid approaches satisfy:
(root cause|primary error|underlying issue).
- If the skill mandates an output shape, assert on it. A skill required to emit a decisive
Recommendation: line can silently stop doing so while the eval still passes.
- Use
file-not-contains/file-not-existsto prove the agent avoided an incorrect action.
Step 7: Write rubric items
Rubric items are judged pairwise (baseline vs. skilled). The overfitting judge classifies each item:
| Classification | Description | Goal |
|---|---|---|
| outcome | Whether the agent reached a correct result — WHAT, not HOW | Target this |
| technique | Whether the agent used a skill-specific procedure | Minimize |
| vocabulary | Whether the agent used the skill's terminology | Avoid |
- Test outcomes, not methods: "Identified the root cause of the build failure", not "Replayed the
binlog using dotnet build /flp".
- Accept any valid approach.
- Never reference the skill by name, and never reuse
SKILL.mdphrasing. - Never reward using the skill — the harness reports activation separately, so a rubric item that
does this measures nothing and inflates the overfit score.
- Do not test knowledge the model already has; it adds no delta.
- Keep each item independently evaluable.
- Do not reward raw volume (test count, report length); judges will compare it when both arms act.
Good:
rubric:
- Correctly identified the missing NuGet package as the root cause of the build failure
- Recognized that downstream failures cascaded from that root cause
- Suggested a concrete fix that resolves itOverfitted:
rubric:
- Replayed the binary log using 'dotnet build /flp:v=diag' # technique
- Measured cold, warm, and no-op build scenarios # vocabulary
- Used the template-comparison skill # rewards activationStep 8: Add constraints sparingly
constraints:
expect_tools: [bash]
reject_tools: [edit, create]
reject_skills: [some-skill]expect_tools: [bash]on an advisory question forces a restore or build and converts an
answer into a timeout with no quality benefit. Only require tools when the task genuinely needs them.
reject_toolsis the right way to keep a read-only stimulus read-only.
Step 9: Add dormancy guards
A dormancy guard proves the skill stays dormant on an off-target request that superficially matches it. Add one per real "when not to use" boundary: wrong input format, out-of-scope request, incompatible project type, wrong framework version, prerequisite absent.
- name: Decline dump analysis request
prompt: |
I already have a .dmp crash dump from my .NET app. Can you help me
analyze it to find the root cause of the crash?
expect_activation: false
graders:
- type: output-matches
config:
pattern: (out of scope|not cover|does not|cannot|only.*collect)
- type: prompt
rubric:
- Stated that dump analysis is out of scope
- Did not open or analyze the dump file
- Did not install analysis tools such as dotnet-dump analyze, lldb, or windbg
- Suggested the correct alternativeNever combine `expect_activation: false` with `constraints.reject_skills`. That forces the skilled arm to run skill-free, so the harness cannot observe whether the target skill hijacks the request. The comparison remains visible as report-only evidence but does not vote in preference; unexpected isolated activation blocks a pass. expect_activation: false alone is the repo convention.Guard rubrics verify three things: recognition (why it does not apply), restraint (no workflow, no file changes, no installs), redirection (the correct next step).
Step 10: Validate
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/<plugin>
python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill-name>For an agent eval, the third command is a no-op: agent.* is outside the experiment's evals: glob. Exercise one by pointing the runner at an experiment file whose glob includes it:
# copy dotnet-skills.experiment.yaml, widen its evals: glob to tests/*/agent.*/eval.yaml
EXPERIMENT_FILE=my-agent.experiment.yaml ./eng/run-skill-evals.sh <plugin>Read the trajectories rather than the verdict — there is no sign-test result for an agent eval.
check_eval_quality.py blocks eleven structural defect classes that can corrupt a result: missing or untracked fixtures, self-contradicting coverage fixtures, empty grader configs, dormancy guards with reject_skills, sub-floor stimulus counts, duplicate YAML keys or stimulus names, and config:/defaults: collisions. Do not add a new eval to eng/eval-quality/underpowered-allowlist.txt — the gate rejects allowlist entries that are new relative to the base branch.
For the official run, submit a PR review containing /evaluate so it binds to the reviewed commit.
Validation Checklist
- [ ] Directory is
tests/<plugin>/<skill-name>/ortests/<plugin>/agent.<agent-name>/ - [ ] Spec uses
stimuli:/graders:, and exactly one ofdefaults:orconfig: - [ ] For a skill eval, at least 5 preference-eligible distinct stimuli exist; dormancy contracts do not count toward this floor (agent evals are exempt)
- [ ] Each stimulus discriminates a different property and has a stable, unique name
- [ ] Prompts never name the skill, the agent, or its vocabulary
- [ ] Every referenced fixture exists and is tracked by
git ls-files - [ ] Every fixture behaves as its stimulus assumes — healthy ones build, deliberately broken ones fail only for the stated reason
- [ ] Every grader has its required
configkey - [ ] Any output shape the skill mandates has a grader
- [ ] Rubric items are outcome-shaped and never reward using the skill
- [ ] Dormancy guards use
expect_activation: falsealone - [ ]
skill-validator checkandcheck_eval_quality.pypass
Common Pitfalls
| Pitfall | Solution |
|---|---|
Writing scenarios: / assertions: | That format no longer loads; use stimuli: / graders: |
Adding defaults: runs: beside an existing config: | Merge into one defaults: block |
| Landing an eval at exactly 5 stimuli | A single tie makes a pass unreachable; size for the effect and tie rate |
Raising runs to clear the floor | Repeats measure reliability for one task; add stimuli |
| Prompt mentions the skill or agent by name | Rewrite as a natural developer request |
| Rubric rewards using the skill | Drop the item — the harness reports activation separately; rubrics measure outcomes |
| Fixture present but ignored by git | Verify with git ls-files; CI setup will fail otherwise |
| Fixture that does not build, or breaks for the wrong reason | Fix the fixture before blaming the skill |
Dormancy guard with reject_skills | Use expect_activation: false alone |
expect_tools: [bash] on an advisory question | Drop it; it causes timeouts, not quality |
| Timeout too short for code generation | Use ~360s; empty output fails every grader |
| Duplicate YAML key left behind by an edit | It overwrites the next stimulus field by field — delete the stray block |
| Duplicate stimulus names | Vally uses names as comparison identity — give every stimulus a stable, unique name |
Direct eval for a disable-model-invocation: true skill | Remove it and cover the reference through consumer outcomes |
| Agent eval sized for the stimulus floor | agent.* evals get no verdict; size them for scenario coverage instead |
Agent eval "run" with ./eng/run-skill-evals.sh | The glob drops it — use a widened EXPERIMENT_FILE |
Agent eval missing environment.skills | Declare the skills the agent routes to, or it cannot invoke them |
environment.skills set in a skill eval | The experiment varies that key and replaces it in every arm; the declaration does nothing |

