Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 30 additions & 3 deletions .agents/skills/create-skill-test/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,17 +1,17 @@
---
name: create-skill-test
description: Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and overfitting avoidance. Do not use for running or debugging existing evals (use improve-skill-quality) nor for skills authoring (use create-skill).
description: Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository. Use when creating skill or workflow-package tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and overfitting avoidance. Do not use for running or debugging existing evals (use improve-skill-quality) nor for skills authoring (use create-skill).
---

# Create Skill Test

Scaffold an evaluation spec (`eval.yaml`) for a skill or agent so it conforms to the Vally schema,
Scaffold an evaluation spec (`eval.yaml`) for a skill, agent, or workflow package so it conforms to the Vally schema,
passes `skill-validator check` and `check_eval_quality.py`, is powerful enough to return a verdict,
and does not overfit to the skill's own wording.

## When to Use

- Creating a new `eval.yaml` for a skill or agent
- Creating a new `eval.yaml` for a skill, agent, or workflow package
- Adding stimuli to an existing eval
- Sizing an eval so the pass gate can actually be reached
- Setting up or repairing fixture files alongside an eval
Expand Down Expand Up @@ -61,6 +61,7 @@ Then locate the target and test directory:
```text
tests/<plugin>/<skill-name>/eval.yaml # skills
tests/<plugin>/agent.<agent-name>/eval.yaml # agents (the agent. prefix disambiguates)
tests/agentic-workflows/<package>/eval.yaml # redistributable gh-aw packages
```

Verify the target exists at `plugins/<plugin>/skills/<skill-name>/SKILL.md` or
Expand Down Expand Up @@ -328,6 +329,32 @@ incompatible project type, wrong framework version, prerequisite absent.
> unexpected isolated activation blocks a pass. `expect_activation: false` **alone** is the repo
> convention.

### Workflow-package scenarios

For a package target, verify `agentic-workflows/<package>/aw.yml`, then read its
entry workflow, local imports, and bundled agents. The native SDK lane evaluates
their real prompt bodies and installed resources against offline fixtures.
Specify collector outputs, revision/tracking evidence, and service responses as
fixture inputs; propose terminal actions in `result.json` rather than pretending
to publish through live GitHub or safe-output tools. Assert the structured result
with deterministic graders. Do not place expected answers in agent-readable
fixtures or staged grader scripts; pass expected values through grader argv.

Prompt expressions are rendered from a flat `workflow-context.json` fixture,
whose keys are exact trimmed expressions and values are strings. Missing context
fails setup. A workflow that correctly chooses noop is still expected-active
decision evidence, not `expect_activation: false` routing evidence. Include
normal, partial, stale, incompatible, missing-evidence, and multi-module cases
where applicable. Keep compilation, helper execution, and actual consumer
publication tests separate: this lane is labeled `workflow-prompt-sdk`, not
end-to-end Actions execution.

```powershell
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate `
agentic-workflows/<package>/aw.yml `
--tests-dir tests/agentic-workflows --runs 1 --verdict-warn-only
```

Guard rubrics verify three things: **recognition** (why it does not apply), **restraint** (no
workflow, no file changes, no installs), **redirection** (the correct next step).

Expand Down
1 change: 1 addition & 0 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
/eng/ @AbhitejJohn @JanKrivanek
/.github/workflows/ @AbhitejJohn @JanKrivanek
/agentic-workflows/ @YuliiaKovalova @JanKrivanek @Evangelink
/tests/agentic-workflows/ @YuliiaKovalova @JanKrivanek @Evangelink

# msbuild
/plugins/dotnet-msbuild/ @dotnet/msbuild @JanKrivanek @YuliiaKovalova
Expand Down
5 changes: 5 additions & 0 deletions .github/workflows/agentic-workflow-validation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ on:
- ".github/workflows/**"
- "agentic-workflows/**"
- "eng/agentic-workflows/**"
- "tests/agentic-workflows/**"
push:
branches: [main]
paths:
Expand All @@ -18,6 +19,7 @@ on:
- ".github/workflows/**"
- "agentic-workflows/**"
- "eng/agentic-workflows/**"
- "tests/agentic-workflows/**"
workflow_dispatch:

permissions:
Expand Down Expand Up @@ -62,3 +64,6 @@ jobs:

- name: Test build-failure operational-value grader
run: python eng/agentic-workflows/test_build_failure_analysis_operational_value.py

- name: Test offline workflow scenario graders
run: python tests/agentic-workflows/test_graders.py
Comment thread
Evangelink marked this conversation as resolved.
70 changes: 63 additions & 7 deletions .github/workflows/evaluation-run.yml
Original file line number Diff line number Diff line change
Expand Up @@ -244,7 +244,43 @@ jobs:
}
}

if ($s) {
if ($p -eq "agentic-workflows") {
# Do not execute discovery helpers from the evaluated checkout.
$root = [IO.Path]::GetFullPath("agentic-workflows")
if (Test-PathHasReparsePoint -allowedRoot $PWD -path $root) {
throw "Workflow collection is missing or contains a reparse point"
}
$packages = @(Get-ChildItem -LiteralPath $root -Directory | Sort-Object Name)
if ($s) {
if ($s -notin $packages.Name) { throw "Unknown workflow package '$s'" }
$packages = @($packages | Where-Object { $_.Name -eq $s })
}
$entries = @($packages | ForEach-Object {
$name = $_.Name
if ($name -notmatch '\A[A-Za-z0-9_-]+\z') {
throw "Invalid workflow package name '$name'"
}
$manifest = "agentic-workflows/$name/aw.yml"
$eval = "tests/agentic-workflows/$name/eval.yaml"
foreach ($path in @($manifest, $eval)) {
if (-not (Test-Path $path -PathType Leaf)) { throw "Workflow package '$name' has no $path" }
if (Test-PathHasReparsePoint -allowedRoot $PWD -path $path) {
throw "Workflow evaluation path contains a reparse point: $path"
}
}
@{
name = "agentic-workflows--$name"
plugin = "agentic-workflows"
target_kind = "workflow"
skills_path = ""
agents_path = ""
package_path = $manifest
eval_path = $eval
}
})
if ($entries.Count -eq 0) { throw "No workflow packages found" }
$json = $entries | ConvertTo-Json -Compress -AsArray
} elseif ($s) {
if ($s.StartsWith("agent.")) {
$agent = $s.Substring("agent.".Length)
if (-not $agent) { throw "Invalid agent target '$s'" }
Expand Down Expand Up @@ -378,6 +414,7 @@ jobs:
ENTRY_SKILLS_PATH: ${{ matrix.entry.skills_path }}
ENTRY_AGENTS_PATH: ${{ matrix.entry.agents_path }}
ENTRY_EVAL_PATH: ${{ matrix.entry.eval_path }}
ENTRY_PACKAGE_PATH: ${{ matrix.entry.package_path }}
ENTRY_MODEL: ${{ matrix.entry.model }}
ENTRY_JUDGE: ${{ matrix.entry.judge }}
run: |
Expand All @@ -399,8 +436,8 @@ jobs:
echo "::error::Invalid matrix value '$val' (must match $name_re, not be '.', and not contain '..')"; exit 1
fi
done
if [ "$ENTRY_TARGET_KIND" != "skill" ] && [ "$ENTRY_TARGET_KIND" != "agent" ]; then
echo "::error::Invalid target_kind '$ENTRY_TARGET_KIND' (must be skill or agent)"; exit 1
if [ "$ENTRY_TARGET_KIND" != "skill" ] && [ "$ENTRY_TARGET_KIND" != "agent" ] && [ "$ENTRY_TARGET_KIND" != "workflow" ]; then
echo "::error::Invalid target_kind '$ENTRY_TARGET_KIND' (must be skill, agent, or workflow)"; exit 1
fi
# Cross-family executor/judge fields (IMPACT-ANALYSIS.md §10) are
# populated when the discover job expands the selected profile.
Expand Down Expand Up @@ -466,6 +503,20 @@ jobs:
if [ "$ENTRY_TARGET_KIND" = "agent" ] && [ -z "$ENTRY_EVAL_PATH" ]; then
echo "::error::Agent matrix entry has an empty eval_path"; exit 1
fi
if [ "$ENTRY_TARGET_KIND" = "workflow" ]; then
package_re='^agentic-workflows/([A-Za-z0-9_-]+)/aw\.yml$'
if [ "$ENTRY_PLUGIN" != "agentic-workflows" ] || ! [[ "$ENTRY_PACKAGE_PATH" =~ $package_re ]]; then
echo "::error::Invalid workflow package path '$ENTRY_PACKAGE_PATH'"; exit 1
fi
package_name="${BASH_REMATCH[1]}"
expected_name="agentic-workflows--$package_name"
if [ -n "$ENTRY_MODEL" ]; then expected_name="$expected_name--$ENTRY_MODEL"; fi
if [ "$ENTRY_EVAL_PATH" != "tests/agentic-workflows/$package_name/eval.yaml" ] ||
[ "$ENTRY_NAME" != "$expected_name" ] ||
[ ${#segs[@]} -ne 0 ] || [ ${#agent_segs[@]} -ne 0 ]; then
echo "::error::Workflow matrix identity does not match package '$package_name'"; exit 1
fi
fi

- name: Checkout skills content
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
Expand All @@ -488,7 +539,7 @@ jobs:
# eval. Otherwise derive the specs from this entry's skills_path so a
# subset leg (per-skill PR entry or shard) reports has_evals accurately
# instead of installing tools or probing a PAT only to exit empty.
if [ "$TARGET_KIND" = "agent" ]; then
if [ "$TARGET_KIND" = "agent" ] || [ "$TARGET_KIND" = "workflow" ]; then
if [ ! -f "$EVAL_PATH" ]; then
echo "::error::Agent eval spec does not exist at '$EVAL_PATH'"
exit 1
Expand Down Expand Up @@ -819,6 +870,7 @@ jobs:
TARGET_KIND: ${{ matrix.entry.target_kind }}
SKILLS_PATH: ${{ matrix.entry.skills_path }}
AGENTS_PATH: ${{ matrix.entry.agents_path }}
PACKAGE_PATH: ${{ matrix.entry.package_path }}
DISPATCH_SKILL: ${{ inputs.skill }}
MODEL: ${{ steps.eval-models.outputs.model }}
JUDGE_MODEL: ${{ steps.eval-models.outputs.judge-model }}
Expand All @@ -836,14 +888,18 @@ jobs:
rm -f "$RUNNER_TEMP/evaluation-copilot-token"
export GITHUB_TOKEN

if [ "$TARGET_KIND" = "agent" ]; then
if [ "$TARGET_KIND" = "agent" ] || [ "$TARGET_KIND" = "workflow" ]; then
# Vally 0.14 has no custom-agent registration field: its public
# EnvironmentConfig exposes skills/files/commands/MCP only, and its
# Copilot executor passes skillDirectories but no customAgents.
# Run the repository's native SDK agent evaluator instead, then
# adapt its evidence into the same schema/result tree as Vally.
AGENT_ARGS=()
for token in $AGENTS_PATH; do AGENT_ARGS+=("$token"); done
if [ "$TARGET_KIND" = "workflow" ]; then
AGENT_ARGS=("$PACKAGE_PATH")
else
for token in $AGENTS_PATH; do AGENT_ARGS+=("$token"); done
fi
if [ ${#AGENT_ARGS[@]} -eq 0 ]; then
echo "::error::No custom-agent paths were supplied for $PLUGIN"
exit 1
Expand Down Expand Up @@ -1198,7 +1254,7 @@ jobs:
run: |
echo "## 🔬 Evaluation Results: $ENTRY_NAME" >> $GITHUB_STEP_SUMMARY
echo "" >> $GITHUB_STEP_SUMMARY
echo "Isolated target vs baseline, judged head-to-head. Skill targets use \`vally compare\`; custom-agent targets use the native SDK evaluator. Each distinct stimulus gives one gate vote; repeated runs report reliability only. A pass needs a complete comparison, an exact one-sided sign test at p ≤ 0.05, and at least a 20% net win. ⚠️ marks an invalid or inconclusive result. 📉 marks a report-only LLM preference loss, not an objective completion regression." >> $GITHUB_STEP_SUMMARY
echo "Isolated target vs baseline, judged head-to-head. Skill targets use \`vally compare\`; custom-agent and workflow-package targets use the native SDK evaluator. Workflow-package results evaluate imported prompts against offline fixtures, not live Actions bootstrap or publication. Each distinct stimulus gives one gate vote; repeated runs report reliability only. A pass needs a complete comparison, an exact one-sided sign test at p ≤ 0.05, and at least a 20% net win. ⚠️ marks an invalid or inconclusive result. 📉 marks a report-only LLM preference loss, not an objective completion regression." >> $GITHUB_STEP_SUMMARY
echo "" >> $GITHUB_STEP_SUMMARY

RESULTS_DIR="artifacts/TestResults/vally/$ENTRY_NAME"
Expand Down
5 changes: 5 additions & 0 deletions .github/workflows/evaluation-workflow-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,8 @@ on:
- "eng/evaluation/test_pr_triage_retry.py"
- "eng/evaluation/find-targets.ps1"
- "eng/evaluation/path-safety.ps1"
- "eng/evaluation/workflow-targets.ps1"
- "eng/evaluation/test_workflow_targets.py"
- "eng/dashboard/**"
- "eng/vally-adapter/**"
- "eng/skill-validator/src/**"
Expand All @@ -49,6 +51,8 @@ on:
- "eng/evaluation/test_pr_triage_retry.py"
- "eng/evaluation/find-targets.ps1"
- "eng/evaluation/path-safety.ps1"
- "eng/evaluation/workflow-targets.ps1"
- "eng/evaluation/test_workflow_targets.py"
- "eng/dashboard/**"
- "eng/vally-adapter/**"
- "eng/skill-validator/src/**"
Expand Down Expand Up @@ -127,4 +131,5 @@ jobs:
- name: Test evaluation workflow behavior
run: |
python eng/evaluation/test_token_failover.py
python eng/evaluation/test_workflow_targets.py
python eng/evaluation/test_pr_triage_retry.py
10 changes: 5 additions & 5 deletions .github/workflows/evaluation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ on:
workflow_dispatch:
inputs:
plugin:
description: "Specific plugin to evaluate (leave blank for all)"
description: "Plugin or agentic-workflows collection to evaluate (leave blank for all)"
type: string
required: false
pr_number:
Expand Down Expand Up @@ -186,7 +186,7 @@ jobs:
$changedFiles = git diff --name-only --diff-filter=ACMR $mergeBase $head

$hasSkillChanges = $changedFiles |
Where-Object { $_ -match '^(?:plugins/[^/]+/plugin\.json$|plugins/[^/]+/skills/[^/]+/|plugins/[^/]+/(?:[^/]+/)*[^/]+\.agent\.md$|tests/[^/]+/[^/]+/)' } |
Where-Object { $_ -match '^(?:plugins/[^/]+/plugin\.json$|plugins/[^/]+/skills/[^/]+/|plugins/[^/]+/(?:[^/]+/)*[^/]+\.agent\.md$|tests/[^/]+/[^/]+/|tests/agentic-workflows/|agentic-workflows/|\.github/graders/)' } |
Select-Object -First 1

# Evaluation pipeline changes need a re-eval. The native custom-agent
Expand All @@ -195,7 +195,7 @@ jobs:
$hasInfraChanges = $changedFiles |
Where-Object {
($_ -match '^eng/vally-adapter/') -or
($_ -match '^eng/evaluation/(?:find-targets|path-safety)\.ps1$') -or
($_ -match '^eng/evaluation/(?:find-targets|path-safety|workflow-targets)\.ps1$') -or
($_ -match '^eng/skill-validator/src/') -or
($_ -match '^dotnet-skills\.experiment\.yaml$') -or
$_ -match '^\.github/workflows/(evaluation|evaluation-run)\.yml$'
Expand Down Expand Up @@ -258,13 +258,13 @@ jobs:
$changedFiles = git diff --name-only --diff-filter=ACMR $mergeBase $head

$hasSkillChanges = $changedFiles |
Where-Object { $_ -match '^(?:plugins/[^/]+/plugin\.json$|plugins/[^/]+/skills/[^/]+/|plugins/[^/]+/(?:[^/]+/)*[^/]+\.agent\.md$|tests/[^/]+/[^/]+/)' } |
Where-Object { $_ -match '^(?:plugins/[^/]+/plugin\.json$|plugins/[^/]+/skills/[^/]+/|plugins/[^/]+/(?:[^/]+/)*[^/]+\.agent\.md$|tests/[^/]+/[^/]+/|tests/agentic-workflows/|agentic-workflows/|\.github/graders/)' } |
Select-Object -First 1

$hasInfraChanges = $changedFiles |
Where-Object {
($_ -match '^eng/vally-adapter/') -or
($_ -match '^eng/evaluation/(?:find-targets|path-safety)\.ps1$') -or
($_ -match '^eng/evaluation/(?:find-targets|path-safety|workflow-targets)\.ps1$') -or
($_ -match '^eng/skill-validator/src/') -or
($_ -match '^dotnet-skills\.experiment\.yaml$') -or
$_ -match '^\.github/workflows/(evaluation|evaluation-run)\.yml$'
Expand Down
12 changes: 12 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -446,6 +446,11 @@ dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate \
plugins/dotnet-msbuild/agents/msbuild.agent.md \
--tests-dir tests/dotnet-msbuild --runs 1 --verdict-warn-only

# Exercise one redistributable workflow package against offline fixtures
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate \
agentic-workflows/msbuild-quality-review/aw.yml \
--tests-dir tests/agentic-workflows --runs 1 --verdict-warn-only

# Run every skill's tests
./eng/run-skill-evals.sh
```
Expand All @@ -462,6 +467,13 @@ Per-skill verdicts are written to `./eval-results/<plugin>/<skill>/results.json`

Tests do **not** run automatically on pull requests. When a PR changes skills, the `pr-status` job posts a pending commit status and a maintainer must trigger the evaluation, binding it to a specific reviewed commit — either by submitting a PR review ("Files changed" → "Review changes") whose body contains `/evaluate` (recommended, no SHA to copy), or by commenting `/evaluate <sha>`. A bare `/evaluate` comment only posts guidance. Results are posted as a PR comment and uploaded as build artifacts.

The same discovery/reporting pipeline covers workflow packages and their specs
under `tests/agentic-workflows/<package>/eval.yaml`. Manual dispatch with
`plugin: agentic-workflows` selects that collection. These results measure
offline workflow decisions/proposals using the real imported prompts and
installed resources; they do not claim live Actions or publication validation.
Keep package compilation and trusted-helper regression tests as separate gates.

The [Skill Value dashboard](https://dotnet.github.io/skills/) provides historical
results for each skill by executor and judge model.

Expand Down
27 changes: 27 additions & 0 deletions agentic-workflows/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,3 +34,30 @@ gh aw update
The installer copies each workflow source and its dependencies into the consumer
repository, then generates the executable `.lock.yml` file there. Generated lock files
are therefore not stored in this distribution directory.

## Scenario evaluation

Every package has a Vally-format spec at
`tests/agentic-workflows/<package>/eval.yaml`. The normal `/evaluate` discovery
includes package sources, resources, shared graders, and these scenarios.
Scheduled evaluations include the collection; manual evaluation dispatch with
`plugin: agentic-workflows` selects only these packages.

For a local run:

```powershell
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate `
agentic-workflows/msbuild-quality-review/aw.yml `
--tests-dir tests/agentic-workflows --runs 1 --verdict-warn-only
```

The native SDK lane loads real local imports and installed package resources,
compares them with a no-workflow baseline, and retains a package-agent arm.
Results are published through the normal pipeline with `skillKind: workflow`
and `evaluationLane: workflow-prompt-sdk`.

These are **offline prompt/decision evaluations**: fixture evidence replaces
collectors and external services, and `result.json` contains proposed actions.
They do not execute Actions bootstrap jobs or publish safe outputs. Compilation,
trusted-helper tests, runtime trace graders, and consumer-repository integration
runs cover different contracts and remain separate evidence.
Loading
Loading