You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Add scenario evaluation for agentic workflow packages - #1268
Redistributable gh-aw packages were outside /evaluate discovery and had no scenario specs. Runtime trace graders did not provide fixture-based A/B coverage or dashboard results. This adds a first-class offline workflow-package evaluation lane to the existing pipeline.
Load individual aw.yml targets, expand local prompt imports, stage resources at their installed paths, and require explicit fixture context. Compare baseline, isolated workflow, and registered package-agent surfaces, retaining proposed result.json output for judging and investigation.
Add 43 distinct stimuli: build-failure analysis (10), MSBuild quality review (10), test-failure analysis (12), and unskip tests (11). Generic structured-output graders receive expectations through evaluator-only arguments rather than model-visible answer maps.
Discover package/resource/eval changes on PRs and include packages in scheduled/manual runs. Preserve native retry/rejudge and result accounting, with skillKind: workflow and evaluationLane: workflow-prompt-sdk in reports and dashboard data.
Update CI, authoring/investigation guidance, package documentation, and ownership of the new eval collection.
Scope: These are offline decision/proposal evaluations using real package prompts and installed resources. They do not execute Actions bootstrap jobs, call live collectors, or publish safe outputs. Package compilation, trusted-helper tests, and runtime trace grading remain separate evidence.
python eng\agentic-workflows\validate_agentic_workflows.py - passed with a session-isolated, checksum-verified gh-aw v0.89.15 compiler; active locks and installed packages validated.
python eng\agentic-workflows\test_build_failure_analysis_operational_value.py - passed, 27 fixtures executing the real runtime grader.
dotnet publish eng\skill-validator\src\SkillValidator.csproj --no-restore -r win-x64 -p:PublishAot=false -o artifacts\workflow-eval-tool --nologo -v:q - passed; smoke used the published layout with Copilot SDK 1.0.11.
Checksum-verified actionlint 1.7.7 with -shellcheck= -pyflakes= on the four changed workflow YAML files - passed.
git diff --cached --check - passed.
Composed smoke: Published evaluator, gpt-5-mini executor and claude-haiku-4.5 judge, four package targets, eight selected positive/boundary scenarios, and 24 SDK arms. Process exit code 0; all required arms and proposal artifacts were produced without scenario execution errors. Native adaptation wrote exactly four expected workflow results with no missing/unexpected evidence, and report rendering was verified. After correcting confirmed instrument defects, replay of the retained proposals passed 19/24 arms, including all eight production-package arms; five genuine baseline/isolated mistakes remain rejected. This is not a statistically powered improvement claim.
Not run / known limitations: The complete 43-stimulus multi-model evaluation and live Actions/publication validation have not run. python eng\eval-quality\selftest_eval_quality.py fails two contained-symlink/cycle cases on this Windows host; both scripts are unchanged from HEAD, and the production quality gate passes.
Checklist
I searched existing issues and pull requests to avoid duplicates.
I kept this pull request focused and avoided unrelated refactors.
I added or updated tests, evals, or documentation when changing skill or agent behavior.
I updated CODEOWNERS when adding or moving owned content.
I updated all marketplace manifests when plugin metadata changed. (N/A: plugin metadata is unchanged.)
I updated eng/known-domains.txt for any new external domains referenced by skill content. (N/A: no new external domains.)
👋 @Evangelink — this PR has 2 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
The reason will be displayed to describe this comment to others. Learn more.
Thanks for the thorough work here. The workflow target model, package staging, adapter integration, and the scenario coverage are all thoughtfully put together.
I found three paths that look worth tightening before merge:
The fixture digest currently depends on operating-system path ordering, and the hosted Linux validate job is red on the partialclass fixture.
A shared grader-only change is interpreted as package graders, which stops discovery instead of selecting all workflow packages.
Shared root contracts under tests/agentic-workflows/, including RESULT_SCHEMA.md, do not currently request evaluation even though they can affect all 43 scenarios.
I left focused inline comments with the evidence and suggested direction. Once these are addressed and the Linux validation is green, this should be in good shape.
The production dashboard cannot reach this branch: generate-benchmark-data.ps1:675 serializes workflow verdicts as skillKind: "invocable". Its activation readers at lines 299, 310, and 792 also select skill telemetry instead of agentActivation*, losing the workflow persona's activation evidence. Update the generator to preserve workflow identity and use agent activation for workflows before rendering these results.
Hunk start line conflicts with renamed method location
The hunk starts at line 3, but Buffer.cs:3 is the opening brace; the renamed method and closing brace are at lines 4–5. This conflicts with both the source snapshot and changed_files.lines: [4], so an agent following the patch's location can produce a correct diagnosis with a suggestion the grader rejects. Start both sides of the hunk at line 4.
This hunk declares seven new lines but contains only six, matching the six-line Generate.targets snapshot. That makes the supplied unified diff malformed, so the scenario presents invalid patch evidence rather than a valid generated-output change. Change the new hunk count to six.
This hunk starts at after-state line 3, but Buffer.cs:3 is {; the renamed method is on line 4. The diff therefore disagrees with both the supplied source and changed_files.lines, so a model following the diff's location can be rejected despite identifying the intended API break. Change the hunk start to line 4 and refresh the fixture's authenticated input digest.
Derive reference eligibility independently from candidate eligibility
Reference eligibility is copied from whole-candidate eligibility. Consequently, partialclass records its closed/completed issue as ineligible merely because the class has structural deferrals. The real resolver calculates issue eligibility independently (IssueResolver.cs:99–109), so this fixture introduces an unrelated tracking-state inconsistency into the class-boundary case. Derive reference eligibility from tracking evidence, retain structural deferrals in decision, and regenerate the manifest/context and digest pins.
Correct malformed diff line count and refresh digest
The hunk declares seven after-state lines, but contains only six, matching the six-line Generate.targets snapshot. This is a malformed diff in a scenario intended to assess generated-output collisions, introducing an unrelated evidence defect that can derail an otherwise correct review. Use +1,6 and refresh the owning eval's authenticated input digest.
This copies candidate eligibility into the tracking-reference eligibility. For partialclass, it marks the closed/completed issue ineligible solely because the class has deferrals. The real resolver keeps these separate (IssueResolver.cs:99-109,64-69), so this fixture adds an unrelated reason to defer and can reward a planner that never recognizes the class boundary. Derive reference eligibility from tracking state independently, then regenerate the manifest and grader digests.
Mark completed issue reference eligibility true independently of candidate deci…
This issue is closed as completed, so its reference eligibility should be true; only the candidate decision should be false because of the partial/nested-class deferrals. The trusted resolver makes that distinction (agentic-workflows/unskip-closed-tests/workflows/unskip-closed-tests-tool/IssueResolver.cs:64-69,99-109). The false reference flag supplies an unintended tracking-based reason to defer, weakening the intended class-boundary case. Regenerate this field with independent reference eligibility and refresh the manifest/context and input digests.
The reason will be displayed to describe this comment to others. Learn more.
This rejects a forbidden phrase even when the answer explicitly denies it. For example, adding This evidence does not establish that all tests passed. to an otherwise valid reason makes the grader fail. This turns all 12 passing test-analysis golden proposals into failures without changing their decisions or findings.
That can also produce a false native_completion_regression when the baseline passes but the isolated answer includes this correct caution, overriding an otherwise positive preference verdict. Could we check unsupported claims semantically rather than banning these phrases anywhere in the text, and add regression coverage for negated and quoted mentions?
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Redistributable gh-aw packages were outside
/evaluatediscovery and had no scenario specs. Runtime trace graders did not provide fixture-based A/B coverage or dashboard results. This adds a first-class offline workflow-package evaluation lane to the existing pipeline.aw.ymltargets, expand local prompt imports, stage resources at their installed paths, and require explicit fixture context. Compare baseline, isolated workflow, and registered package-agent surfaces, retaining proposedresult.jsonoutput for judging and investigation.skillKind: workflowandevaluationLane: workflow-prompt-sdkin reports and dashboard data.Scope: These are offline decision/proposal evaluations using real package prompts and installed resources. They do not execute Actions bootstrap jobs, call live collectors, or publish safe outputs. Package compilation, trusted-helper tests, and runtime trace grading remain separate evidence.
Related issue
N/A
Validation
dotnet test --project eng\skill-validator\tests\SkillValidator.Tests.csproj --no-restore --results-directory artifacts\TestResults\workflow-evals --report-trx- passed, 897 tests.node --test eng\vally-adapter\*.test.mjs eng\dashboard\*.test.js- passed, 227 tests.python tests\agentic-workflows\test_graders.py- passed, 25 tests, including protocol and oracle-leakage regression coverage.python eng\evaluation\test_workflow_targets.py- passed, 6 tests covering PR/manual/scheduled discovery, identity validation, and private-worktree cleanup.python eng\evaluation\test_token_failover.py- passed, 50 existing evaluation-workflow regression tests.python eng\eval-quality\check_eval_quality.py- passed, rerun before commit/push; golden-reference debt remains warning-only.python eng\agentic-workflows\validate_agentic_workflows.py- passed with a session-isolated, checksum-verified gh-aw v0.89.15 compiler; active locks and installed packages validated.python eng\agentic-workflows\test_build_failure_analysis_operational_value.py- passed, 27 fixtures executing the real runtime grader.python eng\agentic-workflows\test_unskip_closed_tests_package.py- passed, 2 installed-layout tests.python agentic-workflows\unskip-closed-tests\tests\run_tests.py- passed, 36 trusted-helper tests.dotnet publish eng\skill-validator\src\SkillValidator.csproj --no-restore -r win-x64 -p:PublishAot=false -o artifacts\workflow-eval-tool --nologo -v:q- passed; smoke used the published layout with Copilot SDK 1.0.11.-shellcheck= -pyflakes=on the four changed workflow YAML files - passed.git diff --cached --check- passed.Composed smoke: Published evaluator,
gpt-5-miniexecutor andclaude-haiku-4.5judge, four package targets, eight selected positive/boundary scenarios, and 24 SDK arms. Process exit code 0; all required arms and proposal artifacts were produced without scenario execution errors. Native adaptation wrote exactly four expected workflow results with no missing/unexpected evidence, and report rendering was verified. After correcting confirmed instrument defects, replay of the retained proposals passed 19/24 arms, including all eight production-package arms; five genuine baseline/isolated mistakes remain rejected. This is not a statistically powered improvement claim.Not run / known limitations: The complete 43-stimulus multi-model evaluation and live Actions/publication validation have not run.
python eng\eval-quality\selftest_eval_quality.pyfails two contained-symlink/cycle cases on this Windows host; both scripts are unchanged from HEAD, and the production quality gate passes.Checklist
eng/known-domains.txtfor any new external domains referenced by skill content. (N/A: no new external domains.)