Repository navigation
ci: every job is hard-cut at 45 minutes — a literal timeout-minutes, gated fleet-wide - #3067
Conversation
…gated fleet-wide Maintainer, 2026-09-02: "hard cut ci runs after 45min — we pay all this." GitHub's default job timeout is 360 minutes. The reusable module-pack lane's `pack` job had no cap; `dotnet test MeshWeaver.Mcp.Test` wedged before its first test on every MeshWeaver.Plugins run; 19 runs sat in_progress at once (11 on main), each holding a billed runner for six hours, and because the repo's required checks `needs:` that job, no Plugins PR could merge and main never reached publish-bake — every satellite starved of a sealed publication. One missing line. - .github/scripts/check-workflow-timeouts.py: every non-`uses:` job in every workflow must carry a LITERAL `timeout-minutes` with 1 <= value <= 45; an expression is refused (a cap that is only evaluable at run time cannot be proven by reading the file); a tree with no workflows, a file with no jobs, or invalid YAML all fail — never a vacuous pass. --self-test proves every case fires on its defect and stays silent on its fix (12 cases). - dotnet-test.yml runs the self-test, then the gate, beside check-workflow-shell.py. - node-repo-validate.yml gains `platform-ref` and fetches the platform's guard at that ref to run it against the CALLER's tree — the compile-check.py centralization, no local copies. - node-repo-publish-bake.yml: the job cap is a literal 45 (the `timeout-minutes` input, default 40, is removed — no caller passed it; measured bake ~20 min). - 24 core jobs gain `timeout-minutes: 45`; alc-unload-probe (120), base-image-acr (60) and flake-repro (60) are lowered to 45 (measured: 24, 7 and 21 min). - AGENTS.md + .claude/skills/ci: the rule, the cost, and the headroom measured today (Plugins portal-host shards 25–30 min, release-images 36–41 min, Education shards ~27 min). Internal CI change: no What's New entry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
There was a problem hiding this comment.
🟢 Approval recommended
The changes consistently apply a clear CI policy and add a self-tested guard that should prevent regressions without introducing skip-trapdoors.
Pull request overview
Enforces a fleet-wide CI policy that every GitHub Actions job must have a literal timeout-minutes with a hard cap of 45 minutes, and adds a guard script + CI gates to prevent regressions.
Changes:
- Added
.github/scripts/check-workflow-timeouts.py(with self-test) and wired it into core CI (dotnet-test.yml) and the reusable satellite validation workflow (node-repo-validate.yml). - Added or lowered
timeout-minutes: 45across workflows/jobs; removed the configurable timeout input fromnode-repo-publish-bake.ymlto keep the cap provable by static inspection. - Documented the 45-minute cap rule in
AGENTS.mdand the/ciskill.
File summaries
| File | Description |
|---|---|
| AGENTS.md | Documents the new hard 45-minute job cap doctrine and guard enforcement. |
| .github/workflows/unlist-orphaned-packages.yml | Adds timeout-minutes: 45 to prevent uncapped hangs. |
| .github/workflows/release-packages.yml | Adds timeout-minutes: 45 to the publish job. |
| .github/workflows/release-images.yml | Adds timeout-minutes: 45 to the images job. |
| .github/workflows/publish-github.yml | Adds timeout-minutes: 45 to the publish job. |
| .github/workflows/publish-compiler-tool.yml | Adds timeout-minutes: 45 to preflight + publish jobs. |
| .github/workflows/plugin-build.yml | Adds timeout-minutes: 45 to preflight. |
| .github/workflows/notify-dependents.yml | Adds timeout-minutes: 45 to the notify job. |
| .github/workflows/node-repo-validate.yml | Fetches the platform guard at platform-ref and runs self-test + repo gate against caller workflows. |
| .github/workflows/node-repo-publish-bake.yml | Removes timeout input and sets a literal timeout-minutes: 45 on the job. |
| .github/workflows/node-repo-module-pack.yml | Adds timeout-minutes: 45 to key jobs to comply with the cap. |
| .github/workflows/hosting-operator.yml | Adds timeout-minutes: 45 to scripts + image jobs. |
| .github/workflows/homebrew.yml | Adds timeout-minutes: 45 to the preflight job. |
| .github/workflows/flake-repro.yml | Lowers timeout from 60 → 45 to align with the cap. |
| .github/workflows/dotnet-test.yml | Adds guard self-test + enforcement step; caps collect-results at 45. |
| .github/workflows/clients.yml | Adds timeout-minutes: 45 to the python-sdk job. |
| .github/workflows/chart-gate.yml | Adds timeout-minutes: 45 to the invariants job. |
| .github/workflows/chart-drift.yml | Adds timeout-minutes: 45 to preflight/discover/matrix jobs. |
| .github/workflows/base-image-acr.yml | Lowers timeout from 60 → 45 to align with the cap. |
| .github/workflows/auto-arm.yml | Adds timeout-minutes: 45 to the arm job. |
| .github/workflows/alc-unload-probe.yml | Lowers timeout from 120 → 45 to align with the cap. |
| .github/scripts/check-workflow-timeouts.py | New guard enforcing literal per-job timeouts (≤45), with self-test and non-vacuous failures. |
| .claude/skills/ci/SKILL.md | Adds the 45-minute cap guidance and checklist item to the CI skill. |
Review details
- Files reviewed: 23/23 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Test Results (shard 0)231 tests 231 ✅ 2m 35s ⏱️ Results for commit e626af0. |
Test Results (shard 3)602 tests 602 ✅ 59s ⏱️ Results for commit e626af0. |
Test Results (shard 1)421 tests 421 ✅ 1m 23s ⏱️ Results for commit e626af0. |
Test Results (shard 5) 3 files 3 suites 1m 12s ⏱️ Results for commit e626af0. |
Test Results (shard 2)539 tests 537 ✅ 1m 37s ⏱️ Results for commit e626af0. |
Test Results (shard 4)1 780 tests 1 780 ✅ 2m 1s ⏱️ Results for commit e626af0. |
Test Results 14 files 14 suites 9m 50s ⏱️ Results for commit e626af0. |
Maintainer, 2026-09-02: "hard cut ci runs after 45min — we pay all this."
What was wrong, measured this morning
GitHub's default job timeout is 360 minutes. The reusable module-pack lane's
packjob had notimeout-minutes;dotnet test MeshWeaver.Mcp.Testwedged before its first test on every MeshWeaver.Plugins run (A total of 1 test files matched the specified pattern.then six hours of silence — e.g. Plugins run 33570175780, job 100064441682, 23:35→05:35Z). 19 runs satin_progressat once, 11 of them onmain, each holding a billed runner until GitHub's cut. The repo's three required checksneeds:that job, so no Plugins PR could merge andmainnever reachedpublish-bake— which is why every satellite's release poll reports "upstream 'plugins' has no SEALED publication". One missing line; I cancelled the 17 hung runs by hand while writing this.Across core, 26 jobs had no cap or a cap above 45 (full inventory in the guard's first run, quoted in the commit).
The rule, and the guard that enforces it
.github/scripts/check-workflow-timeouts.py— every job in every workflow eitheruses:a reusable workflow (exempt in the caller; GitHub ignores the caller's value, and the cap lives on the jobs inside the reusable file, which this same guard gates here) or declarestimeout-minutes: <literal integer>with1 ≤ value ≤ 45. An expression is refused: a cap that can only be evaluated at run time cannot be proven by reading the file. A tree with no workflows, a file with no jobs, or invalid YAML fail — never a vacuous pass.--self-testruns 12 cases and proves each fires on its defect and stays silent on its fix.Wired in two places:
dotnet-test.yml(besidecheck-workflow-shell.py): self-test, then the gate on this tree.node-repo-validate.yml(every satellite'svalidatejob): gainsplatform-ref, fetches the platform's script at that ref and runs it against the caller's tree — thecompile-check.pycentralization; no per-repo copies. A ref it cannot be fetched at fails RED naming it.What changed in the workflows
timeout-minutes: 45addedselect/prepare/pack/verify,collect-results, everypreflight,release-images,notify, …)alc-unload-probe120 → 45 (measured 11–24 min),base-image-acr60 → 45 (5–7 min),flake-repro60 → 45 (18–21 min)node-repo-publish-bake.ymltimeout-minutesinput (default 40) is removed; the job carries a literal 45. No caller passed it (grep across Plugins/Education/Reinsurance/SocialMedia/Manufacturing).Headroom, measured today (last 12 runs per workflow)
Portal hosts (shard 0/2)release-images / imagesEvery course installs … (shard 3/4)publish-bakeMeshWeaver Build and Test,main-cdSatellites
Their own workflows carry jobs above 45 today (Plugins
ci.yml:83260, Educationci.yml:573/129360) and jobs without a cap (everypreflight). Those get fixed in their repos with the lane pin bump that brings this guard; until then the guard cannot reach them (it runs at the pinnedplatform-ref). Recorded in the fleet CI audit.Verification
python3 .github/scripts/check-workflow-timeouts.py --self-test→ 12/12 + no-workflows-dir case.python3 .github/scripts/check-workflow-timeouts.py --root .→61 job(s) checked, 4 reusable-call job(s) exempt, 0 violation(s).check-workflow-shell.py --self-test+ tree run →0 livefindings;check-shared-rules.py --self-test18/18.Internal CI change → no What's New entry.
🤖 Generated with Claude Code
https://claude.ai/code/session_01G4xPnGYEbdu8AVvR5jytyj