Skip to content

Commit dab1dee

Browse files
rodbutterselezarjohntmyersjgarciaoakram
authored
Chore/resync 20260820 (#4)
* refactor(server): normalize compute driver config acquisition (#1974) * refactor(server): remove unused compute runtime constructor parameter Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): normalize compute driver type imports Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): key driver config tables by name Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(server): normalize compute driver config acquisition Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): run gpu workloads from manifest (#1709) * test(e2e): add workload manifest build flow Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): add gpu workload validation tests Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(e2e): build gpu workloads before gpu e2e Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(providers): reserve credential placeholder revisions (#2049) * fix(providers): reserve credential placeholder revisions Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): share placeholder namespace parser Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(providers): cover non-revision env key Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> * fix(CONTRIBUTING): update label format for good first issues (#2056) * fix(helm): generate namespace-aware SANs in certgen and cert-manager templates (#2062) The certgen hook and cert-manager Certificate template hardcoded openshell.openshell.svc.cluster.local in server certificate SANs, breaking deployments in any namespace other than openshell. Use .Release.Namespace in the templates so the SANs match the actual service FQDN regardless of the target namespace. Closes #2060 Signed-off-by: Akram <akram.benaissi@gmail.com> * refactor(core): remove unused extra bind addresses (#2059) Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(mcp): fix granular policy lifecycle examples (#2066) Signed-off-by: Shiju <shiju@nvidia.com> * feat(kubernetes): add combined topology config surface (#2074) * feat(kubernetes): add combined topology config surface Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify topology defaults Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(drivers): reject whitespace in mount fields (#2086) Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(api): remove SandboxTemplate.volume_claim_templates (#2088) The field was added during the Kubernetes driver extraction refactor (#817) as a pass-through mechanism, but was never wired up to a CLI flag, Python SDK helper, or any documentation. The only reachable user path was raw gRPC construction. The Kubernetes driver now always injects the default workspace PVC, removing the branching logic that checked for a user-supplied VCT. Field number 9 is reserved in the proto to prevent reuse. Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(helm): add TLS termination for Envoy Gateway ingress (#2015) The chart's optional Gateway API ingress only rendered a plaintext HTTP listener, so the gateway could not be exposed over TLS. Add an HTTPS listener option that terminates TLS at the Envoy Gateway and forwards plaintext gRPC to the gateway pod. - gateway.yaml renders an HTTPS listener with `tls.mode: Terminate` and `certificateRefs` when `grpcRoute.gateway.listener.protocol=HTTPS`, keeping the default HTTP listener unchanged. Guards fail the render when `certificateRefs` is empty or `server.disableTls` is not true (the chart does not render a BackendTLSPolicy for re-encryption). - values.yaml adds `grpcRoute.gateway.listener.tls.certificateRefs`. - ci/values-gateway-tls.yaml exercises the HTTPS branch in lint/render. - docs/kubernetes/ingress.mdx documents HTTPS setup and clarifies that Envoy Gateway only terminates TLS (no OIDC SecurityPolicy); client identity uses OIDC bearer tokens, with the client-credentials grant for headless agents. - debug-openshell-cluster skill gains HTTPS-ingress troubleshooting rows. - Regenerated the chart README values table. Signed-off-by: Huabing (Robin) Zhao <zhaohuabing@gmail.com> * feat(agents): add manifest-driven gator agent (#1826) * chore(gator): add gator gate skill * chore(gator): add sandbox launcher scaffold * chore(gator): add codex image and docs checks * chore(gator): fold approved provider policy rules * chore(gator): add deterministic reviewer runner * chore(gator): clarify ok-to-test comments * chore(gator): structure launcher harnesses * chore(gator): require e2e for dependabot * chore(gator): add codex refresh profile * chore(gator): wip manifest agent launcher * feat(agents): supervise watch cycles in sandbox * fix(agents): preserve gateway refresh state * fix(gator): continue human response threads * fix(agents): keep watch supervisor retrying * fix(agents): use refreshed Codex credential aliases * fix(gator): avoid misleading gh auth checks * docs(agents): remove architecture build update * fix(gator): use REST-backed GitHub writes * fix(agents): bake immutable agent payloads * fix(agents): upload writable agent workspace * fix(agents): surface gator watch progress * fix(agents): prevent codex stdin hang * fix(agents): align codex subagent input * fix(agents): heartbeat during active cycles * fix(agents): clean up heartbeat sleep * fix(agents): disable gh telemetry in codex harness * fix(agents): reconcile closed gator PRs * fix(agents): query closed gator PR labels separately * fix(agents): tolerate rotated credential placeholders * fix(agents): enforce gator same-sha comment guard Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): scope gator trusted commentary Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(gator): treat reviewer failures as transient Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(agents): refine gator supervised workflow Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(agents): stream codex prompts via stdin Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): clarify trusted gator responses Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * refactor(agents): scope gator PR to scripts Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> * feat(docker,podman): add SELinux label support for bind mounts (#2092) * feat(docker,podman): add SELinux label support for bind mounts The Docker Engine structured Mount API does not support SELinux relabelling (:z / :Z). Move user-supplied bind mounts from the structured `mounts` field to the legacy string-format `binds` field, which does support these options. Add a shared `SelinuxLabel` enum (shared/private) to openshell-core so both Docker and Podman drivers accept an optional `selinux_label` field on bind mount configs. For Docker, labels are appended to the bind string; for Podman, they are pushed to the mount options vec. Signed-off-by: Florian Bergmann <fbergman@redhat.com> * fix(docker): reject missing bind source paths on legacy binds Moving user bind mounts from the structured Mount API to the legacy Binds field changed Docker's behavior for missing source directories: the legacy path silently creates them as empty root-owned dirs instead of erroring. Add an explicit Path::exists() check to preserve the fail-fast behavior operators expect. Signed-off-by: Florian Bergmann <fbergman@redhat.com> --------- Signed-off-by: Florian Bergmann <fbergman@redhat.com> * test(e2e): run rootless podman on ubuntu host (#2119) * test(e2e): run rootless podman on ubuntu host Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): probe rootless capability behavior Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): make capability probe observational Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(policy): accept numeric UIDs for sandbox process identity (#1973) * feat(policy): accept numeric UIDs in sandbox process identity validation Allow run_as_user and run_as_group to be either the literal 'sandbox' or a numeric UID/GID within [1000, 2_000_000_000]. This removes the hard dependency on a baked-in 'sandbox' user in container images, enabling compute drivers to inject resolved UIDs at sandbox creation. Phase 1 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(supervisor): accept numeric UIDs for process identity dropping Allow run_as_user and run_as_group to be numeric UIDs/GIDs, removing the hard dependency on a baked-in 'sandbox' user in container images. Changes: - validate_sandbox_user(): accepts numeric UIDs without passwd lookup (logs OCSF event); keeps passwd check for "sandbox" name; rejects non-numeric non-sandbox strings that fail passwd lookup - prepare_filesystem(): passes numeric UIDs/GIDs directly to chown() instead of requiring a passwd entry - drop_privileges(): resolves numeric UIDs/GIDs directly via UID::from_raw / Gid::from_raw; skips initgroups when target uid matches current euid; uses guard conditions before setgid/setuid calls - session_user_and_home(): falls back to ("{uid}", "/sandbox") for numeric UIDs, avoiding a passwd lookup that will fail Re-exports MIN_SANDBOX_UID and MAX_SANDBOX_UID from openshell-policy so callers have consistent range constants. Phase 2 of #1959. Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-kubernetes): resolve sandbox UID/GID from config or OpenShift SCC annotations Phase 3 of the numeric-UID plan: allow operators to specify explicit sandbox_uid/sandbox_gid in Kubernetes driver config, auto-detect from OpenShift SCC namespace annotations, and propagate resolved values to supervisor container env vars and PVC init container securityContext. Changes: - Add sandbox_uid/sandbox_gid fields to KubernetesComputeConfig - Add SANDBOX_UID/SANDBOX_GID env var constants to openshell-core - Implement resolve_sandbox_identity() to fetch namespace annotations and auto-detect OpenShift SCC UID ranges (sa.scc.uid-range) - Pass resolved UID/GID through SandboxPodParams to pod spec builder - Inject SANDBOX_UID/SANDBOX_GID env vars into supervisor container - Update PVC init container securityContext with resolved UID/GID instead of hard-coded root - Add comprehensive unit tests for resolution logic and annotation parsing (resolve_sandbox_uid, resolve_sandbox_gid, OpenShift SCC annotation parsing) Signed-off-by: Seth Jennings <sjenning@redhat.com> * feat(driver-vm): add configurable sandbox UID/GID and update docs/examples Phase 4 of the numeric-UID plan: replace hardcoded SANDBOX_UID (10001) in VM rootfs preparation with configurable sandbox_uid/sandbox_gid fields. Changes: - Add sandbox_uid/sandbox_gid to VmDriverConfig with serde derives - Pass resolved UID/GID through prepare_sandbox_rootfs_from_image_root to ensure_sandbox_guest_user which writes /etc/passwd/group/gshadow - Update BYOC Dockerfile: remove groupadd/useradd, document runtime UID injection and the ability to skip baked-in sandbox user - Update gateway-config.mdx: document sandbox_uid/sandbox_gid for both Kubernetes (with OpenShift SCC autodetection) and VM drivers - Update sandbox-compute-drivers.mdx: add Sandbox User Identity section explaining numeric UID support across all compute drivers - Update rootfs tests to use non-default UIDs, verify config passthrough Signed-off-by: Seth Jennings <sjenning@redhat.com> * code review changes * fix(supervisor): harden tests for restricted CI container environments Guard tests against CI-specific constraints: root without CAP_SETPCAP, UIDs with no /etc/passwd entry, and restricted /proc access. Signed-off-by: Seth Jennings <sjennings@nvidia.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Seth Jennings <sjennings@nvidia.com> * docs: add Hermes Agent to supported agents table (#2131) * rfc-0006: add driver config passthrough proposal (#1589) * docs(rfc): add driver config passthrough proposal Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): link driver config proposal PR Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): clarify driver config scope Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): clarify driver-local config schemas Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(rfc): clarify driver config extension path * docs(rfc): update driver config baseline * docs(drivers): document bind-mount selinux_label and whitespace rules --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * chore(deps): bump docker/login-action from 4.2.0 to 4.4.0 (#2146) * docs: fix STYLEGUIDE heading to match filename (#2134) * docs(kubernetes): bump cert-manager to v1.20.3 (#2129) * docs: fix article before OpenShell in sync-files (#2133) * docs: warn to redact credentials from log output before sharing (#2124) Add a reminder to the bug report template's Logs field and a new row in the security best-practices Common Mistakes table advising reporters to redact credentials, API keys, and tokens from stack traces before pasting. Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): deliver sandbox JWTs as secrets (#2156) Signed-off-by: Adam Miller <admiller@redhat.com> * chore: remove deprecated --keep flag from docs, scripts, and e2e tests (#2126) * docs: remove deprecated --keep flag from tutorials and examples The --keep flag is deprecated, hidden, and a no-op since sandboxes are kept by default. Remove references from tutorial docs and example READMEs that explain it as a real feature. - Remove --keep from sandbox create commands - Remove --keep explanation text - Clarify that sandboxes are kept by default Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com> * chore: remove deprecated --keep usage from scripts and e2e tests The --keep flag is a deprecated no-op since sandboxes are kept by default. Stop passing it in internal scripts, e2e test scripts, and example demo scripts. Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com> --------- Signed-off-by: Ignas Baranauskas <ibaranau@redhat.com> * chore(deps): bump astral-sh/setup-uv from 8.2.0 to 8.3.0 (#2160) Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.2.0 to 8.3.0. - [Release notes](https://github.com/astral-sh/setup-uv/releases) - [Commits](https://github.com/astral-sh/setup-uv/compare/fac544c07dec837d0ccb6301d7b5580bf5edae39...d31148d669074a8d0a63714ba94f3201e7020bc3) --- updated-dependencies: - dependency-name: astral-sh/setup-uv dependency-version: 8.3.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix(driver-podman): gate Linux-only Path import (#2188) Apply the Linux cfg to the Path import so native macOS lint runs do not report it as unused when the only call site is compiled out. This fixes `mise run rust:lint` on macOS. Signed-off-by: Kris Hicks <khicks@nvidia.com> * docs: fix Docker version format from 28.04 to 28.0 (#2136) * docs: update man page date to 2026 (#2135) * fix(sandbox): acknowledge initial policy revision; expose SDK labels/selectors (#2170) * fix(sandbox): acknowledge initial policy revision The supervisor loaded and enforced a sandbox-scoped policy but never told the gateway which revision it loaded. The policy poll loop seeded itself with the initial revision's hash on its first poll, so `policy_changed` was never true for that revision and `ReportPolicyStatus(LOADED)` — which only ran in the hot-reload branch — was never called. The revision stayed `Pending` and `current_policy_version` stayed 0 even though the sandbox was `Ready` and the policy was effective. This was most visible with sparse policies that get baseline-enriched into a new revision during startup. After the OPA engine is constructed, report the exact sandbox revision the supervisor loaded as LOADED, and seed the poll loop from that revision so it is not re-reported. Report FAILED with the original construction error if engine construction or conversion fails. Only sandbox-sourced revisions (version > 0) whose canonical content matches the loaded policy are acknowledged; global and local-file policies are untouched. Delivery uses the shared bounded retry, is non-fatal on transient failure, and a pending initial acknowledgement is delivered before any newer revision so policy history is never reordered. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * feat(python): expose sandbox labels and selectors The gateway protobuf and CLI already support request-level sandbox labels (`CreateSandboxRequest.name`/`labels`) and selector-based listing (`ListSandboxesRequest.label_selector`), but the public Python SDK dropped them, so Python-created sandboxes could not be found via `openshell sandbox list --selector ...`. Add optional, source-compatible `name`/`labels` to `SandboxClient.create`, `create_session`, and the high-level `Sandbox`, and `label_selector` to `list`/`list_ids`. `SandboxRef` now carries the gateway labels as an immutable mapping (default empty, so `SandboxRef(id, name, status)` still works). Caller-provided label mappings are copied. Attaching the high-level `Sandbox` to an existing sandbox rejects `name`/`labels` since creation metadata cannot change on attach. Template labels remain a separate concept. No protobuf changes are required. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(python): keep SandboxRef hashable and copy high-level labels Excluding the new immutable `labels` field from SandboxRef equality/hash (`compare=False`) preserves the original (id, name, status) identity and keeps the frozen dataclass hashable — a MappingProxyType field would otherwise make `hash(SandboxRef(...))` raise. Also defensively copy caller-provided labels in the high-level `Sandbox` so later caller mutation cannot change what is sent. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): bound initial-policy-ack retries The poll loop retried a pending initial acknowledgement before processing any newer revision, but retried unconditionally forever. A permanently undeliverable ack (e.g. the revision was superseded before it could be reported) would then stall all later policy hot-reloads and provider-env refreshes. Cap the retries; after the bound, give up and resume normal polling so the loop cannot livelock on a stuck acknowledgement. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * test(sandbox): add sparse-policy revision-2 acknowledgement e2e Regression for #2159: create a sandbox with the network-only policy-advisor fixture, which the supervisor enriches with baseline filesystem paths during startup (creating revision 2, superseding revision 1). Assert the effective policy reaches revision 2 and no revision remains Pending once the supervisor acknowledges the load. Adds SandboxGuard::create_keep_with_args to create a kept sandbox with an initial --policy. Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(ci): correct sandbox checks Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * test(cli): serialize mTLS environment access Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): address policy review feedback Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): preserve exact policy acknowledgements Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> * fix(sandbox): preserve local policy overrides Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Kyle Zheng <kyzheng@nvidia.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore(deps): bump astral-sh/setup-uv from 8.3.0 to 8.3.1 (#2191) Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.3.0 to 8.3.1. - [Release notes](https://github.com/astral-sh/setup-uv/releases) - [Commits](https://github.com/astral-sh/setup-uv/compare/d31148d669074a8d0a63714ba94f3201e7020bc3...f98e06938123ccabd21905ea5d0069192241f9f1) --- updated-dependencies: - dependency-name: astral-sh/setup-uv dependency-version: 8.3.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * feat(cli): add --secret-material-env to provider refresh configure (#2178) * feat(cli): add --secret-material-env to provider refresh configure * feat(cli): reject duplicate secret material keys Signed-off-by: Hung Le <hple@nvidia.com> --------- Signed-off-by: Hung Le <hple@nvidia.com> * docs(telemetry): Added first telemetry report for the community (#2190) * docs(telemetry): add community telemetry reports page Add telemetry/README.md to publish aggregate usage trends every two weeks, and link to it from the Telemetry section of the main README. First report covers the July 8, 2026 window. Signed-off-by: Kirit Thadaka <kthadaka@nvidia.com> * docs(telemetry): note telemetry start date (June 1, 2026) Clarify that all-time figures are cumulative from #1433, so readers know when the all-time counts begin. Signed-off-by: Kirit Thadaka <kthadaka@nvidia.com> --------- Signed-off-by: Kirit Thadaka <kthadaka@nvidia.com> * change packit target to new correct copr project (#2185) Signed-off-by: Adam Miller <admiller@redhat.com> * test(supervisor-network): add proxy hostname parser regression tests (#2197) Add regression coverage for parser differentials in the egress proxy's CONNECT hostname handling and OPA wildcard policy matching. Signed-off-by: Shane Utt <shaneutt@linux.com> * fix(tui): route warning logs to status bar instead of stderr (#2210) * fix(tui): route warning logs to status bar instead of stderr tracing::warn/info/debug calls in the TUI crate write to stderr via the global tracing subscriber. In ratatui's alternate-screen/raw-mode, stderr writes corrupt the terminal layout. Error-state sandboxes amplify this as background gRPC polls fail every 2s tick. Replace all 25 tracing calls with app.status_text assignments for direct-access sites, and Vec<String> accumulation for the spawned start_port_forwards task. Add ForwardWarnings event variant to decouple forward warning delivery from the sandbox name in CreateResult, preventing downstream gRPC lookup failures. Closes #2120 Signed-off-by: Ian Miller <milleryan2003@gmail.com> * feat(tui): add ForwardWarnings event variant New event type for non-fatal port-forward warnings during sandbox creation. Keeps warning delivery separate from the sandbox name in CreateResult to avoid corrupting downstream gRPC lookups. Signed-off-by: Ian Miller <milleryan2003@gmail.com> * chore(tui): remove tracing dependency Compile-time guard against reintroducing stderr-writing tracing calls in the TUI crate. 14 other crates retain the dependency. Signed-off-by: Ian Miller <milleryan2003@gmail.com> --------- Signed-off-by: Ian Miller <milleryan2003@gmail.com> * fix(mcp): include tool names in policy logs (#2189) Signed-off-by: Kirit93 <kthadaka@nvidia.com> * chore(deps): bump astral-sh/setup-uv from 8.3.1 to 8.3.2 (#2206) Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.3.1 to 8.3.2. - [Release notes](https://github.com/astral-sh/setup-uv/releases) - [Commits](https://github.com/astral-sh/setup-uv/compare/f98e06938123ccabd21905ea5d0069192241f9f1...11f9893b081a58869d3b5fccaea48c9e9e46f990) --- updated-dependencies: - dependency-name: astral-sh/setup-uv dependency-version: 8.3.2 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * docs(openshift): simplify install steps and add Helm README entries for OpenShift overrides (#2125) Signed-off-by: ChristianZaccaria <christian.zaccaria.cz@gmail.com> * docs(issues): require release and duplicate checks (#2214) Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(core): pin supervisor image tag to gateway version for all drivers (#2070) * fix(core): pin supervisor image tag to gateway version for all drivers The Podman and Kubernetes drivers defaulted the supervisor image to `:latest` via DEFAULT_SUPERVISOR_IMAGE, while the Docker driver already resolved a version-pinned tag. Extract the tag resolution logic into openshell-core so all three drivers use the same OPENSHELL_IMAGE_TAG > IMAGE_TAG > CARGO_PKG_VERSION priority chain. Closes #2068 Signed-off-by: Florent Benoit <fbenoit@redhat.com> * refactor(core): simplify supervisor image tag resolver to slice-based API Remove the Docker driver's wrapper functions and call openshell_core::config::default_supervisor_image() directly. Simplify resolve_supervisor_image_tag to accept &[&str] instead of three separate parameters. Signed-off-by: Florent Benoit <fbenoit@nvidia.com> Signed-off-by: Florent Benoit <fbenoit@redhat.com> --------- Signed-off-by: Florent Benoit <fbenoit@redhat.com> Signed-off-by: Florent Benoit <fbenoit@nvidia.com> * fix(helm): propagate supervisor image overrides (#2216) Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * feat(kubernetes): add sidecar supervisor topology (#2076) * feat(kubernetes): add sidecar supervisor topology Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): avoid similar proxy id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify sidecar topology limits Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): keep sidecar process leaf capless Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): refresh sidecar provider env snapshots Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * test(supervisor): align hot-swap identity regression Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): stage sidecar mtls files before proxy chown Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): simplify sidecar supervisor topology Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(helm): reuse sidecar skaffold values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar iptables helper names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(e2e): harden kube gateway wrapper setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid nft batch rollback on OCP Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules. Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): preserve process identity in sidecar topology Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table. Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): replace sidecar snapshots with control socket Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files. Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * feat(kubernetes): support relaxed sidecar network identity Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy sidecar clippy lint Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): standardize topology naming Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy linux clippy timeout import Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): support kata sidecar on ipv4 pods Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): satisfy linux clippy for sidecar fallback Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(kubernetes): remove stale supervisor topology references Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): enable sidecar binary policy inspection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden sidecar control boundary Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): couple sidecar supervisor lifecycles Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Seth Jennings <sjenning@redhat.com> * feat(kubernetes): support PVC subPath driver config (#2034) * feat(kubernetes): support PVC subPath driver config Signed-off-by: mjamiv <michael.commack@gmail.com> * test(kubernetes): cover writable PVC driver config Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(kubernetes): address PVC subPath review feedback Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(kubernetes): address PVC config review follow-up Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(kubernetes): address PVC review follow-ups --------- Signed-off-by: mjamiv <michael.commack@gmail.com> * fix(network): fail closed when credential placeholders cannot be rewritten (#2162) * fix(network): fail closed when credential placeholders cannot be rewritten When the credential rewriter degrades internally, the proxy forwarded the literal `openshell:resolve:env:<NAME>` placeholder (or its provider alias marker) to the upstream instead of the resolved secret, leaking the reserved token on the wire and causing upstream auth failures (#2161). Two fail-open paths are closed: - secrets: `rewrite_http_header_block` returned the header block verbatim when no `SecretResolver` was available, so the fail-closed marker scan (which ran only on the resolved path) never saw the placeholder. It now scans the header region for reserved markers even with no resolver and returns `UnresolvedPlaceholderError` when one is present. Marker-free traffic still passes through unchanged. - proxy: when TLS was detected on a CONNECT but `tls_state` was `None` (ephemeral CA generation or CA file write failed at startup), the handler fell back to a raw `copy_bidirectional` tunnel, bypassing credential rewrite. Inside the proxy handler `tls_state` is `None` only on CA-init failure (`mode != Proxy` never starts the handler, and `tls: skip` is handled earlier), so it now refuses the connection with a 503 and a High-severity denial event instead of tunneling. The two startup CA-failure logs are raised from Medium to High. Tests: resolver=None with a placeholder in the request line, a header value, and the provider-alias form now fail closed; marker-free passthrough is unchanged; the relay integration test asserts the request is rejected before any byte reaches upstream; and the 503 fail-closed response contract is locked. Signed-off-by: Tony Luo <xialuo@nvidia.com> * fix(network): refuse CONNECT before 200 when TLS termination is unavailable The fail-closed refusal for a terminating CONNECT with no TLS termination state (ephemeral CA init failed) was written after the 200 Connection Established response. Because a CONNECT client only sends its TLS ClientHello after reading the 200, the peek-based TLS detection is inherently post-200, so the 503 landed inside the established tunnel and surfaced to the client as a TLS protocol error rather than a readable status. An 'allowed CONNECT' event was also logged first. Move the decision to a pre-200 gate: query_tls_mode resolves purely from the policy decision + host/port (no peeked bytes), so the route's TLS treatment is known before the tunnel is acknowledged. When TLS state is absent and the route is not tls: skip, write the 503 as the first bytes on the socket, emit the High-severity Denied event, and close. tls: skip routes tunnel raw exactly as before, and no allowed-CONNECT event is emitted on the refusal path. The now-unreachable post-200 branch is kept as defense in depth but no longer writes an in-tunnel 503; it fails closed by dropping the connection instead. Add connection-level regression tests over a real loopback socket: the gate refuses with HTTP/1.1 503 as the first bytes for a terminating route, and writes nothing when TLS termination is present or the route is tls: skip. Signed-off-by: Tony Luo <xialuo@nvidia.com> * fix(network): order the CONNECT TLS-unavailable refusal after SSRF Addresses the gator re-check on #2162. Ordering: the pre-200 fail-closed refusal ran before SSRF/allowed_ips validation, so during CA-init failure an internal-address CONNECT got a 503 tls_termination_unavailable instead of the normal 403 ssrf_denied, weakening operator visibility in degraded state. The SSRF branches now return validated addresses; the refusal runs after that validation (an internal address has already been denied with 403) but still before the upstream connect and before 200 Connection Established. effective_tls_skip is still resolved up front since the refusal consumes it. Tests: add connection-level regressions through the real handle_tcp_connection, driving a CONNECT from a child /bin/bash copy so the /proc process-identity binding resolves it against a permissive policy (the hot-swap test's identity pattern). They assert the first bytes are HTTP/1.1 503 for a terminating route with no TLS state, a 403 (not 503) for an internal address, and no refusal for a tls: skip route. These are gated to Linux at runtime (evaluate_opa_tcp needs /proc); a companion test verifies the OPA policy shape (glob allow, tls mode) on every platform so the precondition is locked where /proc is unavailable. rest.rs: tighten the fail-closed relay test to assert the forwarded buffer is_empty() rather than merely lacking the placeholder/secret. Signed-off-by: Tony Luo <xialuo@nvidia.com> * test(network): keep the CONNECT handler test client fork-free The handler regression tests forked cat to read the proxy reply, so the client socket fd was inherited by a second process with a different binary. The identity resolver correctly denies that as ambiguous shared-socket ownership (the same invariant resolve_process_identity_denies_fork_exec_shared_socket_ambiguity pins), so the tests exercised the deny path instead of the allow path — and on busy CI runners the deny-path /proc fallback scan exceeded the test budget and looked like a hang. The client script now uses only bash builtins (exec, printf, read -d '') so exactly one process owns the socket, and the child is left to exit on EOF instead of being killed mid-read. Signed-off-by: Tony Luo <xialuo@nvidia.com> * test(network): drive the CONNECT handler tests with an in-process client The child-process client (even fork-free) made the handler tests environment-sensitive: on CI runners with a busy or restricted /proc, resolving the child's socket ownership degraded into the whole-/proc fallback scan and a deny, which surfaced as a hang. The client is now an in-process TcpStream and the test policy allows current_exe(), so identity resolution binds the socket to the test process itself in the descendant scan — the same in-process pattern the passing resolve_process_identity tests rely on. The tls: skip test additionally asserts that the handler emitted no DenialEvent at any stage, so it can no longer pass vacuously on a policy or identity deny. Refusal budgets widened to 30s as a belt for slow runners; the refusals themselves return in milliseconds. Signed-off-by: Tony Luo <xialuo@nvidia.com> * test(core): pin the percent-encoded marker no-resolver fail-closed path The no-resolver scan already catches the percent-encoded canonical marker through its decoded pass; this regression pins it: a request line carrying openshell%3Aresolve%3Aenv%3AKEY with no resolver must fail closed with UnresolvedPlaceholderError { location: header }. Signed-off-by: Tony Luo <xialuo@nvidia.com> --------- Signed-off-by: Tony Luo <xialuo@nvidia.com> * fix(server): allow newlines in exec command arguments (#1965) * chore(deps): bump actions/stale from 10.3.0 to 10.4.0 (#2234) * fix(tui): redraw after sandbox shell exits (#2230) * fix(tui): redraw after sandbox shell exits Closes #2229 Render the restored terminal before synchronous gateway refreshes and propagate terminal lifecycle failures from both SSH handoff paths. Signed-off-by: John T. Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): restore input after sandbox shell Discard stale events accumulated around the suspended TUI and let the normal periodic tick refresh state after resuming. Signed-off-by: John T. Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John T. Myers <9696606+johntmyers@users.noreply.github.com> * fix(agents): add confirmation gate to triage-issue batch mode (#2239) Batch triage previously processed all state:triage-needed issues immediately without any confirmation. This led to accidental mass-commenting (29 triage comments on a public repo) when the skill was invoked by mistake. Add a mandatory preview-and-confirm step: the agent must show the issue count and titles, then ask for explicit user confirmation before posting any comments. Single-issue mode is unchanged. Assisted-By: 🤖 Claude Code Signed-off-by: Roland Huß <rhuss@redhat.com> * fix(gator): retry review after draft blocker clears (#2200) Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> * fix(certgen): stage temp dir inside output dir to fix cross-device rename (#2241) generate-certs stages temp files beside --output-dir using dir.with_file_name(). When --output-dir is a container or WSL volume mount, the staging dir lands on a different filesystem and std::fs::rename fails with EXDEV. Move staging inside the output dir so the rename always stays on the same filesystem, preserving atomic replacement. Fixes #2173 Signed-off-by: Grace Smith <grasmith@redhat.com> * refactor(jsonrpc): carry typed inspection errors (#2244) Signed-off-by: Shiju <shiju@nvidia.com> * docs(agents): add gator launch skill (#2203) * docs(agents): add gator launch skill Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): narrow gator launch trigger Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): trim gator launch skill intro Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): harden gator launch examples Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): avoid assuming gator gateway Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(agents): trim gator launch rules Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> * chore(python): lower minimum supported Python to 3.11 (#2247) Debian 12 stable and other LTS environments ship Python 3.11 as the system interpreter. The SDK is pure Python with a bundled native binary, uses no 3.12-only syntax or stdlib APIs, and all dependencies support 3.11. Lower the floor so `pip install openshell` / `uv add openshell` works on 3.11 (security-supported until Oct 2027). Verified with py_compile on CPython 3.11.15 plus ruff and ty. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(policy): keep approved chunk when a mechanistic denial resubmits its endpoint (#2242) * fix(policy): keep approved chunk when a mechanistic denial resubmits its endpoint A mechanistic denial flush for an endpoint already covered by an auto-approved mechanistic chunk flipped that chunk approved -> rejected with no human action. The dedup upsert in put_draft_chunk returns the existing row's id, which aliases onto the approved chunk; the self-reject scan then matched the row against itself and rejected it, while the merged rule stayed enforced — the governance ledger disagreed with the live policy. Guard self_reject_mechanistic_if_already_covered to act only on a still pending effective chunk, and exclude the incoming id from the covering scan. Add a regression test and document the dedup/self-reject invariant. Fixes #2165 Refs NVIDIA/NemoClaw#6329 Signed-off-by: Tinson Lai <tinsonl@nvidia.com> * fix(policy): reject mechanistic chunk via atomic pending compare-and-set The self-reject-when-covered path read a chunk, confirmed it was pending, then issued an unconditional status update to rejected. An approval that committed between the read and the write flipped an already-approved chunk to rejected while its rule stayed merged, recreating the ledger/enforcement mismatch through a concurrent path. Add conditionally_reject_draft_chunk to the policy store: the pending->rejected transition carries a status = 'pending' predicate on the final write and reports whether a row changed. Zero changed rows is a benign no-op, meaning another operation already decided the chunk. Implemented for both SQLite and PostgreSQL. The pending pre-read and the self-exclusion guard stay as defense-in-depth. Adds persistence-level regression tests proving an approved row cannot be conditionally rejected and that an approval racing the reject wins. Signed-off-by: Tinson Lai <tinsonl@nvidia.com> --------- Signed-off-by: Tinson Lai <tinsonl@nvidia.com> * fix(tasks): format all Rust workspaces (#2268) Run cargo fmt for both the root workspace and the standalone e2e Rust workspace from the rust format and format-check tasks. Apply rustfmt to the e2e sources so the expanded formatting check passes. Signed-off-by: Kris Hicks <khicks@nvidia.com> * docs: fix stray bracket in provider create command example (#2275) Key Commands table showed --type [type]] --from-existing with an extra closing bracket. * rfc-0010: gateway interceptors (#1927) * docs(rfc): add gateway interceptors RFC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): clarify gateway interceptors proposal Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): clarify interceptor source of truth Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): refine gateway interceptor proposal Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * docs(rfc): document interceptor order example Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): renumber gateway interceptors RFC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): clarify gateway interceptor service Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): clarify gateway interceptor limits Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): align gateway interceptor config Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): refine gateway interceptor contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): clarify gateway interceptor payload contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): add gateway interceptor describe request Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove interceptor modifies flag Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): shape interceptor evaluation by phase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): simplify interceptor mutation phase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): update interceptor post-commit failure mode Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): accept gateway interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(release-dev): update azure/setup-helm to v5.0.1 (#2274) Dependabot updated the Helm setup action in release-canary but missed the same reference in the release composite action. Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(snap): vendor ssh in openshell snap and remove ssh-keys interface (#2280) Previously, the openshell snap used the ssh-keys interface to get access to the host's ssh binary, which is used for sandbox connect/exec/forward. However, ssh-keys is a privileged interface which also grants access to the public and private ssh keys on the host. As such, it required manual connection in order to be used. This weakened the security sandbox of the snap, and hurt the UX of installing it. This commit changes this by removing the `ssh-keys` interface and instead vendoring the `ssh` binary within the snap. This is safe because OpenShell always invokes the `ssh` binary with `StrictHostKeyChecking=no`, `UserKnownHostsFile=/dev/null`, and `GlobalKnownHostsFile=/dev/null`, and never uses any host credentials or ssh configuration. Openshell only ever access to `~/.ssh/config` to write OpenShell-managed aliases, and this can safely live within the snap sandbox, rather than leaking into the host environment. Signed-off-by: Oliver Calder <oliver.calder@canonical.com> * feat(interceptors): initial gateway interceptor implementation and reference example (#2005) * feat(gateway): add descriptor-driven interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): add service-reflected interceptors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * fix(gateway): harden interceptor evaluation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(interceptors): label metrics and harden governance smoke Signed-off-by: Drew Newberry <anewberry@nvidia.com> * remove on_error: ignore * feat(gateway-interceptors): emit log annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(examples): govern provider profiles in interceptor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): preserve update config annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): support interceptor profile catalogs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * wip * feat(governance-interceptor): sign provider profiles Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(providers): use configured profile sources for refresh updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway-interceptors): add phase-specific evaluation payloads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): compose provider profile sources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): preserve committed responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): reject ambiguous protobuf oneofs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): validate patch candidates per binding Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(gateway-interceptors): use reflected protobuf codec Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(governance-example): canonicalize signed protobuf hashes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): commit policy provenance atomically Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): close signed governance bypasses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): isolate interceptor secrets and authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway): snapshot provider profiles per request Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(gateway): resolve server clippy warnings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): satisfy provider source clippy lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): initialize policy test annotations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(gateway-interceptors): require explicit route allowlist Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(proto): clarify update annotation semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sdk): add openshell-sdk crate (#1862) * feat(sdk): add openshell-sdk crate Additive extraction of the shared async gRPC client core (transport, TLS, OIDC single-flight refresh, edge tunnel, high-level sandbox surface, raw escape hatch) as a new workspace crate. No existing consumers yet; CLI/TUI migration follows in a separate PR. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * refactor(sdk,core): reuse shared JWT exp decoder for refresh deadlines Extract the signature-unverified JWT exp decode out of openshell-core/grpc_client.rs into openshell_core::jwt::parse_exp_secs, and have the openshell-sdk refresh path reuse it to derive a proactive refresh deadline from a bearer JWT when the caller does not advertise expires_at. Addresses review feedback to reuse pre-existing logic rather than reimplement JWT expiry handling per client. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): harden refresh single-flight and redact tokens in Debug Addresses review feedback on the openshell-sdk refresh path. - Make the single-flight cleanup cancellation-safe. The in-flight slot is now cleared by the shared refresh computation itself (epoch-guarded) rather than the leader's post-await code. Previously, if the leader future was dropped (e.g. an FFI caller cancelling its promise) after a follower drove the refresh to completion, the completed future was stranded in the slot and later refresh_now() calls re-joined it, pinning the client to a stale or already-rejected token. Adds a regression test that cancels the leader and asserts the next refresh starts a fresh attempt. - Redact bearer secrets from Debug. RefreshedToken and the oidc RefreshTokenInput/RefreshTokenOutput now use manual Debug impls that omit the access/refresh token fields via finish_non_exhaustive, matching the house style (e.g. SecretResolver, SandboxJwtIssuer). Prevents a stray {:?} or a containing struct's derived Debug from writing tokens to logs. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): fail refresh when the new token can't be encoded as metadata store_bearer now returns an error instead of silently keeping the previous bearer value. The TokenSource commits the refreshed token to its state before the client writes it into the interceptor slot, so a silent drop left the interceptor on the old (expiring) token with no path back to a refresh. Surfacing the error fails the call loudly instead. Adds a unit test covering a token that can't be encoded as gRPC metadata. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): refresh OIDC tokens on raw routes and harden rotation Raw gRPC access never triggered OIDC refresh: a client that only used raw_grpc/raw_inference kept sending the initial bearer until it expired, with no proactive or reactive refresh. Add raw_grpc_fresh and raw_inference_fresh accessors that refresh before returning the client, plus force_refresh for reactive recovery after an Unauthenticated raw RPC. Guard the single-flight refresh commit against a concurrent replace(). The in-flight attempt now records the generation it started from and skips its write when an external replace() has advanced it, so timer or callback driven rotation is no longer clobbered by a slower refresh. Remove TokenSource::snapshot(): it returned an empty string under write contention and had no consumer on the CLI/TUI path. Tests read committed state directly instead. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * docs(sdk): rewrite crate README for consumers Recast the openshell-sdk README as a usable crate README rather than an RFC excerpt. Drop the Responsibilities/Non-responsibilities/Consumers scope-boundary sections and the mTLS migration rationale, folding the useful facts (explicit token, no disk/name resolution, Refresh trait, SdkError mapping) into the intro, a new Auth and refresh section, and Public surface. Remove the dead relative RFC link and status-label prose so the doc renders cleanly wherever it is published. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): address review feedback on OIDC refresh and transport Apply Drew's review notes on PR #1862: - Drop unused `rustls-pemfile` dependency and move `tokio-stream` to dev-dependencies (only used by tests). - Guard OIDC `expires_at` against u64 overflow with `saturating_add`. - Fix stale `#[non_exhaustive]` rationale in `AuthConfig` (the struct `Oidc` variant it described as future already ships). - Stop double-wrapping refresh errors: store the bare refresh-error text so the single `SdkError::auth` wrap happens once at await. - Strip stale CLI porting breadcrumbs from `build_channel` docs, keeping the branch table. - Treat proactive token refresh as best-effort: a transient failure falls through to the request instead of failing an RPC whose current token is still valid, with a regression test. - Collapse `exec`'s inline auth retry into the shared `unary` helper. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): preserve transient/terminal distinction in refresh errors The refresh single-flight collapsed both `RefreshError::Transient` and `RefreshError::Terminal` into a stringified `SdkError::Auth`, so consumers (CLI, TUI, future language bindings) had no machine-readable way to tell a retryable IdP blip from a dead session that needs re-authentication. Carry the `RefreshError` through the shared outcome (kept `Clone` for `Shared`) instead of its rendered text, and map it at the await site to a new `retryable` flag on `SdkError::Auth`. Add `SdkError::auth_retryable` and a `SdkError::retryable()` accessor; transient refresh failures report `true`, every other error `false`. Add classification tests. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> --------- Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): initialize sandbox annotations (#2296) The SDK merged after CreateSandboxRequest and ObjectMeta gained annotations on main, leaving stale struct initializers that prevented the crate and its tests from compiling. Initialize the curated request and mock metadata with empty annotation maps. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(driver-vm): run sandbox supervisor as guest pid 1 (#2299) Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(ci): introduce merge queue (#2024) * feat(ci): introduce merge queue Closes #1946 Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(ci): run GPU E2E for merge groups Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(ci): clarify merge queue GPU gate Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(gateway): add elevated gateway info (#2202) * feat(cli)!: fold gateway metadata into list BREAKING CHANGE: openshell gateway info no longer shows local gateway registration metadata. Use openshell gateway list or openshell gateway list -o json for local registration details. Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(gateway): add elevated gateway info Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(ci): prune snap assets from dev release (#2302) Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat!(openshell-cli): remove openshell policy prove command and z3 dependency (#2318) Remove Z3 from the openshell CLI. Proving is handled by the gateway, so bundling the solver in the client duplicates functionality and complicates portable CLI builds and packaging. Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(server): persist sandbox labels on create (#2306) Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * fix: remove mentions of bundled-z3 in CI and wheel builds (#2322) * fix(gateway): probe Docker socket during driver auto-detection (#2303) Previously, Docker was auto-detected when the CLI was installed or a candidate Unix socket existed. Neither check verified that the Docker API was responsive. A similar check was done when auto-detecting Podman in the past, but was replaced in 1f07bf04 with a probe of candidate Podman sockets instead. This change applies the functional API probing approach introduced for Podman in 1f07bf04 to Docker. It also makes Docker driver initialization use the same socket-selection mechanism as Docker auto-detection instead of Bollard’s local defaults. This means the previously auto-detectable Docker socket paths $HOME/.docker/run/docker.sock and $XDG_RUNTIME_DIR/docker.sock will actually be usable. When no working compute driver can be auto-detected, the gateway exits early with a message saying as much: > configuration error: no compute driver configured and auto-detection found no > suitable driver; set --drivers or OPENSHELL_DRIVERS to kubernetes, podman, > docker, or vm This makes for a better user experience when installing OpenShell without an available supported compute driver. Signed-off-by: Kris Hicks <khicks@nvidia.com> * chore(deps): bump actions/setup-node from 6.4.0 to 7.0.0 (#2289) Bumps [actions/setup-node](https://github.com/actions/setup-node) from 6.4.0 to 7.0.0. - [Release notes](https://github.com/actions/setup-node/releases) - [Commits](https://github.com/actions/setup-node/compare/48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e...820762786026740c76f36085b0efc47a31fe5020) --- updated-dependencies: - dependency-name: actions/setup-node dependency-version: 7.0.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * chore(deps): bump softprops/action-gh-release from 3.0.1 to 3.0.2 (#2288) Bumps [softprops/action-gh-release](https://github.com/softprops/action-gh-release) from 3.0.1 to 3.0.2. - [Release notes](https://github.com/softprops/action-gh-release/releases) - [Changelog](https://github.com/softprops/action-gh-release/blob/master/CHANGELOG.md) - [Commits](https://github.com/softprops/action-gh-release/compare/718ea10b132b3b2eba29c1007bb80653f286566b...3d0d9888cb7fd7b750713d6e236d1fcb99157228) --- updated-dependencies: - dependency-name: softprops/action-gh-release dependency-version: 3.0.2 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * feat(tui): navigate panels via Up/Down arrow overflow at list boundaries (#2287) * feat(tui): navigate panels via Up/Down arrow overflow at list boundaries When at the bottom of a panel's item list, pressing Down/j now moves focus to the next panel instead of being a silent no-op. Likewise, pressing Up/k at the top moves to the previous panel with the cursor on its last item. Empty panels are skipped and the ring wraps around. Closes #2273 Signed-off-by: Varsha Prasad Narsing <vnarsing@nvidia.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(tui): guard Up handlers against stale cursor in empty panels The Down handlers already check whether the list is non-empty before incrementing the cursor, but the Up handlers only checked cursor > 0. When a list becomes empty after a refresh with a nonzero cursor, Up would decrement the stale cursor instead of overflowing to the previous panel. Add the same non-empty guard to all four Up arms. Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * docs(sandboxes): add dashboard keyboard navigation to manage-sandboxes Describe Tab/Shift+Tab panel cycling, Up/Down and j/k boundary overflow, and middle-pane tab switching in the OpenShell Terminal section. Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> --------- Signed-off-by: Varsha Prasad Narsing <vnarsing@nvidia.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782) Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The gateway calls sts:AssumeRole and writes three short-lived credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the provider record; the proxy re-signs requests with SigV4. Adds the aws and aws-s3 provider profiles and a declarative multi-output refresh model (additional_outputs) so one AssumeRole co-mints all three credentials. Signed-off-by: Russell Bryant <rbryant@redhat.com> * rfc-0009: supervisor middleware (#1738) Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(e2e): run VM suite in CI (#2305) * test(e2e): run VM suite in CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): configure KVM permissions directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): flush VM overlay before restart Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: simplify VM test documentation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): include gateway resume in VM run Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm-driver): fixes BYOC sandbox creation failing with ext4-fs write access unavailable (#2150) * feat(supervisor-middleware): add network egress middleware (#2027) Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(gator): require inline review comments (#2346) Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> * fix(cli): preserve symlinks in sandbox upload (#2319) Signed-off-by: lr90 <qiuweimin@matrixorigin.cn> * fix(kubernetes): validate sandbox names against RFC 1123 requirements (#2295) Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> * docs: bump stated Rust MSRV from 1.88 to 1.90 (#2276) * docs: bump stated Rust MSRV from 1.88 to 1.90 Cargo.toml sets rust-version = "1.90" (rust-toolchain.toml pins 1.95.0), so building with the previously documented 1.88 fails Cargo's MSRV check. * docs: bump e2e/rust MSRV to 1.90 * fix: align remaining Rust version fields to 1.90 examples/governance-interceptor/Cargo.toml still had rust-version 1.88. Also bump e2e/rust's prost dependency to 0.14 to match the workspace, since it was on 0.13 in an otherwise standalone crate. * ci: pin docker actions to commit SHA (#2328) login-action and setup-buildx-action used a mutable version tag while every other action in the repo is pinned to a commit SHA. Pin both, and align login-action to the same v4 SHA already used in ci-image.yml. Use the full resolved version in the trailing comment (v3.12.0) to match the more common convention used elsewhere in .github/. * ci(e2e): reuse prebuilt CLI and gateway artifacts (#2311) * ci(e2e): reuse prebuilt CLI artifacts Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(e2e): reuse prebuilt gateway artifacts Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(e2e): reuse prebuilt VM driver artifact Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs: fix broken links and small inconsistencies (#2329) - README: fix github-sandbox tutorial link missing get-started segment - README: replace dead community-sandboxes doc link with the actual repo - README: match supported host list to support-matrix.mdx - architecture/README: list the missing google-vertex-ai-provider doc - SECU…
1 parent 446d9f6 commit dab1dee

1,101 files changed

Lines changed: 348581 additions & 45713 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.agents/skills/build-from-issue/SKILL.md‎

Lines changed: 77 additions & 45 deletions
Large diffs are not rendered by default.

‎.agents/skills/build-openshell-mxc-windows/SKILL.md‎

Lines changed: 341 additions & 0 deletions
Large diffs are not rendered by default.
Lines changed: 230 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,230 @@
1+
# Reference: Windows MSVC maintenance lane
2+
3+
Companion to [SKILL.md](SKILL.md). Use this file for quick lookup while
4+
maintaining the existing build-only Windows MSVC lane.
5+
6+
## Lane Files
7+
8+
| File | Purpose |
9+
|---|---|
10+
| `tasks/windows.toml` | Mise task definitions for `windows:*`. |
11+
| `tasks/scripts/windows-msvc.ps1` | Visual Studio environment discovery, rustup target setup, Cargo invocation, logs, artifact report. |
12+
| `.github/workflows/windows-msvc.yml` | Manual GitHub Actions x64 job and disabled ARM64 scaffold, each with an architecture-specific Rust dependency cache. |
13+
| `architecture/windows-msvc-build.md` | Human-readable design contract. |
14+
15+
## Commands
16+
17+
Use `--skip-tools` for all Windows mise tasks:
18+
19+
```powershell
20+
mise run --skip-tools windows:check:x64
21+
mise run --skip-tools windows:check:arm64
22+
mise run --skip-tools windows:build:x64
23+
mise run --skip-tools windows:build:arm64
24+
mise run --skip-tools windows:test:x64
25+
mise run --skip-tools windows:test:arm64
26+
mise run --skip-tools windows:test:unsupported:x64
27+
mise run --skip-tools windows:test:unsupported:arm64
28+
mise run --skip-tools windows:ci
29+
```
30+
31+
For host-native full validation, detect architecture first:
32+
33+
```powershell
34+
$arch = [System.Runtime.InteropServices.RuntimeInformation]::OSArchitecture
35+
if ($arch -eq [System.Runtime.InteropServices.Architecture]::Arm64) {
36+
mise run --skip-tools windows:check:arm64
37+
mise run --skip-tools windows:build:arm64
38+
mise run --skip-tools windows:test:arm64
39+
mise run --skip-tools windows:test:unsupported:arm64
40+
mise run --skip-tools windows:artifacts
41+
} else {
42+
mise run --skip-tools windows:ci
43+
}
44+
```
45+
46+
The native test tasks reject a target that does not match the host architecture.
47+
Do not report x64 compatibility-under-emulation coverage from an ARM64 run.
48+
49+
The wrapper adds missing rustup targets and clears inherited
50+
`RUSTC_WRAPPER`. It does not install Visual Studio, Rust, Docker, Kubernetes,
51+
Podman, WSL, Hyper-V, or VM tooling.
52+
53+
On Windows, `mise run pre-commit` routes `rust:check`, `rust:lint`, and
54+
`test:rust` through this wrapper for the host-native target. The shared task
55+
definitions retain their existing Unix commands. Only tests for Linux glibc
56+
installer behavior, Linux build-environment shell helpers, and Linux
57+
service/RPM packaging assets skip on Windows. The Windows Clippy command
58+
excludes unsupported runtime packages as top-level targets and allows only
59+
unused imports, dead code, and unused async functions caused by cfg-gated
60+
Windows stubs; other warnings remain errors.
61+
62+
The wrapper limits Cargo to four jobs by default and serializes wrapper-owned
63+
Cargo commands with a host-local mutex. It does not set `CL` or `_CL_` because
64+
`clang-cl` also consumes them and can parse a global `/MP4` option as an input
65+
file.
66+
67+
For ARM64, verify the Visual Studio instance contains the ARM64 MSVC tools,
68+
ARM64 Spectre-mitigated libraries, Clang tools, CMake tools, and a Windows SDK.
69+
Clang supplies host-native `libclang.dll` for `bindgen` and `clang-cl.exe` for
70+
ARM64 crypto dependencies such as `ring` and `aws-lc-sys`. Native ARM64 uses
71+
the normal bundled-Z3 CMake path. An x64-to-ARM64 check/build discovers and
72+
adds host-native Ninja to `PATH`, while the crypto crates select `clang-cl`.
73+
Bundled Z3 uses CMake's Visual Studio ARM64 generator with native MSVC `cl.exe`
74+
because `z3-sys 0.10.9` passes the MSBuild-only `-m` argument. Use a short
75+
`CARGO_TARGET_DIR` if Windows path-length limits are reached.
76+
77+
## Unsupported Driver Rules
78+
79+
Windows is a build target only. These runtimes remain unsupported:
80+
81+
- Docker
82+
- Kubernetes
83+
- Podman
84+
- VM
85+
86+
Rules:
87+
88+
- Keep config/library stubs where the gateway needs them.
89+
- Return clear unsupported errors at runtime.
90+
- Do not build standalone Windows driver binaries.
91+
- Do not add Docker Desktop, WSL, Hyper-V, Podman machine, Podman Desktop, or
92+
VM-backed execution as part of this skill.
93+
94+
Current focused unsupported-contract tests:
95+
96+
```text
97+
windows_builtin_compute_drivers_report_unsupported
98+
```
99+
100+
Run them with the architecture-specific focused task on the native host.
101+
102+
## Cargo Excludes
103+
104+
The Windows wrapper intentionally excludes unsupported runtime packages as
105+
top-level workspace targets for check/test:
106+
107+
```text
108+
--exclude openshell-driver-docker
109+
--exclude openshell-driver-kubernetes
110+
--exclude openshell-driver-kubernetes-secrets
111+
--exclude openshell-driver-podman
112+
--exclude openshell-driver-vault
113+
--exclude openshell-driver-vm
114+
--exclude openshell-sandbox
115+
--exclude openshell-supervisor-network
116+
--exclude openshell-supervisor-process
117+
--exclude openshell-vfio
118+
```
119+
120+
The gateway keeps platform configuration and unsupported-operation contracts
121+
without depending on the Docker, Kubernetes, Podman, sandbox supervisor,
122+
process supervisor, VM, or VFIO runtime crates. The Kubernetes Secrets and
123+
Vault libraries still compile as gateway dependencies; only their standalone
124+
Unix-socket binaries and package-level tests are excluded as top-level targets.
125+
126+
## Common Errors
127+
128+
### Unix imports leak into Windows builds
129+
130+
Symptoms:
131+
132+
```text
133+
unresolved import std::os::unix
134+
unresolved import tokio::net::UnixListener
135+
unresolved import nix::...
136+
```
137+
138+
Fix pattern:
139+
140+
```rust
141+
#[cfg(unix)]
142+
use tokio::net::{UnixListener, UnixStream};
143+
```
144+
145+
Move Unix-only functions into Unix-only modules, or add a Windows stub that
146+
returns an unsupported error.
147+
148+
### Linux-only dependency reaches Windows
149+
150+
Symptoms:
151+
152+
```text
153+
failed to run custom build command for libseccomp-sys
154+
pkg-config could not find libsecret
155+
```
156+
157+
Fix pattern:
158+
159+
```toml
160+
[target.'cfg(target_os = "linux")'.dependencies]
161+
libseccomp = "..."
162+
```
163+
164+
Only gate the dependency if no Windows path should use it.
165+
166+
### ARM64 check fails but x64 passes
167+
168+
Likely causes:
169+
170+
- Native dependency does not support `aarch64-pc-windows-msvc`.
171+
- ARM64 MSVC or Spectre-mitigated libraries are missing.
172+
- Host-native `clang-cl`, Ninja, or CMake is missing during an x64-to-ARM64 build.
173+
- `CL` or `_CL_` injects a global MSVC option such as `/MP4` into `clang-cl`.
174+
- Build script assumes x64 tools.
175+
- Inline assembly or prebuilt artifact lacks ARM64 handling.
176+
177+
Do not skip ARM64 silently. Either fix the target handling or report the exact
178+
blocked dependency.
179+
180+
### Focused tests report many filtered-out tests
181+
182+
This is expected for `windows:test:unsupported:x64`. Cargo runs one named test
183+
and filters the other `openshell-server` tests. Report these as filtered, not
184+
ignored.
185+
186+
## Reporting Counts
187+
188+
Use the log summaries from:
189+
190+
| Log | Count source |
191+
|---|---|
192+
| `test-x86_64-pc-windows-msvc.log` | Full x64 workspace test pass. |
193+
| `test-aarch64-pc-windows-msvc.log` | Full native ARM64 workspace test pass. |
194+
| `test-x86_64-pc-windows-msvc-unsupported-*.log` | Focused unsupported-contract re-runs and filtered counts. |
195+
| `test-aarch64-pc-windows-msvc-unsupported-*.log` | Focused native ARM64 re-runs and filtered counts. |
196+
197+
Separate:
198+
199+
- passed
200+
- failed
201+
- ignored
202+
- filtered out
203+
- cfg-gated zero-test targets
204+
- package-level excludes
205+
206+
Package-level excludes are not printed as ignored tests by Cargo.
207+
208+
## Final Sanity Checks
209+
210+
Before committing Windows-lane changes, choose checks based on the host
211+
architecture:
212+
213+
```powershell
214+
cargo fmt --all
215+
git diff --check
216+
$arch = [System.Runtime.InteropServices.RuntimeInformation]::OSArchitecture
217+
if ($arch -eq [System.Runtime.InteropServices.Architecture]::Arm64) {
218+
mise run --skip-tools windows:check:arm64
219+
mise run --skip-tools windows:build:arm64
220+
mise run --skip-tools windows:test:arm64
221+
mise run --skip-tools windows:test:unsupported:arm64
222+
} else {
223+
mise run --skip-tools windows:check:x64
224+
mise run --skip-tools windows:check:arm64
225+
mise run --skip-tools windows:test:unsupported:x64
226+
}
227+
```
228+
229+
Run the full x64-host `windows:ci` lane when build or test behavior changed and
230+
the host can run that lane natively.

‎.agents/skills/create-github-issue/SKILL.md‎

Lines changed: 34 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -17,22 +17,27 @@ This project uses YAML form issue templates. When creating issues, match the tem
1717

1818
### Bug Reports
1919

20-
Do not add a type label automatically. The body must include an **Agent Diagnostic** section — this is required by the template and enforced by project convention. Apply area or topic labels only when they are clearly known.
20+
Do not add a type label automatically. The body must include a **User Story**, **Problem Statement**, **Impact / Why This Matters**, and **Acceptance Criteria**, followed by bug-specific reproduction steps and environment details. Logs are optional and must be concise and redacted. Apply area or topic labels only when they are clearly known.
2121

2222
```bash
2323
gh issue create \
2424
--title "bug: <concise description>" \
2525
--body "$(cat <<'EOF'
26-
## Agent Diagnostic
26+
## User Story
2727
28-
<Paste the output from the agent's investigation. What skills were loaded?
29-
What was found? What was tried?>
28+
As a <persona>, I want <capability or outcome>, so that <benefit or impact>.
3029
31-
## Description
30+
## Problem Statement
31+
32+
<Summarize what is broken or missing in OpenShell's current behavior and when the issue occurs>
33+
34+
## Impact / Why This Matters
3235
33-
**Actual behavior:** <what happened>
36+
<Explain the consequences for users, the current workaround, and why that workaround is insufficient>
3437
35-
**Expected behavior:** <what should happen>
38+
## Acceptance Criteria
39+
40+
- [ ] <observable outcome that demonstrates the bug is fixed>
3641
3742
## Reproduction Steps
3843
@@ -41,43 +46,54 @@ What was found? What was tried?>
4146
4247
## Environment
4348
44-
- OS: <os>
45-
- Docker: <version>
4649
- OpenShell: <version>
50+
- OS: <os>
51+
- Runtime, deployment, or integration: <relevant details>
4752
4853
## Logs
4954
5055
```
51-
<relevant output>
56+
<optional minimal, redacted output>
5257
```
5358
EOF
5459
)"
5560
```
5661

5762
### Feature Requests
5863

59-
Do not add a type label automatically. The body must include a **Proposed Design** — not a "please build this" request. Apply area or topic labels only when they are clearly known.
64+
Do not add a type label automatically. The body must include a **User Story**, **Problem Statement**, **Impact / Why This Matters**, **Proposed Design**, **Acceptance Criteria**, and **Alternatives Considered**. The proposed design should define the user-facing workflow and externally observable behavior without prescribing internal implementation. Agent investigation is optional. Apply area or topic labels only when they are clearly known.
6065

6166
```bash
6267
gh issue create \
6368
--title "feat: <concise description>" \
6469
--body "$(cat <<'EOF'
70+
## User Story
71+
72+
As a <persona>, I want <capability or outcome>, so that <benefit or impact>.
73+
6574
## Problem Statement
6675
67-
<What problem does this solve? Why does it matter?>
76+
<Summarize the capability or behavior missing from OpenShell today>
77+
78+
## Impact / Why This Matters
79+
80+
<Explain what users must do today, why it is insufficient, and the operational cost, risk, blocked workflow, or adoption barrier>
6881
6982
## Proposed Design
7083
71-
<How should this work? Describe the system behavior, components involved,
72-
and user-facing interface.>
84+
<The desired user-facing workflow and externally observable behavior, without prescribing internal implementation>
85+
86+
## Acceptance Criteria
87+
88+
- [ ] <specific, observable outcome>
7389
7490
## Alternatives Considered
7591
76-
<What other approaches were evaluated? Why is this design better?>
92+
<Other user-facing workflows or behaviors considered and why this approach best satisfies the user story>
7793
7894
## Agent Investigation
7995
80-
<If the agent explored the codebase to assess feasibility, paste findings here.>
96+
<Optional findings from codebase exploration>
8197
EOF
8298
)"
8399
```
@@ -107,6 +123,8 @@ EOF
107123

108124
GitHub built-in issue types (`Bug`, `Feature`, `Task`) should come from the matching issue template when possible, or be set manually afterward. Do not try to emulate them through labels.
109125

126+
Creating an issue does not accept it or queue agent work. Agents never apply `state:accepted`, the `roadmap` label, add issues to the roadmap project, or apply `agent:plan-requested` or `agent:implementation-requested`. Community issues proceed through `triage-issue`; a human accepts technically validated work with `state:accepted` or roadmap placement. The request labels queue work for unattended agents; a user may instead direct an agent to a specific issue.
127+
110128
## Useful Options
111129

112130
| Option | Description |

‎.agents/skills/create-github-pr/SKILL.md‎

Lines changed: 10 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -11,7 +11,7 @@ Create pull requests on GitHub using the `gh` CLI.
1111

1212
- The `gh` CLI must be authenticated (`gh auth status`)
1313
- You must have commits on a branch that's pushed to the remote
14-
- Branch should follow naming convention: `<issue-number>-<description>/<username>`
14+
- For issue-backed work, the branch should follow `<issue-number>-<description>/<username>`. Exempt issue-less changes may use `<description>/<username>`.
1515

1616
## Before Creating a PR
1717

@@ -24,6 +24,10 @@ in the same branch. If the change affects user-facing compute-driver setup,
2424
also update `docs/reference/sandbox-compute-drivers.mdx` or the relevant
2525
deployment docs.
2626

27+
### Check Agent Infrastructure
28+
29+
Use the `sync-agent-infra` skill's maintenance map to identify related skill updates when the branch changes behavior, commands, or development workflows. Run its full consistency check when the branch adds, removes, or renames skills or crates; changes workflow relationships or skill coverage; modifies issue or PR templates; or changes agent cross-references. Resolve any drift before creating the PR.
30+
2731
### Run Pre-commit Checks
2832

2933
Run the local pre-commit task before opening a PR:
@@ -43,7 +47,7 @@ Before creating a PR, verify:
4347
git branch --show-current
4448
```
4549

46-
2. **Branch follows naming convention** - Format: `<issue-number>-<description>/<initials>`
50+
2. **Branch follows naming convention** - Use `<issue-number>-<description>/<initials>` for issue-backed work or `<description>/<initials>` for an exempt issue-less change.
4751

4852
```bash
4953
# Example: 1234-add-pagination/jd
@@ -110,7 +114,7 @@ gh pr create --title "PR title" --body "PR description"
110114

111115
### Link to an Issue
112116

113-
Use `Closes #<issue-number>` in the body to auto-close the issue when merged:
117+
Features, user-visible behavior changes, public API changes, architecture changes, and multi-PR efforts must link an accepted issue. Use `Closes #<issue-number>` in the body to auto-close the issue when merged:
114118

115119
```bash
116120
gh pr create \
@@ -122,6 +126,8 @@ gh pr create \
122126
- Returns 400 instead of 500"
123127
```
124128

129+
Small documentation fixes, mechanical maintenance, and obvious localized bug fixes may omit a separate issue when the PR contains enough context to review the decision and implementation together. In that case, write `No issue required: <brief reason>` in the Related Issue section. Do not use this exception for security fixes; follow `SECURITY.md`.
130+
125131
### Create as Draft
126132

127133
For work-in-progress that's not ready for review:
@@ -153,7 +159,7 @@ PR descriptions must follow the project's [PR template](.github/PULL_REQUEST_TEM
153159
<!-- 1-3 sentences: what this PR does and why -->
154160

155161
## Related Issue
156-
<!-- Fixes #NNN or Closes #NNN -->
162+
<!-- Fixes #NNN / Closes #NNN, or "No issue required: <reason>" for an exempt change -->
157163

158164
## Changes
159165
<!-- Bullet list of key changes -->

0 commit comments

Comments
 (0)