Skip to content

feat(gateway): explain compute driver discovery decisions #3211

Description

@elezar

User Story

As an OpenShell operator, I want the gateway to explain how it selected or rejected local compute drivers, so that I can diagnose startup failures without reproducing them under a debugger.

Problem Statement

Gateway auto-detection probes local driver candidates, such as Docker Unix sockets, but silently treats probe failures as unavailable. When a candidate socket exists but Snap confinement denies access, the gateway eventually reports only that no suitable driver was found and systemd may restart it. Operators cannot tell which candidates were examined, whether a driver was selected, or whether a candidate failed because of a missing socket, an access denial, a timeout, or an unexpected API response.

Impact / Why This Matters

Package and confinement failures are difficult to distinguish from an absent runtime. The current workaround is to manually inspect socket paths, Snap interface connections, and journal output, then infer the cause. That is insufficient for CI and user installations because the decisive probe result is not recorded.

Proposed Design

When gateway driver auto-detection runs, expose diagnostics at debug level that identify each candidate driver and socket probe outcome without logging credentials or other sensitive data. Startup errors should provide a concise summary of the attempted drivers and why no usable driver was selected. A successful selection should state the selected driver and endpoint category in diagnostics.

Acceptance Criteria

  • With debug logging enabled, each local compute-driver candidate records whether it was selected, skipped, or rejected.
  • Docker socket diagnostics distinguish a missing/non-socket path, connection denial, timeout, and an incompatible or unsuccessful API response.
  • Normal startup errors summarize the failed discovery decision without exposing secrets or request contents.
  • A successful auto-detection records the selected driver and does not expose sensitive endpoint data.
  • Coverage verifies representative successful, missing, and permission-denied probe outcomes.

Alternatives Considered

  • Relying only on systemd or Snap logs: these do not report the gateway discovery decision or every candidate it attempted.
  • Treating socket existence as availability: this would hide confinement failures and select unusable drivers.
  • Requiring an explicitly configured driver: useful as a workaround, but does not make default installations diagnosable.

Agent Investigation

Docker discovery currently probes each candidate with a Unix-socket HTTP ping and returns only a boolean result. Failed metadata, connect, write, read, timeout, and response-validation paths are intentionally collapsed into unavailable. This behavior is shared by local API socket discovery and leaves no per-candidate diagnostic trail.

Related: #2869 (comment)

Activity

  1. elezar commented on Sep 7, 2026

    @elezar
    MemberAuthor

    🏗️ build-plan

    Implementation Plan

    Issue type: feat
    Complexity: Medium
    Confidence: High — the behavior is localized to gateway auto-detection, local socket probing, startup diagnostics, and operator docs.
    Minimal scope: Add gateway diagnostics without changing driver selection order, config schema, runtime construction, or remote-driver behavior.
    LSM compatibility: Low impact. SELinux/AppArmor/Snap denials should surface as permission-denied probe outcomes; tests should avoid /proc or fork/exec assumptions.

    Summary

    Gateway auto-detection should explain why each compiled local compute driver was selected, skipped, or rejected. Replace the ComputeDriverRegistration.detect: fn() -> bool hook with a structured detection result, let Docker keep detailed local socket probe outcomes internally, and emit redaction-safe gateway service logs plus concise startup failure summaries for operators.

    The related PR #2869 review comment reinforces the Snap/Docker interface case: a socket can exist while the gateway is denied access, so permission-denied must be distinguishable from Docker absent.

    Scope

    • crates/openshell-core/src/local_api_socket.rs: add a diagnostic probe API with explicit outcomes for missing/non-socket path, permission denied, connect/write/read failure, timeout, empty response, and rejected response. Keep the current first_responsive_socket() wrapper for existing callers.
    • crates/openshell-driver-docker/src/lib.rs: add Docker socket discovery diagnostics and bounded endpoint categories for DOCKER_HOST, /var/run/docker.sock, Docker Desktop, and XDG_RUNTIME_DIR. Expose redacted detail summaries for gateway logs and keep raw socket paths out of log-oriented APIs.
    • crates/openshell-server/src/lib.rs: change the registered detector hook to return ComputeDriverDetectionResult instead of bool. Retain the registry's ordered decision trail so gateway startup can report selected, skipped, and rejected drivers.
    • crates/openshell-server/src/cli.rs: emit stored gateway auto-detection debug logs after tracing is installed. Failed detection still returns a normal startup error with the redacted gateway auto-detection summary.
    • crates/openshell-gateway/src/lib.rs: wire Kubernetes, Podman, and Docker detector functions directly through the structured result API. Docker contributes richer socket details; Kubernetes and Podman provide driver-level summaries.
    • docs/reference/sandbox-compute-drivers.mdx and architecture/compute-runtimes.md: document the gateway/operator diagnostics behavior and clarify that this is gateway service logging, not sandbox telemetry.

    Implementation Steps

    1. Add diagnostic structs/enums in local_api_socket and route the existing boolean helper through them.
    2. Implement Docker-specific candidate categorization and response classification, including unsuccessful status vs incompatible API response.
    3. Replace the compute-driver detection hook's boolean return with a structured detection result.
    4. Store and emit redaction-safe gateway auto-detection decisions after tracing is installed.
    5. Format the no-suitable-driver startup error with driver names and reason buckets, not secrets, endpoint values, or probe request contents.
    6. Update focused docs for operators using OPENSHELL_LOG_LEVEL=debug.

    Test Plan

    • Unit tests: Add openshell-core tests for successful socket probe, missing/non-socket path, timeout/read failure, permission-denied outcome mapping, and rejected response classification. Add openshell-driver-docker tests for Docker response acceptance, Podman/incompatible response rejection, unsuccessful status rejection, and candidate category ordering/redaction. Add openshell-server tests for selected, skipped, rejected, and no-driver summary behavior.
    • Integration tests: N/A; no gRPC protocol or runtime lifecycle behavior changes.
    • E2E tests: N/A for implementation. Final verification should run mise run pre-commit and targeted package-level Rust tests for openshell-core, openshell-driver-docker, and openshell-server; run broader suites if the local toolchain permits.

    Risks & Open Questions

    • Debug logs for failed selection happen before the current full tracing setup unless startup preparation is split. The minimal plan uses the normal startup error for failed discovery and post-setup debug logs for successful detection.
    • Permission-denied socket tests can be platform-sensitive or bypassed by privileged test users. Keep them Unix-only and assert the error classification separately.
    • Do not broaden Docker probing to TCP DOCKER_HOST; unsupported/non-Unix values should be reported as skipped without logging the raw value.

    Documentation Impact

    • Update docs/reference/sandbox-compute-drivers.mdx.
    • Update architecture/compute-runtimes.md because the registered driver detection contract changes from boolean availability to structured gateway diagnostics.
    • No docs/reference/gateway-config.mdx update expected because no gateway TOML fields, driver config keys, defaults, or Helm rendering change.

    Revision 2 — align implementation with gateway/operator diagnostics and structured detector results
    Revision 1 — initial plan

  2. github-actions commented on Sep 22, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  3. lunarwhite commented on Oct 2, 2026

    @lunarwhite
    Contributor

    🏗️ build-plan

    Outcome

    When auto-detection finds no usable driver, the startup error lists every installed driver, its decision, and each probe outcome. When detection succeeds, debug logs show every candidate, and the existing Using compute driver line names the endpoint. Driver selection does not change. One PR closes #3211.

    Findings

    • Every probe failure becomes a boolean. socket_responds() returns false (crates/openshell-core/src/local_api_socket.rs:34-75), Podman's CLI fallback returns None (crates/openshell-driver-podman/src/socket_discovery.rs:160-212), and the registry keeps only detect: Option<fn() -> bool> (crates/openshell-server/src/lib.rs:1279).
    • Detection runs before tracing is installed, and tracing setup needs the selected driver (crates/openshell-server/src/cli.rs:595, :618-635). A failed detection must carry its details in the startup error.
    • Registration keeps source-compatibility shims for custom gateway binaries (lib.rs:1253-1262, :1316-1342), so the API change must be additive.
    • config preflight must not run probes, and it picks the tables to validate through auto_detectable_driver_names() (cli.rs:851-870, lib.rs:1422-1432), which must count the new probe kind.
    • Candidates come from DOCKER_HOST, OPENSHELL_PODMAN_SOCKET, CONTAINER_HOST, $HOME, $XDG_RUNTIME_DIR, and podman output, so output may use only fixed labels.

    Design

    1. Probe outcome in openshell-core: accepted, missing, not a socket, permission denied, connection refused, timed out, I/O error, empty response, unsuccessful (HTTP NNN), or incompatible. Errors are classified by kind at every step, so a denied stat is permission denied and a socket timeout (WouldBlock on Unix) is timed out. File modes, groups, AppArmor or Snap, and SELinux all appear as permission denied; the message does not guess the layer.
    2. Driver report in openshell-core: an ordered list of &'static str labels with outcomes. It has no field for paths, environment values, error text, CLI output, or reply bytes, so redaction holds by construction. Docker and Podman add detection_report(). A non-unix:// DOCKER_HOST and unset variables appear as skipped, the Podman CLI reports not installed, timed out, failed, no socket found, or found, and Kubernetes reports whether KUBERNETES_SERVICE_HOST is set.
    3. Registry in openshell-server: first-party drivers register through a new ComputeDriverRegistration::with_detection_report(..), which replaces the boolean probe; new(..) is unchanged. detect() already runs every probe, so it records selected, skipped (lower priority than <driver>), rejected with its checks, or skipped (not auto-detected).
    4. Output: a failure puts the list in the startup error. On success, the gateway logs one debug! per candidate after tracing is installed and adds the endpoint label to the Using compute driver line (lib.rs:1642). These use plain tracing, not OCSF.
    Error:   × configuration error: no compute driver configured and auto-detection found no usable driver:
      │   kubernetes: rejected (KUBERNETES_SERVICE_HOST is unset)
      │   podman: rejected ($XDG_RUNTIME_DIR/podman/podman.sock: missing; /run/user/$EUID/podman/podman.sock: missing; $HOME/.local/share/containers/podman/machine/podman.sock: missing; podman CLI: not installed)
      │   docker: rejected (/var/run/docker.sock: permission denied; $HOME/.docker/run/docker.sock: missing; $XDG_RUNTIME_DIR/docker.sock: missing)
      │   vm: skipped (not auto-detected)
      │ set --compute-driver <name> or OPENSHELL_COMPUTE_DRIVER=<name>
    

    Out of scope, each needing its own issue: Docker and Podman detect again at construction (crates/openshell-driver-docker/src/lib.rs:850-854, crates/openshell-driver-podman/src/driver.rs:418), and Podman's CLI fallback returns the podman info socket path without probing it (socket_discovery.rs:60-62), so Podman can be selected while its socket is absent.

    Alternatives

    • Defer the error until tracing is installed: restructures startup only to print what the error already says.
    • Probe in config preflight or a separate subcommand: preflight must stay side-effect free, and a subcommand runs in the caller's security context, so it misses the service's confinement denials.
    • Extension points: interceptors, middleware, and providers run after startup and never see driver selection.

    Tests

    • openshell-core: one case per outcome on temporary Unix sockets (missing, regular file, dead listener, silent listener, close after reading the request, HTTP 500, wrong identity, Docker reply), plus a pure EACCES/EPERM mapping test, because root bypasses file modes.
    • openshell-server: decision order, the full no-driver error, and captured debug and info events.
    • Redaction: sentinel inputs such as tcp://user:secret@203.0.113.7:2376 and /home/sentinel-user never appear in output.
    • Preflight validates a report-probe registration without running it.
    • mise run e2e:gateway:no-compute-drivers keeps core and server backend-agnostic. The runtime E2E lanes pin their driver, so none exercises auto-detection.

    Documentation

    • docs/how-it-works/sandboxes/runtimes.mdx: one or two sentences on the per-driver startup error and log_level = "debug".
    • skills/debug-openshell-cluster/SKILL.md: read the per-driver list, and treat permission denied as a socket-access problem, such as group membership or an unconnected Snap interface.
  4. lunarwhite commented on Oct 2, 2026

    @lunarwhite
    Contributor

    Hi @elezar checking are you working on this? or do we still need this feature in the near future?

    I've tried to regenerate the build plan as there are quite a few changes since your last triage and plan. If we don't need it anymore or it has been de-prioritized, please just let me know.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions