Clearance is a recovery-first control plane for trusted Linux CI runners.
Its central safety rule is that a runner must never be reused while previous work may still legitimately be executing. Loss of contact, stale messages, controller failures, and ambiguous execution state are unsafe conditions, not evidence that a runner is free.
Clearance is under active development:
| Component | Implemented responsibility |
|---|---|
| Python reference model, reconciliation classifier, and mechanical checker | Deterministic ownership, fencing, quarantine, abstract cleanup/release semantics, and reconciliation classification, checked by tests. This is verification tooling, not a running service. |
| Controller | Java/Spring Boot job intake, project-scoped idempotency, exclusive runner claims, committed allocation delivery, durable execution results, cancellation, fenced cleanup/release, heartbeat-timeout evaluator with durable runner quarantine, classified quarantine reconciliation, and recovery-mode generation authority with fleet quarantine and generation-gated claims/reports. PostgreSQL is the durable ownership authority; Flyway migrations define the schema. |
| Agent | Standalone Linux Go daemon with durable incarnation/sequence state, execution replay prevention, cgroup v2 execution, workload deadlines, descendant cleanup, workspace scrubbing, physical discovery after agent SIGKILL, and fresh reconciliation observations for quarantined runners. |
| Agent wire contract | Versioned HTTP/JSON behavior with shared Java↔Go compatibility fixtures. |
Job submission does not start execution automatically: tests and the integration harness create initial claims through SchedulerService.claim; there is no general scheduling loop or claim endpoint. Terminal execution retains ownership until current positive cleanup proof permits reuse. Agent-crash recovery can retry interrupted work after verified cleanup. Quarantined runners become reusable only through classified reconciliation with fresh physical evidence. See system architecture and implemented limits for the runtime flow, model boundaries, and unsupported recovery guarantees.
Start with local setup and verification. For a standalone agent, follow agent operation and restart safety. Read security and credential handling before configuring access.
- System architecture — component boundaries, data flow, durable state, and implemented limits.
- Runner ownership model — abstract fencing, quarantine, release, and verification bounds.
- Job intake — API behavior, identities, canonicalization, transactions, and schema invariants.
- Exclusive runner claim — compatibility, ownership transactions, contention, and database invariants.
- Agent API — machine authentication, committed delivery, incarnation rotation, report fencing, and test faults.
- Agent wire contract — wire shapes, compatibility fixtures, and the Java↔Go exchange.
- Agent operation — startup, durable state, restart safety, and failure triage.
- Controller operation — deploy, backup/restore, recovery boot, and triage.
- Configuration reference — environment, flags, and controller settings.
- Agent verification — Linux containment, controller integration, SIGKILL recovery, and bounded diagnostics.
- Security configuration — credentials, transport, and trusted-workload boundaries; report a vulnerability.