Skip to content

Latest commit

 

History

History
463 lines (361 loc) · 17.3 KB

File metadata and controls

463 lines (361 loc) · 17.3 KB

Process Playbook

Operator guide for managing Refarm factory processes. Prefer the refarm runtime commands for daily use; package scripts remain documented here for backend debugging and CI-oriented maintenance.

Architecture overview

The factory runs one backend at a time on port 42000. Two backends exist:

Backend Language Ports Role
tractor Rust :42000 (WS) WASM plugin host, agent, native CRDT sync
farmhand Node.js :42000 (WS CRDT sync) + :42001 (HTTP sidecar) Task orchestration, file queue, effort dispatch

They share port 42000 and must not run at the same time.

  • Use tractor for interactive agent work (REPL, direct prompts, streaming).
  • Use farmhand for batch task dispatch (effort queue, CI smoke, file transport).

Quick reference

pnpm run farm:status        # unified status: both services, ports, artifacts, MODEL
refarm runtime              # selected runtime engine + autostart policy
refarm runtime ensure --wait        # start selected runtime when needed
refarm runtime restart --wait       # stop/start selected runtime and wait
refarm runtime stop                 # stop tracked runtime backends
refarm config set tractor.engine auto      # auto | rust | ts
refarm config set runtime.autostart ask    # ask | always | never
refarm telemetry           # runtime pressure snapshot (queue/in-flight/failures)
pnpm run refarm:telemetry:gate:ci      # strict fail-closed gate (recommended CI policy)
pnpm run refarm:telemetry:gate:strict-all  # enforce all diagnostics (hard mode)
# When checking remote CI results (after local validation):
gh run list --workflow test.yml --limit 5
gh run watch --exit-status
refarm agent install       # manual: force-install agent (normally auto-installed on runtime boot)
pnpm run agent:daemon       # low-level: start tractor in background
pnpm run agent:stop         # low-level: stop tractor
pnpm run farmhand:daemon    # low-level: start farmhand in background
pnpm run farmhand:stop      # low-level: stop farmhand
pnpm run disk:check         # disk usage: target dirs, node_modules, volumes
pnpm run actions:budget:guard:account      # hard Actions guard: monthly net billable quota
pnpm run actions:budget:guard:allocation   # advisory Actions guard: Refarm fairness split
pnpm run actions:budget:guard:modes:json   # discover hard/advisory guard metadata

# Session memory helpers (host-owned)
refarm sessions list        # list known sessions
refarm sessions new         # create and switch active session
refarm sessions fork <id>   # branch from an existing session
refarm sessions use <id>    # session helper to switch active session
refarm tree switch <id>     # timeline-first active-session switch
refarm tree list --json             # read-only session timeline nodes
refarm tree list --limit 5 --json   # bounded session timeline nodes
refarm tree list --scope git --json # read-only git timeline nodes
refarm tree preview <id>            # dry-run fork plan for a session node
refarm tree preview <id> --at <entry> # dry-run fork plan at a historical entry
refarm tree preview <id> --name <branch> # dry-run fork plan with explicit name
refarm tree preview --scope git <commit> # dry-run branch plan for a commit
refarm tree fork --scope git <commit> --name <branch> # create branch without switching
pnpm run refarm:actions:verify # closeout lane for action-readiness changes
pnpm run refarm:tree:verify # closeout lane for tree stabilization changes

CLI runtime controls

refarm owns the daily-driver runtime lifecycle. Use the manual pnpm run agent:* and pnpm run farmhand:* commands below only when debugging the backends directly.

refarm runtime                         # show configured/active engine
refarm runtime --json                  # machine-readable runtime status
refarm runtime ensure --wait           # start the selected runtime when needed
refarm runtime stop                    # stop tracked runtime backends
refarm runtime restart --wait          # restart the selected runtime
refarm config set tractor.engine auto  # prefer Rust tractor, fall back to TS farmhand
refarm config set tractor.engine rust  # require Rust tractor; fail early if missing
refarm config set tractor.engine ts    # force TypeScript farmhand
refarm config set runtime.autostart always  # ask | always | never

runtime.autostart is the canonical autostart key for CLI flows. The legacy farmhand.autostart key still reads and writes the same stored value for compatibility. REFARM_RUNTIME_AUTOSTART is the preferred environment override; REFARM_FARMHAND_AUTOSTART remains a legacy fallback.


Tractor (agent backend)

Start

pnpm run agent:daemon       # background — writes .refarm/tractor.pid + tractor.log
pnpm run agent:start        # foreground — no PID file, Ctrl+C to stop

Prerequisites: tractor binary must exist.

# Build (once, or after Rust changes):
cd packages/tractor && cargo build --release
# Outputs to $CARGO_TARGET_DIR/release/tractor (devcontainer volume)

Stop

refarm runtime stop         # SIGTERM via tracked runtime PID files
# Low-level fallback:
pnpm run agent:stop
# Or kill directly:
kill $(cat .refarm/tractor.pid)

Logs

tail -f .refarm/tractor.log     # background mode only

Check

pnpm run farm:status
# tractor section shows: pid, WS probe result, binary age

Farmhand (task orchestration daemon)

Start

pnpm run farmhand:daemon    # background — writes .refarm/farmhand.pid + farmhand.log
pnpm run farmhand:start     # foreground — no PID file, Ctrl+C to stop

No build step required — runs from source via Node type-stripping with the Farmhand resolver loader (scripts/farmhand-node-loader.mjs).

Stop

refarm runtime stop         # SIGTERM via tracked runtime PID files
# Low-level fallback:
pnpm run farmhand:stop

Logs

tail -f .refarm/farmhand.log    # background mode only

Check

pnpm run farm:status
# farmhand section shows: pid, HTTP sidecar probe, task queue depth

# Direct HTTP probe:
curl -s http://127.0.0.1:42001/efforts/summary | node -e "process.stdin|>JSON.parse|>console.log"
# Runtime-readiness check also validates:
curl -s http://127.0.0.1:42001/sessions | node -e "process.stdin|>JSON.parse|>console.log"

Common scenarios

Scenario 1 — Start fresh for interactive agent work

pnpm run farm:status        # verify nothing is running
refarm runtime ensure --wait
pnpm run agent:repl         # start REPL
# When done:
refarm runtime stop

Scenario 2 — Run a task smoke test

pnpm run farm:status        # verify nothing is running on :42001
# Tests start their own farmhand stub on :42001; a running farmhand will conflict.
pnpm run task:execution:smoke
# Or:
pnpm run task:execution:smoke:agent

If farmhand is already running, stop it first:

refarm runtime stop && pnpm run task:execution:smoke

Scenario 3 — Farmhand for agent task dispatch (batch)

pnpm run farm:status        # verify tractor is not running
pnpm run farmhand:daemon    # starts :42000 (WS) + :42001 (HTTP)
# Dispatch efforts via HTTP:
curl -X POST http://127.0.0.1:42001/efforts -H 'Content-Type: application/json' \
  -d '{"task": {...}, "effort": {...}}'
# Check queue:
curl -s http://127.0.0.1:42001/efforts/summary
# Check session catalog readiness:
curl -s http://127.0.0.1:42001/sessions | node -e "process.stdin|>JSON.parse|>console.log"
# Check rolling pressure window:
curl -s 'http://127.0.0.1:42001/telemetry/window?minutes=30'
pnpm run farm:status
# Stop:
refarm runtime stop

### Scenario 3a — VPN com preparo explícito (sem auto-disparo)

Use este fluxo para evitar tentativas automáticas quando o operador ainda não
está com o celular em mãos.

```bash
# 1) Inspecionar estado da conexão declarada
refarm workspace run refarm vpn-status

# 2) Armar preparo explícito do operador (janela curta)
refarm workspace run refarm vpn-prepare

# 3) Checar prontidão sem efeito colateral
refarm workspace run refarm vpn-ready

# 4) Só então tentar subida segura
refarm workspace run refarm vpn-up-safe

Notas operacionais:

  • vpn-up-safe bloqueia quando não houve preparo explícito recente.
  • vpn-up-safe também bloqueia durante cooldown para evitar rajada de tentativas.
  • vpn-ready é read-only: apenas informa se está pronto para tentar.

Canal reutilizável (além de VPN):

  • O contrato do preparo explícito agora vive no pacote compartilhado packages/operator-state/src/index.ts (helpers de comandos e handoff neutro de atenção).
  • scripts/operator-attention-gate.mjs atua como adaptador fino desse contrato para uso direto em shell.
  • A VPN é apenas um consumidor desse canal (scope = connection-up:<name>).
  • Outras ações sensíveis podem reaproveitar o mesmo contrato, sem acoplar regra de negócio ao workspace rcdc5.
  • O contrato é surface-agnostic (ok, nextAction, nextActions, nextCommand, nextCommands) para beneficiar CLI, web, automações e futuras superfícies sem depender de detalhes do app refarm.

Exemplo genérico:

node scripts/operator-attention-gate.mjs attention:minha-acao --prepare-only --window-ms 120000 --json
node scripts/operator-attention-gate.mjs attention:minha-acao --check-only --json
node scripts/operator-attention-gate.mjs attention:minha-acao --consume-only --json

Fluxo recomendado no CLI (entre dispositivos/processos):

# 1) Armar intenção com perfil reutilizável
refarm intention arm --profile cross-device-handoff --json

# 2) Verificar prontidão explícita da intenção
refarm intention check --profile cross-device-handoff --json

# 3) Projetar esse estado no status base
refarm status --base --attention-profile cross-device-handoff --json

# 4) Consumir intenção após executar a ação sensível
refarm intention consume --profile cross-device-handoff --json

Notas:

  • refarm intention separa o ato de intencionar da execução operacional.
  • Perfis (cross-device-handoff, mobile-ready, operator-sync) reduzem ad-hoc de scope/window e facilitam assimilação de processos recorrentes.
  • refarm status --base continua aceitando override explícito por --attention-scope e --attention-window-ms quando necessário.

Scenario 3b — Canonical local ask flow (daily driver)

refarm runtime ensure --wait
refarm ask "o que é CRDT?"

Notes:

  • Farmhand auto-installs agent from the bundled npm package when it boots. To manually trigger: refarm agent install.
  • refarm ask ... works with the built-in resolver loader.
  • Use refarm telemetry --profile balanced (or conservative/throughput) to watch queue/in-flight pressure and recent failure-rate signals.
  • Use refarm telemetry --strict to fail-closed when diagnostics are present (automation/CI-friendly exit code 2).
  • For automation wrappers that can bootstrap farmhand when needed, use pnpm run refarm:telemetry:gate:ci.
  • To persist gate output artifacts for trend analysis, add --out .artifacts/telemetry/gate-latest.json.
  • Signal meanings + first-response actions live in docs/REFARM_TELEMETRY_RUNBOOK.md.

Scenario 3c — Session-first workflow

refarm sessions new --name "auth-refactor"
refarm ask "planeje os próximos passos"
refarm sessions fork <id-prefix> --name "auth-refactor-alt"
refarm tree preview <id-prefix> --switch
refarm tree switch <id-prefix>
refarm ask --session <id-prefix> "continue deste branch"

Use this when exploring multiple solution branches without losing continuity. --session pins a request to a specific session without switching first. Use refarm ask --new "..." when you explicitly want a fresh conversation: the CLI clears the active pointer, allocates a new session ID, submits the ask with that fresh ID, and persists the same ID only after a successful stream or fallback result. Active-session pointer writes are verified through the shared session-lock.ts helper. The longer-term substrate-agnostic design lives in Refarm Tree Primitive.

Scenario 3d — Timeline-first tree workflow

Use refarm tree when you want renderer-neutral inspection and dry-run readiness before moving state:

refarm tree list --scope all
refarm tree list --scope session
refarm tree preview <session-id-prefix> --switch
refarm tree switch <session-id-prefix>

refarm tree list --scope git --limit 5
refarm tree preview --scope git HEAD --name experiment/refactor
refarm tree fork --scope git HEAD --name experiment/refactor
refarm tree preview --scope git experiment/refactor --switch
refarm tree switch --scope git experiment/refactor

Preview commands are non-mutating. Blocked-but-resolvable previews should return operator-readable readiness (Blocked: ...) and, for JSON output, readyToExecute: false. When a missing value can be supplied later, JSON output may include templates with declared parameters; nextCommands should stay empty until a concrete command is available. Execution remains explicit via tree fork or tree switch. After changing tree contracts or adapter boundaries, run pnpm run refarm:tree:verify before considering the tree slice closed.

Scenario 4 — Port conflict at startup

pnpm run agent:daemon
# ❌  Port 42000 is already bound by PID 4567.
pnpm run farm:status        # identify what's on :42000
refarm runtime stop         # stop tracked runtime backends
# Then retry:
refarm runtime ensure --wait

Scenario 5 — Run agent harness tests

The harness starts its own mock MODEL on a random port — no service conflict possible. It does NOT require tractor or farmhand to be running.

# Build WASM first (outputs to $CARGO_TARGET_DIR):
cd packages/agent && cargo component build --release
# Run harness (serialize with --test-threads=1):
cd packages/tractor && cargo test --test agent_harness -- --ignored --test-threads=1

If you want to also keep tractor running for interactive work: the harness uses an in-memory NativeSync (:memory:) — no port conflict.

Scenario 6 — After disk cleanup, rebuild binaries

After running pnpm run clean:heavy or when CARGO_TARGET_DIR volume is empty:

# Rebuild tractor binary:
cd packages/tractor && cargo build --release
# Rebuild agent WASM:
cd packages/agent && cargo component build --release
# Verify:
pnpm run farm:status        # check ARTIFACTS section

Port assignments

Port Protocol Role Conflict if…
42000 WebSocket CRDT sync (tractor OR farmhand) Both backends running simultaneously
42001 HTTP farmhand HTTP sidecar farmhand + CI smoke stub both running
:0 HTTP mock MODEL (harness tests) OS-assigned — no conflict possible

State files in .refarm/

.refarm/
  tractor.pid      # tractor background PID
  tractor.log      # tractor stdout/stderr (background mode)
  farmhand.pid     # farmhand background PID
  farmhand.log     # farmhand stdout/stderr (background mode)
  .env             # runtime env read by launcher; may differ from Silo identity store
  config.json      # model provider, model, budgets, FS restrictions
  .repl_history    # REPL command history
  tasks/           # FileTransport input queue (farmhand)
  task-results/    # effort outcomes
  task-logs/       # effort NDJSON logs
  task-control/    # retry/cancel signals
  streams/         # stream chunk files
  plugins/         # installed plugin manifests + WASM blobs
    agent/
      plugin.json
      agent.wasm

All .refarm/ contents are gitignored. Root-level integration namespaces such as .project/ or .pi-lens/ are governed separately by ADR-071: use .refarm/ for Refarm-owned local state by default, and require an explicit declaration for other workspace-local namespaces.


Diagnostic checklist

When something is wrong, work top-down:

  1. refarm runtime --json — start here for daily-driver runtime state.
  2. Port conflict? — ss -tlnp | grep '42000\|42001' — identify the PID.
  3. Stale PID file? — process dead but PID file exists → refarm runtime stop then retry.
  4. Binary missing? — ARTIFACTS section in farm:status will tell you what to build.
  5. No credentials? — refarm model doctor --json or refarm sow.
  6. Context/store mismatch? — compare refarm model current --json with refarm runtime status --json and verify SILO_HOME / REFARM_HOME alignment.
  7. Disk full? — pnpm run disk:check → pnpm run clean:light or pnpm run clean:heavy.
  8. WASM/harness fails? — ensure $CARGO_TARGET_DIR is set and agent.wasm is at $CARGO_TARGET_DIR/wasm32-wasip1/release/agent.wasm.

Long-term direction

This playbook is an interim measure. The apps/refarm CLI is evolving to abstract daemon lifecycle, runtime selection, task dispatch, and status into first-class commands. refarm runtime, refarm status, and refarm ask already cover the daily-driver path; use this doc and pnpm run farm:status when you need lower-level backend diagnostics.