Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

nano-longrun-skill

A portable protocol for long-running LLM agent tasks. Three files. Zero dependencies. Works wherever an agent can read and write files.

Drop the skill/ directory into any project. Edit one line in skill/state.json. Bootstrap from Claude, Gemini, Codex, ChatGPT — pick any. The agent advances one atomic task per turn, writes its evidence, and ends. Memory is not trusted. State lives on disk.

# Drop in
cp -R skill /path/to/your/project/

# Bootstrap (pick any CLI)
claude     "Read skill/PROTOCOL.md and follow it."
gemini     "Read skill/PROTOCOL.md and follow it."
codex exec "Read skill/PROTOCOL.md and follow it."

# Or run unattended for hours
AGENT_CMD='claude -p' SKILL_RETRY=3 ./skill/L1/loop.sh

What it does

Spans many turns The agent re-reads the protocol every turn. Memory is intentionally untrusted. State lives in state.json + log.md.
Spans many sessions Kill the process, lose your terminal, swap your laptop — the next invocation continues from the last completed turn.
Spans many CLIs Start in Claude on Monday, resume in Gemini on Tuesday. The protocol is the contract; the CLI is the runtime.
Guards quality Backpressure prevents premature completion. Best-of-N explores alternatives. Adversarial self-review surfaces weaknesses before done is allowed.
Guards drift Hashes of the world (cwd, git HEAD, file checksums) are fenced. External changes halt the run rather than corrupt it.
Stays portable Pure text + JSON. No SDK. No daemon. No vendor lock-in. The skill/ directory is self-contained.

Design philosophy

1. Parasitic host loop. Don't build a scheduler. Every CLI already has one: prompt → response → wait. The agent's job is to re-orient from disk each turn, advance one step, and end. The host's turn-loop is the scheduler. This is what lets the protocol survive on anything that can read and write files.

2. State on disk; memory is not trusted. Each turn cold-starts from three files: PROTOCOL.md, state.json, log.md. The agent never assumes it remembers anything from the previous turn. This is the antidote to context compaction, session timeouts, and CLI swaps — and it's the reason a task started in Claude can finish in Gemini.

3. Self-discipline via invariants, not external hooks. Eleven numbered rules in the protocol carry all the enforcement weight. No host hook, no sidecar daemon, no compiled runtime. The agent's prompt is the contract. This trades hard guarantees for total portability — and pays back in working on every CLI a user might already have.

4. One task per turn. After each task, the turn ends. The next task starts with a fresh re-orient. This forces small blast radius per failure, keeps the audit log line-per-turn parseable by humans and grep, and prevents "agent gets on a roll and drifts" failures common in long runs.

5. Plan quality before work quality. A bad plan misdirects an entire run; the cheapest place to catch a bad plan is before any work starts. Plan-Best-of-N generates alternative strategies; a Plan-Challenger reviews the chosen strategy for gaps; only then does decomposition begin. Catching a strategic flaw at INIT is orders of magnitude cheaper than catching it after fifty turns of drafting.

6. Verification before completion. A task is not done until the agent fills a verification field with concrete, re-readable evidence: file path + size, command + output, test exit code, reviewer verdict. "Looks good" is forbidden. This is the structural antibody against sycophantic self-grading — the agent literally cannot mark done without producing checkable proof.

7. Backpressure against premature completion. Before declaring the run done, the agent must produce a credible adversarial review of its own work — a challenger task whose acceptance is "find ≥3 weaknesses, or explicitly assert none with rationale." The completion counter only advances on a clean pass. This is the antidote to one of the most common long-running-agent failure modes: enthusiastic early stopping.

8. Drift fence over silent corruption. If world_assumptions claim the working directory is X and pwd returns Y, or claim a file's hash is A and it's now B, the run halts rather than write more on top of a stale base. Better to stop and ask a human than to compound silent error across hours of work.

9. Minimum viable substrate. Pure Markdown for the contract, pure JSON for the state, pure text for the log. No SDK. No framework. No installer. The system that runs nowhere runs everywhere. The price of universality is giving up vendor-specific superpowers — accepted deliberately.


Architecture

┌─ L2 — vendor accelerators (delete-safe) ─────────────────────┐
│ Claude Skill, Gemini extension, Codex AGENTS.md,             │
│ Cursor rules, ChatGPT system prompt                          │
└──────────────────────────┬───────────────────────────────────┘
                           │
┌──────────────────────────▼───────────────────────────────────┐
│ L1 — automation envelope (delete-safe)                       │
│ POSIX shell + PowerShell, ~80 lines, zero deps               │
│ Optional: per-turn timeout, retry on transient failures      │
└──────────────────────────┬───────────────────────────────────┘
                           │
┌──────────────────────────▼───────────────────────────────────┐
│ L0 — protocol core (required)                                │
│ PROTOCOL.md + state.json + log.md                            │
│ Pure files. Any CLI that can read and write works.           │
└──────────────────────────────────────────────────────────────┘

Each outer layer is optional. Delete L2 and the system still runs. Delete L1 too and you type continue between turns manually. L0 is the entire protocol — about 400 lines of plain Markdown.


Quality controls

Tunable per run in state.json. Defaults match a quick, light-protection run.

Knob Default What it does
mode "standard" "lightweight" for one-shot tasks; "heavy" for overnight quality runs
max_turns 60 Hard turn-count cap
contest_rounds_required 1 Clean adversarial-review rounds required before done
plan_variants 1 At INIT: generate N candidate plans, judge picks winner (Plan-BoN)
plan_rounds_required 0 Plan-Challenger reviews the chosen strategy before any work begins
per-task variants unset Best-of-N on any task; judge selects winner

Three ready-to-use recipes:

// Fast (~10–20 min, light protection)
{ "mode": "standard", "max_turns": 15, "contest_rounds_required": 0 }
// Standard (~1–2 h, default quality)
{ "mode": "standard", "max_turns": 30, "contest_rounds_required": 1 }
// Heavy (~6–8 h, overnight quality)
{
  "mode": "heavy",
  "max_turns": 150,
  "plan_variants": 2,
  "plan_rounds_required": 1,
  "contest_rounds_required": 2
}

Heavy mode mandates Best-of-N on every drafting task — the protocol enforces this syntactically so quality is not left as an optional recommendation.


What's in the box

nano-longrun-skill/
├── README.md             You are here
├── LICENSE               MIT
├── CHANGELOG.md
└── skill/                The actual package — drop this anywhere
    ├── PROTOCOL.md       Agent contract (~400 lines)
    ├── state.json        Source of truth at runtime
    ├── log.md            Append-only audit
    ├── README.md         Package-level orientation
    ├── L1/               Optional auto-loop (POSIX + PowerShell)
    └── adapters/         Optional vendor glue (Claude / Gemini / Codex / Cursor / ChatGPT)

The skill/ directory is intentionally self-contained. Copy it into any project, treat it like a portable agent runbook.


Non-goals

  • No parallelism — single agent, single state at a time. Best-of-N is serial.
  • No automatic mid-run replan — Plan-Challenger runs once at INIT, not iteratively at runtime.
  • No token-level cost cap — only turn and wall-clock caps. A bloated context within the turn cap still spends.
  • Not a runtime, a contract — the agent must honor the protocol; there is no enforcement daemon.

License

MIT — see LICENSE.

About

A portable protocol for long-running LLM agent tasks. Three files, zero dependencies, works wherever an agent can read and write files.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages