Scaffolding for a way of working in which you own the requirements and the decisions, and agents do the work between them: a planner, a handful of implementers, a research lane that builds and maintains a wiki, two reviewers on different clocks, and one interactive agent that is where you sit. Every Claude instance runs in a container. Coordination goes through a tuple space. Everything a human needs to decide arrives as a file on disk and waits there until you decide it.
The host needs docker, bash, curl and jq. Everything else lives in the image.
Once. Put loop on your PATH. It works on the repository you run it from, the way git
does, so you will be running it from your projects rather than from here:
ln -s "$PWD/bin/loop" ~/.local/bin/loop # or anywhere else on your PATH
Then the token. Interactive — it opens a browser, and you paste the token back:
loop token
That token goes to planner, worker, researcher, summariser and the reviewers — the
unattended agents, which have no way to do an interactive login. interactive (what loop shell gives you) deliberately does not get it: a setup-token token is inference-only, and
Remote Control refuses those. So the first claude you run inside loop shell does a fresh
claude auth login each session — the cost of a full-scope token Remote Control will accept.
Every session:
cd ~/my-project
loop up # tuple space, 4 workers, researcher, summariser — nothing that plans yet
loop shell # ← you live here
loop start # planner + both reviewers, and a planning pass now: when the PRDs are ready
loop stop # planner and reviewers off; workers finish what is queued
loop down
up creates loop/, raw/ and wiki/ in the project and seeds wiki/CLAUDE.md. The project
is the git repository you run it from — its root, from anywhere inside it — because the agents
commit.
More than one project. Each repository gets a loop of its own — containers, tuple space,
serviced — so loop up in two repositories runs two, side by side. loop ls lists them, and
loop -C <dir|name> <command> reaches one from anywhere else. They share the credential, and so
its rate limits (loop up -w 2 is a kind start for a second one), the catalogue services, and
shared/.
The rhythm. Inside loop shell, three things and nothing else:
- Write and argue about PRDs, then
loop startwhen you are happy with them. The most expensive documents in the repository: a bad requirement gets implemented faithfully across a lot of code by agents that will not push back. Take the data model seriously — it decides which workflows are easy and which changes later become awkward. Nothing is built until you start the planner —loop start, fromloop shellor the host — which runs a planning pass there and then, and after that wakes on its own whenever a result, an approved finding or an edited PRD gives it something to do. - Ask research questions —
research "how do mining companies buy occ-health?"— or drop a source straight intoraw/. The summariser hashes that directory every minute and ingests whatever appears. Nothing needs telling. - Triage the inbox. Ask it to go through
loop/feedback/inbox/. Approve → the planner must act on it. Archive with a reason → it stops coming back. This is the only path from a reviewer to the planner, and it runs through you.
Outside, between checkpoints:
loop status # what is running, queue, in flight, backlog, inbox — start here
loop logs planner -f
loop pause / resume # freeze the fleet without stopping containers
The same loop works inside loop shell, so the interactive agent can answer "is the planner
running?" or "why has nothing been reviewed?" from loop status and loop logs, and act on
"stop the planner" itself.
Read loop/PLAN-SUMMARY.md for what the planner queued and why. When an agent does something
dumb, git revert — that is the design, not a workaround.
When something is off:
| An agent keeps doing the wrong thing | edit prompts/*.md — that is where the agents actually live |
| Commits warn about a missing lock | loop reseed |
| Tuple space restarted, queue looks empty | loop requeue |
| An agent asked for a service you do not have | it is in your inbox; add services/<name>/, then loop service start <name> |
backlog … is not shrinking in the summariser log |
a source is defeating it — go and look at it |
| The wiki looks tangled | loop wiki |
The one rule. Everything an agent wants you to decide arrives as a file and waits. Nothing
an agent thinks changes the plan until you move it to approved/. If you find yourself needed
for every turn, something is misconfigured — that is the failure mode this exists to avoid.
Two lanes, joined at the top by you.
RESEARCH LANE BUILD LANE
you: research "<question>" you: PRDs/
│ │
▼ ▼
loop/research/ ──►┌────────────┐ ┌─────────┐ task tuples ┌──────────┐
▲ │ researcher │ │ planner ├──────────────►│ worker ×4│
│ └─────┬──────┘ └─────────┘ └────┬─────┘
│ │ writes ▲ │ commits
│ ▼ │ loop/results/ ▼
│ you ──► raw/ ── immutable └──────────────── the repository
│ ╎ hashed every minute │
│ ╎ hash changed ⇒ run ┌────────┴────────┐
│ ┌─────▼──────┐ 10 passes ┌───────┴─────┐ ┌─────────┴────┐
│ │ summariser │ + a lint │ review-diff │ │ review-drift │
│ └─────┬──────┘ │ 15 min │ │ 4 hours │
│ ▼ └───────┬─────┘ └─────────┬────┘
│ wiki/ ── index.md, log.md, │ │
│ │ CLAUDE.md │ │
└──────────────────┘ gaps become briefs └──► loop/feedback/inbox/ ◄──┘
│
you triage it ────┘
Two arrows matter, and one of them is missing. The dotted one is not an arrow at all — nothing tells the summariser anything; it looks.
No reviewer writes to the planner. Findings land in loop/feedback/inbox/, you read them,
and only what you move to approved/ is binding on the next plan. That is the checkpoint;
everything else runs between checkpoints without you.
Research summarises itself, and nothing has to ring a bell. The summariser hashes raw/
once a minute; when the hash changes it runs. So a researcher finishing its brief triggers a
run, and so does you dragging a clipped article in from a browser — neither knows the
summariser exists. The summariser then writes briefs back into the research queue where it
finds gaps, so the lane feeds itself and you top it up with questions rather than driving it.
| Agent | Clock | Model | Owns | Never touches |
|---|---|---|---|---|
interactive |
you | Opus | PRDs/, RESEARCH-TARGETS.md, the inbox |
wiki/ (except its schema) |
planner |
from loop start: looks every ~1 min, runs when something landed |
Opus | IMPLEMENTATION_PLAN.md, loop/tasks/ |
code, PRDs |
worker ×N |
per task | Sonnet | code and tests | PRDs, the plan |
researcher ×N |
per brief | Sonnet | raw/ (append only) |
wiki/, code |
summariser ×1 |
polls raw/, 1 min |
Sonnet; Opus for the lint | wiki/, all of it |
raw/, code |
review-diff |
from loop start: looks every 15 min, runs when ~10 commits landed |
Opus | loop/reviews/diff/ |
everything else |
review-drift |
from loop start: looks every 4 hours, same rule |
Opus | loop/reviews/drift/ |
everything else |
"Never touches" is what the prompts say, not what the containers enforce — every agent
mounts the whole project — with two exceptions: raw/ is hidden from the build lane, and
no agent has memory (CLAUDE_CODE_DISABLE_AUTO_MEMORY), so each run starts from its prompt
and what is on disk.
The prompts in prompts/ are the real definitions and the thing to edit when an agent does
something dumb. Intervals are REVIEW_DIFF_INTERVAL / REVIEW_DRIFT_INTERVAL in the
environment, seconds.
Models follow the cost of a mistake. Opus where one fans out or lands in your inbox — the
planner's briefs are implemented by every worker at once, and a noisy reviewer spends your
attention. Sonnet where the volume is and a mistake is cheap to catch — a worker has a brief,
tests, a reviewer and git revert behind it. Override per agent from the environment:
PLANNER_MODEL, WORKER_MODEL, RESEARCHER_MODEL, SUMMARY_MODEL, SUMMARY_LINT_MODEL,
REVIEW_DIFF_MODEL, REVIEW_DRIFT_MODEL, INTERACTIVE_MODEL — e.g. WORKER_MODEL=claude-opus-5 loop up. If worker results start coming back FAILED, or review-diff keeps finding work that
went somewhere other than where it was aimed, make the planner's briefs more explicit before
you give the workers a bigger model.
Started, not implied. loop up starts the agents that wait for work — workers,
researchers, the summariser — and not the three on a clock. An agent on a timer that usually
finds nothing to do looks, from outside, exactly like one that is broken, so whether they run
at all is something you say: loop start starts the planner and both reviewers and queues a
planning pass, so the planner's first look always does something; loop stop stops them and
lets the workers finish what is queued. loop status says which state you are in.
Looking is not running. Once started, the three agents on a clock ask a shell script whether the cycle is worth an agent before they start one, and on a quiet repository the answer is usually no. This is the difference between a fleet that costs money while you are asleep and one that does not: the planner alone was fourteen hundred sessions a day to discover, nearly every time, that nothing had happened since the last one.
plan-gateruns the planner when a worker has left a result it has not archived, when you have approved something inloop/feedback/inbox/, when the PRDs change, when a service it asked for becomes available, or when you callreplan. The first three clear themselves: the planner archives what it has accounted for, so "still there" means "not yet read".review-gateruns a reviewer when ten commits have landed since its last high watermark, or when any have and it has been a day — the threshold stops it chasing every commit, the day stops a slow trickle going unreviewed forever. Commits that only touchloop/,wiki/orraw/do not count: a reviewer woken by the summariser's own bookkeeping has nothing to review but the machinery that woke it. Tune withREVIEW_MIN_COMMITS,REVIEW_MAX_AGE_HOURSandREVIEW_IGNORE_PATHS.
Both explain every decision in the container's log, so loop logs review-diff tells you why
it has not run as well as why it has. A gate that is itself broken exits non-zero and the
agent runs anyway: a reviewer that reviews too often is a visible cost, one silently switched
off by a typo in a path is not.
review-diff reads loop/baseline (a commit sha), reviews baseline..HEAD against the
PRDs, the plan and the task briefs, writes a report, moves the baseline. It catches work that
went somewhere other than where it was aimed. That baseline is also what its gate counts from,
which is why the prompt insists on moving it even when the review found nothing.
review-drift takes one PRD area at a time and asks whether the document and the code still
describe the same product. Its output is a recommendation about which side should move —
fix-prd, fix-code, or both — because after a few weeks the code is often the one that
learned something.
Three layers, following karpathy's LLM-wiki pattern:
raw/— sources. Immutable: agents append, and never edit or delete what is already there. A knowledge base that rewrites its own sources cannot be audited, and a contradiction between two sources is content for the wiki rather than a correction to make here. Not in git, and not visible to the build lane. The wiki built from it is committed; the sources stay on disk (back them up if they matter). The planner, workers and reviewers get an empty, read-onlyraw/mounted over the real one, so web pages a researcher fetched never reach an agent that writes code. Only researchers, the summariser andloop shellsee it.wiki/— pages, owned entirely by the summariser.index.mdis the catalog you navigate by;log.mdis the append-only chronological record, with entries prefixed## [YYYY-MM-DD] ingest | …sogrep '^## \[' wiki/log.md | tailis a timeline.wiki/CLAUDE.md— the schema. Page types, naming, linking and the ingest workflow. It is what makes the summariser a disciplined maintainer rather than a generic summariser, and it is the one file inwiki/you should edit yourself. Both of you co-evolve it.
The departure from that pattern is that ingest is automatic. There, you drop in a source
and tell the LLM to process it, one at a time, staying involved. Here nothing tells it
anything: the summariser hashes every file in raw/, hashes the hashes, and compares that
watermark once a minute. A changed watermark means a run. The trade — supervision for
throughput — is paid back through the feedback inbox rather than at every ingest.
Polling rather than a notification is what makes anything a valid producer. A research
agent, you from loop shell, cp from your desktop, a browser extension writing straight into
the directory — none of them has to know the summariser exists, or that there is a tuple space
to reach. The directory is the whole contract.
The watermark is written before a run, not after, so sources that land while the summariser is working change the hash again and earn their own run on the next cycle. And because ten passes is a budget rather than a promise to finish, a run that drains only part of a large backlog comes straight back for the rest — so everything is eventually summarised. If a run ever finishes without the backlog shrinking, it stops and says so once, rather than retrying a source that defeats it every minute forever.
The summariser always runs ten passes and then a lint, however much is waiting. A fixed
budget is what stops one enormous source from starving everything else and what makes a run's
cost predictable; whatever is left is still in the backlog and the next cycle takes it. The
lint is last because contradictions, stale claims and orphans only become visible once a batch
has landed and the wiki is sitting still. wiki-lint does the mechanical half — broken links,
orphans, pages missing from the index, sources nothing cites — and the agent spends its
attention on the half a script cannot check.
Gaps close the loop: when the summariser finds something it cannot write because nobody gathered the material, it writes a research brief, and a researcher goes and answers it.
From loop shell (or on the host as loop research …):
research "How do mining companies procure occupational health services?"
research --slug procurement-chain <<'EOF' # when the question needs real context
# How is occ-health actually bought?
- why: unblocks the [[Procurement]] page
- what would answer it: two or three real tender documents
EOF
research --list |
waiting, in flight, done |
summarise |
ask for a run now, without having added anything |
raw-backlog |
sources in raw/ the wiki has not ingested |
wiki-lint |
the mechanical checks, on demand |
You do not need summarise to get something ingested — just put the file in raw/. Use it
when you want a run without having added anything: a lint pass over a wiki nothing has changed
recently, or another attempt at a backlog that stopped shrinking.
Two derived facts, no bookkeeping files: the watermark says whether to run, and the backlog
says what to do. The backlog is computed from log.md — a source is ingested when the log
names its path — which is why the summarise prompt insists on naming every path it consumed.
One record, nothing to fall out of sync with it.
In this repository — the harness:
bin/loop the only command you run
bin/setup-token.sh mints the credential
bin/serviced.sh stands up services agents ask for, and runs `loop` for the containers
(host-side; it starts containers)
agent/ scripts mounted into every container at /loop/bin
claim-loop.sh the worker and researcher loop (one script, two lanes)
summariser.sh the wiki maintainer's loop
plan-gate, review-gate whether a timed cycle is worth an agent
loop `loop` inside a container — Docker commands via serviced
research, summarise, replan, review-now, wiki-lint, raw-backlog
— yours to run
prompts/ what each agent is
services/ the catalogue of backing services agents may have
secrets/ the token, gitignored
In your project — created by loop up, and mostly committed:
PRDs/ yours
RESEARCH-TARGETS.md what we still need to learn — yours
raw/ sources, immutable, appended by researchers and by you — gitignored
wiki/ the summariser's: pages, index.md, log.md, CLAUDE.md
IMPLEMENTATION_PLAN.md the planner's
loop/PLAN-SUMMARY.md what the planner queued and why — read this first
loop/tasks/ briefs waiting to be claimed
loop/research/ research briefs waiting, and their results
loop/active/ in flight right now
loop/results/ what workers finished or failed at
loop/reviews/ full review reports
loop/feedback/ inbox → approved → archive
loop/services/ what is running, with connection details
loop/STOP exists ⇒ every agent idles
Six tuple shapes, and that is the whole protocol:
| Tuple | Who puts | Who claims |
|---|---|---|
("task", slug, title) |
queue-briefs task, after each planner pass |
a worker |
("research", slug, title) |
research, or the summariser finding a gap |
a researcher |
("lock", "git") |
seeded by loop up |
commit, for the duration of one commit |
("service-request", name, who, why) |
request-service |
serviced, on the host |
("loop-request", id, who, argv) |
loop inside a container |
serviced, which runs bin/loop |
("loop-reply", id, status, output) |
serviced |
the loop that asked |
The rule behind it: the tuple space carries handoffs; disk carries state. A tuple exists
only long enough for one agent to pass something to another. Anything you might want to read
later — briefs, results, reviews, feedback, sources, pages — is a file. That is why the space
being in memory is fine: a restart loses the handoffs, loop requeue puts them back from the
files, and nothing you would have wanted to read was in there anyway.
Summarisation takes this furthest and uses no tuple at all. The summariser polls a hash of
raw/ and recomputes its backlog from log.md, so there is nothing to lose, nothing to
debounce, and nothing that has to know how to reach the space in order to contribute a source.
Its one use of the tuple space is the other direction: putting research briefs into it.
The git lock exists because every agent shares one working tree, and two git commits at once
fail on index.lock for reasons that look like nothing. Agents are told to use commit "msg"
instead, which takes the lock. If commits start warning about a missing lock, loop reseed.
For the same reason commit stages only what its caller owns, never git add -A: in a shared
tree that would put every other agent's half-finished edits into your commit under your name.
What an agent owns is COMMIT_PATHS — set per task from the brief's touches: by
claim-loop.sh, and per agent in docker-compose.yml. loop shell sets none, so a commit
from there still takes the whole tree.
Four workers edit the same checkout. That is how v0.1 worked and it is kept deliberately:
worktrees per agent would mean a merge step, and a merge step means conflict resolution, and
that is a whole machine to maintain for a problem the planner can avoid by not queueing two
tasks that touch the same files. Every brief carries a touches: list and the planner is told
to check new briefs against it. When two agents do collide, you revert — which the article is
right that agents do not mind.
An agent that needs Postgres runs request-service postgres "PRD-03 stores assessments".
serviced sees the request, and:
- in
services/→ starts it onloopnetand copies itsservice.mdintoloop/services/available/, where agents look for connection details. Seconds. - not in
services/→ nothing starts, and the request becomes a feedback item for you.
That asymmetry is the point: the catalogue is the list of what an agent may have, and anything
outside it is a decision. services/README.md says how to add one. The planner is told to
work out its service needs and ask for all of them before queueing implementation tasks, so
this happens while you are looking rather than when a worker is halfway through something.
The catalogue is shared by every loop on the machine: one Postgres, one MinIO, reached by the
same hostname from each. Projects keep their data apart inside them — a database, a bucket, a
key prefix each — and each service.md says how.
Nothing publishes a port to your host. loop forward postgres 5432 when you want a client.
loop up -w 6 -r 2 |
six workers, two researchers |
loop start ["why"] |
start the planner and both reviewers, with a planning pass now — this is what starts the work |
loop stop |
stop them; workers finish what is queued |
loop plan ["why"] |
plumbing: queue a pass for a planner that is already running |
loop research "<q>" |
queue a research question from the host |
loop review diff / drift |
ask for a review now, ahead of what the gate would decide |
loop summarize |
ask for a summarisation run now |
loop wiki |
run wiki-lint |
loop logs planner -f |
watch one agent |
loop pause / loop resume |
freeze the fleet without stopping containers |
loop status |
queue depth, work in flight, unread results, inbox |
loop service list |
what is in the catalogue and what is running |
loop down --all |
stop the services too (they survive a plain down) — for every loop |
loop ls |
every project with a loop, which are up, and their tuple space ports |
loop -C <dir|name> … |
any command, for another project than the one you are in |
Start it by hand the first time round: loop up, then loop shell, and get one PRD to a
state you are happy with before you loop start. The planner will fill a queue from a bad PRD
just as enthusiastically as from a good one — which is the reason it is something you start
rather than something up implies.
loop is there too, for this project, and it is what lets the interactive agent answer for
the loop and steer it. The commands that are only files — plan, research, pause — run in
the container as the same scripts. The ones that need Docker — status, ps, logs,
start, stop, up -w/-r, requeue — go through the tuple space to this loop's serviced,
which runs bin/loop on the host and passes back what it printed. So the container never
holds the Docker socket, and there is one implementation of each command rather than two.
serviced decides what it will run, not the container: anything in the loop can put a
request. down and restart are host-only (they would take down the container asking),
as are logs -f (it never answers), -C, ls and starting or stopping catalogue services
(they reach past this project). If serviced is not running, loop in the container says so
after thirty seconds and takes its request back, so nothing runs later by surprise.
Point it at its own repository and it works:
cd strangeloop && loop up
The harness is then just a project: the planner plans against PRDs you write for it, workers
edit agent/ and prompts/, and the reviewers read the diffs. loop up says as much when it
notices, and spells out what is live and what is not.
The thing that makes this safe is that running containers keep the code they started with.
agent/ and bin/ are read when a process starts, so an agent rewriting claim-loop.sh does
not reach into the loop that is currently executing it — the change waits for loop restart,
which is a decision you make after reading the diff. Self-modification with an explicit
adoption step, rather than a system editing itself out from under its own feet.
| What an agent changed | When it takes effect |
|---|---|
prompts/*.md |
next iteration — they are re-read every pass |
agent/, bin/ |
loop restart |
Dockerfile, docker-compose.yml |
loop restart (it rebuilds) |
services/ |
loop service start <name> |
Two things to know. The project is always a whole repository. Only the project directory is
mounted, so a project nested inside a repo would arrive without its .git and every commit
would fail. loop therefore takes the repository root wherever in it you run, and loop up
refuses a git worktree or a submodule, whose .git points outside the directory. And secrets/ is in the agents' working tree when self-hosting. It is
gitignored so nothing can commit it, and the token is already in their environment, so this
grants nothing new — but it is worth knowing rather than discovering.
No pull requests, no approval before a commit, no per-agent branches. The review happens
after the work lands, and the remedy for bad work is git revert. No retries, no automatic
lease renewal beyond the worker's own heartbeat, no persistence in the tuple space. No metrics
beyond loop status and ts stats. No search over the wiki — index.md is enough at this
scale, and when it stops being enough the answer is a tool like qmd, not an index rewrite. Each of those is a thing that would have to be maintained,
and none of them is what makes this work — the prompts are.