Skip to content

Latest commit

 

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

strangeloop

Scaffolding for a way of working in which you own the requirements and the decisions, and agents do the work between them: a planner, a handful of implementers, a research lane that builds and maintains a wiki, two reviewers on different clocks, and one interactive agent that is where you sit. Every Claude instance runs in a container. Coordination goes through a tuple space. Everything a human needs to decide arrives as a file on disk and waits there until you decide it.

The host needs docker, bash, curl and jq. Everything else lives in the image.

TL;DR

Once. Put loop on your PATH. It works on the repository you run it from, the way git does, so you will be running it from your projects rather than from here:

ln -s "$PWD/bin/loop" ~/.local/bin/loop     # or anywhere else on your PATH

Then the token. Interactive — it opens a browser, and you paste the token back:

loop token

That token goes to planner, worker, researcher, summariser and the reviewers — the unattended agents, which have no way to do an interactive login. interactive (what loop shell gives you) deliberately does not get it: a setup-token token is inference-only, and Remote Control refuses those. So the first claude you run inside loop shell does a fresh claude auth login each session — the cost of a full-scope token Remote Control will accept.

Every session:

cd ~/my-project
loop up        # tuple space, 4 workers, researcher, summariser — nothing that plans yet
loop shell     # ← you live here
loop start     # planner + both reviewers, and a planning pass now: when the PRDs are ready
loop stop      # planner and reviewers off; workers finish what is queued
loop down

up creates loop/, raw/ and wiki/ in the project and seeds wiki/CLAUDE.md. The project is the git repository you run it from — its root, from anywhere inside it — because the agents commit.

More than one project. Each repository gets a loop of its own — containers, tuple space, serviced — so loop up in two repositories runs two, side by side. loop ls lists them, and loop -C <dir|name> <command> reaches one from anywhere else. They share the credential, and so its rate limits (loop up -w 2 is a kind start for a second one), the catalogue services, and shared/.

The rhythm. Inside loop shell, three things and nothing else:

  1. Write and argue about PRDs, then loop start when you are happy with them. The most expensive documents in the repository: a bad requirement gets implemented faithfully across a lot of code by agents that will not push back. Take the data model seriously — it decides which workflows are easy and which changes later become awkward. Nothing is built until you start the planner — loop start, from loop shell or the host — which runs a planning pass there and then, and after that wakes on its own whenever a result, an approved finding or an edited PRD gives it something to do.
  2. Ask research questions — research "how do mining companies buy occ-health?" — or drop a source straight into raw/. The summariser hashes that directory every minute and ingests whatever appears. Nothing needs telling.
  3. Triage the inbox. Ask it to go through loop/feedback/inbox/. Approve → the planner must act on it. Archive with a reason → it stops coming back. This is the only path from a reviewer to the planner, and it runs through you.

Outside, between checkpoints:

loop status             # what is running, queue, in flight, backlog, inbox — start here
loop logs planner -f
loop pause / resume     # freeze the fleet without stopping containers

The same loop works inside loop shell, so the interactive agent can answer "is the planner running?" or "why has nothing been reviewed?" from loop status and loop logs, and act on "stop the planner" itself.

Read loop/PLAN-SUMMARY.md for what the planner queued and why. When an agent does something dumb, git revert — that is the design, not a workaround.

When something is off:

An agent keeps doing the wrong thing edit prompts/*.md — that is where the agents actually live
Commits warn about a missing lock loop reseed
Tuple space restarted, queue looks empty loop requeue
An agent asked for a service you do not have it is in your inbox; add services/<name>/, then loop service start <name>
backlog … is not shrinking in the summariser log a source is defeating it — go and look at it
The wiki looks tangled loop wiki

The one rule. Everything an agent wants you to decide arrives as a file and waits. Nothing an agent thinks changes the plan until you move it to approved/. If you find yourself needed for every turn, something is misconfigured — that is the failure mode this exists to avoid.


The shape of it

Two lanes, joined at the top by you.

  RESEARCH LANE                              BUILD LANE

  you: research "<question>"                 you: PRDs/
          │                                          │
          ▼                                          ▼
    loop/research/  ──►┌────────────┐          ┌─────────┐  task tuples  ┌──────────┐
          ▲            │ researcher │          │ planner ├──────────────►│ worker ×4│
          │            └─────┬──────┘          └─────────┘               └────┬─────┘
          │                  │ writes               ▲                         │ commits
          │                  ▼                      │ loop/results/           ▼
          │   you ──►     raw/  ── immutable        └──────────────── the repository
          │                  ╎ hashed every minute                          │
          │                  ╎ hash changed ⇒ run                   ┌────────┴────────┐
          │            ┌─────▼──────┐  10 passes            ┌───────┴─────┐  ┌─────────┴────┐
          │            │ summariser │  + a lint             │ review-diff │  │ review-drift │
          │            └─────┬──────┘                       │ 15 min      │  │ 4 hours      │
          │                  ▼                              └───────┬─────┘  └─────────┬────┘
          │               wiki/  ── index.md, log.md,               │                  │
          │                  │      CLAUDE.md                       │                  │
          └──────────────────┘ gaps become briefs                   └──► loop/feedback/inbox/ ◄──┘
                                                                                │
                                                              you triage it ────┘

Two arrows matter, and one of them is missing. The dotted one is not an arrow at all — nothing tells the summariser anything; it looks.

No reviewer writes to the planner. Findings land in loop/feedback/inbox/, you read them, and only what you move to approved/ is binding on the next plan. That is the checkpoint; everything else runs between checkpoints without you.

Research summarises itself, and nothing has to ring a bell. The summariser hashes raw/ once a minute; when the hash changes it runs. So a researcher finishing its brief triggers a run, and so does you dragging a clipped article in from a browser — neither knows the summariser exists. The summariser then writes briefs back into the research queue where it finds gaps, so the lane feeds itself and you top it up with questions rather than driving it.

The agents

Agent Clock Model Owns Never touches
interactive you Opus PRDs/, RESEARCH-TARGETS.md, the inbox wiki/ (except its schema)
planner from loop start: looks every ~1 min, runs when something landed Opus IMPLEMENTATION_PLAN.md, loop/tasks/ code, PRDs
worker ×N per task Sonnet code and tests PRDs, the plan
researcher ×N per brief Sonnet raw/ (append only) wiki/, code
summariser ×1 polls raw/, 1 min Sonnet; Opus for the lint wiki/, all of it raw/, code
review-diff from loop start: looks every 15 min, runs when ~10 commits landed Opus loop/reviews/diff/ everything else
review-drift from loop start: looks every 4 hours, same rule Opus loop/reviews/drift/ everything else

"Never touches" is what the prompts say, not what the containers enforce — every agent mounts the whole project — with two exceptions: raw/ is hidden from the build lane, and no agent has memory (CLAUDE_CODE_DISABLE_AUTO_MEMORY), so each run starts from its prompt and what is on disk.

The prompts in prompts/ are the real definitions and the thing to edit when an agent does something dumb. Intervals are REVIEW_DIFF_INTERVAL / REVIEW_DRIFT_INTERVAL in the environment, seconds.

Models follow the cost of a mistake. Opus where one fans out or lands in your inbox — the planner's briefs are implemented by every worker at once, and a noisy reviewer spends your attention. Sonnet where the volume is and a mistake is cheap to catch — a worker has a brief, tests, a reviewer and git revert behind it. Override per agent from the environment: PLANNER_MODEL, WORKER_MODEL, RESEARCHER_MODEL, SUMMARY_MODEL, SUMMARY_LINT_MODEL, REVIEW_DIFF_MODEL, REVIEW_DRIFT_MODEL, INTERACTIVE_MODEL — e.g. WORKER_MODEL=claude-opus-5 loop up. If worker results start coming back FAILED, or review-diff keeps finding work that went somewhere other than where it was aimed, make the planner's briefs more explicit before you give the workers a bigger model.

Started, not implied. loop up starts the agents that wait for work — workers, researchers, the summariser — and not the three on a clock. An agent on a timer that usually finds nothing to do looks, from outside, exactly like one that is broken, so whether they run at all is something you say: loop start starts the planner and both reviewers and queues a planning pass, so the planner's first look always does something; loop stop stops them and lets the workers finish what is queued. loop status says which state you are in.

Looking is not running. Once started, the three agents on a clock ask a shell script whether the cycle is worth an agent before they start one, and on a quiet repository the answer is usually no. This is the difference between a fleet that costs money while you are asleep and one that does not: the planner alone was fourteen hundred sessions a day to discover, nearly every time, that nothing had happened since the last one.

  • plan-gate runs the planner when a worker has left a result it has not archived, when you have approved something in loop/feedback/inbox/, when the PRDs change, when a service it asked for becomes available, or when you call replan. The first three clear themselves: the planner archives what it has accounted for, so "still there" means "not yet read".
  • review-gate runs a reviewer when ten commits have landed since its last high watermark, or when any have and it has been a day — the threshold stops it chasing every commit, the day stops a slow trickle going unreviewed forever. Commits that only touch loop/, wiki/ or raw/ do not count: a reviewer woken by the summariser's own bookkeeping has nothing to review but the machinery that woke it. Tune with REVIEW_MIN_COMMITS, REVIEW_MAX_AGE_HOURS and REVIEW_IGNORE_PATHS.

Both explain every decision in the container's log, so loop logs review-diff tells you why it has not run as well as why it has. A gate that is itself broken exits non-zero and the agent runs anyway: a reviewer that reviews too often is a visible cost, one silently switched off by a typo in a path is not.

review-diff reads loop/baseline (a commit sha), reviews baseline..HEAD against the PRDs, the plan and the task briefs, writes a report, moves the baseline. It catches work that went somewhere other than where it was aimed. That baseline is also what its gate counts from, which is why the prompt insists on moving it even when the review found nothing.

review-drift takes one PRD area at a time and asks whether the document and the code still describe the same product. Its output is a recommendation about which side should move — fix-prd, fix-code, or both — because after a few weeks the code is often the one that learned something.

The research lane

Three layers, following karpathy's LLM-wiki pattern:

  • raw/ — sources. Immutable: agents append, and never edit or delete what is already there. A knowledge base that rewrites its own sources cannot be audited, and a contradiction between two sources is content for the wiki rather than a correction to make here. Not in git, and not visible to the build lane. The wiki built from it is committed; the sources stay on disk (back them up if they matter). The planner, workers and reviewers get an empty, read-only raw/ mounted over the real one, so web pages a researcher fetched never reach an agent that writes code. Only researchers, the summariser and loop shell see it.
  • wiki/ — pages, owned entirely by the summariser. index.md is the catalog you navigate by; log.md is the append-only chronological record, with entries prefixed ## [YYYY-MM-DD] ingest | … so grep '^## \[' wiki/log.md | tail is a timeline.
  • wiki/CLAUDE.md — the schema. Page types, naming, linking and the ingest workflow. It is what makes the summariser a disciplined maintainer rather than a generic summariser, and it is the one file in wiki/ you should edit yourself. Both of you co-evolve it.

The departure from that pattern is that ingest is automatic. There, you drop in a source and tell the LLM to process it, one at a time, staying involved. Here nothing tells it anything: the summariser hashes every file in raw/, hashes the hashes, and compares that watermark once a minute. A changed watermark means a run. The trade — supervision for throughput — is paid back through the feedback inbox rather than at every ingest.

Polling rather than a notification is what makes anything a valid producer. A research agent, you from loop shell, cp from your desktop, a browser extension writing straight into the directory — none of them has to know the summariser exists, or that there is a tuple space to reach. The directory is the whole contract.

The watermark is written before a run, not after, so sources that land while the summariser is working change the hash again and earn their own run on the next cycle. And because ten passes is a budget rather than a promise to finish, a run that drains only part of a large backlog comes straight back for the rest — so everything is eventually summarised. If a run ever finishes without the backlog shrinking, it stops and says so once, rather than retrying a source that defeats it every minute forever.

The summariser always runs ten passes and then a lint, however much is waiting. A fixed budget is what stops one enormous source from starving everything else and what makes a run's cost predictable; whatever is left is still in the backlog and the next cycle takes it. The lint is last because contradictions, stale claims and orphans only become visible once a batch has landed and the wiki is sitting still. wiki-lint does the mechanical half — broken links, orphans, pages missing from the index, sources nothing cites — and the agent spends its attention on the half a script cannot check.

Gaps close the loop: when the summariser finds something it cannot write because nobody gathered the material, it writes a research brief, and a researcher goes and answers it.

Starting it

From loop shell (or on the host as loop research …):

research "How do mining companies procure occupational health services?"

research --slug procurement-chain <<'EOF'      # when the question needs real context
# How is occ-health actually bought?
- why: unblocks the [[Procurement]] page
- what would answer it: two or three real tender documents
EOF
research --list waiting, in flight, done
summarise ask for a run now, without having added anything
raw-backlog sources in raw/ the wiki has not ingested
wiki-lint the mechanical checks, on demand

You do not need summarise to get something ingested — just put the file in raw/. Use it when you want a run without having added anything: a lint pass over a wiki nothing has changed recently, or another attempt at a backlog that stopped shrinking.

Two derived facts, no bookkeeping files: the watermark says whether to run, and the backlog says what to do. The backlog is computed from log.md — a source is ingested when the log names its path — which is why the summarise prompt insists on naming every path it consumed. One record, nothing to fall out of sync with it.

What lives where

In this repository — the harness:

bin/loop            the only command you run
bin/setup-token.sh  mints the credential
bin/serviced.sh     stands up services agents ask for, and runs `loop` for the containers
                    (host-side; it starts containers)
agent/              scripts mounted into every container at /loop/bin
                    claim-loop.sh  the worker and researcher loop (one script, two lanes)
                    summariser.sh  the wiki maintainer's loop
                    plan-gate, review-gate  whether a timed cycle is worth an agent
                    loop           `loop` inside a container — Docker commands via serviced
                    research, summarise, replan, review-now, wiki-lint, raw-backlog
                                   — yours to run
prompts/            what each agent is
services/           the catalogue of backing services agents may have
secrets/            the token, gitignored

In your project — created by loop up, and mostly committed:

PRDs/               yours
RESEARCH-TARGETS.md what we still need to learn — yours
raw/                sources, immutable, appended by researchers and by you — gitignored
wiki/               the summariser's: pages, index.md, log.md, CLAUDE.md
IMPLEMENTATION_PLAN.md   the planner's
loop/PLAN-SUMMARY.md     what the planner queued and why — read this first
loop/tasks/         briefs waiting to be claimed
loop/research/      research briefs waiting, and their results
loop/active/        in flight right now
loop/results/       what workers finished or failed at
loop/reviews/       full review reports
loop/feedback/      inbox → approved → archive
loop/services/      what is running, with connection details
loop/STOP           exists ⇒ every agent idles

Coordination

Six tuple shapes, and that is the whole protocol:

Tuple Who puts Who claims
("task", slug, title) queue-briefs task, after each planner pass a worker
("research", slug, title) research, or the summariser finding a gap a researcher
("lock", "git") seeded by loop up commit, for the duration of one commit
("service-request", name, who, why) request-service serviced, on the host
("loop-request", id, who, argv) loop inside a container serviced, which runs bin/loop
("loop-reply", id, status, output) serviced the loop that asked

The rule behind it: the tuple space carries handoffs; disk carries state. A tuple exists only long enough for one agent to pass something to another. Anything you might want to read later — briefs, results, reviews, feedback, sources, pages — is a file. That is why the space being in memory is fine: a restart loses the handoffs, loop requeue puts them back from the files, and nothing you would have wanted to read was in there anyway.

Summarisation takes this furthest and uses no tuple at all. The summariser polls a hash of raw/ and recomputes its backlog from log.md, so there is nothing to lose, nothing to debounce, and nothing that has to know how to reach the space in order to contribute a source. Its one use of the tuple space is the other direction: putting research briefs into it.

The git lock exists because every agent shares one working tree, and two git commits at once fail on index.lock for reasons that look like nothing. Agents are told to use commit "msg" instead, which takes the lock. If commits start warning about a missing lock, loop reseed.

For the same reason commit stages only what its caller owns, never git add -A: in a shared tree that would put every other agent's half-finished edits into your commit under your name. What an agent owns is COMMIT_PATHS — set per task from the brief's touches: by claim-loop.sh, and per agent in docker-compose.yml. loop shell sets none, so a commit from there still takes the whole tree.

One working tree, on purpose

Four workers edit the same checkout. That is how v0.1 worked and it is kept deliberately: worktrees per agent would mean a merge step, and a merge step means conflict resolution, and that is a whole machine to maintain for a problem the planner can avoid by not queueing two tasks that touch the same files. Every brief carries a touches: list and the planner is told to check new briefs against it. When two agents do collide, you revert — which the article is right that agents do not mind.

Services

An agent that needs Postgres runs request-service postgres "PRD-03 stores assessments". serviced sees the request, and:

  • in services/ → starts it on loopnet and copies its service.md into loop/services/available/, where agents look for connection details. Seconds.
  • not in services/ → nothing starts, and the request becomes a feedback item for you.

That asymmetry is the point: the catalogue is the list of what an agent may have, and anything outside it is a decision. services/README.md says how to add one. The planner is told to work out its service needs and ask for all of them before queueing implementation tasks, so this happens while you are looking rather than when a worker is halfway through something.

The catalogue is shared by every loop on the machine: one Postgres, one MinIO, reached by the same hostname from each. Projects keep their data apart inside them — a database, a bucket, a key prefix each — and each service.md says how.

Nothing publishes a port to your host. loop forward postgres 5432 when you want a client.

Running it

loop up -w 6 -r 2 six workers, two researchers
loop start ["why"] start the planner and both reviewers, with a planning pass now — this is what starts the work
loop stop stop them; workers finish what is queued
loop plan ["why"] plumbing: queue a pass for a planner that is already running
loop research "<q>" queue a research question from the host
loop review diff / drift ask for a review now, ahead of what the gate would decide
loop summarize ask for a summarisation run now
loop wiki run wiki-lint
loop logs planner -f watch one agent
loop pause / loop resume freeze the fleet without stopping containers
loop status queue depth, work in flight, unread results, inbox
loop service list what is in the catalogue and what is running
loop down --all stop the services too (they survive a plain down) — for every loop
loop ls every project with a loop, which are up, and their tuple space ports
loop -C <dir|name> … any command, for another project than the one you are in

Start it by hand the first time round: loop up, then loop shell, and get one PRD to a state you are happy with before you loop start. The planner will fill a queue from a bad PRD just as enthusiastically as from a good one — which is the reason it is something you start rather than something up implies.

From inside loop shell

loop is there too, for this project, and it is what lets the interactive agent answer for the loop and steer it. The commands that are only files — plan, research, pause — run in the container as the same scripts. The ones that need Docker — status, ps, logs, start, stop, up -w/-r, requeue — go through the tuple space to this loop's serviced, which runs bin/loop on the host and passes back what it printed. So the container never holds the Docker socket, and there is one implementation of each command rather than two.

serviced decides what it will run, not the container: anything in the loop can put a request. down and restart are host-only (they would take down the container asking), as are logs -f (it never answers), -C, ls and starting or stopping catalogue services (they reach past this project). If serviced is not running, loop in the container says so after thirty seconds and takes its request back, so nothing runs later by surprise.

Working on strangeloop with strangeloop

Point it at its own repository and it works:

cd strangeloop && loop up

The harness is then just a project: the planner plans against PRDs you write for it, workers edit agent/ and prompts/, and the reviewers read the diffs. loop up says as much when it notices, and spells out what is live and what is not.

The thing that makes this safe is that running containers keep the code they started with. agent/ and bin/ are read when a process starts, so an agent rewriting claim-loop.sh does not reach into the loop that is currently executing it — the change waits for loop restart, which is a decision you make after reading the diff. Self-modification with an explicit adoption step, rather than a system editing itself out from under its own feet.

What an agent changed When it takes effect
prompts/*.md next iteration — they are re-read every pass
agent/, bin/ loop restart
Dockerfile, docker-compose.yml loop restart (it rebuilds)
services/ loop service start <name>

Two things to know. The project is always a whole repository. Only the project directory is mounted, so a project nested inside a repo would arrive without its .git and every commit would fail. loop therefore takes the repository root wherever in it you run, and loop up refuses a git worktree or a submodule, whose .git points outside the directory. And secrets/ is in the agents' working tree when self-hosting. It is gitignored so nothing can commit it, and the token is already in their environment, so this grants nothing new — but it is worth knowing rather than discovering.

What is deliberately not here

No pull requests, no approval before a commit, no per-agent branches. The review happens after the work lands, and the remedy for bad work is git revert. No retries, no automatic lease renewal beyond the worker's own heartbeat, no persistence in the tuple space. No metrics beyond loop status and ts stats. No search over the wiki — index.md is enough at this scale, and when it stops being enough the answer is a tool like qmd, not an index rewrite. Each of those is a thing that would have to be maintained, and none of them is what makes this work — the prompts are.

About

A very opinionated harness for Agentic software engineering

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages