Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
54 changes: 34 additions & 20 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,34 @@
[![Agent Skill](https://img.shields.io/badge/Agent%20Skill-compatible-111827)](https://agentskills.io/)
[![MIT License](https://img.shields.io/badge/license-MIT-2563eb)](LICENSE)

An agent skill for preventing temporary software-development apparatus from becoming permanent
maintenance debt. It governs tests, scripts, harnesses, flags, adapters, caches, compatibility
paths, abstractions, debug surfaces, orchestration, and process documents.

The rule is simple: build whatever investigation requires, then remove it before completion
unless it serves a real continuing consumer, recurring risk, or live transition. Keep the
smallest artefact that meets that need.
An agent skill for one recurring failure of agent-driven development: the apparatus built
*while* solving a task — extra tests, debug flags, replay harnesses, one-off scripts,
compatibility shims, speculative abstractions, process documents — quietly becomes permanent
codebase the moment the task closes, and someone maintains it forever.

The rule the skill enforces is a reversal of burden. While solving, build whatever the
problem demands, freely. When the task closes, everything that remains must name its
continuing need — the consumer that will invoke it, the recurring risk it alone would catch,
or the live transition it carries — and the condition that ends it. "It helped once",
"might be useful", and "someone may depend on it" name no need.

Concretely, after a debugging incident:

| Artefact | Verdict |
|---|---|
| One focused test that would fail if the fixed bug returned | **Keep** — it alone catches a recurrence |
| A test for every rule the changed code enforces, merged into one table | **Keep** — each rule is its own risk |
| Three tests encoding the disproved first theory of the bug | Delete — they guard nothing |
| The synthetic-replay harness that reproduced the incident | Delete — its consumer was the investigation |
| A `DEBUG_*` flag wired into production code | Delete — incident-only surface |
| The investigation notes file | Delete — the commit message carries the story |
| A feature flag whose rollout finished | Delete, or wire an expiry that actually fires |

The skill also covers the harder boundaries: release gates and canaries must guard product
risk, not their own mechanism (a fix for the fixer's fix means the proof sits at the wrong
seam); incident-shipped machinery gets re-justified in calm conditions; scale tests take
their scale from the failure mechanism, not an impressive number; and deletion of artefacts
that predate the task needs consumer evidence, not just a clean grep.

## Install

Expand All @@ -35,19 +56,6 @@ $skill-installer install https://github.com/CodingCossack/anti-machinery/tree/ma

The skill follows the open Agent Skills format and has no harness-specific dependency.

## What it governs

- **Two lives:** temporary investigative machinery is cheap; permanent supported machinery
must name its continuing need.
- **Tests:** retain one proof per distinct recurring risk at the cheapest seam that can catch
its return. Remove duplicate or theory-specific investigation tests.
- **Scale:** derive load from the failure mechanism or a named threshold, not an impressive
arbitrary number.
- **Deletion:** remove new or directly superseded unused artefacts. Treat older unknown
artefacts as candidates until their static, dynamic, and external consumers are checked.
- **Abstraction:** wait for a second concrete consumer to reveal the boundary instead of
building for a predicted future use.

## With Change with Proof

[`change-with-proof`](https://github.com/CodingCossack/change-with-proof) determines what must
Expand All @@ -59,6 +67,12 @@ Use $anti-machinery with $change-with-proof to implement this change, prove it,
temporary or redundant supporting apparatus before completion.
```

## Testing

Behavioural changes are gated by pressure scenarios run against real coding agents on
fixture repositories, with no-skill and paired-skill controls; results are recorded in
[docs/testing.md](docs/testing.md). Run `./scripts/validate.sh` for structural checks.

## Structure

```text
Expand Down
80 changes: 80 additions & 0 deletions docs/testing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# Behavioural testing log

Method throughout: pressure scenarios run as independent `codex exec` agents on
gpt-5.6-sol, medium reasoning effort, with HOME/CODEX_HOME fully isolated so only the
condition's skills are visible; each scenario is a small dependency-free git fixture repo;
ground truth from `git status`/`git diff` after the run. One or two runs per condition —
a directional gate that catches gross regressions, not a statistical eval suite.

## Durable proof seam pressure test, 2026-08-19 (at publication)

Three release-pressure scenarios (a browser canary demanding a permanent two-origin leak
harness; a signed-manifest gate with a distinct decision contract; a retained gate parsing
`status: failed`) run against the published baseline pair, both candidate skills, and
crossed pairs. Result: this skill independently rejected the unjustified high-seam
apparatus (removed canary-specific interception machinery, refused the leak harness) and
independently preserved a focused test of a retained gate's distinct decision contract.
Full table in the change-with-proof repository's `docs/testing.md` ("Durable proof seam
pressure test, 2026-08-19"), recorded there because the scenarios were shared.

## Incident close-out battery, 2026-08-19 (description + boundary revision)

Fixture: a release gate whose `failed → allow` bug is already fixed and committed; the fix
commit also carries classic incident machinery — a synthetic replay harness plus wrapper
test, three disproved-theory tests, a `DEBUG_GATE` production logging surface, an
investigation notes file — and four focused decision tests (passed/failed/expired/unknown).
Task: "close this work out for merge."

| Condition | Machinery removed | Focused proof retained |
|---|---|---|
| No skill | All of it, correctly | All 4 decision tests |
| This skill (previous wording), run 1 | All of it | **Dropped the expiry test** while consolidating ("redundant") |
| This skill (previous wording), run 2 | All of it | **Dropped the expiry test** ("newly added expiry test: unrelated to the changed status rule") |
| This skill (revised wording), 2 runs | All of it | All 4 decision rules, expiry included (2/2) |
| change-with-proof alone | All of it | All 4, plus a counterfactual against the baseline commit |
| Both skills | All of it | All 4 |
| Both skills, final wording of each | All of it | All 4 rules (status rows consolidated into one table test, expiry kept), plus a counterfactual against the pre-fix baseline |

Two findings:

1. **Honest baseline:** gpt-5.6-sol with no skill already performed this obvious close-out
correctly. On easy machinery this skill's marginal value is ~zero; its demonstrated value
is on the harder cases above (the leak-harness scenario, where the baseline pair retained
the machinery and the candidate refused it).
2. **Observed over-pruning, now countered:** alone, the skill twice deleted a live decision
rule's only proof, rationalizing by the test's age and origin. The "Tests are machinery"
section now states the retention criterion explicitly (would this proof alone catch a
real recurrence — never age or origin), and both re-runs kept the proof.

## Trigger battery, 2026-08-19 (description routing, 3 reps × 26 tasks)

Judge: skill router choosing from descriptions only, among this skill, change-with-proof,
and five realistic competitors. Tasks: ten engineering-change positives, ten near-misses,
four apparatus-retention positives (post-fix cleanup, rolled-out flag, adapter-heavy PR,
CI canary pruning), one boundary, one pure-Q&A negative.

| Description | Apparatus positives (12) | Clear false fires (greenfield scaffold ×3, multi-agent design ×3, landing redesign ×3, microservice skeleton ×3) | Co-fires on ordinary coding tasks (30) |
|---|---|---|---|
| Previous ("Use for non-trivial software work — design, planning, …") | 12/12 | 12/12 fired | 30/30 |
| Current ("Use when software work creates, keeps, or approves supporting apparatus…") | 12/12 | 0/12 | 19/30 |

The previous description routed the skill onto essentially all software work, including
greenfield scaffolding where nothing exists to retire. The current one keeps every genuine
retention decision and stays out of work with no apparatus at stake. Co-firing on ordinary
change tasks is intended — such tasks create tests and scripts whose retention this skill
governs — but is now below the indiscriminate 30/30.

## Evidence gaps and traces still needed

- The skill is two days old; there is no end-use trace evidence yet (session-trace mining
on 2026-08-19 found 102 mentions, of which all non-catalog hits were the skill's own
genesis and publication work). Needed: real sessions where the skill was available during
ordinary feature/bugfix work, to measure whether close-out behaviour changes and whether
the co-firing cost is paid back.
- All behavioural results are n=1–2 per condition on one model (gpt-5.6-sol). The
incident-pile ("After the fire") and abstraction-boundary rules have no dedicated
scenario evidence at all — they were carried on argument, not measurement.
- No scenario yet tests the "Deletion needs evidence too" boundary (dynamic/external
consumers of pre-existing artefacts) against this skill specifically; the v2
change-with-proof battery showed no-skill agents already check dynamic registration on
an easy case.
14 changes: 11 additions & 3 deletions skills/anti-machinery/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: anti-machinery
description: Use for non-trivial software work — design, planning, investigation, implementation, debugging, refactoring, migration, testing, review, release, or cleanup — whenever the work may create, retain, or approve supporting apparatus such as tests, scripts, harnesses, reports, flags, adapters, caches, compatibility paths, abstractions, debug surfaces, orchestration, or process documents. Governs what may exist permanently after the task closes. Use alongside change-with-proof when existing behaviour is at stake.
description: Use when software work creates, keeps, or approves supporting apparatus — tests, scripts, harnesses, flags, adapters, caches, compatibility paths, abstractions, debug surfaces, reports, process documents — and above all at task close, review, or cleanup, when deciding what may remain permanent. Apparatus earns permanence only through a continuing consumer, a recurring risk, or a bounded transition it carries. Do not use to decide whether the product change itself is correct or proven; pair with change-with-proof for that.
---

# Anti-Machinery
Expand All @@ -11,13 +11,21 @@ Everything built in service of a task — code structure, tests, scripts, flags,

While solving, build whatever the problem demands — probes, one-off scripts, wide experiments, throwaway harnesses. Spend freely; keep it off the supported path and treat it as scheduled for deletion.

When the task closes, the burden reverses: whatever remains must name its continuing need — the consumer that will invoke it, the recurrence it guards, or the transition it carries for a consumer that still exists, together with the observable condition that ends it. Sunk effort, "might be useful", reviewer comfort, and generalised safety name no need. When the need is real but smaller than the artefact, shrink the artefact to the need.
When the task closes, the burden reverses: whatever remains must name its continuing need — the consumer that will invoke it, the recurrence it guards, or the transition it carries for a consumer that still exists, together with the observable condition that ends it. A stated end condition is not an end condition: wire it to something that will actually fire — a check that fails past the deadline, a scheduled removal, an expiry in the code itself. A comment naming a date will outlive the date. Sunk effort, "might be useful", reviewer comfort, and generalised safety name no need. When the need is real but smaller than the artefact, shrink the artefact to the need.

The same rule governs building: structure justified by a predicted future consumer fails it in advance. Let the second concrete case reveal the boundary rather than guessing it.

## Tests are machinery

Discovery legitimately produces many tests; durable proof needs few. Keep one clear proof per distinct risk that can recur — a contract, a boundary, a fixed bug — at the cheapest seam where the proof would actually fail if the risk returned. A test that re-proves the same rule at an adjacent layer, encodes a disproved theory, or restates the implementation adds maintenance and noise but no failure it alone would catch; it leaves with the investigation that produced it.
Discovery legitimately produces many tests; durable proof needs few. Keep one clear proof per distinct risk that can recur — a contract, a boundary, a fixed bug — at the cheapest seam where the proof would actually fail if the risk returned. A test that re-proves the same rule at an adjacent layer, encodes a disproved theory, or restates the implementation adds maintenance and noise but no failure it alone would catch; it leaves with the investigation that produced it. What stays is decided by one question — would this proof alone catch a real recurrence? — never by age or origin: each rule an artefact enforces is its own risk, and a proof born during an incident, or covering a rule adjacent to the one fixed, is still the only guard that rule has. Dropping a rule's only proof is deletion of proof, not of duplication.

## Gates guard risks, not themselves

Release gates, canaries, and CI checks carry the same burden as any machinery, plus one of their own: they must name not only the recurrence they guard but the ways they can fail while the product is healthy. A gate that blocks for reasons other than its risk is defective machinery — prefer moving the proof to a cheaper seam over repairing the gate in place. A focused test of a gate's distinct decision logic is ordinary proof, not recursion. When a gate needs apparatus that merely re-proves the same assertion or compensates for its own unreliable mechanism — a canary fix that needs a leak-proof harness, a fix for the fixer's fix — stop building. That recursion is evidence the proof sits at the wrong seam or relies on a mechanism too clever to trust; redesign the proof rather than reinforcing it.

## After the fire

Incident pressure defeats the two-lives split: exploration lands directly on supported paths, and the task "closes" with a deploy, so the reversal of burden never runs. When an incident closes, everything it added to permanent paths gets re-justified as if proposed fresh in calm conditions, and the default answer is deletion. Judge the pile, not only the pieces — supporting machinery must stay proportionate to the product it supports, and each individually defensible addition is no defence of the total.

## Scale from the mechanism

Expand Down
Loading