Abstract workflow for implementing a microservices-simulator application aggregate by aggregate, driven by AI agents. Each phase produces a stable artifact that becomes the entry point for the next session — agents are told exactly what to read and what to produce, not overwhelmed up front.
Scope: Sagas transactional model only. TCC is out of scope.
| File | Content |
|---|---|
{App}-domain-model.md — the plain domain |
Entities (§1), relationships incl. composition (§2), rules as standing invariants (§3.1/§3.2), functionalities with Primary Entity / Other Entities (§4) |
{App}-aggregate-grouping.md |
Aggregate partitioning + snapshot value objects (§1), snapshots (§2), technical fields (§2.b), event DAG + consistency policy (§3), events (§4) |
One plain domain, N aggregate groupings: the domain model names no aggregate, and a second grouping
over it must require no edit to it. See HARNESS.md § 5.
applications/{app-name}/
├── pom.xml
└── src/
├── main/java/pt/ulisboa/tecnico/socialsoftware/{app}/
│ ├── commands/
│ │ └── {aggregate}/ ← one subpackage per aggregate
│ ├── enums/ ← domain enums named by more than one aggregate
│ ├── events/ ← shared event classes (published by any aggregate)
│ └── microservices/
│ ├── exception/ ← {App}Exception.java + {App}ErrorMessage.java
│ ├── domain/ ← {App}DomainConstants.java (shared sentinels; only if an aggregate declares one)
│ └── {aggregate}/ ← one subpackage per aggregate
│ ├── {Aggregate}ServiceApplication.java
│ ├── aggregate/
│ │ └── sagas/
│ │ ├── factories/
│ │ ├── repositories/
│ │ └── states/
│ ├── coordination/
│ │ ├── eventProcessing/
│ │ ├── functionalities/
│ │ ├── sagas/
│ │ └── webapi/
│ ├── messaging/
│ ├── notification/
│ │ ├── handling/
│ │ │ └── handlers/
│ │ └── subscribe/
│ └── service/
└── test/groovy/pt/ulisboa/tecnico/socialsoftware/
├── SpockTest.groovy ← package pt.ulisboa.tecnico.socialsoftware (parent, not {app})
└── {app}/
├── BeanConfigurationSagas.groovy
├── {App}SpockTest.groovy
└── sagas/
├── coordination/
│ └── {aggregate}/ ← T4 functionality tests
└── {aggregate}/ ← T1 + T2 (incl. event pub.) + T3 subscription tests
| Layer | Pattern | Example |
|---|---|---|
| Aggregate root | {Aggregate}.java |
Shipment.java |
| Sagas extension | Saga{Aggregate}.java |
SagaShipment.java |
| Saga state enum | {Aggregate}SagaState.java |
ShipmentSagaState.java |
| Factory | Sagas{Aggregate}Factory.java |
SagasShipmentFactory.java |
| Custom repository | {Aggregate}CustomRepositorySagas.java |
ShipmentCustomRepositorySagas.java |
| Service | {Aggregate}Service.java |
ShipmentService.java |
| Command handler | {Aggregate}CommandHandler.java |
ShipmentCommandHandler.java |
| Write functionality | {Operation}FunctionalitySagas.java |
AddShipmentItemFunctionalitySagas.java |
| Read functionality | {Query}FunctionalitySagas.java |
GetOpenShipmentsFunctionalitySagas.java |
| Event subscription | {Aggregate}Subscribes{Event}.java |
ShipmentSubscribesUpdateWarehouseName.java |
| Event handling | {Aggregate}EventHandling.java |
ShipmentEventHandling.java |
| Event handler | {Aggregate}EventHandler.java |
ShipmentEventHandler.java |
| Event processing | {Aggregate}EventProcessing.java |
ShipmentEventProcessing.java |
See docs/concepts/testing.md for the full taxonomy (T1–T4).
| Type | Pattern | Session |
|---|---|---|
| T1 Aggregate | {Aggregate}IntraInvariantTest.groovy |
2.N.a |
| T2 Service | {Aggregate}ServiceTest.groovy — one class per aggregate; also owns event-publication assertions |
2.N.b (read methods; write-method and event-publication cases appended in 2.N.c) |
| T3 Subscription (Inter-Invariant) | {Aggregate}InterInvariantTest.groovy |
2.N.d |
| T4 Read Functionality | {Query}Test.groovy |
2.N.b |
| T4 Write Functionality | {Operation}Test.groovy |
2.N.c |
| T4 Compensation | {Operation}CompensationTest.groovy - one per write op that holds a semantic lock across a later step |
2.N.c |
plan.md lives at applications/{app-name}/plan.md. It is produced by Phase 1 and updated
(checkbox ticked) at the end of every subsequent session. It is the single entry point for every
agent in Phase 2: a Rule Classification table, an Aggregate Implementation Order table
(topological sort), and one Aggregate Details section per aggregate (functionalities, events,
cross-aggregate prerequisites, the per-session file list, a checklist).
Illustrative excerpt (one row of the Implementation Order table):
| # | Aggregate | Upstream deps | Events published | Events subscribed | Sessions |
|---|-----------|--------------|-----------------|-------------------|---------|
| 1 | Warehouse | — | — | — | a b c |Sessions column:
a=domain,b=read functionalities,c=write functionalities,d=event wiring (only when Events subscribed is non-empty).
The full output structure (every section, exact table columns, and generation rules) is
authoritatively defined in
.claude/skills/classify-and-plan/SKILL.md Steps
7-8 — that skill is what generates plan.md, so it owns the shape. Do not restate the template
here; if it changes, edit the skill, not this file.
Before Phase 0. The pipeline starts from the spec pair listed under § "Required Inputs", written
into applications/{app-name}/. Every later phase treats the pair as given: Phase 1 does not
question an aggregate boundary, and no Phase 2 session adds a functionality the domain model omitted.
Write it with docs/templates/domain-model-template.md and
docs/templates/aggregate-grouping-template.md, which
define the section numbers and table shapes the harness parses; /author-spec <pointer to the application being modelled>, which interviews you through the design tree in two parts — the plain
domain first, the grouping only after you sign it off — and writes one file at the end of each; and
applications/trainticket/, a finished pair kept as a worked example. HARNESS.md § 5 and the
templates own the detail.
plan.md does not exist yet. Phase 1 creates it.
One session. No plan.md exists yet. Produces the Maven scaffold, exception classes,
BeanConfigurationSagas.groovy (infrastructure beans only — no domain beans yet), and Spock test
base classes, all produced from the checked-in scaffold templates under
.claude/skills/boot-strap/templates/.
It also creates applications/{app-name}/harness-log.md and is the sole declarer of the run's
self-healing mode, written into that file's header and never rewritten afterwards. --self-healing
turns the mode on; absent, it is off (AGENTS.md § "Harness evolution").
The full procedure — exact files read, every transformation applied, and the complete produced-file
list — is authoritatively defined in
.claude/skills/boot-strap/SKILL.md. Invoke it with
/boot-strap <app-name>.
plan.md does not exist yet. Phase 1 creates it.
One session. plan.md does not exist yet.
{App}-domain-model.md— all sections{App}-aggregate-grouping.md— all sectionsdocs/concepts/rule-enforcement-patterns.md— the pattern taxonomy and classification flowchart
applications/{app-name}/plan.md only, using the structure defined in
.claude/skills/classify-and-plan/SKILL.md (see plan.md — The Job Queue above).
harness-log.md belongs to Phase 0: this phase halts if it is absent rather than creating one,
because creating it here would create it without a declared mode. The agent must:
- Join §4 of domain-model (Primary Entity / Other Entities) against §1 of aggregate-grouping (Entities contained) → the Derived Aggregate Mapping table. Entities the grouping co-locates collapse to one aggregate; an operation left with no other aggregate needs no saga coordination.
- Apply the decision guide to every §3.2 rule → populate the Rule Classification table.
- Topological-sort aggregates by the dependency DAG (§3 of aggregate-grouping) → the Implementation Order table. Aggregates with no upstream deps come first.
- For each aggregate in order, fill the Aggregate Details section: write/read functionalities (split from §4 of domain-model by the derived mapping), events published/subscribed (from aggregate-grouping §4), cross-aggregate prerequisites (P4a rules and P3 DTO-check rules) with their step names, and the full file list per session.
- Set the
dsession checkbox only for aggregates that have a non-empty Events subscribed list.
Any source file, and not the harness-log.md header. Output is plan.md, plus any harness-log rows
this session's own friction produced.
Loop: repeat sessions a → b → c → d for each aggregate in plan.md order.
Each session agent follows these steps:
- Open
plan.md. Find the first unchecked session for the current aggregate. - Read only the docs listed for that session type (below).
- Produce the files listed in the aggregate's file table in plan.md.
- Update
BeanConfigurationSagas.groovy(see per-session instructions). - Tick the checkbox in plan.md before finishing.
Each session type's exact reads, produced files, and BeanConfigurationSagas.groovy updates are
authoritatively defined in its sub-file under .claude/skills/implement-aggregate/ — that sub-file
is what an agent actually executes, so it owns the detail. This table is a one-line orientation
only:
| Session | Name | Sub-file | Adds to BeanConfigurationSagas.groovy |
|---|---|---|---|
| 2.N.a | Domain Layer | session-a.md |
Sagas{Aggregate}Factory, {Aggregate}CustomRepositorySagas |
| 2.N.b | Read Functionalities | session-b.md |
{Aggregate}Service, {Aggregate}CommandHandler, {Aggregate}Functionalities |
| 2.N.c | Write Functionalities | session-c.md |
none — the three beans are registered in 2.N.b; {Op}FunctionalitySagas are per-request objects, not Spring beans |
| 2.N.d | Event Wiring (only if aggregate has subscribed events) | session-d.md |
{Aggregate}EventHandling, {Aggregate}EventHandler, {Aggregate}EventProcessing |
A session's unit of work is a slice: one write functionality (session c) or one subscribed
event (session d), implemented end to end - its own new files, its appends to the session's shared
files, and its own tests.
Each session carries an ordered slice list, emitted into plan.md as sub-checkboxes at Phase 1 and
fixed from then on. Sessions a and b are never sliced. Sessions c and d are sliced one slice
per item only when the item count is greater than 3; at or below the threshold the session is one
implicit slice and carries no sub-checkboxes. The generation rules, the id scheme (2.{N}.{type}{k})
and the ordering rules are owned by
.claude/skills/classify-and-plan/SKILL.md
§ "Step 8.5" - do not restate them here.
Deciding the split at Phase 1 rather than at runtime makes it a reviewable, reproducible artifact of the run: the same spec always slices the same way.
| Command | Scope | Topology |
|---|---|---|
/implement-aggregate [session] |
one session | one agent implements the whole session |
/implement-aggregate-full <N> |
one aggregate, 2.N.a through 2.N.d |
a manager runs the sessions and delegates each slice to a fresh subagent |
Both read the same session-a..d.md sub-files, which remain the single source of truth for how to
implement, and both close a session through
.claude/skills/_shared/session-completion.md,
so they produce the same retro shape and the same one-commit-per-session history.
Slices run strictly sequentially, each with a fresh narrow context, instead of one agent holding a
whole heavy session at once. The manager owns the self-healing gate, the harness-log rows and every
commit; slices report friction and halt on Type 2 (AGENTS.md § "Harness evolution"). Its contract
for a slice is .claude/agents/aggregate-slice.md.
What the choice actually trades:
| Axis | /implement-aggregate |
/implement-aggregate-full |
|---|---|---|
| Scheduled checkpoints | One per session, so four per aggregate. Each ends in a completion report you read before invoking the next. | One per aggregate. The manager runs a through d and reports at the boundary. |
| Unscheduled halts | Type 2 and 2-fw halt and wait for you. |
Identical. A slice halts, the manager surfaces the block verbatim and waits. |
| Token cost | Lower. One agent, one context per session. | Higher. Every slice reads the session sub-file and its concept docs from cold, and the manager holds its own context on top. |
| Context quality | Degrades across a long session as one context accumulates the whole of it. | Higher per unit of work. Each slice starts fresh and narrow, which is the property the topology exists to buy. |
| Failure blast radius | One session. You see the result before the next one starts. | Up to a whole aggregate. A wrong reading can be repeated across four sessions before you next look. |
-full is not unattended. It reduces how often you are asked to start something, not how often
you are asked to decide something: the Type 2 and 2-fw gates are the same gates, they fire on the
same friction, and they wait just as long. Budget for being interrupted either way.
Applications whose plan.md predates slicing are not migrated: /implement-aggregate-full halts
rather than guessing a slice list. Run /implement-aggregate on them, or regenerate the plan.
/review-artifacts is not a phase of this pipeline. It is a static maintenance pass over the harness
itself - docs/, .claude/, AGENTS.md, HARNESS.md - that collects the debris hand-editing and
mid-run repair leave behind: paths that no longer resolve, two files prescribing different things, a
piece of knowledge that lost its single owner, a domain noun leaked in from the application being
generated. It reports; it does not repair.
It is always human-invoked, under both entry points. /implement-aggregate-full stops at the
aggregate boundary and tells the human it is available; it never runs it itself and never continues
to {N+1} either way.
It is expensive. It reads every harness file in full, so it fills a context window fast. Run it in a fresh session, never inline in a Phase 2 session. That cost is why the guidance below is a recommendation rather than a step.
When it is worth running, in descending order:
- After any substantial change to the harness - a refactor, a batch of doc rewrites, a new or retired skill. This is what it is for, and the run that follows reads whatever the edits left behind.
- After each finished aggregate, under self-healing ON (
AGENTS.md§ "Harness evolution"). There, sessions repairdocs/and.claude/skills/under the Type 1 gate, so the artifacts change while they are being read: a fix made in2.{N}.ccan contradict a doc that2.{N+1}.ais about to follow, and the aggregate boundary is the last moment that contradiction is cheap. Act on its Critical and Major findings, in their ownharness:commits, before starting{N+1}. - Occasionally under OFF, if you want the neutral-domain sweep early. Under OFF the harness cannot drift mid-run - the bucket must stay clean and the session commit halts if it is not - so there is nothing new for it to find between aggregates except nouns the run itself could not have introduced. Once at the end of the run is usually enough, and its findings are acted on between runs rather than at a boundary.
Generated automatically at the end of every Phase 2 session, by whichever entry point ran it.
No separate invocation required. After all session files are produced and the plan.md checkbox is ticked, the retro is written to:
applications/{app-name}/retros/retro-{session-id}-{Aggregate}.md
Example: applications/{app-name}/retros/retro-2.3.b-Shipment.md
A single commit covering the implementation files, the retro file and any harness-log.md rows is
then issued automatically, in the message format defined in session-completion.md § "Commit".
The retro template, both assembly topologies and the commit step are owned by
.claude/skills/_shared/session-completion.md.
Under /implement-aggregate the retro is synthesised from the single agent's conversation context;
under /implement-aggregate-full the manager merges the retro fragment each slice returned, which
preserves the same "synthesis from own context" property at the level where the context actually
lives. The file name and schema are identical either way, so /harness-retrospective is unaffected.
| Section | Purpose |
|---|---|
| Files Produced | Audit trail of what was shipped |
| Docs Consulted | Which concept docs were read and whether they were sufficient |
| Skill Instructions Feedback | What worked / what was unclear in the skill sub-file |
| Documentation Gaps | Specific missing or ambiguous content in docs/concepts/ |
| Patterns to Capture | Undocumented patterns discovered during implementation |
| Harness Changes | The harness-log.md row numbers appended this session, and the harness: commit sha of every Type 1 fix |
| One-Line Summary | The single most important finding |