Authoritative guide to test coverage in applications built on the microservices-simulator. Four tiers, layered by what they exercise: the aggregate alone (T1), the service contract plus event publication (T2), consumer-side subscription (inter-invariant) behavior (T3), and saga orchestration (T4).
Single source of truth. This file is the only place tier definitions, required scenarios, assertion-ownership rules, and test code-block templates are stated. Skills and other docs must reference sections here by anchor (e.g.
testing.md § T2 — Service Test) rather than restate the rule text or copy the code blocks. If you find a skill file paraphrasing a rule instead of pointing here, that's drift — fix it by deleting the copy, not by editing both.
| Tier | Name | Test class | Scope | Profile-agnostic? |
|---|---|---|---|---|
| T1 | Aggregate | <Aggregate>IntraInvariantTest |
Intra-invariants (P1): creation happy-path, one violation per non-final P1 rule (EP), boundary on/off-points (BVA). Direct construction + verifyInvariants(). |
yes |
| T2 | Service | <Aggregate>ServiceTest — one class per aggregate, all service methods |
Service contract: change persisted to DB (read back via a fresh UnitOfWork), uniqueness / composite-key guards, not-found paths (Path A SimulatorException / Path B <App>Exception), P3 numeric-guard boundaries. Invoke the *Service bean directly with a UnitOfWork — no saga workflow. Also owns event-publication assertions: per published event type, one payload-asserting case + one negative "does not publish" case. |
yes |
| T3 | Subscription (Inter-Invariant) | <Consumer>InterInvariantTest — only for aggregates with subscribed events |
Consumer side: event received → cached state updated; unrelated event → state unchanged; deletion event → consumer deleted. Trigger publication via a functionality, poll via the consumer's event-handling bean, assert on consumer state. Must not re-assert event-store contents — T2 owns that. | mostly |
| T4 | Functionality | <FunctionalityName>Test + <FunctionalityName>CompensationTest |
Orchestration: saga state-machine traversal (NOT_IN_SAGA → IN_{OP} → NOT_IN_SAGA), semantic-lock acquisition per lock step, P3 guard violations raised through the saga path, compensation on mid-saga failure. |
no (sagas-specific) |
All test classes share the same skeleton, elided from the templates below:
@DataJpaTest @Transactional @Import(LocalBeanConfiguration), extend <AppName>SpockTest, and
declare @TestConfiguration static class LocalBeanConfiguration extends BeanConfigurationSagas {}.
Each fact is asserted in exactly one tier — this is what prevents T2/T4 duplication:
- T4 functionality tests do not assert field-level persistence, uniqueness, or not-found —
those belong to T2. A T4 happy path asserts orchestration outcomes only: the operation
completes, the returned DTO is coherent, and
sagaStateOf(id) == NOT_IN_SAGA. - T4 keeps: lock-acquisition cases (
executeUntilStep) and P3 guard violations that involve cross-aggregate saga coordination. - T2 owns event-publication assertions (merged from old T3): per published event type, one
payload-asserting case + one negative no-publish case, asserted via the
EventServicebean. - T3 owns subscription (consumer-side) assertions: event received → cached state updated / unrelated ignored / deletion → consumer deleted. T3 tests may trigger publication via a functionality but must not re-assert event-store contents — T2 owns that.
- Known coupling:
registerChangedcallsverifyInvariants(), so T2 fixtures can trip P1 rules — T1 and T2 are not perfectly independent layers. Documented, not fixed.
AI assistants given only a coverage metric will read the implementation they just wrote and mirror
it in assertions — tests that are trivially satisfied, break on correct refactors, and give false
security. This checklist is the authoritative smell list, consumed by
.claude/skills/implement-aggregate/.
then:is onlynoExceptionThrown()with no field assertions — flag unless the scenario is explicitly "must not throw."then:checks only non-null / non-empty, or a trivially true condition, never actual values.when:does not call the method under test (bypasses it via a setup helper).- T2: happy path reads back through the same UnitOfWork instance used for the write, or through a fresh one without calling
flushAndClear()first — either way the persistence context still holds the managed write instance, so the assertion never exercises the load path. Read-back isflushAndClear()then a second, fresh UnitOfWork (§ T2 — Service Test). - T3 "ignores unrelated":
originalValuecaptured after the event was processed — the assertion isx == x. Capture ingiven:, before firing. - T3 "reflects event" — carve-out, not a smell: when the event payload re-affirms a value the consumer is already guaranteed to hold and no legal consumer state can differ from it, the payload assertion is trivially satisfied and is nonetheless required, paired with an assertion that the cached publisher version advanced. That pairing is the sanctioned form — flagging it as Fake is itself a Wrong finding. The reachable-contrary-state alternative comes first; see
.claude/skills/implement-aggregate/session-d.md§{Aggregate}InterInvariantTest.groovy.
- Test name / message constant copied verbatim from the implementation without checking
plan.md's rule list — the test validates the implementation deviation instead of catching it. then:mirrors the sequence ofset…calls in the service body — derived from the implementation, not the spec.- T3 deletion-event test puts the post-deletion load in
then:instead of anand:block — outside the exception-capture scope (see § T3). - Not-found exception type contradicts the lookup mechanism — read the service method first (see § T2 Not-Found Paths). Flagging a correct Path B
thrown(<App>Exception)as Fake is itself a Wrong finding. - Test class missing
@Transactional— dirty state bleeds between tests. - P1 intra-invariant violation asserted in a T2/T4 test (belongs in
<Aggregate>IntraInvariantTest).
- Happy-path
then:asserts only fields already set insetup:— at least one asserted field must be a value the operation itself produces. - Removing
unitOfWorkService.registerChanged(aggregate, unitOfWork)from the service would leave the test passing (kill-mutation thought experiment) — add an assertion readable only through the persisted aggregate. - Violation test asserts
thrown(<App>Exception)withoutex.message == <RULE_NAME>— passes on any unrelated bug of that type. - T2: asserts only that "an event exists" (type/count) without asserting the payload fields — a wrong-payload regression slips through.
- Returned DTO assertions cover only a subset of the semantically important fields.
- Ordered-domain boundary rule asserted with only a far-side value — missing the on-point or off-point (see § Choosing Input Values).
- Equivalence Partitioning (EP) — split a rule's input domain into classes treated the same way; test one representative per class (one satisfying value, one violating value).
- Boundary Value Analysis (BVA) — defects cluster at class edges (off-by-one,
<vs<=). For ordered domains, also test the two values straddling the boundary.
For every P1 invariant or P3 numeric guard whose predicate is a comparison on an ordered domain (
<,<=,>,>=,==/!=over a count, timestamp, or collection size), EP's single representative is not sufficient. Add the boundary-straddling pair: on-point — last value that satisfies the rule →notThrown(...); off-point — first value that violates it →thrown(...)andex.message == <RULE>.Routing: P1 boundaries → T1; P3 numeric-guard boundaries → T2.
For categorical invariants — uniqueness, boolean/state freezes, set membership — there is no ordered edge; EP's one-representative-per-class is complete. Do not invent boundary cases.
| Rule shape | Example invariant | On-point (no throw) | Off-point (throws) |
|---|---|---|---|
count > 0 |
NUMBER_OF_ITEMS_POSITIVE |
1 |
0 |
count <= N |
MAX_ITEMS (N=30) |
30 |
31 |
a < b (timestamps) |
START_BEFORE_END_TIME |
end − 1 tick |
start == end |
a >= b (timestamps) |
DISPATCH_BEFORE_START |
firstDispatchTime == startTime |
startTime − 1 tick |
size >= 1 |
MUST_HAVE_ONE_LABEL |
1 |
0 |
Temporal mechanics: the smallest LocalDateTime tick is .minusNanos(1) / .plusNanos(1);
pin both instants explicitly so the on-point is exactly equal. All temporal P1 boundary cases
are T1 direct-aggregate tests — the saga path stamps lastModifiedTime = now() and cannot pin the
on-point. The tick is a T1 construct and stays exactly as written: a fixture that instead crosses a
service boundary and is compared after a read-back is governed by § "Persisted temporal fixtures"
below, and the two rules pull in opposite directions.
When a test instead manufactures a past/future timestamp to trigger a service-level date guard
(a T2 concern, not a P1 boundary), pin it against the same clock the guard under test actually
reads — e.g. pt.ulisboa.tecnico.socialsoftware.ms.utils.DateHandler.now() (fixed UTC), not
LocalDateTime.now() (JVM default timezone). A mismatch between the two can silently fail to
trigger the guard rather than throwing an obvious error.
The two cases above both pin an instant. A third case cannot: where the create path itself stamps
a clock field - DateHandler.now() at creation, or a field derived from it under a P4b construction
invariant - and a P1 invariant orders that field against a caller-supplied one, every aggregate the
production path can build lies on one side of the clock. A state on the far side is unreachable by
construction, not merely inconvenient to set up.
A test that needs such a state - a window that has already elapsed, a deadline already passed - reaches it by creating a short future window and waiting it out:
- Compute the window inside the fixture helper, immediately before the create call
(
start = DateHandler.now().plusSeconds(N)), so the only race is the saga's own latency rather than the time the rest of the test setup took. - Wait by polling the same clock the production code reads (
DateHandler.now().isAfter(end)), not by sleeping a fixed duration. - Give the far-side state its own fixture helper (
create{Adjective}{Aggregate}(...)) that owns both the window and the wait. The plaincreate{Aggregate}helper keeps the signature and defaults its own session mandates - see.claude/skills/implement-aggregate/session-c.md§ "Update{AppClass}SpockTest.groovy". - Assert against the window the helper returns, not against a constant the test declared.
Two shortcuts are forbidden:
- Back-dating a constant. A
PAST_START_TIME/PAST_END_TIMEpair fed to the create path makes the create throw the ordering invariant, and making it not throw means one of the two shortcuts below. - Weakening the invariant - relaxing the P1 rule, or adding a parameter that lets the test supply the stamped clock field - to make the unreachable state reachable. The invariant is the spec; a test that has to break it to run is testing a state the application cannot be in.
If the margin is chosen too small and a stall beats it, creation throws the ordering constant loudly. That is the intended failure mode: the test cannot pass for the wrong reason.
A temporal constant that crosses a service boundary and is later asserted with == after a
read-back must be drawn from the base class's testNow() helper, never from DateHandler.now()
directly:
public static final LocalDateTime {AGGREGATE}_{START_FIELD} = testNow().plusDays(10)LocalDateTime carries nanoseconds. Every SQL timestamp column Hibernate emits for one is
timestamp(6), so a write truncates or rounds the sub-microsecond digits away and the value that
comes back after flushAndClear() is not the value that went in. Whether it differs at all depends
on the host clock's resolution, which the JDK takes from the OS: a host whose clock ticks at 1us
never produces a sub-microsecond draw, so the suite passes; a host with nanosecond resolution turns
every such assertion into a coin flip. The defect is invisible on the machine that writes the test.
testNow() truncates at the source, so the comparison holds on both.
This is not the T1 tick rule, and applying either one in the other's place breaks the suite.
The .minusNanos(1) / .plusNanos(1) boundary fixtures of § Choosing Input Values — EP & BVA are
{Aggregate}IntraInvariantTest values, built on a transient aggregate and never persisted, so
nanosecond resolution costs them nothing. Never truncate a tick. Truncating the result of
.plusNanos(1) collapses the on-point back onto the off-point, and the on-point test that must not
throw starts throwing. Truncate the base instead - the boundary pair survives, because
X.truncatedTo(MICROS).plusNanos(1) is still strictly after X.truncatedTo(MICROS).
Two neighbouring shapes are governed elsewhere and need no truncation:
- A field the create path stamps from the clock (
creationDate,lastModifiedTime) is never asserted==against a fixture constant at all - only!= nullor by ordering. See.claude/skills/implement-aggregate/session-c.md§ "Update{AppClass}SpockTest.groovy", which owns that rule and the re-pinning that makes such a constant clock-relative in the first place. - A window computed to reach a time-gated state (§ "Reaching a time-gated state") is compared
with
isBefore/isAfter, never==, so its resolution is irrelevant.
Before writing any test, locate the plan.md aggregate section for the target aggregate. Its
happy-path postconditions, events-published list, subscribed-events table, and P1/P3 rule list
are the spec — assertions must trace to them, never to the implementation just written (not the
service body, not the EventProcessing class). Write a 1-line // Spec: comment at the top of
each test naming the plan.md section and rule, e.g.
// Spec: plan.md § 5. Shipment - UpdateShipmentNotes; rule SHIPMENT_NOTES_REQUIRED.
The section reference is the ### {N}. {Aggregate} heading classify-and-plan § Step 8 emits - plan.md
has no §n.n numbering, so a §3.5-style citation points at nothing.
If the implementation disagrees (e.g. throws a different message constant than plan.md names), the
implementation is the bug: flag the mismatch, do not adjust the test.
When plan.md is instead silent - it specifies no behaviour for the input under test, such as a
write method called with an out-of-domain target - see
docs/concepts/rule-enforcement-patterns.md § Decision Guide, Step 4, which fixes the behaviour and
requires the resulting constant to be added to plan.md's rule list before the test cites it.
src/test/groovy/pt/ulisboa/tecnico/socialsoftware/
├── SpockTest.groovy ← root Spock marker class; its package is the
│ parent, not <pkg>
└── <pkg>/
├── BeanConfigurationSagas.groovy ← test configuration (infrastructure beans)
├── <AppName>SpockTest.groovy ← base class: @Autowired services + factory helpers
└── sagas/
├── coordination/
│ └── <aggregate>/ ← one dir per primary aggregate
│ ├── <FunctionalityName>Test.groovy (T4)
│ └── <FunctionalityName>CompensationTest.groovy (T4, sibling — not a
│ separate behaviour package)
└── <aggregate>/ ← one dir per aggregate
├── <Aggregate>IntraInvariantTest.groovy (T1)
├── <Aggregate>ServiceTest.groovy (T2)
└── <Aggregate>InterInvariantTest.groovy (T3, consumers only)
Purpose: own the complete P1 matrix (see taxonomy table). All cases are direct-aggregate:
construct, optionally call aggregate mutators, then call verifyInvariants() explicitly — the
saga path cannot pin exact boundary instants. Happy-path fields must trace to the aggregate field
list in plan.md (a constructor field the spec doesn't list is a planning gap to flag, not a field
to copy). P1 Java-final fields need no test coverage in any tier — the compiler enforces
immutability; testing it tests the language, not the domain.
class <Aggregate>IntraInvariantTest extends <AppName>SpockTest {
def "create <aggregate>"() {
when:
def result = new Saga<Aggregate>(/* id, args or dto */)
then:
result.<field> == <expectedValue> // all fields from plan.md's field list
}
def "<aggregate>: <RULE_NAME> violation"() {
given:
def agg = new Saga<Aggregate>(/* valid args */)
// set field(s) to the violating value directly
when:
agg.verifyInvariants()
then:
def ex = thrown(<App>Exception)
ex.message == <RULE_NAME>
}
// Boundary pair per ordered-domain P1 rule — same shape as the violation case:
// on-point: pin the last satisfying value → notThrown(<App>Exception)
// off-point: pin the first violating value → thrown + ex.message == <RULE_NAME>
}Purpose: pin the service contract (see taxonomy table), one class per aggregate covering all
its service methods, invoked directly on the *Service bean with a UnitOfWork — no saga
workflow. No explicit commit is needed: in the sagas profile, registerChanged versions,
invariant-checks, and merges the aggregate inside the service call; the workflow-level
commit(uow) only resets SagaState.
Every read-back calls flushAndClear() first, then reads through a fresh UnitOfWork. Both
halves are required, and neither substitutes for the other:
- A fresh
UnitOfWorkalone is not enough. Under@DataJpaTestthe whole test runs in one transaction, so everyUnitOfWorkin it shares a single persistence context. Hibernate answers the read from the first-level cache and hands back the very instance the write path put there — the assertion then re-reads the object it just built in memory, which is Fake. flushAndClear()(the<AppName>SpockTesthelper:EntityManager.flush()thenclear()) pushes the pending writes to the database and detaches everything, so the next load constructs the aggregate through Hibernate's own instantiation path.
That path is what the read-back is actually there to prove. It is where final fields are set
reflectively, where @Convert converters run, and where a mismatched column mapping or a missing
no-arg constructor first becomes visible — none of which the managed instance would ever exercise.
The rule is unconditional: applying it only to aggregates believed to have final fields makes the
guard depend on a judgement made when the test was written, and it is silently lost the moment the
aggregate changes.
class <Aggregate>ServiceTest extends <AppName>SpockTest {
def "create<Aggregate>: persisted and readable through a fresh UnitOfWork"() {
// Spec: plan.md § <n>. <Aggregate> - Create<Aggregate> postconditions
when:
def dto = <aggregate>Service.create<Aggregate>(/* args */,
unitOfWorkService.createUnitOfWork("create<Aggregate>"))
then: 'read back off a cleared persistence context, through a fresh UnitOfWork'
flushAndClear()
def readBack = <aggregate>Service.get<Aggregate>ById(dto.aggregateId,
unitOfWorkService.createUnitOfWork("check"))
readBack.<field> == <expectedValue>
}
def "<serviceMethod>: <RULE_NAME> violation"() {
// Spec: plan.md § <n>. <Aggregate> - rule <RULE_NAME> (P3 guard / uniqueness)
given:
def existing = create<Aggregate>(/* fixture via base-class helper */)
when:
<aggregate>Service.<serviceMethod>(/* violating args */,
unitOfWorkService.createUnitOfWork("<serviceMethod>"))
then:
def ex = thrown(<App>Exception)
ex.message == <RULE_NAME>
}
// P3 numeric-guard boundaries: on-point (notThrown) / off-point (thrown + message) pairs.
// Not-found per read/mutate method: call with NONEXISTENT_AGGREGATE_ID →
// Path A: thrown(SimulatorException); Path B: thrown(<App>Exception) + message (see below).
}Purpose: pin the publisher side of every event (merged from the former T3 tier). Trigger via a
direct service call with a UnitOfWork — SagaUnitOfWorkService.registerEvent saves the event
immediately via eventService.saveEvent(event) (and marks it published under the local profile).
Assert via the EventService bean (pt.ulisboa.tecnico.socialsoftware.ms.notification.EventService):
per event type, one payload-asserting case (type/count alone is Weak), plus one negative case — an
operation that must not publish leaves the store unchanged. Autowire EventService in the same
<Aggregate>ServiceTest class; append these as their own def methods (not folded into existing
then: blocks).
class <Aggregate>ServiceTest extends <AppName>SpockTest {
@Autowired
EventService eventService
def "<serviceOp> publishes <Xxx>Event with correct payload"() {
// Spec: plan.md § <n>. <Aggregate> - events published by <ServiceOp>
given:
def publisher = create<Aggregate>(/* fixture via base-class helper */)
when:
<aggregate>Service.<serviceOp>(publisher.aggregateId, /* args */,
unitOfWorkService.createUnitOfWork("<serviceOp>"))
then:
def events = eventService.getAllEvents().findAll { it instanceof <Xxx>Event }
events.size() == 1
def event = events[0] as <Xxx>Event
event.publisherAggregateId == publisher.aggregateId
event.<payloadField> == <expectedValue> // every payload field from plan.md
}
def "<nonPublishingOp> publishes no <Xxx>Event"() {
// Spec: plan.md §<n> <Aggregate> — <NonPublishingOp> is not in "Events published"
given:
def publisher = create<Aggregate>(/* fixture via base-class helper */)
def countBefore = eventService.getAllEvents().size()
when:
<aggregate>Service.<nonPublishingOp>(publisher.aggregateId, /* args */,
unitOfWorkService.createUnitOfWork("<nonPublishingOp>"))
then:
eventService.getAllEvents().size() == countBefore
}
}countBefore is captured in given: after the fixture is built — fixture helpers commit
operations that publish events of their own, so a count taken before them measures the fixture, not
the operation under test.
Two distinct not-found paths throw different exception types — read the service method first:
- Path A — primary-key lookup. The service calls
aggregateLoadAndRegisterReaddirectly with an ID; the infrastructure throwsSimulatorException(pt.ulisboa.tecnico.socialsoftware.ms.exception.SimulatorException). - Path B — composite-key lookup. The service first queries a custom repository returning
Optional(e.g.warehouseId + shipmentId) and throws on empty at the service level: expect<App>Exceptionwithex.message == <NOT_FOUND_CONSTANT>.
<Consumer>InterInvariantTest covers the consumer side: event received → cached state updated;
unrelated event → state unchanged; deletion event → consumer deleted. @Scheduled does not
run in @DataJpaTest — call the polling method directly:
<consumer>EventHandling.handle<Xxx>Events(). These tests trigger publication via a functionality
but must not re-assert event-store contents (T2 owns that). If the consumer DTO does not expose a
cached sub-entity field, load the aggregate with the loadForCheck(aggregateId, type) helper the
scaffolded <AppName>SpockTest ships and assert on agg.<subEntity>.<cachedField>:
def agg = loadForCheck(consumer.aggregateId, Saga<Consumer>)
agg.<subEntity>.<cachedField> == <newValue>loadForCheck creates the "check" unit of work and casts, so it is the read-back shape everywhere
a test asserts through the aggregate rather than a DTO — in T3 and T4 alike. Call
unitOfWorkService.aggregateLoadAndRegisterRead directly only where the test needs the unit of work
itself, or registers the read on a unit of work it goes on to use.
Deletion events: when processing calls remove() on the consumer, aggregateLoadAndRegisterRead
filters out DELETED aggregates and throws SimulatorException — the load-and-assert pattern
cannot work. Move the load attempt into an and: block (extending the when: phase's
exception-capture scope, which a then:-placed load would sit outside of), then assert
thrown(SimulatorException).
class <Consumer>InterInvariantTest extends <AppName>SpockTest {
@Autowired
<Consumer>EventHandling <consumer>EventHandling
def "<consumer> <action> on <Xxx>Event"() {
// <action> is a concrete verb: updates, removes, deletes self, anonymizes, invalidates self
given:
def publisher = create<Publisher>(/* args */)
def consumer = create<Consumer>(/* args linked to publisher */)
when: '<publisher> triggers the event'
<publisher>Functionalities.<triggeringOp>(/* args */)
and: 'consumer polls for the event'
<consumer>EventHandling.handle<Xxx>Events()
then: 'consumer cached field is updated'
<consumer>Service.get<Consumer>ById(consumer.aggregateId,
unitOfWorkService.createUnitOfWork("check")).<cachedField> == <newValue>
}
// "ignores <Xxx> event for unrelated entity" twin: create a second publisher,
// capture originalValue in given: BEFORE firing (capturing after is Fake — x == x),
// trigger the op on publisher2, poll, assert consumer.<cachedField> == originalValue.
// Mandatory once per subscribing aggregate — see § "Ordering: the stale-event test" below.
def "<consumer> keeps the newest <field> when a stale <Xxx>Event trails it"() {
given:
def publisher = create<Publisher>(/* args */)
def consumer = create<Consumer>(/* args linked to publisher */)
when: 'the publisher is updated twice before the consumer polls'
<publisher>Functionalities.<triggeringOp>(publisher.aggregateId, <staleValue>)
<publisher>Functionalities.<triggeringOp>(publisher.aggregateId, <newValue>)
and: 'the consumer drains both pending events in one poll'
<consumer>EventHandling.handle<Xxx>Events()
then: 'the newer payload survives the older event that trails it'
<consumer>Service.get<Consumer>ById(consumer.aggregateId,
unitOfWorkService.createUnitOfWork("check")).<cachedField> == <newValue>
}
def "<consumer> is deleted when <Publisher> deletion event is processed"() {
given:
def publisher = create<Publisher>(/* args */)
def consumer = create<Consumer>(/* args linked to publisher */)
when: '<publisher> is deleted'
<publisher>Functionalities.delete<Publisher>(publisher.aggregateId)
and: 'consumer polls for the deletion event'
<consumer>EventHandling.handle<Xxx>Events()
and: 'attempt to load the now-deleted consumer aggregate'
unitOfWorkService.aggregateLoadAndRegisterRead(
consumer.aggregateId, unitOfWorkService.createUnitOfWork("check"))
then: 'DELETED aggregate is not loadable'
thrown(SimulatorException)
}
}One per subscribing aggregate, mandatory. Pick any one field-update event the consumer subscribes
to and prove that a stale event cannot undo a fresher one. The event pipeline makes this reachable
rather than theoretical: findUnprocessedEvents returns the batch timestamp DESC and the whole
batch is handled in that order, so two updates published before a single poll arrive newest-first and
the older payload lands last. events.md § "Reject an event that does not advance the cached
version" owns the guard this test exercises.
The test is per aggregate, not per event type, because the guard is the same code at every cached row of one consumer — one event type discharges it. Choose an event whose payload has two distinguishable values; a re-affirming payload (§ T3 above) cannot show the difference.
Do not substitute a "poll twice and nothing changes" test for it. That one is worth having — it is what proves the version is stamped at all — but it passes whether or not the guard exists, because a stamped version already takes the event out of the eligible set before the second poll. Only two events pending at once reaches the guard.
Every write saga is a finite state machine over its aggregate's SagaState: states —
NOT_IN_SAGA (quiescent) plus one IN_{OP} locked state per operation holding a semantic lock
across steps; transitions — acquire (setSemanticLock(IN_{OP}) drives
NOT_IN_SAGA → IN_{OP}), complete (successful commit drives IN_{OP} → NOT_IN_SAGA),
compensate (mid-saga failure also releases to NOT_IN_SAGA); guards —
setForbiddenStates([...]) blocks a transition when a foreign aggregate is in a listed state.
Testing a write saga means covering its transitions: the happy-path success case covers the full
traversal; one lock-acquisition case per setSemanticLock step pins the intermediate IN_{OP}
state, then resumes to complete; one compensation case per write functionality whose lock is
held across a later step exercises the compensate transition (see § Compensation Test below). Guard
(setForbiddenStates) transitions and Async tests remain deferred — see Appendix; those need a
second saga staged against the first and are out of scope. Everything else is owned elsewhere
(see § Assertion Ownership).
class <FunctionalityName>Test extends <AppName>SpockTest {
def "<functionalityName>: success"() {
given: 'aggregates exist'
// ...
when:
def result = <primary>Functionalities.<functionalityName>(/* args */)
then: 'orchestration outcome only — persistence is asserted in T2'
result.<keyField> == <coherentValue>
sagaStateOf(<aggregateId>) == GenericSagaState.NOT_IN_SAGA
}
// Saga-path P3 guard violations: same shape as the T2 violation case,
// but driven through <primary>Functionalities.<functionalityName>(...).
// One case per saga step that calls setSemanticLock — the acquire transition
def "<functionalityName>: <lockStep> acquires IN_<OP> semantic lock"() {
given:
def uow = unitOfWorkService.createUnitOfWork("<FunctionalityName>")
def func = new <FunctionalityName>FunctionalitySagas(
unitOfWorkService, /* args */, uow, commandGateway)
func.executeUntilStep("<lockStep>", uow)
expect: 'acquire transition: NOT_IN_SAGA → IN_<OP>'
sagaStateOf(<aggregateId>) == <Aggregate>SagaState.IN_<OP>
when:
func.resumeWorkflow(uow)
then: 'traversal completes back to NOT_IN_SAGA'
noExceptionThrown()
}
}Exception — a functionality whose success makes its own aggregate unresolvable. A
delete-shaped operation (one that soft-deletes the primary aggregate, or otherwise leaves it outside
what the unit of work will resolve) has no happy-path case. sagaStateOf(<aggregateId>) loads
through aggregateLoadAndRegisterRead, which throws rather than returning a state once the
aggregate no longer resolves, so the assertion the template mandates cannot run. The substitutes are
all worse: a persistence read-back belongs to T2 (§ Assertion Ownership), and a bare
noExceptionThrown() is the Fake smell named in § Fake/Wrong/Weak.
Such a functionality is fully covered without one: the lock-acquisition case pins the acquire transition, the compensation case pins the compensate transition, and T2 owns the assertion that the operation actually applied — the terminal state, any flag the owning aggregate's invariants require to move with it, and the published event's payload. Record the omission in the T4 file with a one-line comment naming this section, so a reader does not read the gap as missing coverage.
commandGateway is not inherited. Neither SpockTest nor <AppName>SpockTest declares it, so a
lock-acquisition test that constructs its saga directly must declare the field itself:
@Autowired
<Concrete>CommandGateway commandGatewayUse the concrete gateway type registered by BeanConfigurationSagas, not the CommandGateway
interface: the test profile registers more than one implementation, so injection by the interface type
is ambiguous and fails at context startup.
When a write saga acquires a semantic lock and a later step in the same saga throws, the lock
returns to NOT_IN_SAGA automatically — via the core's replay of each aggregate's recorded
pre-lock state on abort (SagaUnitOfWorkService.abortUntilStep → AbortSagaCommand), not a
manually-registered release compensation. Application workflows must not register such a release
(see docs/concepts/sagas.md § Semantic-lock release on abort is automatic). Required for every
write functionality that holds a semantic lock across a later step (a setSemanticLock step with a
dependent step after it); skip it for read-only functionalities and for a functionality whose only
step has no dependents (nothing to compensate). The sagaStateOf(...) == NOT_IN_SAGA assertion
below is the regression guard for lock-lifecycle bugs — it is what caught the currentExecutingStep
re-lock pitfall documented in sagas.md.
The test lets the lock-acquiring step run for real (so the lock is genuinely held and the
compensation genuinely registered), forces the very next step to throw via ImpairmentService,
then makes a three-part assertion: (1) the expected exception propagates — normally
SimulatorException from the injected fault, unless the step throws for an unconditional,
non-fault reason first (an update step whose target fields are P1 final always throws its
immutability constant, for instance — no impairment is needed there, just an added
sagaStateOf(...) == NOT_IN_SAGA assertion on the existing lock-acquisition test); (2)
sagaStateOf(aggregateId) == GenericSagaState.NOT_IN_SAGA (compensation actually ran); (3) a
read-back through the aggregate's read surface shows the mutation never applied.
class <FunctionalityName>CompensationTest extends <AppName>SpockTest {
def setup() {
loadBehaviorScripts()
// ... create fixtures via base-class helpers
}
def cleanup() {
impairmentService.cleanUpCounter()
impairmentService.cleanDirectory()
}
def "<functionalityName>: fault on <nextStep> compensates the lock acquired by <lockStep>"() {
when:
<primary>Functionalities.<functionalityName>(/* args */)
then: 'the injected fault surfaces to the caller'
thrown(SimulatorException)
and: 'compensation released the semantic lock back to NOT_IN_SAGA'
sagaStateOf(<aggregateId>) == GenericSagaState.NOT_IN_SAGA
and: 'the mutation never ran: read-back shows the pre-saga state'
def reread = <primary>Functionalities.<getterName>(<aggregateId>)
reread.<field> == <originalValue>
}
}The read-back uses whatever read coordinator the aggregate actually exposes. The template shows a by-id getter because that is the common shape, but an aggregate may deliberately expose only a filtered list read - and the read surface is fixed by session 2.{N}.b, not by this test. In that case call the list coordinator and select the aggregate under test from the result by its id. Never add a read coordinator, a service method or a command to satisfy this template: an unplanned by-id read is a functionality the domain model did not ask for, and it ships to production code to serve a test.
The ImpairmentService mechanism — read before writing any of these:
-
Already autowired in
<AppName>SpockTest(impairmentServicefield), registered as a bean inBeanConfigurationSagas, with aloadBehaviorScripts()helper on the base test class. CallloadBehaviorScripts()insetup(); callimpairmentService.cleanUpCounter()andimpairmentService.cleanDirectory()incleanup(). -
loadBehaviorScripts()points the impairment directory atsrc/test/resources/groovy/<TestClassSimpleName>/. Put one CSV per saga class you want to fault, named<SagaClassSimpleName>.csv, in that directory. -
CSV format — one
runblock per invocation of that saga class since the counter was last reset:run <stepName>,<fault 0|1>,<delayBefore>,<delayAfter> <stepName>,<fault 0|1>,<delayBefore>,<delayAfter>fault=1makesExecutionPlan.execute()throwSimulatorException("Fault on " + stepName)for that step (simulator/.../coordination/ExecutionPlan.java). Only steps up to and including the faulted one need a row — steps registered after it in the workflow never get checked, since the method throws and unwinds before reaching them.Historical note:
ExecutionPlan.execute()used to check every step's fault flag in registration order before scheduling dependent (non-root) steps for real execution, so only steps with no dependencies ran inline as the check passed over them — a genuine compensation test required the lock-acquiring step to be a root step, or the fault fired before the lock step ever ran, making the test a false positive. This was fixed by wiring the existingcanExecute()method into a topological worklist inexecute(), so a step's fault check now only fires once its full real dependency chain has genuinely completed, at arbitrary depth. Lock-acquiring steps no longer need to be root steps for their compensation to be genuinely testable, at any depth of dependency chain. Always sanity-check a new compensation test by temporarily flipping its fault flag to0and re-running the full suite with logging, not just the exception assertion — confirm the lock-acquiring step'sSTART EXECUTION STEPlog line actually appears before the fault fires. Capturing that log line needs maven's real stdout, which a shell redirect does not reliably give you — use the capture recipe in.claude/skills/_shared/conventions.md§ "Run the test suite" (§ "Inspecting maven output").
The CSV block index is selected purely by "how many times has this saga class been instantiated
since cleanUpCounter() was last called" — not by which test method is running. cleanup()
resets the counter after every Spock feature method, so every test method in a class independently
sees its first saga instantiation as invocation #1 → block 1. This means a compensation-fault case
cannot be added as one more def "..."() inside the existing <FunctionalityName>Test class — the
existing success/lock-acquisition tests need block 1 clean, the compensation test needs block 1
faulty, same block, same file, no way to have both. Put it in its own <FunctionalityName>CompensationTest.groovy,
as a sibling file in the same sagas/coordination/<aggregate>/ package (not a separate
sagas/behaviour/ package) — the resource lookup path only depends on the test class's simple
name, not its package.
With local.messaging.serialize: true (test profile), commands round-trip through Jackson.
CommandResponse.result is typed Object, so a List<XxxDto> stored there deserializes as
LinkedHashMap elements unless MessagingObjectMapperProvider.useForType() returns true for
raw == Object.class, which embeds @class metadata per element. Without that fix, read sagas
returning lists fail with ClassCastException. Irrelevant when the flag is off.
Documented for future work; not part of the current workflow. Compensation tests graduated out of this appendix into core T4 scope — see § Compensation Test above. Guard/forbidden-state transitions and Async tests remain deferred: both require staging two independent sagas against each other (pause one mid-workflow, run the other, resume the first), which is out of scope even though it's deterministic/single-threaded.
Two concurrent operations on overlapping aggregates must either produce a consistent result or
correctly reject one via semantic locks (setForbiddenStates guard transitions). One test per
functionality pair sharing an aggregate with real consistency risk. Includes guard/forbidden-state
coverage, e.g. UpdateWarehouseNameFunctionalitySagas's updateWarehouseNameStep guards against
WarehouseSagaState.READ_WAREHOUSE.
def "concurrent: <op1> step1 → <op2> completes → <op1> resumes"() {
given: 'func1 = new <Op1>FunctionalitySagas(unitOfWorkService, /* args */, uow1, commandGateway)'
when: 'func1.executeUntilStep("<stepBeforeShared>", uow1); <primary2>Functionalities.<op2>(...); func1.resumeWorkflow(uow1)'
then: 'both succeed consistently, or one throws a meaningful exception'
noExceptionThrown() // or: thrown(<App>Exception)
}N concurrent invocations of the same operation must all complete without corrupting state.
def "<functionality>: N concurrent invocations all succeed"() {
given: 'one shared aggregate + N participants'
when: 'fire all concurrently'
participants.collect { p ->
Thread.start { <primary>Functionalities.<functionalityName>(shared.aggregateId, p.aggregateId) }
}*.join()
then: 'flushAndClear(), then read back via a fresh UnitOfWork: <collection>.size() == N'
}