Skip to content

Classify an Orleans placement expiry as a transient delivery failure for the sender - #6460

Closed
systemorph-com[bot] wants to merge 3 commits into
mainfrom
bugfix/systemorph-meshweaver-5731
Closed

systemorph-com[bot] wants to merge 3 commits into
mainfrom
bugfix/systemorph-meshweaver-5731

Conversation

@systemorph-com

Copy link
Copy Markdown
Contributor

What changed

  • RoutingGrain.ClassifyDeliveryException answers ShuttingDown (transient) for a delivery whose message expired before Orleans had placed its target grain (OperationCanceledException: Message expired before placement could complete for grain ...). New predicate RoutingGrain.IsPlacementExpired plus the PlacementExpiredMarker it matches on (narrow: that sentence, on an OperationCanceledException only).
  • RoutingGrain.DeliverToGrainObservable now converts the grain call through ObserveGrainCall. Rx turns a cancelled task into a fresh TaskCanceledException whose text is only "A task was canceled.", which would have hidden Orleans' sentence from every classifier and from the sender's NACK; the exception the task stored is handed over instead.
  • No re-send: the sender is told on the first attempt, so a persistent placement stall cannot pin a routing-pool slot for six full time-to-live waits (the RoutingGrain route dispatches hang at 64 in-flight when _Activity compilation stalls #1172 amplification).
  • src/MeshWeaver.Connection.Orleans/README.md: a short "Placement expiry" section.

Why
Refs #5731: two IMessageHubGrain.DeliverMessage calls to an ephemeral _Activity/compile-* hub died in Orleans placement ("Message expired before placement could complete"). The message reached no activation, so nothing was lost silently - the sender got a NACK - but it was the terminal Failed, and consumers with their own recovery machinery (SynchronizationStream resubscribe, MeshNodeStreamCache transient-owner rule) tear down on Failed and ride out ShuttingDown. Why placement stalled is not evidenced (Loki retention ends before 2026-09-24); this change does not touch Orleans or its log line.

How tested
New PlacementExpiryClassificationTest (pure, no cluster): the expiry reaches the classifier with its text and is not re-sent; the route NACKs ShuttingDown; it is found beside another fault in an aggregate; any other cancellation, and the marker quoted by a non-cancellation, stay Failed; an anti-inert pin checks the marker against the shipped Orleans.Runtime.dll. Not run locally; CI is the build.

Scope note
The classifier and the delivery retry live in src/MeshWeaver.Hosting.Orleans/RoutingGrain.cs and the pinning tests in test/MeshWeaver.Hosting.Orleans.Test/, outside the two folders the brief named (declared on the bug item before touching).


Refs #5731
Bug-Thread: Hosting/Triage/_Thread/bug-systemorph-meshweaver-5731

Opened by Dispatch's control plane on behalf of Essentials/Agent/bug-triage (auto · Provider/OpenRouterEU/anthropic/claude-sonnet-5.5), from the bug's own branch bugfix/systemorph-meshweaver-5731. It is never merged by Dispatch: the merge is the governed activity dev.merge, signed by a person once the checks and the internal review are green.

@github-actions

github-actions Bot commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

Test Results

  3 files    3 suites   8m 2s ⏱️
418 tests 418 ✅ 0 💤 0 ❌
422 runs  422 ✅ 0 💤 0 ❌

Results for commit 938e6c7.

♻️ This comment has been updated with latest results.

@rbuergi

rbuergi commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Verdict: symptom reclassification, not a root-cause fix. Recommend closing unmerged.

Judged against AGENTS.md, "No band-aids: root cause only". Read: the PR body, #5731 with all comments, the diff at 938e6c76. Nothing was built or run.

What it changes. A delivery that expired inside Orleans placement is reported to the sender as ErrorType.ShuttingDown instead of ErrorType.Failed. The cause of the expiry is untouched, and the PR body says so: "Why placement stalled is not evidenced".

Why the classification is not true by construction. ClassifyDeliveryException's own contract, a few lines above the edit, sets the bar:

Only the two conditions that ARE a lifecycle transition by construction qualify: the grain directory mid-handoff, and the host going away. [...] a bare TimeoutException (a target that did not answer across the WHOLE retry budget, i.e. plausibly wedged rather than restarting) must stay terminal, or the answer becomes a resubscribe storm against a hub that never comes back.

A placement expiry is that timeout, one layer earlier: the message waited its whole time-to-live in PlacementService and nothing placed the grain. "No activation saw it" is true, and it is an argument about re-send safety. It says nothing about whether the stall is a transition that ends. The PR concedes the point when it declines to re-send because "a persistent placement stall" would pin a routing-pool slot.

Why the handling is not bounded. The same contract says a consumer told ShuttingDown resubscribes, without a budget. Under a persistent stall each resubscribe is another delivery that waits a full time-to-live in placement. The amplification the PR avoids in the router moves to every consumer.

What the evidence since the PR was opened points at. #5731 was reopened on 2026-10-10 with 7 new expiries; the samples are for messagehub/Governance at 18:11:17Z, logged on memex-portal-deployment-5499f76b76-sz46x. Five minutes later a pod of the same ReplicaSet, 5499f76b76-ltlwc (started 18:10:16Z), was voted dead and stopped (#6432, fold of 18:16:16Z). #6432's reading of 2026-10-10 is that such silos are wedged from their first seconds and answer no grain call until they are evicted, 8 to 20 minutes later. Placement expiries against a silo in that state are a symptom of it, and for those minutes the honest verdict for a sender is the terminal one. I did not establish that these placements were waiting on ltlwc; the correlation is the ReplicaSet and the five-minute window only.

What it would hide. Orleans' own Orleans.Messaging[100071] line is untouched, so the incident log still sees the expiry. What changes is the consumer side: streams that today tear down visibly on Failed would instead keep retrying quietly against a silo that is not coming back.

Root-cause work and where it is tracked.

One part is worth keeping, separately. ObserveGrainCall / OriginalCancellation stop Rx from replacing Orleans' sentence with "A task was canceled.", so the NACK detail and the [ROUTE] warning name what happened. That is diagnostics, changes no verdict, and could land alone with its tests (the text reaches the classifier; the marker is pinned against the shipped Orleans.Runtime.dll).

State of this PR. The only red is Automatic review answered, and its log says the automatic review never landed (0 threads, 0 reviews on this head), so there is no thread to answer. I have not pushed, closed or armed anything.

Review by Claude Opus 5.5 (agent session, tag bot1).

@rbuergi

rbuergi commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Closing on the maintainer's go (2026-10-11): band-aid per AGENTS.md — it reports a placement expiry as ShuttingDown while the cause of the stalled placement is unproven; consumers would resubscribe unbounded against a silo that is not coming back. Root cause is tracked on #6432 / #6395; #5731 stays open as the symptom record. The one diagnostic improvement (keep Orleans' sentence instead of 'A task was canceled.') is being re-landed alone in a small PR. Verdict: #6460 (comment)

@rbuergi rbuergi closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant