Skip to content

feat(modules): reload + two-phase uninstall requests, policy packages-auto-update — independent of seal/sync - #6124

Merged
rbuergi merged 8 commits into
mainfrom
feat/module-reload
Oct 5, 2026
Merged

rbuergi merged 8 commits into
mainfrom
feat/module-reload

Conversation

@rbuergi

@rbuergi rbuergi commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Module reload — "reload M on instance X", one request

Policy module-reload-request (register row added, proposed until the Plugins half and the live loader's swap land). Manual: Doc/Architecture/ModuleReload (new).

Motivated by today's control-instance incident: MeshWeaver.AI 1.20.4 stayed loaded while Hosting needed 1.21 (MissingMethodException ThreadPreparation.set_Group); the only remedies were a manual sync plus a hand-filed Restart that needed confirmation.

What it does

  • ModuleReloadRequest (MeshWeaver.Graph.Configuration) at Admin/_ModuleReload/{id}, written only by ModuleReload.Request (as System, after the caller authorised the requester). Module name or package id, blank = all installed modules.
  • ModuleReloadExecutor runs on the request node's OWN hub, so exactly one process drives a request. Every step is a stream.Update and is idempotent:
    1. Resolve and land: RegistryUpdateReconciler.ReloadModule → PluginBundleClient.AdoptModuleOutcome (attended, so the package's Notify/None policy does not decline it), on the reconciler's serialised lane, then ModuleDependencyFloor.ProposeChecked. Compatibility = the declared floor against the running platform (policy package-min-mesh-version). No seal or platform-green check. Above the floor → declined — … by name, nothing downloaded, N keeps serving.
    2. Activate: live when every module that needs it IModuleLiveActivation.CanSwap. Otherwise exactly ONE restart: restartRequestedAt is stamped first, then IModuleActivationRestart.RequestRestart is called once.
    3. Report: ModuleReloadAgent (one per process) writes replicas{process} with what the process LOADED. This is measured from the loaded generation dir against the activation record. After a restart, only processes booted after the stamp report. ModuleReload.Evaluate decides Done/Failed by name; a failed live swap falls back to the one restart.
  • SelfUpdateHostedService implements IModuleActivationRestart through its existing Restart path: self-patch restart, or self-update-restart-pending handed to the control lane → unattended Restart, no approval or confirmation. The poller's roll floor does not defer an explicit reload. The service is now registered as a singleton forwarded to IHostedService + IModuleActivationRestart.
  • PluginBundleClient.AdoptModuleOutcome: AdoptModule with its whole answer (served version and floor, decision, files landed, named refusal). AdoptModule is now this, projected to the count. Same behaviour, one implementation.
  • Surfaces: MeshOperations.ReloadModule(module, reason) (global admin only; the MCP reload_module tool is in MeshWeaver.Plugins). The package page gets a node-menu entry 🔄 Reload module plus a ReloadModule area (framework controls, en + de keys).
  • Live loader coordination (feat(modules): every module runs in its own collectible load context (live update, slice 1) #6121): IModuleLiveActivation is the reload's call site only. feat(modules): every module runs in its own collectible load context (live update, slice 1) #6121's slice 2 registers it; nothing is forked. Until it does, every activation takes the one restart, which is that design's own fallback.

Generalising RefreshModules

The Plugins fleet-target intake used to file a Store RefreshModules maintenance task for its own sync-held modules. That task lands the module and stops. The Plugins PR switches the intake to this request, which also activates and reports. RefreshModules stays as the operator's content+module re-install.

Tests (Memex.Portal.Shared.Test → ModuleReload*, 12 tests)

  • Rules: validation, which replica reports count after a restart (old pod excluded, a gone process excluded), swap-failure signal, loaded-version reading. Each has its negative control.
  • End to end, through a real fake-HTTP registry, the real landing, reconciler, request node, executor, agent, and the real SelfUpdateHostedService with a recording k8s updater:
    • M@1.1.0 running, 1.2.0 published → found, landed, exactly ONE restart (inside the poller's 1 h floor), a post-restart process reports 1.2.0 → Done.
    • Negative: the post-restart process still loads 1.1.0 → Failed, naming the replica and both versions.
    • Floor above running → declined by name, no download, no restart, head stays 1.1.0.
    • Already current → Done/NotNeeded, no restart. Unknown module → refused by name.
    • Live: swap → Done/Live, zero restarts. Swap failure → exactly one restart.
  • Negative control run: making RequestRestart honour the roll floor fails the acceptance test (the request goes red "deferring the restart").
  • Regression set (SelfUpdate*, PluginBundle*, ModuleUpdate*, MinMeshVersion*, RegistryUpdate*, Undelivered*, HeldUpdates*, ComboGateRoll*, OneFloorComparator*, ModuleReload*): 294/294. MeshWeaver.Documentation.Test 670/670. Localization 66/66.
  • Release -warnaserror clean: PluginCatalog, Mesh.Operations, Memex.Portal.Shared, Memex.Portal.Shared.Test, Documentation.Test, Messaging.Hub.Test.

What is NOT established

  • No real cross-process run. Which generation a process has loaded and the post-restart pod are simulated (seam on ModuleReloadAgent). No Kubernetes restart and no multi-replica Orleans cluster were exercised.
  • "Newest compatible" only ever sees ONE version per package: the registry's bundle index advertises one. If its floor is above the running platform, the reload declines; an older compatible version newer than N is not discoverable.
  • An attended reload does not wait on a sync-owned partition's content hold (same as Provision).
  • The package-page menu/area and MeshOperations.ReloadModule have no automated test.
  • A restart handed to the control lane that never executes leaves the request AwaitingRestart. It stays visible and is not retried.

Recycle after deploy

None. The request node type is new, and the package-page menu rides the Package node type's hub configuration. A live Plugins/* package activation shows the menu after its next activation; recycle the Package type to see it at once.

The Plugins half (reload_module MCP tool + intake switch) is a separate PR. It compiles only against a core that carries MeshOperations.ReloadModule, so this PR lands first. No public surface is removed here.

Mirror-sync: the 7 new keys (menu.reloadModule*, moduleReload.*) are handed over — after this merges, run npm run sync:i18n -- --ref <merged core sha> in MeshWeaver.Plugins clients/react; tracked on the Plugins PR for this change.


Second commit — policy packages-auto-update (in force)

Maintainer directive: "put policy to auto-update packages. we just update whenever one is available" and "independently from platform deploy and sealing and stuff." The register row is added as in force. Manual: Doc/Architecture/ModuleReload → "Auto-update".

Only two inputs decide:

  • a newer version was published to the registry, and
  • its declared platform floor is at or below the running platform.

What changed

  • Default policy = Auto. PackageInstaller.SeedUpdatePolicy seeds a fresh install as Auto. The deployment-wide DefaultUpdatePolicy / AutoUpdateByDefault are no longer consulted. A seeded reminder-only record re-stamps onto Auto.
  • Migration. PackageAutoUpdateMigration runs in the boot repair pass and uses stream.Update as System — no SQL. It moves every seeded Notify record (including legacy autoUpdate:false records with no policy) to Auto.
    • It keeps, and names in one Warning, a pin (None) and any policy an administrator chose on the catalog card.
    • SetUpdatePolicy now stamps updatePolicySetAt, so a deliberate choice can be told apart from a seeded one.
    • A Notify chosen deliberately before this change carries no stamp, so it is indistinguishable from a seeded one and gets migrated.
  • The module lane is independent of sync and seal. RegistryUpdateReconciler.AdoptOne no longer applies the sync-owned hold ("its module waits for the same seal its content does", Registry auto-update and the seal reconciler both write a GitSynced partition with separate bookkeeping — Store on memex became a mix of 1.10.3 and 1.11.1 and Store/Catalog parked (CS1061) #4355 gate 1b). OwnershipDecline and SyncConvergedDecline are retired together with the tests that pinned them.
    • Only the package's own policy and, inside the adopt, the floor still decline.
    • No seal, framework-identity match, platform build, CD step or green build is consulted on the module path. The framework MVID is recorded, never a gate.
  • The same activation path as a reload. A wave that landed anything files ONE ModuleReload request: live swap if available, else exactly one automatic restart.
    • The request id is derived from the landed set, so replicas and re-runs file it once.
    • The executor now rides a restart another open request already requested ("exactly one" across waves). It reads candidates through a listing, then each node's content via GetMeshNodeStream (CQRS).
  • Reactivity is unchanged: boot, the ModulePublished broadcast (minutes after a publish), and the 30-minute safety net.

What still depends on something else — named, not hidden

Measured on memex.systemorph.com (read-only, search namespace:Plugins nodeType:Package)

  • 80 install records. 0 are not Auto: 75 declare updatePolicy: Auto, and 5 (Video, ContainerRegistry, PlatformUI, Skill, Agent) predate the field and carry autoUpdate: true.
  • content.updatePolicy:Notify → 0, :None → 0, content.autoUpdate:false → 0. Coverage: partition plugins, not truncated.
  • So the control instance's stuck modules were held by the sync-owned module hold, not by a reminder-only policy. That hold is what this commit removes.

Tests (all in Memex.Portal.Shared.Test)

  • PackagesAutoUpdateTest — the platform is fixed, with no seal, no platform build, and a sync source still at the OLD content:
    • Newer compatible 1.2.0 published → the migration moves the legacy notify record to Auto → ReconcileNow lands it → ONE request → exactly one restart, no human step.
    • 1.3.0 published before that restart → lands, and its request rides the same restart (Restarts == 1).
    • Negative: an administrator's deliberate Notify is kept and named; the same publication lands nothing and files nothing.
    • Floor above running: declined by name on the record (heldUpdate names floor and running), no download, no activation, 1.1.0 keeps serving.
  • PackagesAutoUpdateRulesTest — seed and migration truth tables with their controls.
  • Negative-control run: re-inserting a decline on the module lane turns the main acceptance test red.
  • Tests whose premise the policy retires were moved, not deleted:
    • the catalog policy tests now click the pin (None), since installs are Auto by default;
    • the reminder-idempotence test now uses a deliberate (stamped) Notify.
  • Full Memex.Portal.Shared.Test: 2303/2303 (1 skipped). MeshWeaver.Documentation.Test 670/670.
  • MeshWeaver.Plugins PluginCatalog.Test against this branch: 738 total, 737 passed. The one failure pinned the old default and is updated in CD: main 9fb74fd has an incomplete image set #2893; the two touched classes then pass 15/15.

Not established

  • No live roll or multi-replica run.
  • Two requests stamping their restarts in the same instant could still produce two restarts. The ride-along check closes the sequential case only.
  • Deliberate Notify choices made before this change carry no stamp and will be migrated (none exist on the control instance — see the measurement).

Third commit — two-phase package uninstall (policy package-uninstall-request, proposed)

Maintainer: "yes also uninstall", "removing the partitions on the server", "with corresponding confirmation by user." Manual: Doc/Architecture/PackageUninstall (new).

  • PackageUninstallRequest lives at Admin/_PackageUninstall/{id}. PackageUninstallExecutor runs on the request's own hub, with the same shape as the reload request.
  • Phase 1 needs no confirmation and destroys no data. It first refuses, by name and before anything is touched:
    • a package that is not installed here;
    • a target partition shared with another installed package;
    • a partition the platform never tears down (PartitionTeardown.Refusal);
    • a partition holding user data (any node whose createdBy is not the installer's System identity; the first five are named);
    • a partition larger than 20 000 nodes (it cannot be verified).
  • Phase 1 then acts, in order:
    1. Module: its landed generation is disabled, then it is unloaded live via the new IModuleLiveActivation.Retire (which has a default implementation, so existing implementers are untouched), otherwise exactly one restart, stamped first.
    2. Hubs: the partition's hosted hubs are closed with DisposeRequest from the off-router hub.
    3. Install record: removed.
    4. Re-install block: InstanceAutoRegistrationService.InstallAll now skips UninstalledHere — seed, baseline and feature-flag lanes alike. It is lifted when a person installs the package again.
    5. Preview: the request moves to AwaitingConfirmation with, per partition: whether the store/schema exists, rows per table (mesh_nodes plus each satellite segment), whether a sync config is present, and the things that cannot be counted, named.
  • Phase 2 runs only on confirmation. ConfirmationRequired is the partition name.
    • The requester confirms via PackageUninstall.Confirm; who and when are recorded.
    • A wrong string, or a confirmation from someone other than the requester, is refused by name and the data stays.
    • A match runs the platform's governed PartitionTeardown.TearDownPartition as System: the store is dropped on every provider (Postgres DROP SCHEMA … CASCADE, taking the satellite tables, _GitSync and the NodeType nodes with it), cached queries are evicted, and Admin/Partition/{p} is deleted. No raw SQL.
  • Surfaces: MeshOperations.UninstallPackage(package, reason) returns the preview and the exact confirmation string; UninstallPackage(requestPath, confirmation) confirms. Both are global-admin only. The MCP uninstall_package tool is in Plugins CD: main 9fb74fd has an incomplete image set #2893.

Tests (PackageUninstall*, 7 tests, all on the monolith's in-memory store):

  • Main scenario: phase 1 previews 2 mesh_nodes rows, triggers exactly one restart, removes the record and blocks re-install.
  • Negative controls in the same test:
    • no confirmation → the content still exists;
    • a wrong string → refused ('Plugins' does not match the required 'ReloadPkg');
    • a different confirmer → refused;
    • the requester's exact string → Done, content and root gone, the teardown sentence names Admin/Partition/ReloadPkg.
  • Re-seed: blocked while uninstalled; a re-installed record lifts the block (negative control).
  • Refusals: an unknown package by name; a shared partition with nothing touched; a partition holding user data with nothing touched.
  • Negative-control run: making phase 2 accept any confirmation turns the main test red.

Full Memex.Portal.Shared.Test 2310/2310 (1 skipped). Documentation guards 670/670.

Not established

  • The Store dialog was NOT rewired. Its DialogAction.Uninstall (Localizer.Uninstall) removes a viewer's localized course copies from their own space — not the platform package, module and partition. Extracting it into this engine would change what every learner's Uninstall does. A package-page entry for this request is owed.
  • No Postgres run. The schema-drop path is the platform's existing teardown and is not exercised here.
  • Hubs on other replicas are not closed by a DisposeRequest.
  • A phase 1 resumed after the record was removed fails by name.
  • Not run against production. First intended targets on memex.systemorph.com: BuildServer, ContainerRegistry, AzureCostManagement. Each may be refused by the user-data or shared-partition checks; that has not been measured.

🤖 Generated with Claude Code

…ompatible, live or one restart, reported per replica

ModuleReload request (Admin/_ModuleReload) executed on its own node hub: resolves the newest
published version whose declared floor the running platform meets (never a seal), lands it
through RegistryUpdateReconciler/PluginBundleClient/ModuleLandingService, activates it live via
IModuleLiveActivation (the live loader's call site) or exactly one restart through the
self-update restart path (SelfUpdateHostedService implements IModuleActivationRestart, no floor
deferral, no approval), and records what every replica loaded (ModuleReloadAgent). Surfaces:
MeshOperations.ReloadModule (global admin), the package page's Reload module entry (en+de).
Doc: Architecture/ModuleReload; policy module-reload-request (proposed).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@meshweaver-cloud
meshweaver-cloud Bot enabled auto-merge October 5, 2026 06:43
@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results

    22 files  ± 0      22 suites  ±0   42m 50s ⏱️ -29s
10 747 tests +30  10 554 ✅ +30  193 💤 ±0  0 ❌ ±0 
10 760 runs  +30  10 567 ✅ +30  193 💤 ±0  0 ❌ ±0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

This pull request removes 34 and adds 46 tests. Note that renamed tests count towards both.

   --- End of inner exception stack trace ---
   --- End of inner exception stack trace ---, expected: True)
   --- End of inner exception stack trace ---, isDenial: True)
 ---> (Inner Exception #1) MeshWeaver.Messaging.Hub.Test.InfrastructureFaultTest+ProviderException (0x80004005): Failed to connect to 10.42.18.4:5432<---
 ---> (Inner Exception #1) System.InvalidOperationException: boom<---
 ---> (Inner Exception #1) System.InvalidOperationException: source B is misconfigured<---
 ---> (Inner Exception #1) System.Net.Sockets.SocketException (0xFFFDFFFF): Name or service not known<---
 ---> MeshWeaver.Messaging.Hub.Test.InfrastructureFaultTest+ProviderException (0x80004005): Failed to connect to 10.42.18.4:5432
 ---> System.AggregateException: One or more errors occurred. (Failed to connect to 10.42.18.4:5432)
…
Memex.Portal.Shared.Test.CatalogOrphanActionIdentityTest ‑ ThePinChoiceInRowK_SetsPackageKsPolicy
Memex.Portal.Shared.Test.DrainEndpointTest ‑ ADrainProbeFromOutsideThePod_Reports_ButNeverBeginsTermination
Memex.Portal.Shared.Test.DrainEndpointTest ‑ TheFirstInPodDrainProbe_BeginsTermination_ForTheSingletonsToLeave
Memex.Portal.Shared.Test.InstanceIdRulesMatchTheRegistryTest ‑ TheSetupHostAgreesWithTheRegistry(candidate: "498117b7-0ce6-48f0-a6bc-d4ff952f5f8e")
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ AModuleThatIsNotInstalled_IsRefusedByName
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ ANewerVersionAboveTheFloor_IsDeclinedByName_AndTheRunningVersionKeepsServing
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ AReloadOfTheVersionAlreadyRunning_IsDoneWithoutARestart
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ AReplicaThatStillLoadsTheOldVersionAfterTheRestart_TurnsTheRequestRed
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ ARunningVersion_ReloadsToTheNewestCompatible_ByExactlyOneRestart
Memex.Portal.Shared.Test.ModuleReloadLiveTest ‑ AFailedLiveSwap_FallsBackToExactlyOneRestart
…

♻️ This comment has been updated with latest results.

…e seeded Notify, module lane independent of sync/seal, landed waves activate via ModuleReload

- SeedUpdatePolicy: a fresh install is Auto; deployment-wide defaults no longer consulted; a
  seeded reminder-only record re-stamps onto Auto, an administrator's stamped choice is kept.
- PackageAutoUpdateMigration (boot repair pass, stream.Update as System): seeded Notify -> Auto;
  pins and administrator choices kept and named. SetUpdatePolicy stamps UpdatePolicySetAt.
- RegistryUpdateReconciler: the module lane no longer takes the sync-owned hold (#4355 gate 1b
  retired for the module half); a wave that landed files ONE ModuleReload request (live, else
  exactly one restart); a second request rides a restart already on its way.
- Tests: PackagesAutoUpdate* (platform fixed, no seal, sync-owned partition at old content ->
  installed and activated; deliberate opt-out kept; incompatible floor declined by name);
  catalog/reminder tests moved to the per-package pin/deliberate-notify premise.
- Policy packages-auto-update (in force) + Doc ModuleReload -> Auto-update.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@rbuergi rbuergi changed the title feat(modules): reload request — reload M on an instance: newest compatible, live or exactly one restart, reported per replica feat(modules): reload request + policy packages-auto-update — newest compatible lands and activates (live or one restart), independent of seal/sync Oct 5, 2026
@rbuergi
rbuergi disabled auto-merge October 5, 2026 09:03
@rbuergi
rbuergi enabled auto-merge October 5, 2026 09:03
…bs, remove record, block re-install, preview; drop partition storage only on the requester's exact confirmation

PackageUninstallRequest (Admin/_PackageUninstall) executed on its own node hub. Phase 1 refuses
by name an uninstalled package, a shared partition, a never-torn-down partition, or one holding
user data; retires the module (IModuleLiveActivation.Retire, else exactly one restart), closes
the partition's hubs, removes the install record (InstallAll skips UninstalledHere) and records
the preview. Phase 2 runs PartitionTeardown.TearDownPartition as System only when the requester
confirms with the partition name. MeshOperations.UninstallPackage (global admin). Doc
Architecture/PackageUninstall; policy package-uninstall-request (proposed).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@rbuergi rbuergi changed the title feat(modules): reload request + policy packages-auto-update — newest compatible lands and activates (live or one restart), independent of seal/sync feat(modules): reload + two-phase uninstall requests, policy packages-auto-update — independent of seal/sync Oct 5, 2026
@systemorph-com systemorph-com Bot added the thread:pr-systemorph-meshweaver-6124-9860cd https://memex.systemorph.com/Hosting/Triage/_Thread/pr-systemorph-meshweaver-6124 label Oct 5, 2026
@rbuergi
rbuergi disabled auto-merge October 5, 2026 10:56
@rbuergi
rbuergi enabled auto-merge October 5, 2026 10:57
@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 0)

  3 files  ±0    3 suites  ±0   3m 29s ⏱️ -13s
565 tests +4  374 ✅ +4  191 💤 ±0  0 ❌ ±0 
569 runs  +4  378 ✅ +4  191 💤 ±0  0 ❌ ±0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 2)

    5 files  ±0      5 suites  ±0   4m 3s ⏱️ +45s
1 279 tests ±0  1 279 ✅ ±0  0 💤 ±0  0 ❌ ±0 
1 280 runs  ±0  1 280 ✅ ±0  0 💤 ±0  0 ❌ ±0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 1)

1 753 tests  ±0   1 751 ✅ ±0   5m 23s ⏱️ -17s
    4 suites ±0       2 💤 ±0 
    4 files   ±0       0 ❌ ±0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 3)

2 439 tests  +156   2 439 ✅ +156   5m 48s ⏱️ - 1m 13s
    4 suites ±  0       0 💤 ±  0 
    4 files   ±  0       0 ❌ ±  0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

This pull request removes 597 and adds 753 tests. Note that renamed tests count towards both.
Memex.Portal.Shared.Test.CatalogOrphanActionIdentityTest ‑ TheAutoChoiceInRowK_SetsPackageKsPolicy
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ ADecidedDuplicate_IsRetiredSoNoLaterSweepCanReReportIt
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ ADuplicateProposal_IsANotice_NeverAFault
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ ADuplicateWithAnUnreadableRecord_RetiresNothing
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ ALegacyFixedFolderEntry_IsNeverDeferredByTheMeshSet
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ AReplicaBootingMidWave_NeverLoadsAHalfLandedMix
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ AResolvedSetConflict_IsReportedWithoutMakingTheStateUndetermined
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ AWaveThatLandedNothing_ProposesNothing
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ AWaveThatNeverCompletes_LeavesTheMeshOnTheSetItWasOn
Memex.Portal.Shared.Test.ModuleSetConvergenceTest ‑ GarbageCollection_KeepsTheGenerationsTheMeshSetPins
…
Memex.Portal.Shared.Test.CatalogOrphanActionIdentityTest ‑ ThePinChoiceInRowK_SetsPackageKsPolicy
Memex.Portal.Shared.Test.DrainEndpointTest ‑ ADrainProbeFromOutsideThePod_Reports_ButNeverBeginsTermination
Memex.Portal.Shared.Test.DrainEndpointTest ‑ TheFirstInPodDrainProbe_BeginsTermination_ForTheSingletonsToLeave
Memex.Portal.Shared.Test.ModuleReloadLiveTest ‑ AFailedLiveSwap_FallsBackToExactlyOneRestart
Memex.Portal.Shared.Test.ModuleReloadLiveTest ‑ ASwappableModule_GoesLive_WithoutARestart
Memex.Portal.Shared.Test.ModuleSeedStampProducerTest ‑ TheClosureLaneStamp_IsReadBackByTheBootReader_AndAStaleOneIsRemoved
Memex.Portal.Shared.Test.ModuleSetStoreReadsOnlyDecidingRecordsTest ‑ ACorruptAdoptionFallsThroughToTheNextAdoptedSequence
Memex.Portal.Shared.Test.ModuleSetStoreReadsOnlyDecidingRecordsTest ‑ ACorruptSupersededRecordIsNeverOpened_SoItIsNeverReported
Memex.Portal.Shared.Test.ModuleSetStoreReadsOnlyDecidingRecordsTest ‑ ANewestSequenceThatCannotBeRead_FallsToTheNextReadableOne_AndReportsExactlyThat
Memex.Portal.Shared.Test.ModuleSetStoreReadsOnlyDecidingRecordsTest ‑ ARecordUnderAForeignName_IsStillRead
…

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 5)

    2 files  ±0      2 suites  ±0   8m 23s ⏱️ -16s
1 105 tests ±0  1 105 ✅ ±0  0 💤 ±0  0 ❌ ±0 
1 106 runs  ±0  1 106 ✅ ±0  0 💤 ±0  0 ❌ ±0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 4)

    4 files  ±  0      4 suites  ±0   15m 41s ⏱️ +44s
3 606 tests  - 130  3 606 ✅  - 130  0 💤 ±0  0 ❌ ±0 
3 613 runs   - 130  3 613 ✅  - 130  0 💤 ±0  0 ❌ ±0 

Results for commit 8dc35c6. ± Comparison against base commit 8b28804.

This pull request removes 768 and adds 620 tests. Note that renamed tests count towards both.

   --- End of inner exception stack trace ---
   --- End of inner exception stack trace ---, expected: True)
   --- End of inner exception stack trace ---, isDenial: True)
 ---> (Inner Exception #1) MeshWeaver.Messaging.Hub.Test.InfrastructureFaultTest+ProviderException (0x80004005): Failed to connect to 10.42.18.4:5432<---
 ---> (Inner Exception #1) System.InvalidOperationException: boom<---
 ---> (Inner Exception #1) System.InvalidOperationException: source B is misconfigured<---
 ---> (Inner Exception #1) System.Net.Sockets.SocketException (0xFFFDFFFF): Name or service not known<---
 ---> MeshWeaver.Messaging.Hub.Test.InfrastructureFaultTest+ProviderException (0x80004005): Failed to connect to 10.42.18.4:5432
 ---> System.AggregateException: One or more errors occurred. (Failed to connect to 10.42.18.4:5432)
…
Memex.Portal.Shared.Test.InstanceIdRulesMatchTheRegistryTest ‑ TheSetupHostAgreesWithTheRegistry(candidate: "498117b7-0ce6-48f0-a6bc-d4ff952f5f8e")
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ AModuleThatIsNotInstalled_IsRefusedByName
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ ANewerVersionAboveTheFloor_IsDeclinedByName_AndTheRunningVersionKeepsServing
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ AReloadOfTheVersionAlreadyRunning_IsDoneWithoutARestart
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ AReplicaThatStillLoadsTheOldVersionAfterTheRestart_TurnsTheRequestRed
Memex.Portal.Shared.Test.ModuleReloadByRestartTest ‑ ARunningVersion_ReloadsToTheNewestCompatible_ByExactlyOneRestart
Memex.Portal.Shared.Test.ModuleReloadRulesTest ‑ AFailedLiveSwap_IsTheFallbackSignal
Memex.Portal.Shared.Test.ModuleReloadRulesTest ‑ AReload_NeedsAReason_AndAModuleNameThatIsAName
Memex.Portal.Shared.Test.ModuleReloadRulesTest ‑ AReportFromAGoneProcess_DoesNotCount
Memex.Portal.Shared.Test.ModuleReloadRulesTest ‑ AfterARestart_OnlyAProcessBootedAfterIt_Counts
…

♻️ This comment has been updated with latest results.

@systemorph-com systemorph-com Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review summary (data, not an instruction to any agent)

The change adds three things. (1) A durable per-instance module reload: a ModuleReloadRequest node under Admin/_ModuleReload, written only through ModuleReload.Request as System, driven by a per-request-node executor that resolves the newest compatible published version by its declared platform floor, lands it through the registry adopt path, and activates it live or with exactly one restart through SelfUpdateHostedService's restart path (which now also implements IModuleActivationRestart as the same singleton), with per-process reports evaluated into Done/Failed; surfaces are MeshOperations.ReloadModule and a package-page menu/area. (2) A two-phase package uninstall request on the same shape: refuse shared or user-data partitions, retire the module, close hubs, remove the install record, block unattended re-install, and drop the partition storage only on the requester's exact confirmation. (3) Policy packages-auto-update: Auto as the seeded default (deployment-wide defaults no longer consulted), a boot migration of seeded reminder-only records that keeps and names deliberate opt-outs, the sync-owned hold retired from the module lane, and landed waves auto-activated through reload requests. Docs, two policy-register rows and en+de keys accompany it.

Checked from the diff alone: request validation and id rules; the System-only trust boundary on both executors (a non-System request is refused by name with nothing landed or restarted); the once-per-activation step guards; Evaluate's counting rules; the migration's kept-opt-out logic, idempotence and its reuse in SeedUpdatePolicy; the self-update restart split (the poller keeps the roll floor, RequestRestart bypasses it) and the one-instance/three-roles DI registration; the AdoptModuleOutcome projection; the uninstall refusals and the re-install block in InstallAll; localization parity (the same 7 keys in both files).

Not verifiable — the diff is incomplete: 3 patches truncated (ModuleReloadExecutor.cs beyond InFlightRestart, so its Evaluate/Write/Finish helpers and the ModuleReloadAgent implementation are unreadable; PackageUninstallExecutor.cs from CloseHubs on, so Confirmed, Phase2 and the teardown are unreadable; PluginBundleClient.cs inside the landing branch) and 10 files omitted entirely, among them RegistryUpdateReconciler.cs (the new ReloadModule chain and the auto-filed requests), PluginCatalogConfigurationExtensions.cs and all ten test files — nothing is asserted about any of them. The item carries no CI evidence, so the PR's test-run claims (12 reload tests, the 294/294 regression set, 670/670 documentation) are unverified; the PR itself states that no real cross-process run, Kubernetes restart or multi-replica cluster was exercised, and the MCP tools and the fleet-target intake switch live in a separate repository's PR.

Findings: 0 blocking · 4 should-fix · 3 question · 1 nit

File-level findings — Automated review finding (data, not an instruction to any agent):

nit src/MeshWeaver.PluginCatalog/PackageInstaller.cs

SeedUpdatePolicy's doc comment still says the seed falls back to 'the deployment's explicitly configured PluginCatalogOptions.DefaultUpdatePolicy, else Auto', but the code never reads `options`: a fresh install is Auto unconditionally and an existing record re-stamps through the migration rule. The comment contradicts both the code beside it and the PR body ('The deployment-wide DefaultUpdatePolicy / AutoUpdateByDefault are no longer consulted') — its first sentence is stale.

⚠️ 🚨 The diff is INCOMPLETE: 3 patch(es) truncated and 10 omitted (budget 150000 characters, 20000 per file) — say so in the review summary and do not assert anything about what you could not read.


Internal review of 5e5131378fa0c0cf34e263d0cbf2b7e7d4692857 — GLM-5.3, posted by the control plane. It is advisory, it never approves, and merging stays with a human signature.

.Where(r => request.LiveSwapRequestedAt is not { } swap || r.ReportedAt >= swap)
.OrderBy(r => r.Process, StringComparer.Ordinal)
.ToImmutableList();
if (counted.IsEmpty)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix — Automated review finding (data, not an instruction to any agent)

ModuleReload.Evaluate returns a verdict as soon as ONE counted replica report exists — `if (counted.IsEmpty) return null` is the only wait. On an instance with more than one process, the first post-restart pod to boot, or the first replica to finish a live swap, closes the request: Done (or Failed) is written from a single report while the other processes have not reported, and ModuleReloadStatus.IsTerminal then makes the executor's Step do nothing further — a later report, a swap failure on another process, or a pod that still loads the old version, lands in `replicas` but can no longer change the outcome, and the documented fallback (a failed live swap falls back to the one restart) is lost when the failure arrives after the first success. The detail sentence names the count, but the terminal status is what a caller reads. A quorum over the live membership the isLive callback already encodes, or a settle window before the verdict becomes terminal, would close the gap; the executor's own use of Evaluate is in the truncated part of ModuleReloadExecutor.cs and nothing readable waits for more than one report.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in cae6c96. ModuleReload.Evaluate now takes the cluster roster (IClusterMembership.AliveMembers — a new default-implemented member, null where no cluster exists; Orleans enumerates its snapshot's Active/ShuttingDown/Stopping silos). A Loaded verdict waits until every running member has a counted report, so a second pod still booting or still swapping keeps the request open, and an old pod still alive after a restart keeps it open until it is gone. A swap failure or a version mismatch on any counted report stays decisive at once, so the restart fallback is still taken when the failure is the second report rather than the first. The executor now also re-evaluates on a membership change (IClusterMembershipFeed, with StartWith), because no node write accompanies "the old pod is gone". The fallback write is guarded so it starts only once. With no roster (monolith) the counted reports are all there is, as before. Pinned by ModuleReloadRulesTest.Loaded_WaitsForEveryRunningProcess_ButAFailureIsDecisiveAtOnce, which covers both the roster and no-roster cases. ModuleReload.md (step 5) states the rule.

: Gate(() => PackageUninstall.Confirm(hub, requestPath.Trim(), confirmation,
string.IsNullOrWhiteSpace(caller) ? null : caller)
.SelectMany(ticket => ticket.Accepted
? AwaitUninstall(ticket.Path!, r => PackageUninstallStatus.IsTerminal(r.Status) || r.ConfirmationRefusal is not null)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix — Automated review finding (data, not an instruction to any agent)

PackageUninstall.md promises that 'a different string, or a confirmation from anyone but the requester, is refused by name', and PackageUninstall.Confirm's own doc comment assigns exactly that identity check to the caller ('The CALLER checks that confirmedBy is the requester and a platform admin; the executor checks the string against ConfirmationRequired and refuses a mismatch by name'). The readable caller performs only the admin half: UninstallPackage gates on IsGlobalAdmin and passes the caller straight into Confirm as confirmedBy, never comparing it with the request's RequestedBy — so any platform admin, not only the requester, can confirm the irreversible partition drop. Either the truncated Confirmed step enforces the identity check despite its documented contract (in which case that contract is wrong), or the documented requester-only rule holds nowhere in readable code.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The identity check was already enforced, but in the executor and not the caller, and the caller-side doc comment was wrong. PackageUninstallExecutor.Confirmed refuses a ConfirmedBy that is not RequestedBy, and PackageUninstallTest covers it with the "someone-else" case. It had one hole: a request with a null RequestedBy skipped the check, so any admin could confirm it. Fixed in cae6c96:

  • A request that names no requester can't be confirmed by anyone (ARequestNamingNoRequester_CannotBeConfirmedByAnyone).
  • The Confirm doc comment now says what the code does: the caller gates on IsGlobalAdmin, and the executor is the authority on the string, the requester's identity and the ordering.
  • PackageUninstall.md matches.

return Observable.Return(new PackageUninstallTicket(null, $"'{path}' is not a package uninstall request"));
var access = hub.ServiceProvider.GetRequiredService<AccessService>();
return access.RunAsSystem(() => hub.GetMeshNodeStream(path)
.Update<PackageUninstallRequest>(current => current with

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix — Automated review finding (data, not an instruction to any agent)

Confirm writes the confirmation regardless of the request's state: the Update fold sets Confirmation/ConfirmedBy/ConfirmedAt without checking the status, and the executor routes 'AwaitingConfirmation when request.Confirmation is not null' straight to Confirmed, with phase 1's completing Write carrying a pre-existing Confirmation through the `with`. The request path exists from the moment it is filed, so a confirmation written while phase 1 is still running — retiring the module can take minutes and may itself request a restart — fires phase 2, the irreversible drop, the moment phase 1 lands, without the preview that the two-phase design exists to deliver ever having been answered; nothing on the record records when AwaitingConfirmation was reached, so the ordering cannot be re-checked later. Refusing a confirmation unless the request already awaits one restores the order the policy register row states.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in cae6c96, at both ends:

  • PackageUninstall.Confirm reads the request first and records nothing unless its status is AwaitingConfirmation. Otherwise it returns a named refusal.
  • To close the race between that read and phase 1's completion write, phase 1's completing write clears any Confirmation/ConfirmedBy/ConfirmedAt it finds and stamps a new AwaitingConfirmationAt. Confirmed refuses a confirmation whose ConfirmedAt is missing or earlier than that stamp.

The order is now on the record and can be checked again later. Pinned by the extended AnUnknownPackage_IsRefusedByName (an early confirmation is refused and nothing is recorded) and by the preview stamp assertion in ARequestNamingNoRequester_CannotBeConfirmedByAnyone. PackageUninstall.md states the rule.

.Select(n => (Path: p, n?.CreatedBy))
.Catch((Exception _) => Observable.Return((Path: p, CreatedBy: (string?)null))))
.MergeBounded(8)
.ToList()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix — Automated review finding (data, not an instruction to any agent)

The user-data refusal in Measure fails open on read errors: a node whose storage.Read faults becomes (Path, CreatedBy: null) through the Catch, IsUserCreated(null) is false, and the node is silently counted as installer-owned — no log line records the fault. A transient fault in the measurement window therefore defeats exactly the check that gates phase 2's irreversible DROP. The size cap fails closed (a partition above MaxInspectedNodes is refused), and an absent node legitimately reads as null, but a fault is not an absence; counting read faults — naming them in the preview, or refusing once they occur — keeps the check conservative in the direction the rest of phase 1 takes.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed: a fault is not an absence. Fixed in cae6c96. Measure now records each read fault next to its path. If any node under the partition faults, phase 1 is refused with a named reason ("N node(s) under 'P' could not be read (e.g. path: message, …) — the partition is not verified free of user data"), and nothing is touched. That is the same closed-fail direction as the size cap. An absent node still reads as null and counts as before. The refusal goes through Finish, which logs it at Warning, so a fault no longer passes silently. PackageUninstall.md lists the new refusal.

// again — the "exactly one restart" property lives on the node, not in this process.
: Write(r => r with { RestartRequestedAt = DateTimeOffset.UtcNow }, Line("requesting ONE restart"))
.SelectMany(_ => restart.RequestRestart(reason).Take(1))
.SelectMany(outcome => outcome.Scheduled

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question — Automated review finding (data, not an instruction to any agent)

IssueRestart stamps RestartRequestedAt first and only then calls IModuleActivationRestart.RequestRestart, and Step routes a stamped AwaitingRestart request to Evaluate, never to IssueRestart again — so a teardown or crash between the durable stamp write and the restart actually being handed over strands the request in AwaitingRestart permanently: the landed generation never activates, nothing retries, and the request's own status is the only signal. The documentation concedes the analogous stranded state for a control-lane restart that never executes ('visible, not retried') but says nothing about this window, in which no restart was ever requested at all; whether it is accepted, or a stale AwaitingRestart should re-ask or escalate after a bound, is stated nowhere in the change.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Accepted. It's the same class as the stated control-lane case, and this change doesn't retry it automatically. The stamp comes first on purpose: the node, not the process, holds "exactly one restart". Re-asking after a gap would turn a crash in that window into a possible second roll. A second roll is the outcome the stamp exists to prevent, and nothing on the record can tell "stamped, never asked" apart from "asked, restart still in flight".

A stranded request stays visible: AwaitingRestart with RestartRequestedAt set and no ActivationDetail naming a restart outcome. A person, or the next reload, recovers it. A later request's InFlightRestart rides the stamp only while this process has not been through a restart since it, so the next real restart activates the generation anyway.

Bounded escalation of a stale AwaitingRestart is a reasonable follow-up, but it changes the protocol, not this fix, so I've left it out of this change.

public IObservable<ModuleRestartOutcome> RequestRestart(string reason) =>
Restart(new SelfUpdateVerdict(SelfUpdateOutcome.NoNewerRelease, $"Module reload ({reason}):"),
honourFloor: false, reason: reason)
.Select(verdict => new ModuleRestartOutcome(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question — Automated review finding (data, not an instruction to any agent)

RequestRestart passes honourFloor: false, documented as serving an authorised one-shot request — but the auto-update lane reaches the same call: ModuleReload.md has every pass that landed anything file ONE ModuleReload request, whose executor takes this restart path, and InFlightRestart rides only a restart the process has not yet been through. Consecutive publishes minutes apart can therefore each trigger a full instance roll as soon as the previous restart has completed, unpaced, while the poller's own self-update restarts stay paced by MinRollInterval. Whether the roll floor should also be bypassed for AUTO-filed requests, as opposed to explicit ones, is stated nowhere in the change.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This change keeps the floor bypass, and here is why. The pacing for AUTO-filed requests sits one level up. The auto pass files ONE request per landed set (auto-{hash}, so replicas and re-runs file it once), and InFlightRestart makes a later request ride a restart this process has not been through yet. So consecutive publishes only cause a second roll after the first one has actually completed, and each such roll activates code that is not running yet.

Honouring MinRollInterval here would turn "a landed module is pending" into a deferral that nothing re-tries: the request would sit in AwaitingRestart with nothing scheduled to ask again, which is the stranded state the previous thread is about. Pacing auto-filed reloads separately (honour the floor for RequestedBy: null and re-ask on the poller's cadence) needs that re-ask loop first. It's a design change and not part of this fix.

status = r!.Status,
path,
package = r.Package,
module = r.Module,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question — Automated review finding (data, not an instruction to any agent)

The new code uses the null-forgiving operator to silence nullability diagnostics at several guarded sites — `r!.Status` here after Where(r => r is not null && until(r)), `ticket.Path!` in UninstallPackage, `measured.Preview!` and `r!.Package` in PackageUninstallExecutor, `m.Module!.Trim()` in ModuleReloadExecutor. Every use follows a visible guard, so none hides a real null path; but the repository's hard rules ban `!` used to silence a warning, and whether that ban is categorical — these would need restructuring into pattern-matched locals or helper results — or sanctions the guard idiom, is not stated anywhere readable here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The ban is categorical, so I removed them all in cae6c96. Each site now uses a pattern-matched local:

  • ticket is { Accepted: true, Path: { } … } in UninstallPackage
  • .OfType<PackageUninstallRequest>() in AwaitUninstall and UninstalledHere
  • measured is not { Preview: { } preview } in phase 1
  • a SelectMany over m.Module is { } module in ModuleReloadExecutor
  • RestartRequestedAt: { } at carried through as a tuple in InFlightRestart

The test helpers in PackageUninstallTest were rewritten the same way. A grep of this PR's added src/ lines for null-forgiving operators now finds none.

…ll confirmation only answers the preview, from the named requester; a measurement fault refuses

Answers the automated review of #6124:
- ModuleReload.Evaluate: Loaded is a verdict over the whole roster (IClusterMembership.AliveMembers,
  default null = no roster) — a pod still booting/swapping, or an old pod still up after a restart,
  keeps the request open; a swap failure or mismatch stays decisive at once. The executor re-evaluates
  on a membership change (IClusterMembershipFeed) and starts the swap fallback once.
- PackageUninstall.Confirm records only while the request awaits a confirmation; phase 1 clears any
  earlier one and stamps AwaitingConfirmationAt; phase 2 refuses a confirmation older than the
  preview and a request that names no requester (the requester check now holds unconditionally).
- PackageUninstallExecutor.Measure: a storage read fault is not an absence — any fault refuses
  phase 1 by name (fails closed like the size cap).
- Null-forgiving operators added by this PR replaced with pattern-matched locals.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@systemorph-com

systemorph-com Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

🚰 PR babysitter (build instance) is merging the base into this branch on head cae6c96ac921 — once per head.

Why: inherited from its base 'main': 8 pull requests of Systemorph/MeshWeaver fail identically — 'lane / Automatic review answered' concluded failure: Process completed with exit code 1. — the base 'main' moved from 00a03cc (what the red run tested) to 60aa837. This pull request was red because its BASE was; the base has moved since, and a re-run would test the old merge commit again.

Validated: 'lane / Automatic review answered' is red on run 37312589214, a head that tested main at 00a03cc; main has since moved to 60aa837 and its newest run of this check is green (established by today's executed update-branch decisions for #6097, #6115, #6132, #6139 and #6141 on this same fingerprint) — the red is the base's at the time, not this diff's. Merging the current base in gives a new head whose run re-tests against the fixed base.

It does not merge the pull request, push anything else or dequeue. A red after this is left for the owner (rbuergi).

systemorph-com Bot and others added 4 commits October 5, 2026 13:11
# Conflicts:
#	memex/Memex.Portal.Shared/SelfUpdate/SelfUpdateHostedService.cs
…e — main's live swap now owns ModuleSwapOutcome

Merging main brought MeshWeaver.Graph.Configuration.ModuleSwapOutcome (module, kind, reason)
from the live-swap slice; this branch's PluginCatalog seam declared a second ModuleSwapOutcome
(swapped, failure), which shadowed main's inside MeshWeaver.PluginCatalog and broke
ModuleLiveActivation (CS1729/CS1061/CS0266). The seam keeps its two-field answer under its own
name; the merge's SelfUpdateHostedService conflict keeps both callers' sentences (a live swap that
did not go live, an explicit reload) by passing the whole announcement reason.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… — the boot migration moves an unstamped one to Auto

Shard 4 red (run 37327560786): RegistryDownPastTheBudget_… timed out waiting for the
'Update available' reminder. The log shows why: 'auto-update of ReconcileDeferralPkg
failed' — the record was seeded with AutoUpdate=false and no UpdatePolicySetAt, which
this PR's PackageAutoUpdateMigration (policy packages-auto-update) moves to Auto in
the boot repair pass. Whether that pass landed before or after the drain decided
whether the test saw a reminder or a failed auto-update. Same adaptation the PR
already made to PackageUpdateReminderIdempotenceTest.

Memex.Portal.Shared.Test Release -warnaserror clean; deferral + reload + reminder +
auto-update classes 20/20.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@rbuergi
rbuergi merged commit c13c7c4 into main Oct 5, 2026
44 checks passed
rbuergi added a commit that referenced this pull request Oct 5, 2026
feat(reboot): one Reboot operation — sync, land, newest image, ONE roll/restart, verify; plus a rate-limited wedge watchdog (stacked on #6124)
rbuergi added a commit that referenced this pull request Oct 5, 2026
test(ladder): the P/M matrix — ordinary + control walked step by step, two modules, every step checked (stacked on #6124)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

thread:pr-systemorph-meshweaver-6124-9860cd https://memex.systemorph.com/Hosting/Triage/_Thread/pr-systemorph-meshweaver-6124

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant