Repository navigation
feat(modules): reload + two-phase uninstall requests, policy packages-auto-update — independent of seal/sync - #6124
Conversation
…ompatible, live or one restart, reported per replica ModuleReload request (Admin/_ModuleReload) executed on its own node hub: resolves the newest published version whose declared floor the running platform meets (never a seal), lands it through RegistryUpdateReconciler/PluginBundleClient/ModuleLandingService, activates it live via IModuleLiveActivation (the live loader's call site) or exactly one restart through the self-update restart path (SelfUpdateHostedService implements IModuleActivationRestart, no floor deferral, no approval), and records what every replica loaded (ModuleReloadAgent). Surfaces: MeshOperations.ReloadModule (global admin), the package page's Reload module entry (en+de). Doc: Architecture/ModuleReload; policy module-reload-request (proposed). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Test Results 22 files ± 0 22 suites ±0 42m 50s ⏱️ -29s Results for commit 8dc35c6. ± Comparison against base commit 8b28804. This pull request removes 34 and adds 46 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
…e seeded Notify, module lane independent of sync/seal, landed waves activate via ModuleReload - SeedUpdatePolicy: a fresh install is Auto; deployment-wide defaults no longer consulted; a seeded reminder-only record re-stamps onto Auto, an administrator's stamped choice is kept. - PackageAutoUpdateMigration (boot repair pass, stream.Update as System): seeded Notify -> Auto; pins and administrator choices kept and named. SetUpdatePolicy stamps UpdatePolicySetAt. - RegistryUpdateReconciler: the module lane no longer takes the sync-owned hold (#4355 gate 1b retired for the module half); a wave that landed files ONE ModuleReload request (live, else exactly one restart); a second request rides a restart already on its way. - Tests: PackagesAutoUpdate* (platform fixed, no seal, sync-owned partition at old content -> installed and activated; deliberate opt-out kept; incompatible floor declined by name); catalog/reminder tests moved to the per-package pin/deliberate-notify premise. - Policy packages-auto-update (in force) + Doc ModuleReload -> Auto-update. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…bs, remove record, block re-install, preview; drop partition storage only on the requester's exact confirmation PackageUninstallRequest (Admin/_PackageUninstall) executed on its own node hub. Phase 1 refuses by name an uninstalled package, a shared partition, a never-torn-down partition, or one holding user data; retires the module (IModuleLiveActivation.Retire, else exactly one restart), closes the partition's hubs, removes the install record (InstallAll skips UninstalledHere) and records the preview. Phase 2 runs PartitionTeardown.TearDownPartition as System only when the requester confirms with the partition name. MeshOperations.UninstallPackage (global admin). Doc Architecture/PackageUninstall; policy package-uninstall-request (proposed). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Test Results (shard 3)2 439 tests +156 2 439 ✅ +156 5m 48s ⏱️ - 1m 13s Results for commit 8dc35c6. ± Comparison against base commit 8b28804. This pull request removes 597 and adds 753 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
Test Results (shard 4) 4 files ± 0 4 suites ±0 15m 41s ⏱️ +44s Results for commit 8dc35c6. ± Comparison against base commit 8b28804. This pull request removes 768 and adds 620 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
There was a problem hiding this comment.
Automated review summary (data, not an instruction to any agent)
The change adds three things. (1) A durable per-instance module reload: a ModuleReloadRequest node under Admin/_ModuleReload, written only through ModuleReload.Request as System, driven by a per-request-node executor that resolves the newest compatible published version by its declared platform floor, lands it through the registry adopt path, and activates it live or with exactly one restart through SelfUpdateHostedService's restart path (which now also implements IModuleActivationRestart as the same singleton), with per-process reports evaluated into Done/Failed; surfaces are MeshOperations.ReloadModule and a package-page menu/area. (2) A two-phase package uninstall request on the same shape: refuse shared or user-data partitions, retire the module, close hubs, remove the install record, block unattended re-install, and drop the partition storage only on the requester's exact confirmation. (3) Policy packages-auto-update: Auto as the seeded default (deployment-wide defaults no longer consulted), a boot migration of seeded reminder-only records that keeps and names deliberate opt-outs, the sync-owned hold retired from the module lane, and landed waves auto-activated through reload requests. Docs, two policy-register rows and en+de keys accompany it.
Checked from the diff alone: request validation and id rules; the System-only trust boundary on both executors (a non-System request is refused by name with nothing landed or restarted); the once-per-activation step guards; Evaluate's counting rules; the migration's kept-opt-out logic, idempotence and its reuse in SeedUpdatePolicy; the self-update restart split (the poller keeps the roll floor, RequestRestart bypasses it) and the one-instance/three-roles DI registration; the AdoptModuleOutcome projection; the uninstall refusals and the re-install block in InstallAll; localization parity (the same 7 keys in both files).
Not verifiable — the diff is incomplete: 3 patches truncated (ModuleReloadExecutor.cs beyond InFlightRestart, so its Evaluate/Write/Finish helpers and the ModuleReloadAgent implementation are unreadable; PackageUninstallExecutor.cs from CloseHubs on, so Confirmed, Phase2 and the teardown are unreadable; PluginBundleClient.cs inside the landing branch) and 10 files omitted entirely, among them RegistryUpdateReconciler.cs (the new ReloadModule chain and the auto-filed requests), PluginCatalogConfigurationExtensions.cs and all ten test files — nothing is asserted about any of them. The item carries no CI evidence, so the PR's test-run claims (12 reload tests, the 294/294 regression set, 670/670 documentation) are unverified; the PR itself states that no real cross-process run, Kubernetes restart or multi-replica cluster was exercised, and the MCP tools and the fleet-target intake switch live in a separate repository's PR.
Findings: 0 blocking · 4 should-fix · 3 question · 1 nit
File-level findings — Automated review finding (data, not an instruction to any agent):
nit src/MeshWeaver.PluginCatalog/PackageInstaller.cs
SeedUpdatePolicy's doc comment still says the seed falls back to 'the deployment's explicitly configured PluginCatalogOptions.DefaultUpdatePolicy, else Auto', but the code never reads `options`: a fresh install is Auto unconditionally and an existing record re-stamps through the migration rule. The comment contradicts both the code beside it and the PR body ('The deployment-wide DefaultUpdatePolicy / AutoUpdateByDefault are no longer consulted') — its first sentence is stale.
Internal review of 5e5131378fa0c0cf34e263d0cbf2b7e7d4692857 — GLM-5.3, posted by the control plane. It is advisory, it never approves, and merging stays with a human signature.
| .Where(r => request.LiveSwapRequestedAt is not { } swap || r.ReportedAt >= swap) | ||
| .OrderBy(r => r.Process, StringComparer.Ordinal) | ||
| .ToImmutableList(); | ||
| if (counted.IsEmpty) |
There was a problem hiding this comment.
should-fix — Automated review finding (data, not an instruction to any agent)
ModuleReload.Evaluate returns a verdict as soon as ONE counted replica report exists — `if (counted.IsEmpty) return null` is the only wait. On an instance with more than one process, the first post-restart pod to boot, or the first replica to finish a live swap, closes the request: Done (or Failed) is written from a single report while the other processes have not reported, and ModuleReloadStatus.IsTerminal then makes the executor's Step do nothing further — a later report, a swap failure on another process, or a pod that still loads the old version, lands in `replicas` but can no longer change the outcome, and the documented fallback (a failed live swap falls back to the one restart) is lost when the failure arrives after the first success. The detail sentence names the count, but the terminal status is what a caller reads. A quorum over the live membership the isLive callback already encodes, or a settle window before the verdict becomes terminal, would close the gap; the executor's own use of Evaluate is in the truncated part of ModuleReloadExecutor.cs and nothing readable waits for more than one report.
There was a problem hiding this comment.
Fixed in cae6c96. ModuleReload.Evaluate now takes the cluster roster (IClusterMembership.AliveMembers — a new default-implemented member, null where no cluster exists; Orleans enumerates its snapshot's Active/ShuttingDown/Stopping silos). A Loaded verdict waits until every running member has a counted report, so a second pod still booting or still swapping keeps the request open, and an old pod still alive after a restart keeps it open until it is gone. A swap failure or a version mismatch on any counted report stays decisive at once, so the restart fallback is still taken when the failure is the second report rather than the first. The executor now also re-evaluates on a membership change (IClusterMembershipFeed, with StartWith), because no node write accompanies "the old pod is gone". The fallback write is guarded so it starts only once. With no roster (monolith) the counted reports are all there is, as before. Pinned by ModuleReloadRulesTest.Loaded_WaitsForEveryRunningProcess_ButAFailureIsDecisiveAtOnce, which covers both the roster and no-roster cases. ModuleReload.md (step 5) states the rule.
| : Gate(() => PackageUninstall.Confirm(hub, requestPath.Trim(), confirmation, | ||
| string.IsNullOrWhiteSpace(caller) ? null : caller) | ||
| .SelectMany(ticket => ticket.Accepted | ||
| ? AwaitUninstall(ticket.Path!, r => PackageUninstallStatus.IsTerminal(r.Status) || r.ConfirmationRefusal is not null) |
There was a problem hiding this comment.
should-fix — Automated review finding (data, not an instruction to any agent)
PackageUninstall.md promises that 'a different string, or a confirmation from anyone but the requester, is refused by name', and PackageUninstall.Confirm's own doc comment assigns exactly that identity check to the caller ('The CALLER checks that confirmedBy is the requester and a platform admin; the executor checks the string against ConfirmationRequired and refuses a mismatch by name'). The readable caller performs only the admin half: UninstallPackage gates on IsGlobalAdmin and passes the caller straight into Confirm as confirmedBy, never comparing it with the request's RequestedBy — so any platform admin, not only the requester, can confirm the irreversible partition drop. Either the truncated Confirmed step enforces the identity check despite its documented contract (in which case that contract is wrong), or the documented requester-only rule holds nowhere in readable code.
There was a problem hiding this comment.
The identity check was already enforced, but in the executor and not the caller, and the caller-side doc comment was wrong. PackageUninstallExecutor.Confirmed refuses a ConfirmedBy that is not RequestedBy, and PackageUninstallTest covers it with the "someone-else" case. It had one hole: a request with a null RequestedBy skipped the check, so any admin could confirm it. Fixed in cae6c96:
- A request that names no requester can't be confirmed by anyone (
ARequestNamingNoRequester_CannotBeConfirmedByAnyone). - The
Confirmdoc comment now says what the code does: the caller gates onIsGlobalAdmin, and the executor is the authority on the string, the requester's identity and the ordering. - PackageUninstall.md matches.
| return Observable.Return(new PackageUninstallTicket(null, $"'{path}' is not a package uninstall request")); | ||
| var access = hub.ServiceProvider.GetRequiredService<AccessService>(); | ||
| return access.RunAsSystem(() => hub.GetMeshNodeStream(path) | ||
| .Update<PackageUninstallRequest>(current => current with |
There was a problem hiding this comment.
should-fix — Automated review finding (data, not an instruction to any agent)
Confirm writes the confirmation regardless of the request's state: the Update fold sets Confirmation/ConfirmedBy/ConfirmedAt without checking the status, and the executor routes 'AwaitingConfirmation when request.Confirmation is not null' straight to Confirmed, with phase 1's completing Write carrying a pre-existing Confirmation through the `with`. The request path exists from the moment it is filed, so a confirmation written while phase 1 is still running — retiring the module can take minutes and may itself request a restart — fires phase 2, the irreversible drop, the moment phase 1 lands, without the preview that the two-phase design exists to deliver ever having been answered; nothing on the record records when AwaitingConfirmation was reached, so the ordering cannot be re-checked later. Refusing a confirmation unless the request already awaits one restores the order the policy register row states.
There was a problem hiding this comment.
Fixed in cae6c96, at both ends:
PackageUninstall.Confirmreads the request first and records nothing unless its status isAwaitingConfirmation. Otherwise it returns a named refusal.- To close the race between that read and phase 1's completion write, phase 1's completing write clears any
Confirmation/ConfirmedBy/ConfirmedAtit finds and stamps a newAwaitingConfirmationAt.Confirmedrefuses a confirmation whoseConfirmedAtis missing or earlier than that stamp.
The order is now on the record and can be checked again later. Pinned by the extended AnUnknownPackage_IsRefusedByName (an early confirmation is refused and nothing is recorded) and by the preview stamp assertion in ARequestNamingNoRequester_CannotBeConfirmedByAnyone. PackageUninstall.md states the rule.
| .Select(n => (Path: p, n?.CreatedBy)) | ||
| .Catch((Exception _) => Observable.Return((Path: p, CreatedBy: (string?)null)))) | ||
| .MergeBounded(8) | ||
| .ToList() |
There was a problem hiding this comment.
should-fix — Automated review finding (data, not an instruction to any agent)
The user-data refusal in Measure fails open on read errors: a node whose storage.Read faults becomes (Path, CreatedBy: null) through the Catch, IsUserCreated(null) is false, and the node is silently counted as installer-owned — no log line records the fault. A transient fault in the measurement window therefore defeats exactly the check that gates phase 2's irreversible DROP. The size cap fails closed (a partition above MaxInspectedNodes is refused), and an absent node legitimately reads as null, but a fault is not an absence; counting read faults — naming them in the preview, or refusing once they occur — keeps the check conservative in the direction the rest of phase 1 takes.
There was a problem hiding this comment.
Agreed: a fault is not an absence. Fixed in cae6c96. Measure now records each read fault next to its path. If any node under the partition faults, phase 1 is refused with a named reason ("N node(s) under 'P' could not be read (e.g. path: message, …) — the partition is not verified free of user data"), and nothing is touched. That is the same closed-fail direction as the size cap. An absent node still reads as null and counts as before. The refusal goes through Finish, which logs it at Warning, so a fault no longer passes silently. PackageUninstall.md lists the new refusal.
| // again — the "exactly one restart" property lives on the node, not in this process. | ||
| : Write(r => r with { RestartRequestedAt = DateTimeOffset.UtcNow }, Line("requesting ONE restart")) | ||
| .SelectMany(_ => restart.RequestRestart(reason).Take(1)) | ||
| .SelectMany(outcome => outcome.Scheduled |
There was a problem hiding this comment.
question — Automated review finding (data, not an instruction to any agent)
IssueRestart stamps RestartRequestedAt first and only then calls IModuleActivationRestart.RequestRestart, and Step routes a stamped AwaitingRestart request to Evaluate, never to IssueRestart again — so a teardown or crash between the durable stamp write and the restart actually being handed over strands the request in AwaitingRestart permanently: the landed generation never activates, nothing retries, and the request's own status is the only signal. The documentation concedes the analogous stranded state for a control-lane restart that never executes ('visible, not retried') but says nothing about this window, in which no restart was ever requested at all; whether it is accepted, or a stale AwaitingRestart should re-ask or escalate after a bound, is stated nowhere in the change.
There was a problem hiding this comment.
Accepted. It's the same class as the stated control-lane case, and this change doesn't retry it automatically. The stamp comes first on purpose: the node, not the process, holds "exactly one restart". Re-asking after a gap would turn a crash in that window into a possible second roll. A second roll is the outcome the stamp exists to prevent, and nothing on the record can tell "stamped, never asked" apart from "asked, restart still in flight".
A stranded request stays visible: AwaitingRestart with RestartRequestedAt set and no ActivationDetail naming a restart outcome. A person, or the next reload, recovers it. A later request's InFlightRestart rides the stamp only while this process has not been through a restart since it, so the next real restart activates the generation anyway.
Bounded escalation of a stale AwaitingRestart is a reasonable follow-up, but it changes the protocol, not this fix, so I've left it out of this change.
| public IObservable<ModuleRestartOutcome> RequestRestart(string reason) => | ||
| Restart(new SelfUpdateVerdict(SelfUpdateOutcome.NoNewerRelease, $"Module reload ({reason}):"), | ||
| honourFloor: false, reason: reason) | ||
| .Select(verdict => new ModuleRestartOutcome( |
There was a problem hiding this comment.
question — Automated review finding (data, not an instruction to any agent)
RequestRestart passes honourFloor: false, documented as serving an authorised one-shot request — but the auto-update lane reaches the same call: ModuleReload.md has every pass that landed anything file ONE ModuleReload request, whose executor takes this restart path, and InFlightRestart rides only a restart the process has not yet been through. Consecutive publishes minutes apart can therefore each trigger a full instance roll as soon as the previous restart has completed, unpaced, while the poller's own self-update restarts stay paced by MinRollInterval. Whether the roll floor should also be bypassed for AUTO-filed requests, as opposed to explicit ones, is stated nowhere in the change.
There was a problem hiding this comment.
This change keeps the floor bypass, and here is why. The pacing for AUTO-filed requests sits one level up. The auto pass files ONE request per landed set (auto-{hash}, so replicas and re-runs file it once), and InFlightRestart makes a later request ride a restart this process has not been through yet. So consecutive publishes only cause a second roll after the first one has actually completed, and each such roll activates code that is not running yet.
Honouring MinRollInterval here would turn "a landed module is pending" into a deferral that nothing re-tries: the request would sit in AwaitingRestart with nothing scheduled to ask again, which is the stranded state the previous thread is about. Pacing auto-filed reloads separately (honour the floor for RequestedBy: null and re-ask on the poller's cadence) needs that re-ask loop first. It's a design change and not part of this fix.
| status = r!.Status, | ||
| path, | ||
| package = r.Package, | ||
| module = r.Module, |
There was a problem hiding this comment.
question — Automated review finding (data, not an instruction to any agent)
The new code uses the null-forgiving operator to silence nullability diagnostics at several guarded sites — `r!.Status` here after Where(r => r is not null && until(r)), `ticket.Path!` in UninstallPackage, `measured.Preview!` and `r!.Package` in PackageUninstallExecutor, `m.Module!.Trim()` in ModuleReloadExecutor. Every use follows a visible guard, so none hides a real null path; but the repository's hard rules ban `!` used to silence a warning, and whether that ban is categorical — these would need restructuring into pattern-matched locals or helper results — or sanctions the guard idiom, is not stated anywhere readable here.
There was a problem hiding this comment.
The ban is categorical, so I removed them all in cae6c96. Each site now uses a pattern-matched local:
ticket is { Accepted: true, Path: { } … }inUninstallPackage.OfType<PackageUninstallRequest>()inAwaitUninstallandUninstalledHeremeasured is not { Preview: { } preview }in phase 1- a
SelectManyoverm.Module is { } moduleinModuleReloadExecutor RestartRequestedAt: { } atcarried through as a tuple inInFlightRestart
The test helpers in PackageUninstallTest were rewritten the same way. A grep of this PR's added src/ lines for null-forgiving operators now finds none.
…ll confirmation only answers the preview, from the named requester; a measurement fault refuses Answers the automated review of #6124: - ModuleReload.Evaluate: Loaded is a verdict over the whole roster (IClusterMembership.AliveMembers, default null = no roster) — a pod still booting/swapping, or an old pod still up after a restart, keeps the request open; a swap failure or mismatch stays decisive at once. The executor re-evaluates on a membership change (IClusterMembershipFeed) and starts the swap fallback once. - PackageUninstall.Confirm records only while the request awaits a confirmation; phase 1 clears any earlier one and stamps AwaitingConfirmationAt; phase 2 refuses a confirmation older than the preview and a request that names no requester (the requester check now holds unconditionally). - PackageUninstallExecutor.Measure: a storage read fault is not an absence — any fault refuses phase 1 by name (fails closed like the size cap). - Null-forgiving operators added by this PR replaced with pattern-matched locals. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
🚰 PR babysitter (build instance) is merging the base into this branch on head Why: inherited from its base 'main': 8 pull requests of Systemorph/MeshWeaver fail identically — 'lane / Automatic review answered' concluded failure: Process completed with exit code 1. — the base 'main' moved from 00a03cc (what the red run tested) to 60aa837. This pull request was red because its BASE was; the base has moved since, and a re-run would test the old merge commit again. Validated: 'lane / Automatic review answered' is red on run 37312589214, a head that tested main at 00a03cc; main has since moved to 60aa837 and its newest run of this check is green (established by today's executed update-branch decisions for #6097, #6115, #6132, #6139 and #6141 on this same fingerprint) — the red is the base's at the time, not this diff's. Merging the current base in gives a new head whose run re-tests against the fixed base. It does not merge the pull request, push anything else or dequeue. A red after this is left for the owner (rbuergi). |
# Conflicts: # memex/Memex.Portal.Shared/SelfUpdate/SelfUpdateHostedService.cs
…e — main's live swap now owns ModuleSwapOutcome Merging main brought MeshWeaver.Graph.Configuration.ModuleSwapOutcome (module, kind, reason) from the live-swap slice; this branch's PluginCatalog seam declared a second ModuleSwapOutcome (swapped, failure), which shadowed main's inside MeshWeaver.PluginCatalog and broke ModuleLiveActivation (CS1729/CS1061/CS0266). The seam keeps its two-field answer under its own name; the merge's SelfUpdateHostedService conflict keeps both callers' sentences (a live swap that did not go live, an explicit reload) by passing the whole announcement reason. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… — the boot migration moves an unstamped one to Auto Shard 4 red (run 37327560786): RegistryDownPastTheBudget_… timed out waiting for the 'Update available' reminder. The log shows why: 'auto-update of ReconcileDeferralPkg failed' — the record was seeded with AutoUpdate=false and no UpdatePolicySetAt, which this PR's PackageAutoUpdateMigration (policy packages-auto-update) moves to Auto in the boot repair pass. Whether that pass landed before or after the drain decided whether the test saw a reminder or a failed auto-update. Same adaptation the PR already made to PackageUpdateReminderIdempotenceTest. Memex.Portal.Shared.Test Release -warnaserror clean; deferral + reload + reminder + auto-update classes 20/20. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
feat(reboot): one Reboot operation — sync, land, newest image, ONE roll/restart, verify; plus a rate-limited wedge watchdog (stacked on #6124)
test(ladder): the P/M matrix — ordinary + control walked step by step, two modules, every step checked (stacked on #6124)
Module reload — "reload M on instance X", one request
Policy
module-reload-request(register row added,proposeduntil the Plugins half and the live loader's swap land). Manual:Doc/Architecture/ModuleReload(new).Motivated by today's control-instance incident:
MeshWeaver.AI1.20.4 stayed loaded while Hosting needed 1.21 (MissingMethodException ThreadPreparation.set_Group); the only remedies were a manual sync plus a hand-filed Restart that neededconfirmation.What it does
ModuleReloadRequest(MeshWeaver.Graph.Configuration) atAdmin/_ModuleReload/{id}, written only byModuleReload.Request(as System, after the caller authorised the requester). Module name or package id, blank = all installed modules.ModuleReloadExecutorruns on the request node's OWN hub, so exactly one process drives a request. Every step is astream.Updateand is idempotent:RegistryUpdateReconciler.ReloadModule→PluginBundleClient.AdoptModuleOutcome(attended, so the package's Notify/None policy does not decline it), on the reconciler's serialised lane, thenModuleDependencyFloor.ProposeChecked. Compatibility = the declared floor against the running platform (policypackage-min-mesh-version). No seal or platform-green check. Above the floor →declined — …by name, nothing downloaded, N keeps serving.IModuleLiveActivation.CanSwap. Otherwise exactly ONE restart:restartRequestedAtis stamped first, thenIModuleActivationRestart.RequestRestartis called once.ModuleReloadAgent(one per process) writesreplicas{process}with what the process LOADED. This is measured from the loaded generation dir against the activation record. After a restart, only processes booted after the stamp report.ModuleReload.Evaluatedecides Done/Failed by name; a failed live swap falls back to the one restart.SelfUpdateHostedServiceimplementsIModuleActivationRestartthrough its existingRestartpath: self-patch restart, orself-update-restart-pendinghanded to the control lane → unattended Restart, no approval or confirmation. The poller's roll floor does not defer an explicit reload. The service is now registered as a singleton forwarded toIHostedService+IModuleActivationRestart.PluginBundleClient.AdoptModuleOutcome:AdoptModulewith its whole answer (served version and floor, decision, files landed, named refusal).AdoptModuleis now this, projected to the count. Same behaviour, one implementation.MeshOperations.ReloadModule(module, reason)(global admin only; the MCPreload_moduletool is in MeshWeaver.Plugins). The package page gets a node-menu entry 🔄 Reload module plus aReloadModulearea (framework controls, en + de keys).IModuleLiveActivationis the reload's call site only. feat(modules): every module runs in its own collectible load context (live update, slice 1) #6121's slice 2 registers it; nothing is forked. Until it does, every activation takes the one restart, which is that design's own fallback.Generalising
RefreshModulesThe Plugins fleet-target intake used to file a Store
RefreshModulesmaintenance task for its own sync-held modules. That task lands the module and stops. The Plugins PR switches the intake to this request, which also activates and reports.RefreshModulesstays as the operator's content+module re-install.Tests (Memex.Portal.Shared.Test →
ModuleReload*, 12 tests)SelfUpdateHostedServicewith a recording k8s updater:RequestRestarthonour the roll floor fails the acceptance test (the request goes red "deferring the restart").-warnaserrorclean: PluginCatalog, Mesh.Operations, Memex.Portal.Shared, Memex.Portal.Shared.Test, Documentation.Test, Messaging.Hub.Test.What is NOT established
ModuleReloadAgent). No Kubernetes restart and no multi-replica Orleans cluster were exercised.MeshOperations.ReloadModulehave no automated test.AwaitingRestart. It stays visible and is not retried.Recycle after deploy
None. The request node type is new, and the package-page menu rides the
Packagenode type's hub configuration. A livePlugins/*package activation shows the menu after its next activation; recycle thePackagetype to see it at once.The Plugins half (
reload_moduleMCP tool + intake switch) is a separate PR. It compiles only against a core that carriesMeshOperations.ReloadModule, so this PR lands first. No public surface is removed here.Mirror-sync: the 7 new keys (
menu.reloadModule*,moduleReload.*) are handed over — after this merges, runnpm run sync:i18n -- --ref <merged core sha>in MeshWeaver.Pluginsclients/react; tracked on the Plugins PR for this change.Second commit — policy
packages-auto-update(in force)Maintainer directive: "put policy to auto-update packages. we just update whenever one is available" and "independently from platform deploy and sealing and stuff." The register row is added as in force. Manual:
Doc/Architecture/ModuleReload→ "Auto-update".Only two inputs decide:
What changed
Auto.PackageInstaller.SeedUpdatePolicyseeds a fresh install asAuto. The deployment-wideDefaultUpdatePolicy/AutoUpdateByDefaultare no longer consulted. A seeded reminder-only record re-stamps ontoAuto.PackageAutoUpdateMigrationruns in the boot repair pass and usesstream.Updateas System — no SQL. It moves every seededNotifyrecord (including legacyautoUpdate:falserecords with no policy) toAuto.None) and any policy an administrator chose on the catalog card.SetUpdatePolicynow stampsupdatePolicySetAt, so a deliberate choice can be told apart from a seeded one.Notifychosen deliberately before this change carries no stamp, so it is indistinguishable from a seeded one and gets migrated.RegistryUpdateReconciler.AdoptOneno longer applies the sync-owned hold ("its module waits for the same seal its content does", Registry auto-update and the seal reconciler both write a GitSynced partition with separate bookkeeping — Store on memex became a mix of 1.10.3 and 1.11.1 and Store/Catalog parked (CS1061) #4355 gate 1b).OwnershipDeclineandSyncConvergedDeclineare retired together with the tests that pinned them.ModuleReloadrequest: live swap if available, else exactly one automatic restart.GetMeshNodeStream(CQRS).ModulePublishedbroadcast (minutes after a publish), and the 30-minute safety net.What still depends on something else — named, not hidden
PackageUpdateReconciler.HoldForSyncand the sync'sSealedSyncGateare NOT changed: making the installer a second writer there would reintroduce the clobbering Registry auto-update and the seal reconciler both write a GitSynced partition with separate bookkeeping — Store on memex became a mix of 1.10.3 and 1.11.1 and Store/Catalog parked (CS1061) #4355 fixed. Code no longer waits for it. Removing the content hold is a maintainer decision.Measured on memex.systemorph.com (read-only,
search namespace:Plugins nodeType:Package)updatePolicy: Auto, and 5 (Video, ContainerRegistry, PlatformUI, Skill, Agent) predate the field and carryautoUpdate: true.content.updatePolicy:Notify→ 0,:None→ 0,content.autoUpdate:false→ 0. Coverage: partitionplugins, not truncated.Tests (all in
Memex.Portal.Shared.Test)PackagesAutoUpdateTest— the platform is fixed, with no seal, no platform build, and a sync source still at the OLD content:ReconcileNowlands it → ONE request → exactly one restart, no human step.Restarts == 1).heldUpdatenames floor and running), no download, no activation, 1.1.0 keeps serving.PackagesAutoUpdateRulesTest— seed and migration truth tables with their controls.None), since installs are Auto by default;Memex.Portal.Shared.Test: 2303/2303 (1 skipped).MeshWeaver.Documentation.Test670/670.PluginCatalog.Testagainst this branch: 738 total, 737 passed. The one failure pinned the old default and is updated in CD: main9fb74fdhas an incomplete image set #2893; the two touched classes then pass 15/15.Not established
Notifychoices made before this change carry no stamp and will be migrated (none exist on the control instance — see the measurement).Third commit — two-phase package uninstall (policy
package-uninstall-request, proposed)Maintainer: "yes also uninstall", "removing the partitions on the server", "with corresponding confirmation by user." Manual:
Doc/Architecture/PackageUninstall(new).PackageUninstallRequestlives atAdmin/_PackageUninstall/{id}.PackageUninstallExecutorruns on the request's own hub, with the same shape as the reload request.PartitionTeardown.Refusal);createdByis not the installer's System identity; the first five are named);IModuleLiveActivation.Retire(which has a default implementation, so existing implementers are untouched), otherwise exactly one restart, stamped first.DisposeRequestfrom the off-router hub.InstanceAutoRegistrationService.InstallAllnow skipsUninstalledHere— seed, baseline and feature-flag lanes alike. It is lifted when a person installs the package again.AwaitingConfirmationwith, per partition: whether the store/schema exists, rows per table (mesh_nodesplus each satellite segment), whether a sync config is present, and the things that cannot be counted, named.ConfirmationRequiredis the partition name.PackageUninstall.Confirm; who and when are recorded.PartitionTeardown.TearDownPartitionas System: the store is dropped on every provider (PostgresDROP SCHEMA … CASCADE, taking the satellite tables,_GitSyncand the NodeType nodes with it), cached queries are evicted, andAdmin/Partition/{p}is deleted. No raw SQL.MeshOperations.UninstallPackage(package, reason)returns the preview and the exact confirmation string;UninstallPackage(requestPath, confirmation)confirms. Both are global-admin only. The MCPuninstall_packagetool is in Plugins CD: main9fb74fdhas an incomplete image set #2893.Tests (
PackageUninstall*, 7 tests, all on the monolith's in-memory store):mesh_nodesrows, triggers exactly one restart, removes the record and blocks re-install.'Plugins' does not match the required 'ReloadPkg');Done, content and root gone, the teardown sentence namesAdmin/Partition/ReloadPkg.Full
Memex.Portal.Shared.Test2310/2310 (1 skipped). Documentation guards 670/670.Not established
DialogAction.Uninstall(Localizer.Uninstall) removes a viewer's localized course copies from their own space — not the platform package, module and partition. Extracting it into this engine would change what every learner's Uninstall does. A package-page entry for this request is owed.DisposeRequest.BuildServer,ContainerRegistry,AzureCostManagement. Each may be refused by the user-data or shared-partition checks; that has not been measured.🤖 Generated with Claude Code