You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
feat(server): add HA capacity metrics, scoped locks, and graceful drain
Multi-replica gateways route correctly, but they exposed no capacity
signals, dropped every supervisor session at once on shutdown,
serialized all cross-object mutations fleet-wide on one PostgreSQL
advisory lock, and polled one full sandbox record per watched sandbox
per second on every replica.
Graceful drain: on SIGTERM the gateway closes supervisor admission and
reports "draining" on /readyz and /health (/healthz stays 200).
Gateways with PostgreSQL and a peer endpoint (every Helm install on an
external database) keep the listener open for a 3-second propagation
delay, then close their sessions at most 100 ms apart within 12
seconds; each session keeps serving relays and heartbeats until its
slot. The existing shutdown then runs with up to 10 seconds of
cleanup, 25 seconds in total. A draining owner fails peer relays for
sessions it no longer holds immediately, and disconnect cleanup skips
the endpoint-status lock when a replacement session exists.
Supervisors reset their reconnect backoff after an accepted session
and jitter retries, and a new session receives RelayOpen only after
SessionAccepted is queued.
Capacity metrics: a gateway_metrics module exports per-replica held
supervisor sessions, draining state, pending relays against the 256
and 32 caps, relay rejections, expiries and claim latency, peer
request rate, outcome (with the owner's code before the remap to
UNAVAILABLE) and latency, mutation-lock waits and timeouts, and
watch-poller cost. Labels are bounded, series are zero-initialized,
new latency metrics are histograms, and existing duration summaries
are unchanged.
Scoped mutation locks: hierarchical intention locks on a global key,
a key per workspace, and a key per sandbox replace the single
fleet-wide key. Global settings and platform profiles hold the global
key exclusively; provider and workspace-profile writes hold their
workspace key exclusively; every sandbox mutation, admin or
supervisor, holds only its sandbox key exclusively. Keys are taken in
a process-local RwLock table, then as PostgreSQL advisory locks in
ascending order on one connection from a dedicated 4-connection lock
pool, with one 10-second deadline. Timeouts return UNAVAILABLE with
reason MUTATION_LOCK_TIMEOUT instead of INTERNAL. The global key keeps
its legacy value, so mixed-version rollouts stay mutually exclusive.
DeleteProvider now takes its workspace guard, the settings mutex is
removed, and startup endpoint-status reconciliation takes one sandbox
guard at a time instead of holding the global key for the whole scan.
Watch polling: Store::get_resource_versions reads only id and
resource_version (PostgreSQL id = ANY($2), SQLite IN, 1000 ids per
statement), and the cross-replica poller makes one batched call per
tick with unchanged notification semantics.
Helm: the default termination grace period rises from 5 to 30 seconds,
and an optional autoscaling/v2 HorizontalPodAutoscaler (disabled by
default) omits spec.replicas from the workload when enabled. Chart
validation covers the effective maximum replicas and resource requests
or limits for utilization targets, and the certgen hook pods no longer
match the gateway selector.
Tests and docs: mise run test:rust:postgres runs the ignored
PostgreSQL-backed tests (batched lookups, cross-store lock exclusion,
cancellation, pool bounds, a drain-rate envelope); a Kubernetes HA e2e
test verifies that supervisor sessions move off gateway pods during a
rollout; unit and integration tests cover relay saturation, 5000
watched sandboxes, drain pacing, and lock scopes. The HA guide, a new
gateway metrics reference, the architecture notes, the API errors
reference, and the cluster debugging skills describe the drain
lifecycle, capacity signals, autoscaling, and PostgreSQL connection
sizing.
Part of #3528
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
|`deploy/helm/openshell/ci/values-high-availability.yaml`| HA test overlay (`replicaCount: 2` with external PostgreSQL Secret) |
461
+
|`deploy/helm/openshell/ci/values-autoscaling.yaml`| Render-only overlay for the optional gateway HorizontalPodAutoscaler (helm lint and helm-unittest) |
Copy file name to clipboardExpand all lines: TESTING.md
+18Lines changed: 18 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -48,6 +48,23 @@ mise run test:rust # cargo test --workspace
48
48
49
49
Rust validation checks tracked Cargo lockfiles; run `mise run rust:lockfiles:check` to check them directly. If one is stale, refresh it with Cargo using its adjacent manifest, review the diff, and commit the update.
50
50
51
+
### PostgreSQL-backed tests
52
+
53
+
Tests that need a real PostgreSQL server, such as advisory-lock concurrency
54
+
across two stores, are `#[ignore]`d and named `postgres_*`. Run them with:
55
+
56
+
```shell
57
+
mise run test:rust:postgres
58
+
```
59
+
60
+
The task starts a disposable PostgreSQL container with Docker or Podman
61
+
(set `CONTAINER_ENGINE` to choose), runs the tests one at a time, and removes
62
+
the container. Each test works in its own temporary schema. To use your own
63
+
disposable database, set `OPENSHELL_TEST_POSTGRES_URL`. Never point it at a
64
+
database that a running gateway uses: the tests take fleet-wide advisory locks.
65
+
CI does not run these tests; the Kubernetes HA e2e suite covers PostgreSQL end
66
+
to end.
67
+
51
68
### Native Windows validation
52
69
53
70
Use `mise run --skip-tools pre-commit` with the existing Rust/MSVC toolchain.
@@ -403,6 +420,7 @@ Available task variants:
403
420
|---|---|
404
421
|`e2e:kubernetes`| Default Rust e2e against Helm-deployed gateway |
405
422
|`e2e:kubernetes:db`| All database backend scenarios (SQLite + external PostgreSQL) |
423
+
|`e2e:kubernetes:ha-rebalancing`| Two gateway replicas behind Envoy with external PostgreSQL: scale, pod deletion, rollout drain and session redistribution, and file sync during pod rolls |
0 commit comments