Skip to content
harsha-moparthyPublic

About

Zero-downtime PostgreSQL schema migration: expand/backfill/verify/cutover/contract with cross-node xmin-fenced verification. Zero-dependency Go.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

pgshift — zero-downtime schema migration for PostgreSQL

A phase-gated expand→backfill→verify→cutover→contract migrator in zero-dependency Go, measured against a real primary + streaming replica: it widens a write-hot 2-million-row integer column to bigint holding its ACCESS EXCLUSIVE lock 3–4 ms for a 2.6–8.5 ms write stall — bounded by one lock_timeout even under lock contention — where the naive ALTER TABLE ... TYPE it replaces stalls every writer for 1,145 ms it cannot avoid. Verification runs across the replication lag boundary fenced on xmin: on a correct migration under deliberately induced lag it fenced 20,000 lagging-row differences and reported zero mismatches, where the naive single-pass verifier raised 20,000 false alarms; against three fault-injected triggers it caught every real divergence. Rollback is demonstrated from all four pre-contract phases, each verified from pg_catalog and row checksums. An adversarial review of the verifier found a false negative that would have certified corrupt data — it is bug #1, fixed and pinned by a regression test.


Why this project exists

Changing a large table's column type without blocking production writes is unglamorous, permanently in demand, and newly urgent because AI-assisted development ships schema changes faster than review can reason about their lock behaviour. The naive statement — ALTER TABLE t ALTER COLUMN c TYPE bigint — takes an ACCESS EXCLUSIVE lock and rewrites the entire table under it. Every reader and writer blocks for the whole rewrite. On the 2-million-row table measured here that is 1.1 seconds of total outage; on a real hundred-million-row table it is minutes. This is what PlanetScale, gh-ost, and pt-online-schema-change were built to avoid, and it is what the database platform companies sell.

The safe alternative is the expand-and-contract dance: add a new column, dual-write to it via a trigger, backfill the old data in throttled batches, verify the two agree, atomically swap the names, and only much later drop the old column. Each step is individually simple. The engineering is in the parts that are silent when done wrong:

  • Dual-write correctness. A trigger that handles INSERT but not UPDATE looks perfect in a by-hand test and diverges the moment a real row is updated.
  • Verification under concurrent writes — "the subtle part." The spec names this, and it is subtle for a reason that is not the obvious one. See below.
  • The cutover lock window. The rename is microseconds; the wait for its lock is what takes production down, because every query queues behind a queued DDL lock request.
  • Rollback that actually preserves data after cutover, not just in principle.

This project is built so each of those failure modes is demonstrated — with a control arm that exhibits the failure and a treated arm that does not — rather than asserted.

The one thing that makes verification hard

Checksumming a column against its shadow is trivial on a quiet table. On a write-hot one the naive worry is a torn read; the real problem is comparing two values captured at different instants.

The operationally correct place to verify is against a read replica — pt-table-checksum exists precisely so verification does not add a full-table scan to the primary's load. But a streaming replica lags: a row updated on the primary a moment ago holds its new value there and its old value on the replica, purely because replication has not caught up. A verifier that compares the two directly reports that row as a mismatch. Do that on every recently-written row of a busy table and it cries corruption continuously — and an alarm that always fires gets silenced, and then a real divergence ships behind it.

pgshift fences on xmin, the transaction id stamped on each tuple version, which PostgreSQL replicates physically — so the same row version has the same xmin on both nodes:

The rule is symmetric, and the symmetry is the part that is easy to get wrong: no comparison is evidence about the dual-write unless both nodes are showing the same tuple version.

  • same version, values agree → proven equal. Cleared.
  • same version, values differ → the replica has replayed the very version the primary shows, and they still disagree. No lag can explain that. A genuine dual-write defect, and the verdict is final.
  • different versions → nothing is proven, in either direction. Re-checked after the replica catches up.

Applying the fence only to the "differ" case leaves a false negative that I shipped and then had to fix: a stale replica value that happens to equal the primary's unchanged source clears a row that is diverged on the primary right now. See bug #1. The verification experiment measures both halves of the design.

Zero dependencies, on purpose

go.mod has no require block. The PostgreSQL v3 wire protocol client — startup, SCRAM-SHA-256 with mutual-auth verification, simple and extended query, COPY, ErrorResponse with SQLSTATE, the out-of-band cancel request — is hand-rolled in internal/pgwire, and so are the latency recorder and the workload harness.

The decisive reason is measurement. This tool's central claims are per-statement latency distributions attributed to migration phases, and a driver sits exactly on that measurement path: it owns pooling, statement caching, and automatic re-preparation when DDL invalidates a cached plan — which this tool triggers, so a driver would charge its own re-prepare round trip to the migration phase and the number would measure the driver. Owning the wire means the exact statement text, the bytes, and the timing boundaries are all code in this repo.

The SCRAM implementation is pinned to the RFC 7677 test vector (not just "it connected"), and includes the server-signature verification step that is easy to omit because the connection works without it — omitting it means authenticating against a server that cannot prove it knows the password.

Measured results

Measured on an Apple M4 Pro (14 cores, 48 GB), Go 1.26.5, PostgreSQL 16.14, macOS, against a local primary + streaming replica over loopback with fsync=on (deliberately — see honesty checks). Every number reproduces with make bench and its raw JSON is committed under results/. The workload is 16 connections at a 3,000 writes/sec offered rate (60% UPDATE, 20% INSERT, 20% SELECT) against the table being migrated, throughout every phase.

Headline: the same migration, naive vs pgshift, under identical load

Widen amount integer → bigint on a 2,000,000-row table.

naive ALTER ... TYPE (control) pgshift (treated)
write stall (max gap between commits) 1,145 ms 2.6 – 8.5 ms (uncontended)
...worst case the full table rewrite, unavoidable one lock_timeout (200 ms), then it yields and retries
ACCESS EXCLUSIVE lock held ~1,148 ms 3 – 4 ms
client-visible errors 0 0
final column type (checked in pg_catalog) bigint bigint
data equal to source (independent checksum) — yes

The control arm is not handicapped: it runs the fastest naive form there is — a plain TYPE change with an implicit cast — on the same table, under the same workload, from the same starting state. Its 1,145 ms stall is what the naive approach costs when it is doing its best, and it is not avoidable by tuning: the table has to be rewritten under the lock.

pgshift's stall has two regimes, and both are reported because they answer different questions:

  • Uncontended (the common case): 2.6 – 8.5 ms across four runs, with the lock acquired on the first attempt every time and held 3–4 ms. The rename is a few catalog updates; there is nothing else in the lock window.
  • Contended: bounded by one lock_timeout. If a long-running transaction holds a conflicting lock, the cutover's request times out after cutover.lock_timeout_ms (200 ms here) with SQLSTATE 55P03 and retries, rather than queueing every subsequent query behind a waiting DDL. So the worst a single attempt can cost writers is one lock_timeout, and between attempts the table is completely free. Earlier in this project's development — when the cutover mistakenly compiled a PL/pgSQL function inside the lock and so held it 114 ms — that regime was the common one, and the measured stall was 206 ms. That is the number the fix eliminated, and it is what the bound looks like when it is actually reached.

Both figures are from committed artifacts; the 200 ms yield behaviour is independently reproducible with a held ACCESS SHARE lock (see RUNBOOK).

Per-phase write latency (treated arm, ms)

Latency is timed around the client's full round trip — connection, lock wait, commit — because that is what an application pays, and attributed to the phase in effect when each operation completed.

phase p50 p99 p99.9 max gap
before (baseline) 0.17 0.33 0.75 1.4
expand 1.65 20.44 — 14.8
backfill 0.22 1.21 3.92 19.5
verify 0.21 1.49 1.83 2.5
cutover 5.82 12.55 — 8.5
after 0.20 0.50 2.24 2.9

Reading this honestly: the backfill is nearly invisible to writers — a p99 of 1.21 ms against a 0.33 ms baseline, and a 19.5 ms worst-case gap which is one chunk's row locks. Verify is invisible too (p99 1.49 ms), because its scan runs against the replica, which is the operational reason to verify there. Cutover is the one phase writers feel, and here it cost a p50 of 5.82 ms and a p99 of 12.55 ms — the brief lock plus the reverse-trigger swap. Expand's 20 ms p99 is over only 29 operations, so it is a handful of samples rather than a distribution.

The earlier, pre-fix run of this same benchmark showed a 659 ms max-gap outlier in verify (against a p99.9 of 127 ms) and a 206 ms cutover gap. The verify outlier was an autovacuum/checkpoint interaction during a longer read pass, not a lock the migration took; it did not reproduce once the cutover stopped holding its lock for 114 ms. Reporting that it happened matters more than that it went away — a single 659 ms gap on a busy table is worth knowing about.

Verification: 3 real faults caught, 20,000 lag-induced false alarms avoided

The verifier is run across the replication lag boundary with lag deliberately induced (replay paused on the standby, then rows mutated on the primary — a faithful model of a standby that cannot keep up). Each fault mode's data is judged three ways: an independent intra-row checksum (the ground truth, using neither verifier), the re-check verifier (fenced on xmin), and a single-pass control (no fence — what most tools ship).

trigger independent checksum re-check verifier single-pass control
correct equal equal (0 mismatches, 20,000 lagging rows fenced) 20,000 FALSE ALARMS
skip UPDATE diverged 20,000 mismatches ✓ 20,000
lossy cast diverged 20,000 mismatches ✓ 20,000
drop 6/7 of keys diverged 2,857 mismatches ✓ 20,000

The top row is the point. On a correct migration under replication lag, the single-pass verifier reports 20,000 mismatches — every lagging row, a flood of false corruption alarms that would abort a healthy migration. The re-check verifier recognises all 20,000 as version skew via xmin, waits for the replica to converge, and reports clean. The bottom three rows confirm it still catches genuine divergence — it is not simply silent. The xmin fence buys the gap between "clean" and "20,000 false alarms," measured, not asserted.

The fence is applied symmetrically, to equality as well as inequality, and a dedicated regression test pins that: accepting value equality across two different tuple versions is a false negative, and a likely one. See bug #1 below.

Backfill throttle: responds to replication lag, holds it lower

The backfill throttles on replication lag. Because a laptop's primary and replica share disks and replay about as fast as they write, the lag is induced (replay paused mid-backfill) so the throttle's response is measured rather than hoped for. Both arms measure lag; only the throttled arm acts on it.

arm throttle back-offs extra sleep max observed lag duration
throttled (target 100 ms) 18 3,440 ms 1,143 ms 9,258 ms
unthrottled (measures only) 0 0 1,423 ms 6,909 ms

The throttled backfill backs off 18 times once lag crosses the target, holding the peak lag ~20% lower — at the cost of taking ~34% longer. That is the trade the throttle exists to make: a backfill is a WAL generator, and a standby that falls far enough behind stops being a failover candidate for the duration of the migration. A back-off here means a sleep above the configured floor — distinguished from the floor sleeps every backfill does between chunks, which an earlier version conflated.

Rollback from every phase, verified from the catalog

Each drill runs the cycle to a target phase, rolls back, and judges the result by inspecting pg_catalog and row checksums — never the tool's own report.

rolled back from schema clean (catalog) data preserved
expanded ✓ ✓ (source byte-identical)
backfilled ✓ ✓ (source byte-identical)
verified ✓ ✓ (source byte-identical)
cutover ✓ ✓ (post-cutover write preserved by the reverse trigger)

The cutover row is the real guarantee: a write made after cutover survives the rollback, because the reverse dual-write trigger kept the old column current. Roll back a cut-over migration without that and every write since cutover is silently lost — reverted to data as of the cutover instant.

Bugs found and fixed

Each was found by distrusting a number or a claim, and each is now pinned by a test.

1. The verifier certified corrupt data when a stale replica value coincided with the primary's source — the worst bug in the project. The xmin fence was applied only to inequality: the comparison did if equal { continue } before ever checking whether the two nodes were on the same tuple version. The reproduction is mundane, which is what makes it dangerous:

v1: (source=10, shadow=10)   correct; replica replays it
v2: (source=10, shadow=99)   primary updates; broken trigger diverges the shadow
                             replica still on v1, so it reports shadow=10

The verifier compares the primary's source (10) against the replica's shadow (10), calls it equal, skips the row, and reports Equal=true on a row that is diverged on the primary right now — so cutover proceeds on corrupt data. On any low-cardinality column (a status enum, a small integer, a boolean) a stale shadow coincides with the current source constantly, so this is not a corner case. I reproduced it directly in psql before fixing it. The rule is symmetric and now stated as such: no comparison is evidence unless both nodes show the same tuple version — same version + agree = cleared, same version + differ = confirmed, different version = nothing proven, re-check. Pinned by TestFenceAppliesToEqualityNotJustInequality, which asserts both that the row is not cleared while versions differ and that it is confirmed once the replica catches up.

2. A proven mismatch was downgraded to a suspect and could be "resolved" away. The first pass computed xminMatch — i.e. proved a same-version disagreement — and then threw that verdict away (suspects[k] = s.pRow, dropping the flag), sending the row back through the re-check loop. If the workload then overwrote that row with a correct value, the next round saw agreement and cleared it: a defect that was demonstrably real when observed reported as clean. A same-version disagreement is a statement about a tuple version that existed, so no later round can overturn it; proven rows are now recorded immediately in a separate set and reported terminally.

3. Mismatch samples showed innocent rows instead of the offenders. The report sampled 10 keys from every row that failed to clear, in random map order. With 1,000 lagging rows and one real defect, the operator saw the actual offender with about 1% probability — and the rows they did see had differing xmins, directly contradicting the definition of a mismatch the report claimed to show. Deleted keys could appear too, rendering as empty strings. Samples now come only from the proven set.

4. transient_diffs counted round-observations, not rows. It is the headline evidence number for "the lag path was exercised," and it was incremented on every re-check round — so with 40 rounds configured, one persistently lagging row could contribute up to 41. The reported figure was 80,000 for 20,000 rows. Now counted on the first pass only: 20,000, which is the number of rows it claims to be.

5. The cutover held its lock 40× longer than necessary. The first working cutover held ACCESS EXCLUSIVE for 114 ms, and the write stall was 751 ms against a 200 ms lock_timeout — worse than the tool's own bound. The cause: the reverse trigger's PL/pgSQL function was created inside the locked transaction, so compiling a function was happening under the lock every writer was queued behind. CREATE FUNCTION needs no table lock, and PL/pgSQL resolves column names at execution time, so the function can be pre-created against a column layout that does not exist yet. Moving it outside the lock cut the hold to 3–4 ms and made the uncontended stall single-digit milliseconds. The lesson generalises: the lock window should contain only what genuinely requires the lock.

6. A crash between the cutover commit and its state save was unrecoverable. SaveState runs in its own transaction after the swap commits. A connection drop in that window left a cut-over schema with a state row still reading verified — and then re-running cutover retried the first rename, hit 42701 duplicate_column (non-retryable, hard failure), while contract simultaneously refused because the recorded phase was below cutover. Manual surgery was the only way out. rollback already trusted the catalog over the state row for exactly this reason; cutover now does too, detecting the cut-over layout and reconciling the bookkeeping instead of re-attempting the rename. Pinned by TestCutoverReconcilesStateAfterCrash, which simulates the crash by rewinding the state row.

7. A row keyed exactly math.MinInt64 would never have been verified. The chunk cursor was seeded at math.MinInt64 with a key > $1 predicate, so a row at that exact key was skipped by every pass — and a diverged row there would have returned Equal=true. The first chunk now uses >=. Unreachable in practice with an identity key; "unreachable in practice" is what every silent skip is called beforehand, and the fix is one comparison operator.

8. Re-check-on-change, as first designed, defended against nothing. The verifier originally compared source and shadow in the same row on the same node, with xmin-based re-checking for rows that changed mid-scan. Measured on a live table it reported 0 changed rows — and it always would, because two columns of one tuple are read in one MVCC snapshot and are internally consistent by construction. The "correct under concurrent writes" claim was vacuous: the re-check machinery could never fire. The fix was to move verification to where the subtlety actually lives — across the primary/replica lag boundary — where a naive comparison genuinely produces false positives and the xmin fence is genuinely load-bearing. A benchmark reporting zero for the number that proves your mechanism works is the mechanism telling you it is pointed at the wrong problem.

9. A ten-byte message could have sized a 2 GB allocation. readMessage reads a length off the wire and allocates a body buffer from it — a length the peer controls. Without a cap, a corrupted or hostile length is a make([]byte) of that size, and the process dies of OOM while the connection looks healthy. Fixed with a 64 MiB cap checked before allocation, and an undersize check (a length below 4 is int(length)-4 = a negative size). Both are pinned by FuzzReadMessage, which ran 2.86M executions clean.

10. NULL would have compared equal to a missing value. The intra-row verifier's predicate started as shadow <> cast(source). NULL <> NULL is NULL, which is not true, so a <> predicate silently drops every row where either side is NULL — including the un-backfilled rows where the shadow is NULL and the source is not. A verifier written that way reports "no mismatches" on a table full of un-backfilled rows. Fixed to IS DISTINCT FROM, which treats NULL as a comparable value.

11. The throttle "back-off" count conflated floor sleeps with real backoffs. The first throttle experiment reported "200 throttle sleeps" for both arms — but that counted the 5 ms floor pause every backfill does between chunks, not the adaptive backoff. The throttled and unthrottled arms looked identical. Fixed by counting only sleeps above the floor as back-offs, which is the number that actually distinguishes "the throttle engaged" from "the backfill paused its configured minimum." Now the throttled arm shows 18 back-offs and the unthrottled arm 0, which is the real signal.

12. INSERT's row count would have been read as its OID. CommandComplete for INSERT is INSERT <oid> <rows> — the count is the last field, not the second as it is for UPDATE/DELETE. A parser taking the second field reports the OID (usually 0) as the rows affected. Fixed to take the last field; pinned by a table test.

13. The lag signal, read but not acted on, reported nothing. The throttle originally only read replication lag when target_lag_ms > 0. That meant the control arm (target 0) could not report the lag it let accumulate — the exact comparison the control exists to make. Fixed so the signal is always measured when wired, and only acted on when a positive target is set. Now "max observed lag" is a real number in both arms.

Controls and honesty checks

  • fsync stays on. The headline claims are write-latency numbers; fsync=off would remove the disk from the critical path and flatter every one of them.
  • The control arm is not handicapped. It runs the fastest naive migration there is, same table, same load, same starting state. A manufactured-worse baseline would make the comparison worthless.
  • Every verdict comes from an independent judge. Post-migration equality is a checksum computed by a separate query; the final type is read from pg_catalog; rollback cleanliness is a catalog inspection; the fault experiment decides "is this data actually diverged" with an intra-row checksum that uses neither verifier. Nothing reads the migrator's own success report to decide whether it succeeded.
  • The stall is a max-gap, not a quantile. A migration that blocks all writers for a second and then drains the queue has an excellent p99 — the blocked statements are few and each ran fast once unblocked. The gap between consecutive commits is what an application experiences as downtime, so it is measured directly. Both are reported.
  • Quantiles are exact order statistics, subsampling disclosed. Nearest-rank over a retained reservoir, never interpolated from buckets; count and retained are both in every artifact.
  • Lag is induced deliberately, and said to be. Natural lag on one laptop is a couple of milliseconds and would leave the throttle and the cross-node verifier untested. Pausing replay is a faithful model of a standby that cannot keep up, and the experiments say plainly that they induce it.
  • SCRAM is pinned to the RFC vector. Not "it connected" — the client-final message and server signature are checked against RFC 7677 §3, so a subtle HMAC or base64 error is caught rather than being consistent-with-itself-and-wrong.
  • Numbers that came out worse than theory are reported as measured. The verify phase's 659 ms max-gap outlier, the throttle's ~34% duration cost, the 20,000 false alarms the naive verifier produces — all stated, none smoothed.

Honest limitations

  1. One column per migration. The expand-and-contract dance generalises to multi-column and whole-table copies, but the interesting engineering (dual-write correctness, cross-node verification, the cutover window) is identical on one column and easier to hold to an exact standard. Multi-column is a stated non-goal here, not a half-built feature.
  2. Integer chunk keys only. The backfill chunks on an integer primary key. A text or uuid key needs collation-aware ordering, where a disagreement between the server's collation and the tool's comparison silently skips rows at chunk boundaries — so the tool refuses such keys at preflight rather than risk it.
  3. Measured on loopback. The primary and replica share one machine's disks and network, so absolute latencies and the natural replication lag do not include real network delay. The shapes transfer — the cutover stall is bounded by lock_timeout regardless, and the verification logic is topology-independent — but a production run would see larger absolute lag and a genuinely useful throttle.
  4. The production latency signal is a benchmark input. The backfill can throttle on production write p99, and the harness feeds it its own measured p99. A real deployment would wire this to its APM or pg_stat_statements; the plumbing is a one-function interface, but that integration is not written.
  5. No CREATE INDEX CONCURRENTLY for the shadow. The shadow column inherits no indexes. A migration of an indexed column would need the index rebuilt on the shadow (concurrently, so as not to lock), which this tool does not do — it migrates the column's type, not its indexes.
  6. contract does not reclaim disk. DROP COLUMN is catalog-only; the space is returned only by a later VACUUM FULL or pg_repack, both table rewrites. The tool says so rather than claiming the space back.
  7. Single-node verification is weaker, and labelled. Without a replica the verifier compares intra-row in one snapshot — correct and complete for one node, but it cannot exercise the lag path. VerifyResult.Mode records which ran, so a single-node "verified" is never mistaken for the cross-node guarantee.

Tech stack

Component Choice
Language Go 1.26, zero external dependencies (stdlib only)
Wire protocol hand-rolled PostgreSQL v3: startup, SCRAM-SHA-256 (mutual, RFC-7677-pinned), simple + extended query, COPY, cancel request; message reader fuzzed
Migration expand (shadow column + BEFORE-trigger dual-write) → throttled backfill → cross-node verify → rename cutover → contract, all phase-gated with persisted state in the database
Verification primary-source vs replica-shadow, fenced on xmin; single-node intra-row fallback; single-pass control arm
Throttle multiplicative-increase/decrease on replication lag and write-p99, both bounded
Safety short lock_timeout with yield-and-retry on all DDL; SET LOCAL for scoped GUCs; catalog-based preflight guards; type allowlist (a type name cannot be quoted)
Measurement exact-quantile latency recorder with max-gap stall detection; write-hot workload harness with per-phase attribution
CI GitHub Actions: gofmt, vet, go test -race, fuzz — offline; plus a job that stands up a real primary+replica with the repo's own script and runs the integration + fault + rollback drills

Quickstart

Needs postgresql@16 (initdb, pg_ctl, pg_basebackup on PATH — the cluster script finds the Homebrew/apt locations itself). No other install, no network.

make cluster-up     # initdb a primary (5514) + streaming replica (5515), SCRAM auth
make build
source .env.example # or set PGSHIFT_* yourself

# Run a migration, phase by phase (each is a separate, resumable command):
./bin/pgshift preflight -plan configs/plan.example.json   # checks the catalog, no changes
./bin/pgshift expand    -plan configs/plan.example.json
./bin/pgshift backfill  -plan configs/plan.example.json
./bin/pgshift verify    -plan configs/plan.example.json -json
./bin/pgshift cutover   -plan configs/plan.example.json
./bin/pgshift status    -plan configs/plan.example.json
# rollback is available until you contract:
./bin/pgshift rollback  -plan configs/plan.example.json -confirm

Reproduce every number in this README:

make bench          # brings up the cluster, runs all five experiments, writes results/

Run the offline correctness suite (no database):

make test           # unit + plan validation + SCRAM RFC vector
make race           # the same under -race (the important one)
make fuzz           # the wire message reader

Repository layout

pgshift/
├── cmd/
│   ├── pgshift/         # the phase-gated CLI (preflight/expand/backfill/verify/cutover/contract/rollback/status)
│   └── shiftbench/      # the evidence: control vs treated, faults, throttle, rollback drills
├── internal/
│   ├── pgwire/          # hand-rolled PostgreSQL v3 client; SCRAM (RFC-pinned); message reader fuzzed
│   ├── migrate/         # the phase machine, dual-write triggers, cross-node verifier, lag probe, rollback
│   ├── workload/        # write-hot harness with per-phase latency attribution; the table seeder
│   ├── metrics/         # exact-quantile recorder with max-gap stall detection
│   └── config/          # connection config from the environment; password by variable name, never value
├── scripts/
│   └── cluster.sh       # initdb primary + pg_basebackup streaming replica, entirely under run/
├── configs/             # example plan
└── results/             # committed benchmark JSON — every number above reproduces from here

The claim, stated narrowly

On one laptop against a local PostgreSQL 16 primary and streaming replica: pgshift widens a write-hot 2M-row integer column to bigint holding its lock 3–4 ms for a 2.6–8.5 ms write stall, bounded by one lock_timeout under contention, where the naive ALTER stalls every writer for 1,145 ms it cannot avoid; its cross-node verifier, fenced on xmin in both directions, reports zero mismatches on a correct migration under induced lag that makes a single-pass verifier raise 20,000 false alarms, and catches every one of three injected dual-write faults; and rollback preserves data from all four pre-contract phases, each checked against pg_catalog and row checksums. It does not claim real-network latencies, multi-column migrations, index rebuilds, or that contract reclaims disk — those limits are listed above.

The "Bugs found" list is the part I would read first. Seven of its thirteen entries are bugs in the verifier and the measurement rig rather than in the migration logic — including one that certified corrupt data as equal, and one where a benchmark reporting 0 for its own key metric revealed that the mechanism it was testing could never fire. A verifier is a claim about detecting something that must never happen, so it is exactly the component whose bugs are invisible until someone attacks it deliberately. I would rather that class of error keep being caught by the harness and by review than by a reader.

About

Zero-downtime PostgreSQL schema migration: expand/backfill/verify/cutover/contract with cross-node xmin-fenced verification. Zero-dependency Go.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages