Fast, snapshottable PostgreSQL for local development on Apple silicon — load a fresh dump into a spare instance and swap it in without downtime.
pg_restore of a large dump can take ~90 minutes (in my case, having expensive GIN indexes),
and the dev database is
unreachable the whole time. This repo runs two persistent Apple container
machines, vpg-a and vpg-b, each hosting one PostgreSQL 17 backend on
its own Incus + copy-on-write XFS snapshot store. One machine is active
(serving your app); the other is staging (where the next dump loads).
promote swaps them.
- Snapshots in seconds — XFS reflink (copy-on-write) checkpoints and rollbacks, not multi-gigabyte directory copies.
- Zero-downtime dump loads — import into staging while active keeps serving,
then
make pg.promoteswaps roles without copying data. Clients use a stable endpoint with fixed role ports:127.0.0.1:5442(active) /:5443(staging). - On-demand disk reclaim —
make pg.staging.rebuilddelete+recreates the staging machine to free its grown macOS disk, while active keeps serving. - Low power at runtime, thanks to Apple
container.
-
Apple silicon running macOS 26.
-
Apple's
containerCLI (1.1+) and Go — via Homebrew:brew install container go
macOS client
│ 127.0.0.1:5442 (active) / :5443 (staging) — stable, never changes
▼
socat client proxy (host launchd agents, internal/socatproxy) — re-pointed on promote
│ maps each role port to whichever machine holds that role
├───────────────────────────────┬───────────────────────────────┐
▼ ▼
┌──── vpg-a <ip>:5432 ────┐ ┌──── vpg-b <ip>:5432 ────┐
│ Incus host │ │ Incus host │
│ pg-dev-a (PostgreSQL 17)│ │ pg-dev-b (PostgreSQL 17)│
│ eth0:5432 ─proxy dev─▶ 127.0.0.1:5432 (in-container) │
│ XFS store + reflink snaps│ │ XFS store + reflink snaps│
└──────────────────────────┘ └──────────────────────────┘
Each Apple machine is a persistent outer Linux environment with its own Incus
daemon and one backend, exposed on the machine's eth0 by a single Incus proxy
device on the backend container (no separate proxy container, no static-IP
pinning). The active/staging role is a host-side pointer
(var/active-machine), not something a machine knows. The repository stays on
macOS, visible at the same /Users/... path through each machine's home mount,
so .env, the pointer, and exports live outside the machines.
A resident daemon, pgdevd, runs inside each machine under systemd and
serves an HTTP/JSON API on that machine's eth0 (port 5440, bearer token in
var/agent-token). The host CLI, pgdev, holds one client per machine and
routes by role:
pgdev (macOS) ── HTTP/JSON ──▶ pgdevd (vpg-a | vpg-b) ──▶ Incus socket + XFS store
Each daemon serves exactly its one backend (slot-implicit API); active/staging
is decided host-side. So promote is purely host: flip the active pointer,
re-point the client proxy — no data moves, no daemon call. up provisions each
machine's backend from a golden pg-dev-base image (PostgreSQL installed once,
then incus published) and adds its eth0 proxy device; status/ip/
refresh fan out over both. Each daemon owns its own single-mutation lock,
write-ahead journal, and crash recovery; machine setup (XFS store + Incus
topology) is pgdevd bootstrap, run as its unit's ExecStartPre. make pgdevd
builds both binaries (one git-stamped version); pgdev agent deploy (run by
make start, --machine a|b|both) delivers the prebuilt daemon and its config
machine-local (never read over the home mount at runtime), restarts the
unit, and confirms the GET /v1/version handshake. The only remaining
container exec passthroughs are the interactive shell/logs in
scripts/pg-dev-local (the host picks the active/staging machine); psql is a
plain local psql against the 127.0.0.1 endpoint.
A backend's nested incusbr0 address is not routed to macOS. Each backend is
instead exposed on its own machine's eth0:5432 by a single bind=host Incus
proxy device on the backend container, connecting to PostgreSQL on the
container's 127.0.0.1 — so the connect target never drifts and there is no
static-IP pinning. Do not use a backend's 10.x address from macOS.
Each machine's eth0 IP is an unpinnable bootpd DHCP lease that can change
(the whole /24 too) after a macOS reboot or machine recreation, and the two
leases drift independently. So the client endpoint is decoupled from them: two
per-user launchd agents run socat (internal/socatproxy), owning
127.0.0.1:5442 (active) / :5443 (staging) and relaying each to whichever
machine currently holds that role. The ports are offset from 5432 so a local
PostgreSQL isn't shadowed. Clients always use 127.0.0.1:5442 / :5443
— permanent, identical on every Mac.
socat cannot re-point itself, so every promote and IP change reconciles the
agents: rewrite the plist, bootout, wait for the port to actually free, load it
again, then verify the live process really dials the intended target. That
reconcile is driven from var/pgdev.db (SQLite) and serialized by a flock, so
two concurrent commands can't leave a stale mapping behind. Sessions on the
demoted machine are dropped by the reload — reconnect to land on the new
database.
Install it once with make proxy.install (make proxy.status /
proxy.uninstall manage it); after that promote/refresh keep it in step
automatically and never touch launchd if it isn't installed. The first run also
needs a one-time macOS Local Network grant — see
macOS Security.
PG_PROXY_HOSTNAME sets the hostname printed in psql/.pgpass lines (default
host.docker.internal, so the endpoint also resolves from sibling
containers/k3d; use 127.0.0.1 for host-only). PG_CLIENT_BIND widens the
listener bind (default 127.0.0.1; set 0.0.0.0 only if a sibling container
can't reach the Mac's loopback — it exposes the dev backend on every interface).
No connection pooler: each port is a per-connection TCP passthrough, so
CREATE/DROP DATABASE, LISTEN/NOTIFY, prepared statements, advisory locks
and parallel pg_restore behave like direct connections. Promoting reloads the
proxy, which drops existing sessions (reconnect); the role ports don't change.
Apple 1.1's recommended container-machine kernel lacks the Btrfs, ZFS, and
DM-thin stack needed by Incus's optimized snapshot backends. Incus therefore
falls back to dir, whose snapshots are full directory copies—not practical
for repeated multi-gigabyte PostgreSQL checkpoints.
pgdevd bootstrap (each daemon's ExecStartPre) instead creates a sparse XFS
loop filesystem (140 GiB by default) inside its machine's root disk, mounts it,
and configures the Incus storage/network/profile. Apple's 1.1 boot examples
show a 512 GiB root device, but that size is not a documented compatibility
guarantee. PostgreSQL data for each slot is mounted from the XFS filesystem,
and snapshot commands use reflink copies. Creating a checkpoint is fast and
consumes additional blocks only as the live dataset and snapshots diverge.
Snapshots cover PostgreSQL data, not the disposable Ubuntu container root, and
are per machine (a hard reset discards that machine's). Every snapshot stops
PostgreSQL cleanly before cloning the data and starts it again
afterward. Restoring an older snapshot retains the original workflow's
timeline semantics: snapshots newer than the target are shown and deleted
after confirmation (interactively you get a [Y/n] prompt, but non-interactively
you must pass force=1, e.g. make pg.restore name=foo force=1, because stdin
cannot answer prompts through the machine transport).
The XFS size is a logical ceiling; the backing file is sparse. Apple container
CLI 1.1 does not offer a machine disk-size flag. PG_DATA_DISK_SIZE is only
used at first creation; to grow it later, raise the value in .env and the next
make start will expand the sparse XFS store online via xfs_growfs. Shrinking
the store is not supported.
This is a local-development tool for one trusted developer and laptop. It is
not replication, automatic failover, a zero-downtime service, or a hardened
multi-user deployment. PostgreSQL durability features such as fsync are
deliberately disabled for import speed. Snapshots are checkpoints on the same
physical disk, not backups.
make deps # also auto-creates .env from .env.example (defaults work; edit if you like)
make start # builds the image, creates both machines (vpg-a, vpg-b) and provisions their backends
make pg.statusThe first make start builds an Ubuntu 26.04 machine image (systemd, Incus, jq,
XFS tools) and creates both machines; later starts reuse them. It then installs
PostgreSQL 17 in each machine's nested Ubuntu 24.04 container (several minutes)
for any slot that has no backend yet — so make start is also the way back from
make pg.staging.purge. A slot that already has a backend is left untouched.
make pg.up is the same provisioning step on its own. Run make proxy.install
once for the stable 127.0.0.1 endpoints
(see Networking); make start keeps them re-pointed afterwards.
The first connection also needs a one-time macOS Local Network grant — see
macOS Security.
Status prints endpoints similar to:
.pgpass lines:
host.docker.internal:5442:*:<PG_USER>:<PG_PASSWORD>
host.docker.internal:5443:*:<PG_USER>:<PG_PASSWORD>
psql commands:
active: psql --host=host.docker.internal --port=5442 --username=<PG_USER> --dbname=<PG_DB>
staging: psql --host=host.docker.internal --port=5443 --username=<PG_USER> --dbname=<PG_DB>
Put the printed lines in ~/.pgpass. The endpoint host is permanent — it does
not change across reboots or machine recreation, so saved connection strings keep
working. PG_PROXY_HOSTNAME sets that host (default host.docker.internal; use
127.0.0.1 for host-only access).
Port 5442 always means the active/current dataset. Port 5443 always means the opposite staging dataset.
make start
make pg.status
psql -h host.docker.internal -p 5442 -d "$PG_DB"
make pg.logsLoad a fresh dump without blocking the active database:
# 1. Reset staging to a clean start: soft (reflink, instant) or, to also reclaim
# macOS disk from prior imports, hard (delete+recreate the staging machine).
make pg.staging.reset # soft
# make pg.staging.rebuild # hard reset — reclaims disk; active untouched
# 2. Import through the staging port on the stable endpoint.
pg_restore --host=host.docker.internal --port=5443 --dbname="$PG_DB" \
--jobs=4 your-dump.pgdump
# 3. Verify and checkpoint staging.
psql -h host.docker.internal -p 5443 -d "$PG_DB" -c '\dt'
make pg.staging.snapshot name="$(date +%Y-%m-%dT%H-%M-%S)_dump_import"
# 4. Swap roles. Open connections reconnect; host and ports stay the same.
make pg.promotemake pg.promote requires both backends running (start staging with make pg.staging.start if stopped). If the new data is bad, make pg.promote again
immediately points :5442 back to the previous machine and its untouched data.
Unprefixed commands operate on the active physical slot:
make pg.snapshot name="$(date +%Y-%m-%dT%H-%M-%S)_before-migration"
make pg.restore name=<snapshot>
make pg.restore-last
make pg.snapshotsThe staging slot has the parallel command family:
make pg.staging.snapshot name=<snapshot>
make pg.staging.restore name=<snapshot>
make pg.staging.restore-last
make pg.staging.reset # soft reset (reflink)
make pg.staging.rebuild # hard reset: recreate the machine, reclaim disk
make pg.staging.stop
make pg.staging.startforce=1 replaces a same-named snapshot:
make pg.snapshot name=before-test force=1Snapshot names may contain letters, digits, dots, underscores, and hyphens.
make status # endpoints, roles, states, IPs, timelines
make status/incus # Incus versions/resources/list
make pg.ip # both machines' IPs and endpoints
make pg.refresh # re-discover both machine IPs, re-point the client proxy
make machine.status # Apple machine JSON (both)
make machine.shell # shell into the active machine (slot=a|b to pick)
make pg.shell # shell in the active backend; pg.staging.shell for staging
make pg.logs # tail active PostgreSQL logs; pg.staging.logs for stagingSnapshots and the XFS data trees disappear with the Apple machine. There is no
built-in full-setup export/import; before a destructive step, dump anything you
need over the client ports with ordinary pg_dump (e.g. pg_dump -h 127.0.0.1 -p 5442 …) and restore it with pg_restore into the staging port afterward.
make stop # stop both machines; keep all data
make system.stop # also stop Apple's container services
make pg.down # delete each machine's backend container AND its XFS data tree
make delete # delete BOTH machines and everything inside them
make recreate # delete both, rebuild, and start freshmake delete, make recreate, and make pg.down are destructive and hit
both machines. To reclaim macOS space while staying live, prefer make pg.staging.rebuild — it deletes+recreates only the staging machine (freeing
its sparse image) and never touches active. make recreate is the full nuke.
On macOS 15 Sequoia and later (incl. 26 Tahoe), Local Network Privacy gates
any connection to the machines' 192.168.64.0/24 subnet, so the client proxy
needs that permission — granted once, and only once.
Until it is granted, the proxy accepts your client on :5442/:5443 but cannot
reach the backend (EHOSTUNREACH), and psql reports "server closed the
connection unexpectedly." Grant it under System Settings → Privacy & Security
→ Local Network (enable the socat entry), then restart the running agents so
they re-evaluate it — macOS caches the decision at process start, so a proxy that
was already running when you ticked the box stays blocked:
make proxy.uninstall && make proxy.installThe grant then sticks across rebuilds. It is keyed to the binary's
code-signing identity, and the proxy runs Homebrew's socat — one stable,
already-signed binary at a fixed path that this repo never rebuilds. (The
removed Go forwarder was the opposite: make re-signed it on nearly every
build, so each rebuild looked like a new program — a fresh prompt and another
duplicate row in the Local Network list.)
If a client hangs after the grant, check make proxy.status and the socat logs
in var/<prefix>-socat-<role>.log.
Disk space — the sparse-VM-disk trap (important): the Apple container
machine's root disk (vdb) is a sparse image on macOS that only grows.
Blocks written inside the guest (the XFS PostgreSQL store, pg_restore output,
Incus image layers) are added to the macOS-side image and are not returned to
macOS when you delete them — Apple's runtime does not compact the disk on
discard/TRIM, and there is no supported per-machine size cap. So a large restore
can silently ratchet your Mac's free space to zero. When the next guest write
then fails, the loop-backed XFS store shuts down mid-write with an I/O error and
PostgreSQL drops into recovery (you'll see Input/output error on data files and
XFS … Filesystem has been shut down).
Controls:
make disk— show macOS free space, the Apple container storage footprint, the guest root-disk usage, and the.xfsstore's actual (physical) size.make disk.check— a fail-fast pre-flight (macOS free space vsDISK_MIN_FREE_GB, default 40 GiB) that gatespg.upand the staging restore/reset commands. SetDISK_MIN_FREE_GBto at least the size of the dump you're about to restore (a restore can grow the image by roughly the DB size). Note the client-sidepg_restoreitself runs outside make, so this guards the step right before it, not the copy.- Reclaim a bloated VM disk by deleting the machine (the only dependable
shrink —
container system pruneonly touches the CLI's image/build cache). Everyday:make pg.staging.rebuildfrees the staging machine's image while active keeps serving. Full:make recreate(dump anything you need first). - Cap future growth (strategic): relocate the bulky data onto a dedicated
APFS volume with a quota (
diskutil apfs addVolume … -quota …) mounted into the machine, so the payload can't consume the whole macOS volume. Not wired up here yet — seeissues/.
Repository location: The repository must live under your macOS home directory.
The Apple container machine only mounts $HOME (home-mount), and every make
target executes the repo's scripts inside the machine via that mount. A repo
outside $HOME fails with a raw missing-directory error.
systemd status: systemctl is-system-running inside the machine reports
degraded permanently. This is expected and harmless—Apple's guest kernel has
no loadable-module support, so systemd-modules-load can never succeed.
See .env.example. The main settings are:
MACHINE_PREFIX(defaultvpg→vpg-a/vpg-b),MACHINE_CPUS,MACHINE_MEMORY— per machine (both share the Mac, so memory is per-machine);PG_DATA_DISK_SIZE— first-creation XFS logical size (per machine);PG_BACKEND_PORT— port each backend is exposed on (5432);PG_CLIENT_ACTIVE_PORT,PG_CLIENT_STAGING_PORT— host loopback ports the client proxy listens on (5442/5443);PG_PROXY_HOSTNAME— hostname printed in psql/.pgpass lines (defaulthost.docker.internal;127.0.0.1for host-only);PG_PROXY_DEBUG— debug-level structured logging for the tracking DB and the proxy reconcile (on by default);PG_LOG_COLOR(auto/always/never) for its coloring —autocolors only a terminal and honorsNO_COLOR, so redirected runs stay clean and fully dated — andPG_LOG_FORMAT(text/json) to swap the colored rendering for machine-readable JSON.
The lifecycle logic lives in the Go binaries — pgdev on the host, pgdevd in
the machine — not in the script. The make targets are thin, tab-completable
wrappers that call those binaries and cover the Apple-machine chores around them
(make start/delete/disk). The residual scripts/pg-dev-local now holds
only the interactive shell/logs passthroughs, pending their move to SSH.
