Skip to content

Architecture (as built)

How pgoverlay actually works — written from the code rather than from the original design spec, so where the two disagree this document follows the code. For the file-by-file map of which package does what, see the code tour; for why each major structural choice was made, see the design decisions.

Components

            pgb (CLI) ────────────────┐
            │ local mode:             │ server mode: REST + bearer token
            │ embeds the engine       ▼
            │                ┌─ branchd ────────────────────────────────┐
            │                │  REST API :7070   (+ embedded web UI)    │
            │                │  pgproxy :6432    (wire-protocol router) │
            │                │  reconcile loop   (reap, repair, GC)     │
            │                └──────────────┬────────────────────────────┘
            ▼                               ▼
        ┌─ engine ──────────────────────────────────┐
        │  sagas: create/from-branch/reset/recover/   │
        │  destroy; diff; masking                     │
        │  cow.Planner: overlay | zfs | csi (names,    │
        │  entrypoints, zfs argv, clone plans)        │
        └───────┬──────────────────────┬────────────┘
                ▼                      ▼
        registry (SQLite)      runtime.Driver
        states + journal       ├─ DockerDriver (containers, volumes)
        sources, branches,     └─ KubeDriver   (pods; hostPath dirs on one
        layers, mask scripts,                   storage node, or CSI PVCs)
        tokens

One engine, two frontends: the CLI embeds it directly (local mode), branchd serves it over REST and the CLI becomes a thin API client (server mode). The registry is SQLite — a single writer. branchd runs as one replica by default; with leader election (High availability) extra replicas stand by and only the leader writes. Local mode must not run against a registry a branchd is using.

The CoW mechanism (overlay backend, default)

pgb source add runs pg_basebackup in a one-shot helper container, streaming the source cluster into a named volume — that volume is the read-only lower layer for every branch. pgb branch create makes one empty volume for the branch's writes and starts a stock postgres container whose entrypoint assembles an OverlayFS mount inside the container:

 ┌─ branch container (CAP_SYS_ADMIN) ──────────────────────────┐
 │   PGDATA = /pgoverlay/merged   ← overlayfs mount             │
 │                ▲                                            │
 │     ┌──────────┴───────────┐                                │
 │     │ upper+work (writes)  │  volume: pgoverlay-br-pr-1-rw   │
 │     ├──────────────────────┤                                │
 │     │ lower (read-only)    │  volume: pgoverlay-src-main ────┼─▶ shared by
 │     └──────────────────────┘  (settled pg_basebackup copy)  │   all branches
 │   entrypoint.sh: mount overlay → lazyrw self-test + probe → │
 │     LD_PRELOAD=liblazyrw → exec docker-entrypoint.sh        │
 └─────────────────────────────────────────────────────────────┘

Mounting in-container (not on the host) is what makes the same code work on Colima/macOS and bare Linux. Postgres boots on the merged view. A settled seed (the default, below) was shut down cleanly, so the branch starts without recovery; an unsettled one (--seed-settle=off) gets ordinary WAL crash recovery, as if the machine power-cycled at backup time.

One non-obvious flag is load-bearing: the entrypoint execs postgres with -c recovery_init_sync_method=syncfs. The default (fsync) opens every data file read-write before recovery, and a read-write open of a lower-layer file forces a full OverlayFS copy-up — i.e. a complete copy of the database into the branch's rw layer. syncfs replaces that per-file pass with one syscall and copies nothing up; it's what makes branch creation O(1) in data size (measured in Benchmarks).

The same rule applies after startup. Postgres opens every relation segment read-write, for reads too (md.c), so on its own the first query that touches a table would copy each of its segment files (up to 1 GiB) whole into the branch's rw layer (Reads copy up too). Three pieces keep a branch's cost to what it writes:

  • The lazyrw shim (internal/cow/lazyrw), an LD_PRELOAD library the entrypoint loads into the branch's Postgres, opens relation and transaction-status files read-only and reopens a file read-write on its first write-class call, which is when OverlayFS copies it up. Reads copy nothing; a write copies the file it touches once. The builds (glibc and musl, x86_64 and aarch64) are committed with their SHA-256 sums, embedded with go:embed, and installed into /pgoverlay/rw/lazyrw/ by the same helper that writes the entrypoint (base64 in its environment, checked after decoding), on create, reset, both freeze volumes and recover. Before exporting LD_PRELOAD the entrypoint self-tests the kernel (a read-only descriptor opened before a copy-up must see the data written after it, which OverlayFS guarantees since Linux 4.19) and probes the build against the image's postgres; it pins io_method=worker on PG 18 and later. It writes lazyrw, eager or off to /pgoverlay/rw/cow-mode, and the engine reads that after every readiness wait, logs it (a warning for eager while --lazyrw=on) and exports pgoverlay_branch_cow_mode (internal/engine/cowmode.go). --lazyrw=off turns it off; the setting travels as PGOVERLAY_LAZYRW in the branch's environment, so a branch picks a change up on its next start.
  • Seed settle (internal/pgctl/settle.go) makes reads write nothing: see below.
  • Copy-up detection (internal/cow/fsprobe.go, internal/engine/cowfs.go): on a filesystem that reflinks (XFS reflink=1, btrfs) the kernel turns a copy-up into an extent clone, which is block-level copy-on-write. branchd probes which one it has at startup, exports pgoverlay_cow_copyup_mode, sets an XFS copy-on-write extent size hint on a volume root it manages, and in clone mode counts branch usage as exclusive bytes with the embedded pgoverlay-du instead of du -sb. --volume-root DIR (docker) creates every volume as a bind volume under DIR (internal/runtime/docker_volumeroot.go), so any host with such a disk gets clones without moving Docker. See Core concepts.

Every branch container gets CAP_SYS_ADMIN and apparmor=unconfined for the mount, on Docker whatever the backend. Branch ports are published on the Docker host's 127.0.0.1: branchd picks a free loopback port and pins it, and containers restart with the daemon (unless-stopped), so the address survives Docker and host restarts.

Seeding: basebackup vs dump

The seed itself has two methods. pg_basebackup (default) is a physical, crash-consistent copy: fast at size, but it requires REPLICATION privilege and a physical replication connection, which managed providers (Supabase, Neon, RDS, Cloud SQL) don't offer — and, left as it is, it makes every branch replay WAL on first start, as if the machine power-cycled at backup time (the settle below does that once instead). --via dump is a logical copy: a helper container runs initdb into the seed volume and pipes pg_dump from the remote into it, needing only a normal user, optionally scoped to schemas (--dump-schema). It is slower for large databases (full SQL restore, index rebuilds), but the resulting layer is a clean-shutdown cluster, so branches skip crash recovery entirely. Either way the seed is just a data dir in the source layer — everything downstream (overlay/zfs/csi branching, refresh generations, masking) is identical.

After either method the seed is settled (--seed-settle, default freeze), on every backend. For a basebackup seed a helper on the branch image starts Postgres once on the copy (socket only, with the source's preload libraries, archiving, TLS and other settings that cannot start in a throwaway container overridden), which completes the backup's recovery; runs VACUUM (FREEZE, ANALYZE) on every database and then freezes pg_statistic once more; checkpoints; switches to a fresh WAL segment; and stops it cleanly. It then trims that segment to a hole after the shutdown checkpoint (the file reads the same) and removes unused segments after it, so a branch's first WAL write copies about 1 MiB, not 16 MiB. A dump seed runs the same VACUUMs before its own clean stop, and the same trim. Every helper runs as the image's postgres user. Branches then start from a clean shutdown, and a read in a branch has no hint bits to set, nothing to prune and no anti-wraparound VACUUM due, so it writes nothing and, with the shim, copies nothing. recover skips the VACUUM and off skips the step. The settle is part of the seed: it runs under the seed's heartbeat and fails the seed only if the copy cannot start or stop cleanly (a failed VACUUM is a warning).

A basebackup of a standby is made safe to boot: the seed deletes the standby and recovery signal files and strips the recovery settings, and the branch entrypoints override the settings that would tie a branch to the source host (port, listen and socket addresses, hba_file, archiving, synchronous standbys). The seed helper also checks the data directory's major version against the image and fails a mismatched seed. Each source can name its own image (--image) for extensions, locales or libc the stock postgres:<major> image lacks.

The zfs backend (experimental)

branchd --cow zfs --zfs-dataset tank/pgoverlay swaps the layer mechanics: sources seed into datasets (<prefix>/src-<name>-gN), branch create is zfs snapshot + zfs clone (block-level CoW — no copy-up, no shim, no overlay assembly; the entrypoint shrinks to perms + pid cleanup + exec), and zfs commands run in privileged helpers with /dev/zfs mapped in. Same engine, same sagas — cow.Planner decides what the driver is asked to do. Details and verification walkthrough: ZFS backend.

Sagas and states

Branch rows move through a journaled state machine (legalBranch in internal/registry/registry.go); every transition is a compare-and-swap written together with its journal row, and records the actor:

stateDiagram-v2
    [*] --> creating
    creating --> ready
    creating --> failed
    ready --> resetting: reset, freeze or csi quiesce, reconcile restart
    ready --> destroying
    resetting --> ready
    resetting --> failed
    failed --> resetting: reset, or recover onto the existing data
    failed --> destroying
    destroying --> destroying: failed attempt (journaled note), retried
    destroying --> destroyed
    destroyed --> [*]

failed is not terminal: pgb branch reset re-clones a failed branch (for a failed create, a retry), and pgb branch recover restarts it on its recorded volumes with no re-clone — the way back for a branch that failed with its data intact, such as a freeze parent interrupted by a crash. A destroy that fails part-way leaves the row in destroying with a journaled note, and running the destroy again (by hand, or reconcile's retry_destroy) re-runs the idempotent teardown. A resetting branch that is destroyed is first moved to failed. Sources have a smaller machine: seeding → ready | failed; a failed seed is replaced by the next source add of the same name.

Provisioning is a saga: each step (layer create, entrypoint install, container start, readiness wait, masking) registers a compensation, and any failure unwinds them in reverse — no orphaned containers, volumes, or datasets. While a saga runs it heartbeats its rows (every min(--stuck-timeout/4, 30s)), so reconcile can tell a slow operation from an abandoned one. Masking scripts (per-source ordered SQL, stored in the registry) run inside the fresh branch via psql over the local socket after readiness and before the branch is marked ready, so a branch never serves unmasked data; a failing script fails the branch. Failure reasons are stored without Postgres DETAIL/CONTEXT lines (which can quote row data) and capped at 1 KiB.

The reconcile loop runs one pass at startup and then every --reconcile-interval. It reaps expired branches; fails creating/resetting branches and seeding sources that have made no progress for --stuck-timeout (so a crash less than that ago is repaired by a later pass, not the startup one); retries destroys stuck in destroying; restarts ready branches whose container or pod is gone and records moved addresses; removes orphaned containers and finished helpers; and garbage-collects unreferenced layers and volumes older than --stuck-timeout. Every destructive step is re-checked against the registry just before it runs. The action list is in Troubleshooting.

Source generations

pgb source refresh seeds a new layer (...-g2, -g3, …) and bumps the source's generation. Existing branches keep the layer they were cloned from; new branches use the new one. An old generation is garbage-collected when the last branch referencing it is destroyed. Branches never follow the source — a branch is a point-in-time snapshot by design.

Per-branch credentials

By default a branch inherits its source's credentials: the data directory is a byte-for-byte clone, so the same role/password just works. branchd --rotate-branch-credentials (chart: rotateBranchCredentials) switches to per-branch passwords: on every branch create and reset the engine generates a 32-hex crypto/rand secret and applies it inside the branch — ALTER ROLE … WITH PASSWORD over the same local-socket psql path masking uses, after masking and before the branch is marked ready — then stores it on the branch row (registry v7), encrypted with AES-256-GCM under a dedicated at-rest key (secret.key in the state directory, or PGOVERLAY_SECRET_KEY; see Security). The API returns it as password (omitted in inherit mode), pgb connect embeds it in the printed DSNs, and pgb branch ls never shows it. A password the configured key cannot decrypt is reported as password_unavailable instead of failing the row; a reset mints a new one. Branch-from-branch children get their own password; the parent's freeze/quiesce restart deliberately does not re-rotate (its data already carries its password), and neither does a recover. A reset rotates again — the old branch password stops working. A destroyed branch's row keeps no password.

The trade-off: with rotation a leaked branch DSN exposes only that branch, never the production-shaped source. But static-credential preview flows break — Vercel-style env templating (one DATABASE_URL template with the branch name substituted per preview) relies on every branch accepting the same source password, so those flows need inherit mode (the default).

Branch diff

DiffBranch answers "what changed in this branch?" without ever touching the source: it provisions an internal throwaway branch (diff-<6 hex>) from the target's own recorded base — whatever a reset of the target would re-provision onto, never the source's current generation. Both instances are then dumped in-container over the local socket (pg_dump --schema-only --no-owner --no-acl, plus a table-statistics query), the unified diff is computed host-side (a linear-space Myers diff), and the throwaway is destroyed through the normal branch-destroy path. The throwaway is a regular registry row with a one-hour TTL, so if branchd dies mid-diff the reaper (or reconcile) cleans the stray.

What "the base" is depends on the backend:

  • overlay: the pinned source-generation volume and frozen-layer chain, so the baseline is exactly what the branch started from;
  • zfs and csi, for a branch created from another branch: the parent's live dataset or PVC. The diff (like a reset) therefore compares against the parent's current state, and changes the parent made after the fork show up reversed. On csi the parent is quiesced around the clone (CHECKPOINT, stop, clone, restart) because cloning an in-use PVC is not crash-safe, and a child whose parent was destroyed can no longer be diffed or reset.

Row counts are planner estimates (pg_class.reltuples); a table never analyzed is counted exactly when its heap is 64 MiB or less, and otherwise reported as unknown. Deltas show direction and magnitude, not an audit. Because the default seed settle runs ANALYZE, every table in a settled seed has an estimate, so a change too small to trigger the branch's autovacuum ANALYZE shows as no change until the table is analyzed in the branch; tables the branch created are always counted exactly (while small). Tables are keyed by schema and name. Optional sampling (?data=N, at most 500 rows per table) returns branch-only rows of grown tables by primary key.

Proxy routing

Per-branch host ports are annoying, so branchd bundles a wire-protocol router. Clients connect to :6432 with the branch name suffixed to the database name:

 psql "dbname=postgres@pr-42" ──► pgproxy reads the startup message,
                                  resolves pr-42 → container host:port
                                  (registry lookup), rewrites dbname back
                                  to "postgres", then splices bytes

Everything after the startup message — including SCRAM authentication — is relayed untouched between client and branch. With --pg-tls-cert/--pg-tls-key the router answers SSLRequest with 'S' and terminates TLS before reading the startup message (sslmode=require works); without certs it answers 'N' as before. The REST API gets the same treatment via --api-tls-*.

While relaying the backend's startup response the router records its BackendKeyData, so a CancelRequest (psql's Ctrl-C, a driver's cancel) is forwarded to the backend of the one live session holding that key; unknown or ambiguous keys are dropped. The map is per branchd process, so with several replicas a cancel must reach the replica carrying the session. The router is also an unauthenticated surface, so it bounds each connection's startup (first byte in 2 s, startup in 10 s, backend ReadyForQuery in 30 s), caps connections (256) and per-IP startups (64), closes sessions idle in both directions for 15 minutes, and refuses every routing failure with the same message. After a failed dial it re-reads the branch's address once (rate-limited) in case the branch moved. Details in Security.

Branch-from-branch: the layer DAG

pgb branch create child --from-branch parent snapshots a running branch. Overlay mode can't snapshot a live upper dir atomically, so the engine freezes it: CHECKPOINT on the parent, stop its container, the parent's current rw volume becomes an immutable layer row in the registry, the parent restarts on a fresh rw volume layered above it, and the child starts on its own fresh rw volume with the same lower chain. Both now share the frozen layer read-only:

 parent:  [new upper]──┐
                       ├──► frozen layer (old upper) ──► source data
 child:   [new upper]──┘         (refcounted)

PGOVERLAY_LOWERS carries the chain newest-first; frozen layers contribute their /upper subdir, the source its /data. Layers are garbage-collected when the last branch whose chain references them is destroyed — destroying the parent first leaves the child (and the layer) intact. ZFS and CSI modes skip the freeze machinery entirely: they snapshot/clone at the block or volume level (ZFS: zfs snapshot + clone, with no parent interruption; CSI: PVC clone after a brief parent CHECKPOINT and stop, then the parent restarts).

Overlay chains only grow: each fork of a parent adds a layer, a reset keeps the chain, and there is no compaction yet. Branch-from-branch is refused (403) once the parent's chain reaches --max-layer-depth (default 100).

Kubernetes: storage-node and CSI models

The kube driver has two storage strategies. hostPath (default) maps "volumes" to subdirectories of a data root (default /var/lib/pgoverlay) on one designated storage node; helpers are one-shot pods and branches are plain pods, all pinned with nodeName, branch pods carrying SYS_ADMIN (and unconfined seccomp and AppArmor) for the overlay mount. hostPath branches run the same entrypoint as Docker branches, lazyrw shim included (the install helper carries the builds in its environment, which on Kubernetes travels in the helper's short-lived Secret), and a data root on XFS or btrfs gives clone copy-up. csi (--kube-storage csi) makes every volume a PVC and every branch a PVC clone (dataSource, or VolumeSnapshot+restore when a snapshot class is configured): branch pods need no SYS_ADMIN, no node pin — they schedule anywhere, which is the multi-node payoff. The trade-off: clone CoW economics belong to the CSI driver (instant on EBS/Ceph/zfs-localpv, full copy on naive drivers). Either way: no CRDs, no operator — branchd is a normal Deployment with a namespace-scoped Role: pods create/delete/get/list/watch, pods/exec, pods/log, and Secrets create/delete (helper pods read their environment, including a seed's source password, from a short-lived Secret instead of the pod spec), plus PVCs and VolumeSnapshots in csi mode and Leases and pod patch with leader election. Helper pods carry the pgoverlay.instance label and an ownerReference to branchd's own pod, so Kubernetes garbage-collects them if branchd dies mid-seed.

What runs where

concern mechanism
data files only ever touched inside containers (helpers/entrypoints)
host Go code pure control plane: registry, sagas, driver API calls
seeding pg_basebackup or pg_dump helper, runs as the image's postgres uid:gid (999:999 Debian, 70:70 Alpine); then the settle helper on the branch image, as the same user
copy-on-write in the branch the lazyrw shim, preloaded into the branch's Postgres only (overlay backend)
copy-up probe a helper with a branch container's privileges, once at branchd startup
masking, credential rotation, diff dumps exec into the branch (as postgres on Docker, as root on Kubernetes)
disk usage du -sb helper on the rw layer, or pgoverlay-du (exclusive bytes) when copy-up clones (zfs: zfs list -o used)
web UI single static page, go:embed, no build toolchain
GitHub App separate pgoverlay-github service driving the REST API