ZFS backend (experimental)
Status: experimental. The zfs backend is unit-tested (command/plan construction, engine saga with a fake driver) and has a gated integration test (
PGOVERLAY_ZFS_IT=1), but it has not been exercised in CI on this project's development machine — Colima/macOS has no ZFS in the VM. This page contains everything needed to verify it by hand on a real Linux host. The default OverlayFS backend remains the supported path.
Why a second backend
OverlayFS copies up whole files. Since v1.0.0 the overlay backend's
lazyrw shim makes that happen on a file's first write rather than its first
read-write open, so reads copy nothing; but on ext4 the first write to a
table or index segment still copies the whole segment (up to 1 GiB), and a
branch grows by every segment it writes. Where the volumes sit on XFS
(reflink=1) or btrfs, the overlay backend's copy-up becomes a clone and
writes are block level too
(concepts). ZFS does
copy-on-write at the block level on its own: a clone shares blocks with
its origin snapshot and pays only for the blocks it actually changes, no
matter which file they live in or how Postgres opens them, and it needs
neither the shim nor a reflink filesystem. If you already run ZFS, and
especially if your branches write into many large tables, the trade is
usually worth it.
Three consequences worth knowing:
- ZFS clones have no copy-up, so the
recovery_init_sync_method=syncfsflag the overlay entrypoint needs (see benchmarks — The fix) isn't load-bearing here. The zfs entrypoint keeps it anyway for parity — it is simply harmless on ZFS. - ZFS branches run without the lazyrw shim (
--lazyrwis ignored): a read of a clone never copies anything, so there is nothing for it to do. - Seed settle applies to zfs seeds as to overlay ones (
--seed-settle, defaultfreeze): the seed dataset is recovered, frozen, analyzed and cleanly shut down before its first snapshot, so branches start without WAL replay and their reads write no hint bits into the clone. - The branch entrypoint shrinks to perms + stale-pid cleanup + exec
(
internal/cow/entrypoint_direct.sh, shared with the CSI backend — both hand the container a ready-made writable clone): there is nothing to assemble, the clone is the writable data directory.
How it works
With branchd --cow zfs --zfs-dataset <prefix>, layers become ZFS datasets
under <prefix> instead of docker/kube volumes:
| pgoverlay object | overlay backend | zfs backend |
|---|---|---|
| source generation N | volume pgoverlay-src-<name>[-gN] |
dataset <prefix>/src-<name>-gN |
| branch writable layer | volume pgoverlay-br-<name>-rw |
dataset <prefix>/br-<name> (clone) |
| branch create | empty rw volume + in-container overlay mount | zfs snapshot <src>@br-<name> + zfs clone |
| branch destroy | remove rw volume | zfs destroy -r clone, then the snapshot |
| branch usage | du -sb on the rw volume |
zfs list -Hp -o used <clone> |
- Seeding (
pg_basebackup) targets the dataset's mountpoint, bind-mounted into the seed helpers. The backend assumes default mountpoints (/<dataset>— noaltroot, no custommountpoint=). - zfs commands run in privileged one-shot helpers (the pinned alpine
helper image,
runtime.UtilityImage: alpine 3.24 by digest,--privileged, host/dev/zfsmapped in; on kube, a privileged pod). The helper installs the zfs userland at run time (apk add zfs), so it needs network access and an alpinezfspackage version compatible with the host's zfs kernel module. - Branch containers bind-mount the clone's mountpoint at
/pgoverlay/rwand run withPGDATA=/pgoverlay/rw/data. No overlay assembly; a branch of a settled seed starts from its clean shutdown, and with--seed-settle=offit runs WAL crash recovery on first boot, exactly as the overlay backend does. They still run withCAP_SYS_ADMINand AppArmor unconfined: the runtime drivers give every branch container those settings, whatever the backend (see Security). - Destroys are idempotent: an already-absent dataset/snapshot doesn't fail the destroy (so a half-created, failed branch stays destroyable), but a destroy that fails with the target still present — e.g. a busy clone — is surfaced as an error.
The registry stores dataset names in the same columns that hold volume names
on overlay. Do not switch an existing PGOVERLAY_HOME between backends —
seed sources freshly under the backend you intend to use.
Requirements
- Linux host where the docker daemon runs (the pool must be visible to the kernel that runs your containers — on macOS that means inside the VM, which Colima does not provide; this is why the IT is skipped there).
- An imported zpool and a dataset prefix pgoverlay owns, e.g.
tank/pgoverlay. /dev/zfspresent on the host (zfs kernel module loaded).- Default mountpoints for everything under the prefix.
- Outbound network from helper containers (
apk add zfs).
Manual verification walkthrough
On a Linux box with docker and ZFS installed (a file-backed pool is fine for testing):
$ truncate -s 10G /var/tmp/pgoverlay-pool.img
$ sudo zpool create tank /var/tmp/pgoverlay-pool.img
$ sudo zfs create tank/pgoverlay
$ ls -l /dev/zfs
crw-rw-rw- 1 root root 10, ... /dev/zfs
Start a demo source and branchd in zfs mode (same demo source as the README quickstart):
$ docker run -d --name demo-src -e POSTGRES_PASSWORD=secret postgres:17 \
-c wal_level=replica -c max_wal_senders=4
$ docker exec demo-src sh -c 'until pg_isready -U postgres; do sleep 1; done'
$ docker exec demo-src psql -U postgres \
-c "CREATE TABLE t(i int); INSERT INTO t SELECT generate_series(1,100000);"
$ docker exec demo-src sh -c \
'echo "host replication all all scram-sha-256" >> "$PGDATA/pg_hba.conf"'
$ docker exec demo-src psql -U postgres -c "SELECT pg_reload_conf();"
$ SRC_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' demo-src)
$ export PGOVERLAY_TOKEN=$(openssl rand -hex 16)
$ ./bin/branchd --cow zfs --zfs-dataset tank/pgoverlay
2026/06/10 12:00:00 REST API listening on :7070
Seed a source — expect a new src-main-g1 dataset holding a data/ dir:
$ AUTH="Authorization: Bearer $PGOVERLAY_TOKEN"
$ curl -H "$AUTH" -d "{\"name\":\"main\",\"host\":\"$SRC_IP\",\"port\":5432,
\"user\":\"postgres\",\"password\":\"secret\"}" localhost:7070/v1/sources
{"name":"main","state":"ready",...}
$ zfs list -r tank/pgoverlay
NAME USED AVAIL REFER MOUNTPOINT
tank/pgoverlay ... ... ... /tank/pgoverlay
tank/pgoverlay/src-main-g1 ... ... ... /tank/pgoverlay/src-main-g1
$ sudo ls /tank/pgoverlay/src-main-g1
data
Create a branch — expect a snapshot + clone and a near-instant create:
$ curl -H "$AUTH" -d '{"name":"pr-1","source":"main"}' localhost:7070/v1/branches
{"name":"pr-1","state":"ready","port":<P>,...}
$ zfs list -r -t all tank/pgoverlay
NAME USED AVAIL REFER MOUNTPOINT
tank/pgoverlay/src-main-g1@br-pr-1 0 - ... -
tank/pgoverlay/br-pr-1 ... ... ... /tank/pgoverlay/br-pr-1
Verify isolation and usage (block-level CoW: the clone's used stays small
after small writes — no whole-segment copy-up):
$ psql "host=localhost port=<P> user=postgres password=secret" \
-c "DELETE FROM t WHERE i > 50000"
$ docker exec demo-src psql -U postgres -c "SELECT count(*) FROM t" # still 100000
$ curl -H "$AUTH" localhost:7070/v1/branches/pr-1/usage
{"bytes":<N>} # ≈ `zfs list -Hp -o used tank/pgoverlay/br-pr-1`
Destroy — expect the clone and snapshot to be gone:
$ curl -H "$AUTH" -X DELETE localhost:7070/v1/branches/pr-1
$ zfs list -r -t all tank/pgoverlay
NAME USED AVAIL REFER MOUNTPOINT
tank/pgoverlay ... ... ... /tank/pgoverlay
tank/pgoverlay/src-main-g1 ... ... ... /tank/pgoverlay/src-main-g1
The same flow is automated as TestZFSEndToEndBranching:
$ PGOVERLAY_ZFS_IT=1 PGOVERLAY_ZFS_DATASET=tank/pgoverlay \
go test ./internal/engine/ -run TestZFSEndToEnd -count=1 -v
Cleanup: zpool destroy tank && rm /var/tmp/pgoverlay-pool.img.
Kubernetes note
The zfs backend follows the same storage-node model as overlay: the
zpool must live on the designated storage node, every zfs helper runs there
as a privileged pod, and branch pods mount the clone mountpoints via
hostPath (type Directory — a missing mountpoint fails the pod visibly
rather than starting on an empty dir). The Helm chart does not wire --cow
flags yet; run branchd with --cow zfs --zfs-dataset ... via a chart fork or
manual Deployment edit if you want to try it in-cluster.
Known limitations
- Experimental: no CI coverage on real ZFS; verify with the walkthrough above.
- Helper containers install the zfs userland at run time (network required; alpine package vs host kernel-module compatibility is on you).
--cowis a branchd flag; local-modepgb(no--server) is overlay-only.- One backend per
PGOVERLAY_HOME— don't mix. - A parent with live children cannot be destroyed or reset (their clones depend on snapshots of its dataset); both are refused up front.
Branch-from-branch is not a limitation here — it is the backend's best
feature. CreateBranchFrom snapshots the parent's clone and clones that
(zfs snapshot <parent>@br-<child> + zfs clone), so unlike the overlay
backend there is no freeze: the parent is never checkpointed, stopped or
restarted, and no layer rows are written. Covered by
TestZFSCreateBranchFromSnapshotsParentClone against the fake driver; like
everything else on this page, unverified on real ZFS.
One consequence to know: a child's recorded base is the parent's live
dataset, not the snapshot taken at the fork. Resetting the child, or running
pgb diff on it, snapshots the parent again, so the reset returns the child
to the parent's current state and the diff compares against it: changes
the parent made after the fork show up reversed in the child's diff.