Benchmarks
Measured with hack/benchmark.sh on 2026-06-10, against
pgoverlay's Docker runtime with the OverlayFS copy-on-write backend. All numbers
are real, single-machine measurements — no extrapolation.
Results
| Database size | pgbench scale | Seed time | Branch create (p50 of 5) | Branch rw overhead after create | rw after updating 1% of rows |
|---|---|---|---|---|---|
| 1.00 GiB (1,074,124,467 B) | 68 | 20 s | 1.90 s | 33.1 MiB (34,748,732 B) | 1.04 GiB (1,113,471,292 B) |
| 5.00 GiB (5,370,484,403 B) | 342 | 41 s | 1.89 s | 33.1 MiB (34,748,732 B) | 5.20 GiB (5,586,819,388 B) |
Individual branch-create runs (seconds):
- 1 GiB: 2.934, 2.323, 1.851, 1.901, 1.791 → p50 1.901
- 5 GiB: 2.062, 1.890, 1.796, 1.890, 1.918 → p50 1.890
Branch creation is now independent of database size — ~1.9 s p50 at both 1 GiB and 5 GiB — and a fresh branch costs 33.1 MiB of disk (recycled WAL segments written during crash recovery plus overlay bookkeeping), not a copy of the dataset. That was not true before 2026-06-10 (see Before the fix). What a branch costs after it starts serving queries is a different number: see Reads copy up too for how it was before v1.0.0, and True copy-on-write (v1.0.0) for how it is now.
One number needs honest framing: the rw layer after the 1% UPDATE probe is now ≈ the full dataset size, where the old table showed only +10–179 MiB. The old probe was measuring growth of a layer that creation had already copied up in full; the new probe pays the copy-up at first write instead. OverlayFS copies up whole files, and Postgres heap/index segments are files of up to 1 GiB — so a write that touches a segment copies that entire segment into the rw layer, and this bulk UPDATE + CHECKPOINT ended up copying up essentially the whole pgbench dataset. The cost moved from create time to first use, and, as the next section shows, before v1.0.0 "use" included reads: a branch converged on the size of the tables it touched, not only the ones it wrote. (This table predates v1.0.0's seed settle and lazyrw shim. Writes still copy whole segments on ext4, so the 1% UPDATE column would not change much there; on XFS or btrfs it would be a fraction of it.)
10 GiB was not run in this pass; the script's disk check skips any size that
doesn't fit (each size needs ~2.2× the target free in the Docker VM) and sizes
are configurable: BENCH_SIZES_GIB="1 5 10" hack/benchmark.sh.
Reads copy up too
Fixed in v1.0.0 for the default setup: see True copy-on-write (v1.0.0). This section is kept as it was measured, because it is the evidence for that work, and because a branch whose kernel or image cannot run the lazyrw shim (eager mode) still behaves this way.
An earlier version of this page said that branches which mostly read stay
thin. That was wrong, and the pre-v1 review measured why (issue
#13). Setup: commit
c55cab2, Docker Engine 29.6.2 on Linux amd64 (ext4), PostgreSQL 17, a
489 MB table that was VACUUM (FREEZE)d on the source before seeding, so
62,500 of its 62,528 pages were all-frozen and setting hint bits could not
explain any growth. The rw layer was measured with the same du -sb helper as
above.
| Branch state | rw layer |
|---|---|
| fresh branch | 33.1 MiB |
after a 1-row UPDATE on a small table |
33.5 MiB |
after SELECT count(*) on a tiny table + CHECKPOINT |
34.6 MiB |
right after SELECT count(*) on the 489 MB table, before any checkpoint |
523.1 MiB |
The file that appeared in the rw layer, base/5/16384, was 488.5 MiB: the
whole table, copied by one read-only query.
Why. It is the SyncDataDirectory mechanism from The fix,
on the normal read path. PostgreSQL's storage manager opens every relation
segment read-write, whatever the query does: src/backend/storage/smgr/md.c
opens with O_RDWR | PG_BINARY in REL_14_STABLE (line 494), and through
_mdfd_open_flags() (line 146) in REL_17_STABLE. OverlayFS decides on
copy-up at open() time, from the open flags, before any byte is written: a
read-write open of a file that exists only in a lower layer copies the whole
file into the upper layer first. So the first query of any kind that touches
a table copies each segment it opens (up to 1 GiB per segment), and pays for
that copy inline. A small table costs almost nothing, which is why the 1-row
UPDATE above barely moved the layer.
What it means.
- A fresh branch still costs about 33 MiB and still starts in about 2 s; creation is unaffected.
- A branch grows by the size of every table and index file it opens, read or write, up to roughly the size of the data it touches. A test suite that scans a few large tables will copy those tables once per branch.
- The first query to touch a large table is slow on a new branch, because it waits for the copy.
syncfscannot help here: the open flags are how stock Postgres reads. Avoiding the copy needs either a filesystem that clones instead of copying, or changing the flags Postgres's opens reach the kernel with, which is what v1.0.0 does.
What the advice was. Before v1.0.0 this page pointed read-heavy branches
of large databases at the block-level backends (zfs, Kubernetes
csi mode). Since v1.0.0 the overlay
backend reads without copying, so that advice now applies only to branches
that write into many large tables on a host whose volumes cannot clone; and
a branch reported as eager by pgoverlay_branch_cow_mode should be fixed
(Troubleshooting) rather than
sized for.
True copy-on-write (v1.0.0)
Issue #49 set the bar:
on the default setup (plain Docker on an ext4 host, stock postgres:14 to
18 images) a read-only workload adds approximately nothing to a branch, a
write copies as little as the mechanism allows, create time stays
independent of database size, and no new privileges.
What changed
- The lazyrw shim, on by default on the overlay backend (Docker and
Kubernetes hostPath). An
LD_PRELOADlibrary in the branch's Postgres opens table files read-only and reopens a file read-write on its first write, which is when OverlayFS copies it (how it works). - Seed settle, on by default. Each new seed is recovered,
VACUUM (FREEZE, ANALYZE)d and cleanly shut down once, with its last WAL segment trimmed, so branches start without WAL replay, their reads have no hint bits to set, and their first WAL write copies about 1 MiB (how it works). - Block level where the filesystem allows it. Where the volumes sit on
XFS (
reflink=1) or btrfs, OverlayFS copy-up is an extent clone. branchd detects it, counts usage as exclusive bytes, and--volume-rootputs the volumes on such a disk on any Docker host (how it works).
The evaluation
Every option in #49 was measured on one host before anything was built: Linux 7.0 (amd64), Docker 29.6.2 with volumes on ext4, PostgreSQL 17, and a table of about 500 MB (1.5M rows, 488 MB of heap plus a 32 MB primary key) frozen on the source. Branches started with crash recovery, as they did then; latencies are with a warm page cache, so the first query's extra time is the copy.
| Mechanism | SELECT count(*) adds |
First query | A write copies | Verdict |
|---|---|---|---|---|
| OverlayFS on ext4 (before v1.0.0) | +555 MiB | 20.8 s | the whole file, on open | replaced |
OverlayFS with metacopy=on |
the same bytes | 11.3 s | the whole file, on open | rejected: the kernel copies data on any write-mode open |
| fuse-overlayfs | the same bytes | 15.8 s | the whole file, on open | rejected |
| lazyrw shim + OverlayFS, ext4 | +0 bytes | 277 ms | the touched segment (up to 1 GiB), once | default |
OverlayFS on XFS reflink=1 or btrfs |
~0 | ~10 ms per file | blocks | automatic where the volumes reflink |
| Managed loopback XFS pool on ext4 | ~0 | 208 ms | blocks | deferred: fsync ran at 0.4 to 0.6x of ext4, and it needs a privileged loop device lifecycle |
| btrfs snapshots, dm-thin, qcow2 over nbd | ~0 | about 0.2 s | blocks | rejected: privileged, fragile set-up, extra kernel modules, the slowest fsync |
Per-file FICLONE backend, no overlay |
0 | blocks | later: needs a reflink filesystem, and create time grows with the file count |
The raw OverlayFS numbers behind the first rows (bytes of the branch's upper
directory, du -sb): 34,735,718 B fresh, 582,461,030 B after the
SELECT count(*), and byte for byte the same with metacopy=on and with
fuse-overlayfs; the first SELECT took 19.4 s and the second 116 ms in that
run. With the shim the upper directory stayed at 16.2 MiB across the read.
Two findings from the evaluation shaped the implementation:
- An unsettled seed still copies on read. With a seed that was not
frozen, the shim's first
SELECTstill copied 68 MiB (PG 14 and 17) and 119 MiB (PG 18): reading rows sets their hint bits, which is a real write. pgoverlay'spg_basebackupseed is exactly such a copy, hence seed settle. - Select-only throughput is unchanged.
pgbench -Sgave 27,511 tps in a branch with the shim and 27,007 tps on plain Postgres on ext4. TPC-B numbers from that host were not usable: it was shared with other workloads, and plain ext4 TPC-B fell from 925 to 122 tps within 40 minutes. The pgbench runs below replace them: one on a host that was not quiet either, and one on dedicated GitHub-hosted runners.pg_test_fsync(8 kB write plusfdatasync, interleaved) is the comparison that held: 670/757/621 ops/s on ext4, 429/275/251 on XFS on a loop device, 170/117/95 on btrfs on a loop device, which is why the loop pool is deferred.
Measured on the implementation
The integrated tree (v1-cow at 65372f9), built with make build and run
with default settings on the same host: a postgres:17 source with a
200 MiB table that was never vacuumed, seeded with --seed-settle=freeze,
once with Docker's own volumes (ext4) and once with --volume-root on an
XFS reflink=1 disk. Branch usage is what pgb branch ls --usage reports.
| Check | ext4 (Docker volumes) | XFS --volume-root |
|---|---|---|
| seed including settle | 22.8 s (settle 6.6 s) | 17.9 s (settle 6.2 s) |
| branch create | 2.9 s | 3.5 s |
copy-up probe (pgoverlay_cow_copyup_mode) |
copy |
clone, 16 KiB cowextsize hint set |
cow-mode |
lazyrw |
lazyrw |
| seq scan of the table after settle | 25,408 pages read, 0 dirtied | same |
first read pass (count, EXPLAIN, pg_relation_size, full read), then CHECKPOINT |
data +0.24 MiB, WAL +16 MiB | same |
| second read pass | +0 bytes | +0 bytes |
1-row UPDATE of a small table |
+0.08 MiB | +0.08 MiB |
1-row UPDATE of the 200 MiB table |
+200 MiB (its segment, once) | +0.02 MiB exclusive (du -sb: +200 MiB) |
the same read with --lazyrw=off |
+201.5 MiB, cow-mode off |
same |
At that commit the first read pass still copied a constant amount, not
proportional to the data read: 224 KiB of pg_statistic (hint bits on the
rows ANALYZE wrote after pg_statistic itself was vacuumed) plus 8 KiB each
of pg_xact, pg_subtrans and pg_multixact, and the whole current 16 MiB
WAL segment the first time the branch wrote WAL (a clone on XFS). Both were
fixed before release; see the next section. Reset, branch-from-branch (the
child reads the parent's writes; the parent does not see the child's) and
destroy behaved the same in both runs.
Other measurements from the pull requests that built it (same host):
- a read of a frozen 196 MB table in a branch created by the engine: +0 bytes
in about 200 ms with the shim, +207 MB in 3.5 s with
--lazyrw=off(#53); PG 18 and apostgres:17-alpine(musl) branch also came uplazyrwwith reads adding 0 bytes; - a 1-row
UPDATEof a ~56 MiB table: +58 MiB with Docker's volumes on ext4, +44 KiB of exclusive bytes with--volume-rooton XFSreflink=1and on btrfs (#51); - a branch of a settled seed starts with no redo, and a sequential scan of a never-vacuumed table dirties 0 buffers, on PG 14, 16-alpine, 17 and 18 (#52).
After the release fixes: the torture suite
#56 removed the two
constants above and made Alpine sources seedable;
#57 added the
copy-on-write torture suite (internal/engine/cow_it_test.go) and ran the
version matrix with the shim. Measured on v1-cow at a77c260, same host,
ext4 Docker volumes, postgres:17, a 226.6 MB frozen table and a 226.6 MB
table that was never vacuumed:
| Check | Result |
|---|---|
read pass on either table (count, EXPLAIN, pg_relation_size, full scan), then CHECKPOINT |
+0 bytes of data; a second pass +0; neither the table, its index nor pg_statistic is copied (before #56: pg_statistic, 224 KiB) |
the same read with --lazyrw=off |
+258,129,920 bytes |
1-row UPDATE of the 226.6 MB table |
+226,598,912 bytes, the touched segment, in 12 to 18 s including the checkpoints that write it; a second UPDATE of it +0 |
1-row UPDATE of a small table |
+90,112 bytes |
TRUNCATE / DROP of a seed table |
42 ms / 69 to 102 ms, +1.29 MB of catalog pages, no table data |
CREATE INDEX (8.24 MB index) |
data +9.47 MB |
VACUUM FULL (255.8 MB rewrite) |
data +256.2 MB; the old file is not copied |
CREATE DATABASE, FILE_COPY and WAL_LOG |
both work |
TPC-B killed with SIGKILL, then a restart by reconcile |
balances add up, pg_amcheck --heapallindexed clean, still lazyrw |
| branch-from-branch | the child reads the parent's writes; its own reads +0 |
| reset | reads afterwards under 2 MiB |
--volume-root on XFS reflink=1 |
probe clone; read +0; a 1-row UPDATE of a 113 MB table +40,960 bytes of exclusive usage (du -sb: +113 MB) |
version matrix: 14, 15, 16, 17, 18 and 17-alpine (musl) |
all lazyrw, every read +0 bytes |
The WAL segment. A branch appends its WAL to the segment that holds the
seed's shutdown checkpoint, so its first WAL write used to copy that whole
16 MiB segment. Measured for #56 on ext4 Docker volumes with postgres:17,
5 fresh branches per variant: the first one-row INSERT in a branch took 46
to 119 ms (median 56 ms), the second about 1 ms. The settle, and the dump
helper unless --seed-settle=off, now switch to a fresh segment just before
the clean stop, so the shutdown checkpoint sits at the start of a zero-filled
segment, and turn everything after its first two pages into a hole. The file
stays byte for byte the same (the tail is checked to be zeros and the copy
compared with cmp before the rename), so crash recovery is unaffected; the
torture suite's SIGKILL and pg_amcheck runs use such seeds. Unused future
segments after it are removed too (a dump seed had carried 4 to 15 of them,
up to 240 MiB). With GNU tools and with busybox (the Alpine seeds are trimmed
as well):
| first write in a branch | the branch's copy of the segment | |
|---|---|---|
| before | 46 to 119 ms (median 56) | 16,384 KiB allocated |
after, pg_basebackup seed |
12 to 27 ms | about 1,024 KiB allocated |
| after, dump seed | 20 to 24 ms | about 1,024 KiB allocated |
Branch usage counts apparent bytes (du -sb), so it still shows that segment
as 16 MiB. wal_recycle=off was measured as an alternative and not adopted:
a settled seed has a single segment, which the branch's first write has
already copied, so there is no copy-up left for it to save; it only keeps
fewer recycled segments (16 to 32 MiB less after 45 MiB of WAL) at the cost of
zero-filling every new one. It stays available as the experimental
--wal-recycle=off.
What a branch of a settled seed still writes, whatever it reads: about 1 MiB
allocated for that WAL segment, 8 KiB each of pg_xact, pg_subtrans,
pg_multixact and pg_control, and about 190 KB of relcache init files.
Other timings from the same runs: a seed plus settle of 2 × 226 MB took 38 to
41 s (the settle alone 4.7 to 8.6 s on a 250 MB seed, with the extra
pg_statistic VACUUM), a branch create 2.4 to 3.2 s, and a branch-from-branch
9.4 s.
Throughput and the first-write stall
Measured with
hack/bench-cow.sh
(#58), driving a real
pgb in local mode (and branchd, for the copy-up probe). Every
configuration gets the same data: postgres:17, pgbench scale 50 (about
750 MiB) plus stall_big, a 949 MiB table in a single segment, seeded and
settled with freeze:
- A: plain Postgres on a Docker volume (ext4), no branch; the baseline.
- B: a branch with
--lazyrw=off, which copies a file when it is opened, as releases before v1.0.0 did. - C: a branch with the shim on Docker's volumes (ext4); the default.
- D: a branch with the shim and
--volume-rooton XFSreflink=1.
Each round creates the branch, runs pgbench select-only (-S) and TPC-B at
-c 8 -j 4 -T 60 once cold (right after create) and once warm (after the cold
run), and then times the first one-row UPDATE of stall_big. Rounds run
every configuration once, interleaved, so that anything else happening on the
host falls on all of them. It ran twice: on dedicated GitHub-hosted runners,
which decided the release gate, and before that on a shared host, where the
write half of the gate could not be decided. Cells are medians, with the
range (min to max) where the spread matters.
"Allocated and not shared" is what pgoverlay-du counts: on XFS the bytes a
branch does not share with the seed, which is what branch usage reports in
clone mode; on ext4, where nothing is shared, the allocated size.
On a quiet host: GitHub-hosted runners
The bench-cow
workflow ran hack/bench-cow.sh on two fresh GitHub-hosted VMs that ran
nothing else (run 36541991870,
commit a1507f4: v1-cow at 6df469b plus the workflow and the harness
changes described below):
- amd64:
ubuntu-24.04, 4 vCPUs (AMD EPYC 9V45), 15,989 MiB of RAM, Linux 6.17.0-1022-azure, Docker 28.0.4, Docker's volumes on ext4 on the VM's NVMe disk. A, B, C and D. - arm64:
ubuntu-24.04-arm, 4 vCPUs (Neoverse-N2), 15,947 MiB, the same kernel and Docker, ext4. A, B and C, with the shim's aarch64 glibc build.
Five rounds on each. The order rotated every round (A B C D, then B C D A,
and so on), so no configuration always ran right after another one's
copy-up. The 1-minute load average was 6 to 9 during the runs: that is the
benchmark itself, 8 pgbench clients and 8 backends on 4 vCPUs (it was 2.6
before the first round). The UPDATE of stall_big is followed by a timed
CHECKPOINT, after an untimed one that flushes the TPC-B runs, because with
the shim the page is written out, and the segment copied, there and not in
the UPDATE (see the first-write item below).
D's --volume-root is the only XFS a hosted runner can have: a sparse 24 GiB
file on the same ext4 NVMe filesystem as Docker's volumes, attached as a loop
device with direct I/O and formatted mkfs.xfs -m reflink=1. The copy-up
probe reported mode=clone on it (the 64 MiB probe file used 0 bytes) and
mode=copy on ext4. D's throughput is XFS on a loop device on ext4: two
filesystems and a loop device on every write and fsync, not an XFS disk.
amd64:
| A: plain | B: eager | C: lazyrw, ext4 | D: lazyrw, XFS on a loop device | |
|---|---|---|---|---|
| branch create (s; A: container start) | 0.14 | 1.37 | 1.36 | 1.49 |
branch after create (MiB, du -sb) |
n/a | 1.3 | 0.4 | 0.3 |
after a select-only run (MiB, du -sb) |
n/a | 768.0 | 16.5 | 16.5 |
| after a select-only run (MiB, allocated and not shared) | n/a | 753.1 | 1.6 | 0.5 |
after TPC-B (MiB, du -sb) |
n/a | 1,821.4 | 1,812.9 | 1,806.7 |
| after TPC-B (MiB, allocated and not shared) | n/a | 1,821.4 | 1,812.9 | 1,818.3 |
| select-only tps, cold | 126,069 (121,218-129,754) | 115,788 (115,648-119,876) | 124,938 (122,669-126,613) | 122,551 (120,991-124,535) |
| select-only tps, warm | 119,174 (117,673-120,951) | 118,064 (115,678-119,835) | 116,833 (115,758-120,181) | 117,029 (114,787-117,367) |
| TPC-B tps, cold | 10,392 (10,359-10,927) | 10,499 (10,303-10,567) | 9,865 (9,687-10,171) | 7,834 (7,429-8,037) |
| TPC-B tps, warm | 11,781 (11,495-11,828) | 11,823 (11,649-11,949) | 11,648 (11,522-11,911) | 9,175 (8,985-9,358) |
first one-row UPDATE of stall_big (s) |
0.05 | 3.95 (3.89-5.20) | 0.05 | 0.05 |
the CHECKPOINT after it (s) |
0.08 | 0.26 | 5.51 (4.20-5.79) | 0.13 (0.12-0.15) |
arm64:
| A: plain | B: eager | C: lazyrw, ext4 | |
|---|---|---|---|
| branch create (s; A: container start) | 0.17 | 1.45 | 1.44 |
after a select-only run (MiB, du -sb) |
n/a | 768.0 | 16.5 |
after TPC-B (MiB, du -sb) |
n/a | 1,783.4 | 1,779.1 |
| select-only tps, cold | 40,169 (38,505-40,687) | 36,292 (35,297-37,390) | 38,846 (37,238-41,553) |
| select-only tps, warm | 37,840 (35,724-38,666) | 36,824 (35,585-39,721) | 37,326 (35,802-40,747) |
| TPC-B tps, cold | 4,287 (3,975-4,533) | 4,158 (3,691-4,303) | 4,075 (3,957-4,115) |
| TPC-B tps, warm | 4,601 (4,414-4,680) | 4,644 (4,357-4,936) | 4,721 (4,380-4,911) |
first one-row UPDATE of stall_big (s) |
0.06 | 4.06 (3.56-4.20) | 0.06 |
the CHECKPOINT after it (s) |
0.09 | 0.21 | 4.09 (3.31-4.34) |
Within a round the configurations ran minutes apart on the same VM, so the ratio of two of them in one round cancels most of whatever drifts on the host. Ratios of warm throughput, median (min to max) over the 5 rounds:
| amd64 select-only | amd64 TPC-B | arm64 select-only | arm64 TPC-B | |
|---|---|---|---|---|
| C / B (the shim's cost) | 1.006 (0.970-1.018) | 0.984 (0.974-1.021) | 1.026 (0.975-1.061) | 1.014 (0.990-1.017) |
| C / A | 0.989 (0.957-1.008) | 0.995 (0.984-1.013) | 1.005 (0.946-1.086) | 1.026 (0.939-1.101) |
| B / A (the overlay's cost) | 0.987 (0.975-1.015) | 1.012 (0.986-1.016) | 0.980 (0.934-1.058) | 1.009 (0.944-1.112) |
| D / C | 0.998 (0.974-1.007) | 0.789 (0.771-0.796) | n/a | n/a |
Warm TPC-B round by round on amd64:
| round | order | A | B | C | D | C / B |
|---|---|---|---|---|---|---|
| 1 | A B C D | 11,781 | 11,858 | 11,588 | 9,145 | 0.977 |
| 2 | B C D A | 11,632 | 11,823 | 11,522 | 9,175 | 0.974 |
| 3 | C D A B | 11,495 | 11,649 | 11,648 | 8,985 | 1.000 |
| 4 | D A B C | 11,828 | 11,666 | 11,911 | 9,242 | 1.021 |
| 5 | A B C D | 11,807 | 11,949 | 11,752 | 9,358 | 0.984 |
Release gate: passes on both architectures. v1.0.0's gate for #49 is select-only throughput within noise of the eager branch and plain Postgres, and TPC-B within 5% of the eager branch once warm.
- Select-only, warm: passes. amd64: C is 1.0% below B on medians and 0.6% above it paired (1.006), inside a noise band of ±3.5% (A's and B's own spread across rounds). Against A, C is 2.0% lower on medians and 1.1% paired (0.989), the same as B (B / A 0.987): that is the overlay's cost, with or without the shim. arm64: C is 1.4% above B and 1.4% below A.
- TPC-B, warm: passes. amd64: C 11,648 tps against B 11,823, 1.5% lower on medians; paired C / B 0.984, and no round below 0.974. B's own warm TPC-B varied by 2.5% from round to round. arm64: C 4,721 against B 4,644, 1.7% higher; paired 1.014 (0.990 to 1.017).
The 9.4% gap on the shared host below was the other workloads' I/O: with nothing else on the machine, B's warm TPC-B spread fell from 48.6% of its median to 2.5%, and C stayed within 2.6% of B in every paired round. No change to the shim was needed.
What else it shows:
- Reads copy nothing, as on the shared host: a select-only run grew B by 768 MiB and C by 16.5 MiB (1.6 MiB allocated; the rest is the branch's trimmed WAL segment). D grew by 0.5 MiB of exclusive bytes.
- The first write to a 949 MiB segment, with the checkpoint timed. B
pays in the
UPDATE(3.95 s: opening the segment copies it). C'sUPDATEtakes 0.05 s and itsCHECKPOINT5.51 s: the copy happens when the page is written out. D'sCHECKPOINTtakes 0.13 s, because there the copy-up is a clone. - Clones save the pages that are not written. D's cold TPC-B run is
about 470,000 transactions, each updating a random one of the 5 million
pgbench_accountsrows, so it writes nearly every page of the table; the rest of the 1.8 GiB is what the run added (WAL, new row versions,pgbench_history), which no seed shares. After it, D owns as much as C. On the shared host's spinning disk the same run managed 53 tps, and D owned 49 MiB of 690. - The copy moved from the first read to the first write. Cold
select-only, B is 7% below C (it copies
pgbench_accountsduring the reads); cold TPC-B, C is 6% below B (it copies it at the first write). Once warm, both are the same. - D: select-only throughput is C's (0.998), and warm TPC-B is 21% lower.
That is the loop device and the second filesystem under every
fsync(the evaluation'spg_test_fsyncon XFS on a loop device also ran at 0.4 to 0.6 of ext4); it says nothing about XFS on a real disk, which was not measured.
On a shared host (the first run, kept as measured)
Built with make build from v1-cow at a77c260. The rounds ran A, B, C, D
in the same order every time, and the stall timed only the UPDATE. Two runs
were pooled: 6 samples each for A, B and C, 3 for D.
The host was shared, and not quiet. A ThinkPad with Linux 7.0 x86_64,
4 × Core i3-7100U at 2.4 GHz, 7,305 MiB of RAM and Docker 29.6.2, which ran
a k3s node, a kind cluster and other containers throughout: the 1-minute load
average was 3 to 16 during the runs, recorded per run below. It also has two
storage tiers. Docker's volumes (A, B, C) are on an SSD with ext4; the XFS
disk behind D's --volume-root is a spinning disk. D's throughput
measures that disk, not reflink, and cannot be compared with A, B and C;
D is in the table for usage and the stall. The copy-up probe of the real
branchd reported mode=copy on ext4 (the 64 MiB probe file used 64 MiB)
and mode=clone on the XFS disk (0 bytes used, 64 MiB shared).
| A: plain | B: eager | C: lazyrw, ext4 | D: lazyrw, XFS (spinning disk) | |
|---|---|---|---|---|
| branch create (s; A: container start) | 1.77 (0.54-1.84) | 2.72 (2.55-2.82) | 2.80 (2.53-3.29) | 3.41 (3.30-3.45) |
branch after create (MiB, du -sb) |
n/a | 1.3 | 0.4 | 0.3 |
after a select-only run (MiB, du -sb) |
n/a | 768.0 | 16.5 | 16.5 |
| after a select-only run (MiB, allocated and not shared) | n/a | 753.0 | 1.6 | 0.5 |
after TPC-B (MiB, du -sb) |
n/a | 1,118.3 | 1,008.2 | 689.6 |
| after TPC-B (MiB, allocated and not shared) | n/a | 1,110.3 | 1,008.3 | 49.2 |
| select-only tps, cold | 24,617 (21,222-26,396) | 12,726 (11,506-17,065) | 22,700 (21,527-26,725) | 92 (86-94) |
| select-only tps, warm | 26,422 | 25,967 | 26,107 | 137 |
| TPC-B tps, cold | 801 (412-1,135) | 635 (518-846) | 391 (215-641) | 53 |
| TPC-B tps, warm | 1,051 (825-1,130) | 927 (695-1,145) | 840 (647-1,096) | 55 |
first one-row UPDATE of stall_big (s) |
0.19 | 46.36 | 0.16 | 0.23 |
| 1-minute load average during the runs | 4.51 (3.06-7.37) | 6.03 | 5.09 | 4.78 |
What it showed:
- Reads copy nothing. A select-only run grew the eager branch by
768 MiB, all of
pgbench_accountsand its index, copied on the read's read-write open; it grew C by 16.5 MiB, of which 1.6 MiB is allocated. The difference is the branch's WAL segment, 16 MiB long but mostly a hole since the settle trims it. D grew by 0.5 MiB of exclusive bytes. - Clones are counted as clones. After TPC-B, D's branch is 690 MiB by
du -sbbut owns 49 MiB;duover-counts it 14 times. On ext4 the two numbers agree for C (1,008 MiB): a copy is a copy. - Select-only throughput is unchanged once warm: C is 1.2% below A and 0.5% above B, within the noise of 6 samples. Cold, B pays for its copies inside the reads (12,726 tps, about half of C's 22,700).
- The copy moved from the first read to the first write. Cold TPC-B is where C pays it: 391 tps against B's 635, which had paid during the select-only run.
- The first write to a 949 MiB segment. B's
UPDATEtook 46 s, because opening the segment copied it. C's and D's took 0.16 s and 0.23 s, but that does not mean C's copy was fast: theUPDATEchanges the page in shared buffers, and the file is first written, and copied, when Postgres writes the page out, at the next checkpoint or when a backend or the background writer evicts it. A separate check on the same host (v1-cowat87a7ec2, a 446 MiB single-segment table, load average 3 to 5) showed it: with the shim theUPDATEtook 0.18 s and the branch grew by 16 MiB (its WAL segment), thenCHECKPOINTtook 20.1 s and the branch grew by 455 MiB; with--lazyrw=offtheUPDATEtook 21.0 s and theCHECKPOINT0.22 s. On ext4 the stall is still there, once per segment of at most 1 GiB, but it lands on whatever writes the page out; on XFS or btrfs the copy is a clone. See Troubleshooting. (The quiet-host run above times thatCHECKPOINTtoo.)
The release gate on this host: select-only passed (C 26,107 tps against A 26,422, -1.2%, and B 25,967, +0.5%). Warm TPC-B was not cleanly met and was inconclusive: C's median of 840 tps was 9.4% below B's 927, but the two ranges overlapped by about 85% (C 647 to 1,096, B 695 to 1,145), the ratio of C to B within a round swung from 57% to 158%, and the 87 tps gap was about a fifth of the spread within either configuration. The quiet-host run above settled it: the gap was the other workloads' I/O on the same SSD.
Running it
On any Linux host with a local Docker engine:
make build
BENCH_ROUNDS=5 BENCH_VOLUME_ROOT=/path/on/xfs-or-btrfs BENCH_OUT=results.md \
BENCH_JSON=results.jsonl hack/bench-cow.sh
Without BENCH_VOLUME_ROOT, D is skipped. The other knobs are
BENCH_DURATION, BENCH_CLIENTS, BENCH_JOBS, BENCH_SCALE,
BENCH_STALL_ROWS, BENCH_CONFIGS and BENCH_ROTATE. The table ends with
the per-round ratios and the gate's verdict: PASS, FAIL (which needs the
medians, the paired median and at least 80% of the paired rounds below 0.95),
or INCONCLUSIVE.
On GitHub-hosted runners, run the bench-cow workflow from the Actions tab
(inputs: rounds, duration, scale), or push a branch named bench/<anything>,
which runs it with the defaults. Each job records the host, puts the table and
the raw per-run JSON in the job summary, and uploads them as the
bench-cow-amd64 and bench-cow-arm64 artifacts.
Not measured
The numbers above are from linux/amd64 on kernel 7.0 (the shared host) and
from GitHub-hosted amd64 and arm64 VMs on kernel 6.17 (the pgbench gate run).
arm64 also runs in CI: the shim's C tests and the copy-on-write torture suite
pass on ubuntu-24.04-arm with the aarch64 glibc build. Docker Desktop,
Colima and OrbStack run recent kernels and arm64 or amd64 images, so the shim
is expected to work there, but they have not been measured; Kubernetes
hostPath runs the same entrypoint and install path as Docker and was not
measured either. Neither was XFS or btrfs on a real disk under pgbench: the
only reflink throughput here is a spinning disk and a loop device.
The fix
One flag. The branch entrypoint
(internal/cow/entrypoint.sh) now hands off
with:
exec docker-entrypoint.sh postgres -c recovery_init_sync_method=syncfs
Before replaying WAL, Postgres syncs the data directory
(SyncDataDirectory). With the default recovery_init_sync_method=fsync it
opens every data file read-write to fsync it — and on OverlayFS a
read-write open of a lower-layer file triggers a full copy-up, so the
pre-recovery sync pass alone copied the entire dataset into the branch's rw
layer. syncfs replaces the per-file pass with a single syncfs() call on
each filesystem containing data, which opens nothing read-write and copies
nothing up.
Why this is safe here: syncfs syncs the whole filesystem the data
directory lives on — a superset of what the per-file fsync pass covers — so
the durability guarantee the sync pass exists for (no dirty pages from before
recovery lingering unflushed) is preserved, and WAL crash-recovery semantics
are unchanged. The documented trade-offs of syncfs (errors on unrelated
files on the same filesystem can be reported, or I/O errors missed on some
kernels) are irrelevant for pgoverlay branches: they are disposable dev/test
databases whose filesystem is the branch's own overlay mount. syncfs is
Linux-only and requires Postgres 14+; branch containers are always Linux and
always postgres:14+, which the entrypoint comments.
Before the fix: branch creation scaled with data size
Everything in this section describes commit 2468425 and earlier — it is
kept because the diagnosis is the evidence for the fix. Same script, same
machine, 2026-06-10, before the entrypoint change:
| Database size | pgbench scale | Seed time | Branch create (p50 of 5) | Branch rw overhead after create | rw growth after updating 1% of rows |
|---|---|---|---|---|---|
| 1.00 GiB (1,074,124,467 B) | 68 | 8 s | 7.66 s | 1.05 GiB (1,123,601,062 B) | +10.2 MiB (+10,665,984 B) |
| 5.00 GiB (5,370,484,403 B) | 342 | 29 s | 61.85 s | 5.05 GiB (5,420,035,330 B) | +179.1 MiB (+187,850,752 B) |
Individual branch-create runs (seconds):
- 1 GiB: 7.710, 11.223, 7.658, 6.386, 6.382 → p50 7.658
- 5 GiB: 69.889, 66.125, 61.394, 59.976, 61.853 → p50 61.853
These numbers contradicted the naive expectation for a copy-on-write system ("creation time is independent of data size"), and the diagnosis was:
A branch starts as a stock postgres:17 container whose PGDATA is an
OverlayFS mount: the seeded source volume read-only below, an empty rw volume
on top. Because the seed is a pg_basebackup, the branch's first boot is
crash recovery. Before replaying WAL, Postgres fsyncs every file in the data
directory (SyncDataDirectory, recovery_init_sync_method=fsync, the
default), and it opens each regular file read-write to do so. On OverlayFS, a
read-write open of a lower-layer file triggers a full copy-up — so this
pre-recovery sync pass copied the entire dataset into the branch's rw layer.
The WAL replay itself is trivial (redo done ... elapsed: 0.00 s in the
branch logs); essentially all of the create time was this copy-up — ~140 MB/s
for the 1 GiB dataset (which fits the VM's page cache) and ~87 MB/s for the
5 GiB one (which doesn't).
Two measurements pinned the mechanism down:
- The rw layer right after creation was ≈ the full database size (1.05 GiB for a 1.00 GiB database; 5.05 GiB for a 5.00 GiB one — the excess over the database size is recycled WAL segments and overlay bookkeeping).
- A control run identical except for
-c recovery_init_sync_method=syncfs(onesyncfs()call instead of per-file fsync) finished recovery with the rw layer at 16 KiB — no copy-up at all.
That control run is exactly the fix that now ships; the Results table above is the same benchmark re-run with it in place. The old write-amplification numbers (+10.2 MiB / +179.1 MiB for a 1% UPDATE) measured growth of an already-fully-copied-up rw layer — WAL plus dirtied heap pages only — and are not comparable to the post-fix probe, which pays whole-file copy-up at first write (see the note under Results).
Branch-from-branch
Measured 2026-06-11 (Phase 5), one quick pass at ~1 GiB: a 1.12 GiB source
(1M rows × ~1 KB, pg_database_size = 1,200,805,555 B), branch b1 created,
1% of rows updated on b1 + CHECKPOINT, then b2 created from b1
three times (create/destroy cycles):
| Database size | create-from-branch (p50 of 3) | individual runs |
|---|---|---|
| 1.12 GiB | 4.79 s | 8.36 s, 4.79 s, 3.86 s |
Branch-from-branch on the overlay backend is a freeze: the parent is
checkpointed and stopped, its rw volume becomes an immutable layer, and both
the parent and the child restart over that frozen layer chain — so a
create-from-branch pays two Postgres startups (parent restart + child start)
instead of one, roughly 2× a plain branch create (~1.9 s p50 above). No data
is copied; the first (slowest) run replays the 1% UPDATE's WAL in both
instances, later runs freeze a near-empty rw layer. The ZFS backend
snapshots+clones the parent's dataset instead — no freeze or parent restart
at all. The CSI backend clones the parent's volume after a CHECKPOINT and
a brief stop of the parent, which is then restarted.
Methodology
Everything below is what hack/benchmark.sh does; run it yourself with
make build && hack/benchmark.sh (it needs a local Docker engine: branch
ports are published on the engine host's 127.0.0.1). The read-path
measurement in Reads copy up too was taken by hand, with
the same du -sb helper.
Source database. A throwaway postgres:17 container with
-c wal_level=replica -c max_wal_senders=4 and a
host replication all all scram-sha-256 line appended to pg_hba.conf
(the stock image has no remote replication entry). Data is generated with
pgbench -i -q -s <scale> followed by CHECKPOINT; the size reported is
pg_database_size(current_database()).
Scale calibration. A scale-10 database is initialized first and bytes-per-scale derived empirically: an empty database is 7,689,907 B, scale 10 is 164,665,011 B → 15,697,510 B (≈15.0 MiB) per scale unit including indexes. Target scales are computed from that: 68 → 1.00 GiB, 342 → 5.00 GiB.
Seed time. Wall time of pgb source add (1 s resolution), which runs
pg_basebackup from the source container into the copy-on-write source
volume.
Branch create p50. Five cycles of
pgb branch create benchwork --from benchsrc / pgb branch destroy benchwork;
each time is parsed from the CLI's ready in <duration> line (the CLI stops
the clock when Postgres in the branch accepts connections). p50 is the median
of the five.
Branch rw overhead. du -sb of the branch's rw volume (upper + work
dirs) via a one-shot alpine:3.21 container, immediately after create — the
same measurement GET /v1/branches/{name}/usage and pgb branch ls --usage
perform where copy-up copies data (on a volume root where it clones, branchd
reports exclusive bytes instead; see
usage accounting).
Write amplification. On the still-running branch:
UPDATE pgbench_accounts SET abalance = abalance + 1 WHERE aid <= <scale*1000>
(exactly 1% of rows), then CHECKPOINT, then the same du -sb; the table
reports the resulting rw size.
Disk check. Free space is read from df -Pk / inside a one-shot
container (a container's root is an overlay on the filesystem Docker keeps
volumes on); sizes needing more than ~2.2× target + 2 GiB margin are skipped.
Hardware
- Host: Apple M1 Pro, 16 GiB RAM, macOS 26.5.1.
- VM: Colima (macOS Virtualization.Framework), aarch64, 4 CPUs, 8 GiB
RAM, Ubuntu 24.04.2, kernel 6.8.0-64-generic, Docker 28.4.0 (overlayfs
storage driver). Docker volumes live on a VM-local ext4 data disk
(
/dev/vdb1, 98 GiB, mounted at/var/lib/docker). - Images:
postgres:17(17.10) for source and branches,alpine:3.21for measurement helpers.
Caveats
- Virtualization overhead. Everything runs inside a Colima VM; bare-metal Linux with local NVMe will be faster across the board. (Pre-fix, the copy-up-bound create times tracked the VM's disk throughput — ~87 MB/s effective at 5 GiB; post-fix, create time is no longer disk-bound.)
- virtiofs is not in the data path. Colima mounts the macOS home over virtiofs, but all measured I/O hits Docker named volumes on the VM-local ext4 disk — these numbers do not include virtiofs penalties (nor its cache benefits).
du -sbis apparent size, not allocated blocks; sparse files (rare inPGDATA) would overcount, and overlay/ext4 metadata is not charged.- Seed time has 1 s resolution (wall-clock around the CLI call); branch create times come from the CLI's own millisecond-rounded duration. Seed times are not comparable between the before/after tables (different cache state, same mechanism).
- Postgres runs with stock image defaults (128 MB
shared_buffers, 1 GBmax_wal_size); a tuned source would seed faster. - An idle kind (Kubernetes-in-Docker) control-plane container from another test setup was present in the same VM during the runs; it was not serving any workload.