Upgrading to v1.0
What changes when you move a deployment from a v1.0.0-rc build to v1.0.0.
Coming from pgbranch (v1.0.0-rc.3 or older)? That was a rename with no
automatic migration; read
the pgbranch → pgoverlay rename
first.
Before you upgrade
- Back up the state directory (
PGOVERLAY_HOME, in the chart<dataRoot>/stateor the persistence PVC):pgoverlay.dband, if present, its-walfile. The registry migrates in place on the first start (schema v11 to v15; the new migrations only add indexes and columns), and an older binary refuses to open the upgraded file (registry schema is newer than this build), so the backup is your way back. - Check
PGOVERLAY_TOKENis at least 16 characters. branchd now refuses to start with a shorter one. Generate a new one withopenssl rand -hex 16if needed, but see the next point. - Do not change
PGOVERLAY_TOKENin the same restart as the upgrade if you use--rotate-branch-credentials. Rotated passwords used to be encrypted under a key derived from the token; the first v1.0 start re-encrypts them under a new dedicated key (secret.keyin the state directory, generated on that start) and needs the old token to read them. Change the token on a later restart; from then on it no longer affects stored passwords. Back upsecret.keywith the registry from now on.
Copy-on-write
v1.0.0 changes what a branch copies on the overlay backend (#49): reads copy nothing, a write copies the file it touches once (only its changed blocks on XFS or btrfs), and new seeds are settled. Nothing needs configuring, and no registry migration is involved, but existing branches and seeds do not change by themselves.
- Existing branches keep copying eagerly. A branch created by an earlier
build has the old entrypoint in its writable layer, so its Postgres runs
without the lazyrw shim and every table it opens is copied, as before.
branchd reports it as
eagerinpgoverlay_branch_cow_mode, with the reason "the branch's entrypoint predates lazyrw" in its log. It moves to the shim only when something installs the current entrypoint (a restart by reconcile or by the Docker daemon keeps the old one, and nothing migrates branches automatically):pgb branch reset NAME(discards the branch's writes);pgb branch recover NAME, for afailedbranch (keeps them);- creating a branch from it (
--from-branch): the freeze restarts the parent on a fresh writable layer with the current entrypoint, and its data carries over in the frozen layer.
- Existing seeds are not settled. Branches of a seed taken before v1.0.0
still replay the backup's WAL on first start, and their first reads set
hint bits, which are writes: with the shim, those copy the segments they
touch.
pgb source refresh NAMEtakes a new, settled generation for new branches; existing branches keep the generation they were created from. - Seeding takes longer. The settle adds roughly one read of the database
plus a write of its unfrozen pages to every
source addandrefresh, in an extra helper on the branch image. It can fail a seed that used to succeed when the source's configuration cannot start in a container (anincludeof a file outside the data directory, a setting the image does not know); every branch of that seed would have failed the same way. See Troubleshooting.--seed-settle=offrestores the old behaviour. - Seeds belong to the image's
postgresuser. The seed helpers look the user up in the branch image (id -u postgres,id -g postgres) instead of assuming uid 999: 999:999 on the Debian images, 70:70 on the Alpine ones. Alpine sources, which could not be seeded before, now work, and their branches run the musl build of the shim; an image without apostgresuser cannot be seeded (Troubleshooting). - Diff row counts of seeded tables are estimates. The settle runs
ANALYZE, so a small table that was never analyzed on the source is no longer counted exactly; a small change to it shows once the branch has analyzed it (usage).--seed-settle=recoverkeeps the old counting. - New settings, all defaulting to the new behaviour (
pgbin local mode reads the same environment variables; details in the reference):- branchd
--lazyrw=on|off(PGOVERLAY_LAZYRW, Helmcow.lazyrw); --seed-settle=freeze|recover|off(PGOVERLAY_SEED_SETTLE, HelmseedSettle);--volume-root(PGOVERLAY_VOLUME_ROOT, docker only) and--xfs-cowextsize;- the experimental
--wal-recycle.
- branchd
- New metrics:
pgoverlay_branch_cow_mode{mode}andpgoverlay_cow_copyup_mode{mode}; alert onmode="eager"(Observability). - Privileges are unchanged. Branch containers keep exactly the
capabilities they had. branchd runs one more helper at startup, the copy-up
probe, with a branch container's privileges, and it creates and removes two
temporary
pgoverlay-probe-*volumes (reconcile collects them if branchd dies first). - Kernel. Copy-free reads need Linux 4.19 or later where branches run.
On older kernels branches work as before and report
eager.
Behaviour changes
GitHub App branch names. Branches created by pgoverlay-github are now
named gh-<repo-key>-pr-<number> (or gh-<repo-key>-<ref> in git-branch
mode), where the repository key is six hex characters of the SHA-256 of the
lowercased owner/name. Branches named gh-pr-<n> by an earlier build are
not destroyed when their pull request closes; destroy them by hand or let
their TTL expire, and update anything that hard-codes a name. The connect
helpers derive the new names: pgoverlayconnect needs Options{Repo, PR} or
Options{Repo, Ref}, and pgoverlay-connect 0.2.0 (npm) needs {repo, pr}
or {repo, ref}. Details: GitHub App.
REST API (all additive for well-behaved clients, see REST API):
- Request bodies are decoded strictly: an unknown field (a typo such as
ttlforttl_seconds) is a400, and bodies over 1 MiB are a413. GET /v1/branches/{name}/diffneeds the operator role and the leader.- New statuses:
422for a failed seed or masking script (with the tool's output),503when leadership moves or branchd shuts down mid-operation,504when a branch operation outlives--stuck-timeout. All three mean the operation was rolled back. - A failed destroy answers
409(something still uses the branch's volume),502(the container runtime is unreachable) or500with the cause it journaled, instead of a bare500; the branch staysdestroyinguntil a destroy succeeds. - Token names must be lowercase (
^[a-z0-9][a-z0-9._-]{0,62}$) androotis reserved. Existing tokens keep working. - New:
POST /v1/branches/{name}/recover;Branch.password_unavailable,Branch.proxy_host/proxy_port,Source.imageandCreateSourceRequest.image. Duplicate names return409naming the holder instead of SQLite's text, and not-found errors name the missing object.
CLI. pgb source add rejects an empty --host, --user, --database
or --pg-version, an out-of-range --port, and a --via other than
basebackup or dump. --server must be an http(s):// URL. pgb doctor
exits 2 (not 1) when it cannot compute the plan. pgb connect refuses a
branch that is not ready. New commands: pgb branch recover,
pgb source clear-mask, pgb version; new flags: --image and
--no-password on source add, --proxy-host/--proxy-port on connect.
Seeding. A basebackup source whose major differs from --pg-version now
fails at seed time instead of producing branches that cannot start. Seeding
from a standby works (the recovery settings are stripped). Branches no longer
inherit the source's port, listen and socket settings, logging collector,
archiving or synchronous standby names (Troubleshooting).
Docker runtime. Branch ports are chosen by pgoverlay and pinned, branch
containers restart with the daemon (unless-stopped), and reconcile restarts
branch containers that are removed or stopped. Execs inside branches run as
the postgres OS user. ssh:// docker contexts are refused with an
explanation instead of failing obscurely.
Web UI. The token is kept per tab (sessionStorage), so paste it again in
a new tab; reset asks for confirmation.
Helm chart.
shutdownTimeout(default 60 s) sets branchd's drain budget and the pod'sterminationGracePeriodSeconds(the budget plus 30 s).- With
replicaCount > 1orleaderElection.enabled, the API Service routes to the leader through thepgoverlay.leaderpod label, and the Role gainspods patch. The Role also gainssecrets create/deletefor helper environments. - In csi mode with persistence,
nodeis no longer needed; replicas co-locate through a pod affinity. networkPolicy.sourceEgressnow applies to the seed helper pods; branch pods may only reach DNS.- Give ghook an operator token with
ghook.apiTokenSecret; the chart's NOTES warn while it falls back to the admin token. - Label the namespace for Pod Security (
privilegedfor hostpath,baselinefor csi with persistence); the NOTES print the level.
Images and the Action. Release images are multi-arch (linux/amd64,
linux/arm64) from v1.0.0. The Action gains proxy_host, proxy_port,
proxy_database, user and (in rotation mode) password outputs and a
proxy_host input; connect through the proxy_* outputs. action@v1 moves
to v1.0.0 when it is released.