Skip to content

Observability

branchd exposes Prometheus metrics and a real readiness endpoint on the REST API port. Both sit outside the bearer-token auth: scrapers and kubelet probes don't authenticate, and neither endpoint leaks secrets.

Endpoints

Path Auth Purpose
/healthz none Liveness — the process is up. Used by the Deployment's livenessProbe.
/readyz none Readiness — 200 only when the registry is reachable and the container driver responds (a cheap ListManaged); 503 otherwise. Used by the readinessProbe.
/metrics none Prometheus exposition over branchd's private registry.

Metrics

Metric Type Labels Meaning
pgoverlay_branches_total gauge state Branches by state (reported from the registry on scrape).
pgoverlay_sources_total gauge state Sources by state.
pgoverlay_branch_op_duration_seconds histogram op Branch operation latency (op = create|reset|destroy|from_branch|diff).
pgoverlay_branch_op_errors_total counter op Failed branch operations.
pgoverlay_masking_duration_seconds histogram — Time applying a source's masking scripts inside a branch.
pgoverlay_reaper_runs_total counter — Reconcile passes that applied a plan (the TTL-reaping half; counted once per apply).
pgoverlay_reaper_reaped_total counter — Expired branches destroyed by reconcile (reap actions are counted here, not below).
pgoverlay_reconcile_runs_total counter — Reconcile passes.
pgoverlay_reconcile_actions_total counter action Reconcile actions taken (fail_stuck|fail_stuck_source|retry_destroy|restart_branch|update_endpoint|remove_orphan_container|remove_orphan_helper|gc_layer|gc_volume); see what each does.
pgoverlay_compensation_failures_total counter kind Saga compensations or failure transitions that themselves failed (kind = transition|undo|cleanup). Each one may have left a resource behind for reconcile to collect.
pgoverlay_inflight_ops gauge — Branch operations currently in flight.
pgoverlay_leader gauge — 1 on the replica that is leader (accepts mutations, runs reconcile), 0 on followers; always 1 without --leader-elect.
pgoverlay_leader_transitions_total counter — Times this replica gained or lost leadership.
pgoverlay_disk_bytes_free gauge — Free bytes on the measured filesystem (read via statfs on every scrape; see below).
pgoverlay_disk_bytes_total gauge — Total bytes on the measured filesystem.
pgoverlay_branch_cow_mode gauge mode Ready overlay branches by the copy-on-write mode their Postgres started in: lazyrw (the shim is active: reads copy nothing, a file is copied into the branch on its first write), eager (it is not, although --lazyrw=on; see Troubleshooting), off (--lazyrw=off) or unknown (not read yet or unreadable). Overlay backend only.
pgoverlay_cow_copyup_mode gauge mode Overlay backend: 1 for what an OverlayFS copy-up costs where the volumes live, as probed at startup, 0 for the other modes. clone (XFS reflink=1, btrfs: extents are shared, block-level copy-on-write), copy (data is copied), unknown (not probed yet, or the probe failed). See copy-up mode.

What the disk gauges measure. They statfs one path on every scrape, and that path is not always where branch data lives:

Setup Path measured by default Is branch data there?
Docker runtime PGOVERLAY_HOME (~/.pgoverlay) No. Branch and seed volumes are Docker volumes under the engine's data root (/var/lib/docker/volumes, inside the VM on Colima or Docker Desktop, or on another machine for a remote engine). The gauges cover the registry's filesystem only, unless both happen to be the same filesystem
Docker runtime with --volume-root the volume root, when it is a directory on branchd's machine; else PGOVERLAY_HOME Yes when branchd runs on the Docker host. A volume root on a remote Docker host cannot be measured from branchd
Kubernetes hostpath, chart layout (PGOVERLAY_HOME inside --kube-data-root) the mounted state directory, <dataRoot>/state Yes when the state directory is the hostPath (persistence off, the default in hostpath mode), since it sits on the data root's filesystem. With persistence.enabled=true it is the registry PVC instead
Kubernetes hostpath, branchd running on the storage node itself --kube-data-root Yes
Kubernetes csi not emitted Each branch is its own PVC; watch the CSI driver's capacity metrics

--disk-root <path> overrides the choice: point it at a path on the filesystem that holds the branch data (for Docker, the engine's data root when branchd runs on the Docker host; for hostpath with persistence, a mount of the data root). branchd logs the path it measures at startup (disk gauges measure the filesystem of ...).

Alerts worth having

- alert: PgoverlayCompensationFailures   # leaked resources to look for
  expr: increase(pgoverlay_compensation_failures_total[15m]) > 0
  labels: { severity: warning }
- alert: PgoverlayNoLeader               # every write is refused
  expr: max(pgoverlay_leader) == 0
  for: 1m
  labels: { severity: critical }
- alert: PgoverlayLeaderFlapping
  expr: increase(pgoverlay_leader_transitions_total[15m]) > 4
  labels: { severity: warning }
- alert: PgoverlayBranchOpsFailing
  expr: increase(pgoverlay_branch_op_errors_total[15m]) > 0
  labels: { severity: info }
- alert: PgoverlayBranchesCopyEagerly   # reads copy whole tables into branches
  expr: pgoverlay_branch_cow_mode{mode="eager"} > 0
  for: 15m
  labels: { severity: warning }

A compensation failure means a saga's cleanup did not complete; the next reconcile passes usually collect what it left, and branchd's log names the resource. The leader alerts matter only with --leader-elect (see High availability).

Running out of disk (ENOSPC)

On the overlay backend every branch shares one filesystem (the Docker data root or --volume-root, or the storage node's data root), so a full disk is a fleet-wide, not per-branch, failure. A write copies the whole file it touches (a table segment, up to 1 GiB) into the branch the first time; where copy-up clones (pgoverlay_cow_copyup_mode{mode="clone"}), only the blocks it rewrites take new space. Branches in eager mode (pgoverlay_branch_cow_mode{mode="eager"} or off) grow on reads too: there the first time Postgres opens a table file, OverlayFS copies it whole into the branch (Reads copy up too), so a test suite that scans large tables fills the disk faster than its writes suggest.

  • Overlay copy-up fails. The first write to a file in any branch (the first open, in eager mode, even for a read) must copy the whole file up into that branch's upper layer; with no free space the copy-up returns ENOSPC and the query — and often the whole transaction — fails.
  • Postgres write failures across all branches. WAL/heap writes in every running branch start failing; branches may refuse to accept writes or shut down their backends.
  • Registry failures. When the SQLite registry lives on the same filesystem (Kubernetes hostpath without persistence), a full disk can also break branch bookkeeping (create/destroy state transitions), turning a space problem into a control-plane problem.

These surface as confusing Postgres errors with no obvious common cause. pgoverlay_disk_bytes_free explains them when it measures the right filesystem (see the table above); on Docker, watch the engine host's disk directly, or run branchd on that host with --disk-root pointing at the Docker data root. pgb branch ls --usage shows which branches hold the space.

Warn well before the disk is full (10% free) so there is time to reap branches or grow the volume:

- alert: PgoverlayStorageRootLow
  expr: pgoverlay_disk_bytes_free / pgoverlay_disk_bytes_total < 0.10
  for: 5m
  labels: { severity: warning }
  annotations:
    summary: "pgoverlay storage root <10% free"
    description: >
      The filesystem branchd measures (the branch data root, when
      configured as described in docs/observability.md) is nearly full.
      Overlay copy-up and Postgres writes will start failing across every
      branch. Reap branches (TTL/--max-branches) or grow the volume.

If you prefer an absolute floor (e.g. on a fixed-size PV), alert on pgoverlay_disk_bytes_free < 5e9 (5 GiB) instead of the ratio.

Scraping

The Helm chart annotates the branchd pod for Prometheus pod-discovery:

metadata:
  annotations:
    prometheus.io/scrape: "true"
    prometheus.io/port: "7070"   # = .Values.api.port
    prometheus.io/path: /metrics

If you run a Prometheus CRD / PodMonitor instead of annotation-based discovery, point it at the API port and /metrics. Outside Kubernetes, scrape http://<branchd-host>:7070/metrics directly.