Files
felhom.eu/REPORT-ep0-datastore-volume-2026-07-27.md
T
admin ad8057c4e3 docs(ep0): relocate the felhom-offsite PBS datastore onto the 100 GB volume
Supervised runbook execution. No code change, no version bump.

felhom-offsite moved from ep0's 40 GB root disk (/srv/pbs-felhom) to a
dedicated 100 GB Hetzner Cloud Volume (/mnt/pbs-datastore, ext4 -m 0,
by-id fstab, relatime). Datastore NAME unchanged, so the PBS-DR
descriptors, per-box storage ids, ACLs and namespaces are untouched.

Capacity: 37.2 GB -> 98 GB total, 28.9% -> 13% used, headroom to the 80%
warn 19 GB -> ~65 GB. This CLEARS the R-82 Phase 0 P0.3 STOP. Per-tenant
encryption still precludes cross-customer dedup, so the slope is
unchanged - the volume buys runway, not a better cost model.

Verified: byte totals and chunk counts identical (9748), 7/7 snapshots
across all three namespaces, backup:backup ownership, clean itemised
dry-run, full verify job TASK OK with 0 errors, and a restore round-trip
(source_tier pbs, pass true, mount_parity ok, clean teardown).

Nothing deleted - the original 13 GB stays at /srv/pbs-felhom as the
rollback until a new weekly backup lands. GC deliberately not run.

Three findings recorded:
- the `scratch` datastore points at a non-existent path (pre-existing;
  now logs ENOENT every start) - operator decision
- the runbook's S6 guard test proves the wrong proposition:
  RequiresMountsFor re-mounts rather than refusing, so the test only
  bites when the device is genuinely unavailable (re-run that way, and
  the refusal was observed) - amendment recommended
- S11: storage box u629193 has no live backup path, BUT ep0 carries an
  enabled sshfs mount unit against it that must be removed before the
  box is deleted

Deviations: the volume arrived pre-formatted and mounted; S8 ran on
demo-felhom rather than demo-hp (no SSH key for demo-hp); the window was
contended by a stale in-memory 10-minute restore-test cadence whose
config had already been reverted on disk.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ARoadHBf8rHoscfiqeVZn
2026-07-27 09:36:11 +02:00

5.1 KiB
Raw Blame History

REPORT — ep0 PBS datastore relocated onto the 100 GB volume (2026-07-27)

Class: supervised operational run (RUNBOOK execution). No code changed. No version bump. Written as REPORT-<topic>.md per the parallel-session rule — the shared REPORT.md was not touched.

Full record with all evidence: documentation/runbooks/RUNBOOK-ep0-datastore-volume-2026-07-27.md


Outcome: DONE and verified

felhom-offsite now lives on a dedicated 100 GB Hetzner Cloud Volume instead of ep0's 40 GB root disk.

Before After
Path /srv/pbs-felhom (root disk) /mnt/pbs-datastore (volume)
Datastore total 37.2 GB 98 GB
Used 28.9 % 13 %
Headroom to the 80 % warn 19 GB ≈65 GB
Additional customers before warn ≈2 ≈713

Datastore name unchanged — the PBS-DR descriptors, per-box storage ids, ACLs and namespace layout that R-39/R-82 made self-healing are untouched.

Window: 06:58 → 07:19 UTC (PBS down 07:00 → 07:17). Nothing was deleted.

Acceptance evidence

Gate Result
Copy integrity 13,242,207,822 = 13,242,207,822 B · 9,748 = 9,748 chunks · 7 = 7 snapshots · backup:backup · itemised dry-run 0 lines
Snapshot counts per ns demo-felhom 2=2, demo-felhom-01 3=3, demo-hp 2=2
atime semantics rw,relatime,discardrelatime present, noatime absent (GC correctness)
Verify job TASK OK, 3/3 groups, forced re-verification of every snapshot, 0 errors
§6 mount guard refusal observedJob … failed with result 'dependency'; mountpoint stayed empty
§8 restore round-trip source_tier: pbs, pass: true, mount_parity: ok, clean teardown, 12m1s

Three findings the operator should act on

  1. scratch datastore is configured at a path that does not exist (/srv/pbs-scratch). Pre-existing, not caused here, but now logs ENOENT on every PBS start. This is the PRIME RISK shape ("reports fine, is not there") already live in the config. Decision needed: remove the stanza or create the directory.

  2. The runbook's §6 acceptance test proves the wrong proposition. RequiresMountsFor is a mount-first ordering guarantee, not a refusal — systemd silently re-mounts an unmounted volume and PBS then starts safely. The test only bites when the device is genuinely unavailable, which is how it was re-run and passed. Amendment recommended in the runbook record.

  3. §11 — storage box u629193 is NOT simply unused. No live backup path references it (no datastore, no restic repo, no fstab, no known_hosts pin; R-17 already deleted u629193-sub1), but ep0 carries an enabled, currently-mounted sshfs unit mnt-pbs\x2dstoragebox.mount/mnt/pbs-storagebox, holding spike leftovers. Disable and remove that unit before deleting the box, or ep0 logs a failed mount every boot. The deletion is the operator's console click.

Deviations from the runbook as written

  • The volume arrived already formatted and mounted by Hetzner at /mnt/HC_Volume_106469259 (§2 assumed neither). Operator ruled: reformat + repath. The 5 % reserve was reclaimed (-m 0).
  • §8 ran on demo-felhom, not demo-hp — DooPlex holds no SSH key for demo-hp (the G1 gap). Same tier, same relocated datastore, larger archive.
  • The window was contended by a stale 10-minute restore-test cadence on demo-felhom: the config had already been reverted to 3.5 days on disk, but the cadence is read once at daemon start and NRestarts=0. Restarting the agent applied it (cadence=84h0m0s). The in-flight test was allowed to finish rather than aborted.

Process errors made during this run (recorded deliberately)

  • rsync -aHAX OOM-killed ep0 (3.7 GB RAM, no swap). Cause: a PBS .chunks/ tree pre-creates all 65536 shard dirs → 75,341 inodes, and -H retains the whole inode map. -H was dropped only after proving no hardlinks exist (-links +1 → 0; max link count → 1); PBS references chunks by digest, never by hardlink.
  • /usr/bin/time -v is not installed on ep0 → exit 127, rsync never ran, and a | grep … || true wrapper swallowed it while printing a success-looking line.
  • rsync --version | head -1 reported a working rsync 3.4.1 as "missing" — the §12 pipe-into-head trap, fourth recorded instance in this project. Both fixed by capturing the command's own $?.

Deferred

  1. Old copy retained at /srv/pbs-felhom (13 GB, 9,748 chunks) as the rollback. Rollback is a two-line datastore.cfg revert. Reclaim only after a new weekly offsite backup lands on the volume, with explicit go-ahead.
  2. GC not run — now unblocked by the round-trip, but left for a separate deliberate run. No GC schedule is configured on this PBS at all.
  3. Hub PBS-DR capacity gauge not re-read — the hub is a ClusterIP service and was not reachable from DooPlex. If it still reports 37.2 GB it is reading the wrong filesystem: a real bug.
  4. ep0 has no swap (temporary 4 GB file removed; box left as found). Worth a small permanent swapfile — outside this runbook's scope.