Files
felhom.eu/documentation/audits/DRILL-version-travel-2026-09-26.md
T
2026-09-27 14:15:01 +02:00

5.3 KiB
Raw Blame History

DRILL — 2026-09-26/27: a backup's data and its version travel together; the next PostgreSQL apps; adventurelog; leftovers

Brief: "a backup's data and its version always travel together; the next PostgreSQL apps; adventurelog; small leftovers". Architecture read before any claim: 07-backup-architecture.md §6 (now §6.6 added), 09-update-architecture.md §3 decisions 13–42, §5.3, §6.4 parts 4, 5, 10. Evidence: audits/version-travel-2026-09-26/.

Not done, or changed

  • One controller release (v0.275.0), not two, with Parts A, D2–D4 and R-699 (found live during Part B, filed and fixed). Live proofs ran on 0.275.0-rc1 (identical except R-699); 0.275.0 itself ran the adventurelog box proof.
  • The session paused ~17 h (usage limit) between the build and the live proof. In that gap the box's own night ran the engine step on docmost, which became the shape-2 live proof — observed, not pressed.
  • R-691 (2) NOT built — Load does not look at the off-site copy. (1) built and red-proofed, not live-proven (no nextcloud kept folder existed on 9202).
  • Part B moved 2 apps (paperless-ngx, tandoor). calcom and claper were not attempted (no seed route in the fixtures).
  • adventurelog's health fix is a DROP of the override, not /usr/bin/node: no single node path works on both versions (measured). It rides the image move in the same commit, as asked.
  • Off-site restore NOT exercised live (9202 has no off-site target; the boxes that do are fenced). Unit-tested.
  • The conversion-copy RELEASE on a real 18 dump not observed — due on 9202 after tonight's 00:30 data run.
  • D1 took THREE agent releases. Instead of waiting 6 h for the scheduled cycle, the scheduler's own read-only verdict (-selftest=restore-test-due) was read on demo-hp: the golden was skipped, but the pick fell to a leftover archive of guest 9100, deleted in August. Fixed in v0.136.0 (the guest must exist) — whose lookup then met PVE's 403 for a guest outside the agent's pool and made the local tier UNKNOWN (my regression, read the same way); fixed in v0.137.0, verified read-only on demo-hp with the pre-release binary before the release. Each red-proofed and signed-delivered to both hosts.
  • Part E (night watch) not done.
  • D4 is not ruled: option A is built and in force (STATUS).

Claims in the brief, judged

  1. CaptureRecoveryUnit rewrites compose/ and image_pins before the data is refreshed — TRUE (measured: the refresh 2 min after the update rewrote both; the dumps stayed).
  2. A restore replays the database's raw volume tar before the SQL dump — TRUE (restore_unit.go; live: the 16 datadir tar replaced the volume first).
  3. The 18 definition over a 16 tar does X — REFUSED by the engine, and the app was left DOWN: postgres:18 refused the old layout, the replay timed out, the start failed. Not "a fresh cluster with the SQL".
  4. A restored version older than the ladder would jump — TRUE for a person's press, FALSE for the automatic leg (LegSkipOlderThanLadder already refuses). The page now says the box will not update it by itself.
  5. 07 says nothing on which version a restore brings back — TRUE for 07 (only 09 §5.3 said the restore pins to the unit's compose). 07 §6.6 now says it.

A — the data and its version (R-696)

A1 spike (A1/README.md) → A2–A4 built (stamps, the definition kept with its data, restore guard, off-site definition, data times on all three tiers, release needs a dump on the new major) → red-proofs RP1–RP6 (A5-redproofs/) → live (A5-live/README.md): shape 2 in the box's own night (restore offered 00:30 = the data; restore back whole at 16 with the sentence „A(z) docmost visszaállt a(z) 2026-09-27 02:30-i mentésből, a(z) docmost:0.96.0, postgres:16-alpine, redis:7-alpine verzióra. Elérhető frissítés: a doboz lépésenként hozza naprakészre."; climb again 55 s); shape 1 (back whole at 0.95.0, "No pending database migrations"; climb 46 s). A7: 42 of 42 ladder digests resolve; filed R-698.

B, C, D

B: B/README.md — paperless-ngx → 18 and tandoor → 17, both venues, engine gate ALLOWED, catalog 46affe0, 53b84ca. C: C/README.md — adventurelog v0.13.0, catalog 06ea7da. D1 agent v0.135.0 → v0.137.0 (signed delivery to both hosts); D2 R-695; D3 R-691 (1); D4 R-694 (per-app measurement, 6 of 7 keep the login in their data).

Teardown

  • Machine: bench LXC 9401 created and destroyed (hostname checked), template removed. 9202: docmost, tandoor and adventurelog removed through the product; paperless-ngx (standing) now on PostgreSQL 18 like the live catalog, its kept 16 copy waiting for the release; back on the LIVE catalog (controller.yaml.pre-vt0926), controller 0.275.0. Drill repo reset to live main (f1a7d6c).
  • Host: demo-hp pct list 9201 + 9202 only; local and nvme-scratch as before the bench.
  • Hub: global floor 0.275.0 (declared MinAgent 0.131.0); six signed agent_update jobs (0.135.0, 0.136.0, 0.137.0 × 2 hosts); no customer or appliance created.

Register

Before: 339 rows / 682,159 B (the brief's count). Opened R-697, R-698, R-699; closed R-655, R-689, R-694, R-695, R-696, R-699 (moved to CLOSED-ITEMS.md); narrowed R-691 (to its (2)) and R-463 (3 of 11 moved). After: 336 rows / 678,058 B.