Files
felhom.eu/REPORT-golden-kept-offsite-2026-09-28.md
T
2026-09-28 15:56:50 +02:00

5.4 KiB
Raw Blame History

REPORT — 2026-09-28: weekly golden + Day-0 vouch, kept data from the off-site copy, claper/calcom, demo-hp space, night read, first live off-site restores

Brief: "the weekly golden, the Day-0 vouch, use my kept data from the off-site copy, two more PostgreSQL fixtures, demo-hp's restore-test space, the night watch, and the first live off-site restore" (revised 2026-09-28). Architecture read before any claim: 07-backup-architecture.md §6.5, §6.6; 06-offsite-connectivity.md; 03-host-agent.md (restore-test storage); 09-update-architecture.md §3 decisions 35–42. Operator change mid-session (15:14): "finish today" → Part E and D2 were done in the day, not overnight (below).

The Parts

Part Step State Note
A 1 read the manifest (rollback) done agent 0.132.0, golden 0.258.0, min_agent 0.131.0
A 2 bake golden 0.276.0 done, with a deviation attempt 1 picked the arm64 template; stopped by CC before pct create finished, nothing published. Attempt 2: all markers, leak grep 0 (control 1), 404 → 200, anonymous download sha equal, VM back to virgin
A 3 vouch (3 fields) done R-120 gate refused golden 0.258.0 first (negative control); agent=0.137.0 golden=0.276.0 min_agent="0.131.0" read back
A 4 Day-0 test install done installer 1.28.0, agent + golden fetched through the manifest and sha-verified, controller 0.276.0 healthy, claim gate armed; not claimed; scratch customer deleted by the hub's own cascade
A 5 golden-currency gate done green, not waived (record documentation/tests/golden-0.276.0-2026-09-28/); waiver file left as it is (to 2026-10-04)
A 6 agent CHANGELOG line done felhom-agent 5c68c86; no floor move in Part A
B 1–3 code, tests, red-proofs done controller v0.277.0; RP1–RP3
B 4 live Tier 1/Tier 2 regression on 9202 Tier 1 done; Tier 2 not runnable 9202 has one drive
B 5 release + floor done floor 0.277.0 at 10:11 (before 01:30), both boxes healthy
B — changed: a second release, v0.278.0 R-704 blocked Part E; floor 0.278.0 at 15:36
C calcom done memory fix first (R-703, 768M OOM at every start → 1536M); PG 16 → 18 proven on bench + box; catalog 037f956
C claper done PG 16 → 17 proven on bench + box; catalog 4a249b9; found R-702 (default admin)
D 1 R-701 (a) done — not enough 20.3 → 26.6 GiB free; 31 needed; 14:13 cycle refused again
D 2 night read changed the night 27/28 was read (logs, hub, agent journals), not tonight's; demo-felhom's off-site/update legs not readable
E 1 setup done nextcloud on demo-hp, seeded, joined the off-site copy
E 2 (a) off-site restore done — changed timing the snapshot came from the page's "run now" (15:40), not the night; seed back, marker gone
E 2 (b) kept data from off-site done choice named "távoli mentés, 2026-09-28 15:40"; seed + files back
E 3 teardown done app removed with its data; app list equal to before; verification copy deleted by the product

Claims in the brief that turned out wrong (or right)

  • "0.276.0's MinAgent is 0.131.0" — right.
  • "The R-120 gate accepts golden 0.276.0" — right (it refused 0.258.0 and accepted 0.276.0).
  • "The Day-0 test install leaves no hub record" — wrong. It left a host, a vaulted break-glass credential, a claim code, a WireGuard peer (10.77.0.5, synced toward ep0) and reports. The customer DELETE cascade removed the hub side; the ep0 peer removal was not observed (ep0 untouched by rule).
  • "The pairing code is shown" — wrong for this path. The one-liner install shows no pairing code (that is the ISO path); what shows is the controller's claim gate (dashboard not yet claimed).
  • "pct fstrim frees enough for R-701" — wrong. It freed 6.3 GiB; the pool is 53.9 GiB and the guest holds ~26 GiB, so 31 GiB free cannot be reached by trimming.
  • "paperless-ngx runs on demo-hp" — right; it converted to PostgreSQL 18 in the night 27/28 (04:22 CEST, rows equal, 0 documents).
  • "An app joins the off-site copy by a per-app switch" — right (POST /backup/offbox/toggle, app_backup.<app>.offbox).

Evidence

  • Golden, vouch, Day-0, D1, D2: documentation/audits/evidence-golden-0276-2026-09-28/
  • Part B + E: documentation/audits/kept-offsite-2026-09-28/ (redproofs/, live/, E/)
  • Part C: documentation/audits/pg-calcom-claper-2026-09-28/

Rows

Opened: R-702 (claper default admin, P1), R-703 (calcom OOM — closed the same day), R-704 (leftover holds — fixed in 0.278.0, WATCHING), R-705 (no "run the night now"), R-706 (verification copy survives removal). Closed: R-691, R-703. Updated: R-463, R-687, R-701. Register 337 → 342 rows. unproven.py --summary: not walked 35 of 55 (unchanged; the capability map was not edited).

Teardown, three layers

  • Machines: drill VM reverted to virgin (qemu gone); bench LXC 9401 created and destroyed twice; nextcloud, calcom, claper removed through the product on 9201/9202; 9202 pointed back to the live catalog; drill catalog = live.
  • Hosts: demo-hp local-lvm trimmed (kept); template cache files removed; no storage added.
  • Hub: scratch customer drill-g0276 deleted by its cascade; manifest now golden 0.276.0 / agent 0.137.0; floor 0.278.0. nextcloud's off-site snapshots stay in demo-hp's repository (removal never touches off-site history, R-474).