Files
felhom.eu/documentation/audits/DRILL-night-2026-09-26.md
T
2026-09-26 04:42:47 +02:00

9.8 KiB
Raw Blame History

DRILL — 2026-09-25 evening → 26 night: the box converts a PostgreSQL major (docmost first); kept data; the night watch

Brief: "the box converts a PostgreSQL major for its first app (docmost); a reinstall over kept data asks the household…". Architecture read before any claim: 09-update-architecture.md §3 decisions 13–36, §3b Q5, §6.1a, §6.4 part 10; 07-backup-architecture.md §6; 08-alarm-ladder.md. Evidence: audits/night-2026-09-26/ (part0, A, B, C, D, E, F, G, tools).

Not done, or changed

  • Two controller releases, not one — as the brief allowed: v0.273.0 (the conversion; floor raised at 11:32Z so the live catalog could move docmost before the night) and v0.274.0 (kept data; 12:30Z).
  • Part E was built by a parallel helper session in its own git worktree (the brief's parts B–D and E touch different files); I reviewed it, found one defect live (R-692) and fixed it before the release, and ran the whole E5 proof myself.
  • The live proofs of Part D ran on 0.273.0-rc1, not the released image: the release adds only the recovery line's wording (found live in case (c)) and R-687's two small items.
  • docmost moved with its own memory limit raised 384M → 512M (decision 39): the bench marked the step memory_tight, and the move gate refuses a tight step whose limit does not move. The mark stays at 512M (80.4 %) — Node sizes its heap from the limit (R-693).
  • The bench was a throwaway LXC 9401 on demo-hp, created and destroyed tonight (the first create picked an ARM template and failed to start; destroyed, name checked first, recreated with amd64). 9202 could not be the bench: the harness's container names collide with the installed docmost and it ends with down -v.
  • Part F3 (adventurelog, R-655): measured, not moved. The answer to "skip or pre-seed download-countries" is in the row; the move needs a choice between two options with different costs.
  • My own tool error in Part A4: a bare docker compose up -d in docmost's stack dir (no .env there — the controller passes the env) started docmost without its secrets; it crash-looped and the box stopped it after 7 restarts (decision 28 working). Started again with the product's Start; the seed read back.
  • The kept pre-conversion copy was released too early on demo-hp (R-696, found in Part G): the release trusted the unit's refresh time, and the unit's dump was from PostgreSQL 16. Low harm tonight; not fixed (a third release).

Claims in the brief, judged

  1. postgres:18 moved its datadir and refuses or warns at /var/lib/postgresql/data — TRUE, and stronger: it REFUSES (exit 1) even an EMPTY volume there (A/A1-images.txt).
  2. The safety dump is pg_dump per database with --no-owner --no-privileges, not pg_dumpall — TRUE. For docmost both routes rebuild an identical database (one role, one database); the box uses pg_dumpall (decision 38).
  3. Restore brings the old datadir back so the old engine starts — TRUE, by hand (A4) and by the product three times (Part D cases a–c).
  4. The R-487 restore of a removed app picks up the kept drive files — FALSE on 0.272.0: nextcloud came back with no env, no database, its files mounted on the guest's root disk — the unit on the DATA drive was never found (R-690, fixed in v0.274.0, proven live).
  5. nextcloud needs a rescan for files newer than its backup — TRUE (occ files:scan --all, 0.6 s); now the template's after_load:.
  6. Only demo-hp runs docmost — TRUE.
  • Also measured: R-657's install LOOP did not recur on 0.272.0 — the install succeeded and the old files were silently unreachable; the choice replaces both shapes.

Part 0 — part0/README.md

Decisions 35, 36 recorded; D3 in STATUS. Last night: demo-felhom's leg did opengist in 20 s at 04:15; its whole-box backup ran at 07:29, three hours after the leg — nothing waited; demo-hp's leg had nothing and reported "steps": null (fixed); no whole-box backup was due there. Found: R-689 — demo-hp's restore test picks the golden template as "the newest settled archive" and fails every 6 h.

Part A — the spike (A/README.md)

Target 18 (decision 37); load from pg_dumpall with zero tolerated errors (decision 38); the check (owners, encodings, roles, extensions, per-table rows) in 0.5 s; the undo after emptying proven by hand; space: the dump is 0.2 % of the datadir — the box's bound is the volume's size × 1.25 + the 2 GB floor.

Part B — controller v0.273.0 (B/)

internal/stacks/pgconvert.go; 17 tests; 9 red-proofs each seen failing (mark missing → refused; marker re-checked before emptying; counts compared; PG_VERSION checked; a cut-off dump; a restart during converting undone; the old copy kept; the release wired; the space check).

Part C — the catalog (C/)

Harness v4 (converts on the bench, moves an 18 mount, writes the mark only when both venues converted); the engine gate's proof clause with 5 decoys + a red-proof; the postgis family judged. Bench: negative control C3 failed; docmost run 1 (384M) converted in 25.0 s, proven, memory_tight 90.9 %; run 2 (512M) converted in 11.0 s, proven, 80.4 %; abort (16 images on the 18 datadir) refuses, as expected — the product's route back is the undo, not that.

Part D — the live proof, the floor, the move (D/)

case (9202, drill catalog, 0.273.0-rc1) phases end PG_VERSION after seed
happy path safety-dump → pulling → copying → converting 10 s → starting → verifying done, 42.1 s 18 read back
(a) the load fails (adminpack, gone in 17+) … → converting → undoing undone, 43.6 s 16 read back
(b) unhealthy on 18 (drill probe :3999) … converting → starting → verifying (90 s) → undoing undone, 148.1 s 16 read back
(c) SIGKILL 1 s after the volume was emptied … converting → (restart) → undoing undone, 40 s after the restart 16 read back

Floor 0.273.0 (then 0.274.0) read back from the hub; both demo boxes healthy within ~25 s. docmost moved in the live catalog at ~14:40 CEST (afd3a60): the engine gate printed ALLOWED … proven on both venues and carries the conversion mark; the Hungarian copy gate caught the RAM line's change on the first push (freeze updated on purpose).

Part E — kept data (E/)

E1 answers above (claims 4, 5). Built: the install choice (409 kept_data_choice), the „Megőrzött adatok" page, the read-only view, the drive-full naming, after_load:. E5 on 9202 (0.274.0-rc2), both languages, all passed: ask (use off without a database copy) → start fresh (126 MB renamed to kept/nextcloud/<date>/) → seed, a file, a unit, a file after it → remove keeping data → use (loaded from the unit on the DATA drive; rescan ran; account + both files back; a never-written file stays invisible) → start fresh again → Load refused while a leftover occupies the folder → Delete refused for a wrong name and for an unlisted path → Delete → Load (files moved back, database loaded, rescan; account + both files back) → a write into the view: Read-only file system. R-692 found and fixed live (two leftovers named „Filebrowser").

Part F

F1 (R-687 [] + the taken step's log): v0.273.0, 2 red-proofs. F2 (R-688): hub v0.125.0 — the delete dialog lists the Cloudflare items to remove by hand; live preview read for demo-hp. F3 (R-655): measured, not moved (see the row).

Part G — the night watch

Read-only on both demo boxes (G/). demo-hp converted docmost BY ITSELF in the automatic leg: off-site copy 04:14–04:18 → leg 04:18:25–04:19:20 CEST → CONVERTED 16 → 18 in 9.898s — the check is equal (2 database(s), 50 table(s), 73 row(s)), PG_VERSION 18 → DONE in 54s; hub summary done=1, steps: [docmost 55.1 s]. Content read back (read-only SELECTs, before 12:41Z / after 02:40Z): users 1 = 1, workspaces 1 = 1, spaces 1 = 1, pages 4 = 4; PG_VERSION 16 → 18; front door 200; all three containers healthy. The whole-box backup was due since the evening and held for its window [04:30, 08:30); it started at 04:30:45, eleven minutes after the leg ended — nothing had to wait (R-687 item 4 still unobserved). demo-felhom: db-dump 02:30, Tier 2 03:30, off-site 04:15, leg nothing to do — hub reads "steps": [] (R-687's fix, live). 9202's release (B5) worked: 13:30:39Z, one hour after a backup whose dump (11:57) came after the conversion (11:13). demo-hp's release was premature (R-696): at 02:30:16Z it removed the 16 copy citing "Tier 1 at 02:20:16Z" — the unit's manifest refresh; its docmost-postgres.sql is 02:15:01Z and says Dumped from database version 16.15.

Teardown

  • Machine: 9202 — docmost and nextcloud removed through the product; this session's kept folders deleted through the Kept-data page (the older romm/paperless leftovers stay as found); back on the LIVE catalog (controller.yaml from the saved copy), controller 0.274.0; one empty folder recreated by the file browser's stale bind removed by hand after the loop was understood (R-695). The drill repo reset to live main. demo boxes: nothing touched but the two floors (read-only otherwise); docmost on demo-hp now runs PostgreSQL 18 — the intended outcome.
  • Host: bench LXC 9401 destroyed (name checked), its template removed; demo-hp pct list 9201 + 9202 only.
  • Hub: global floor 0.274.0 (MinAgent 0.131.0); hub 0.125.0; no customer or appliance created.

Register

Before: 334 rows / 674,532 B. Opened R-689, R-691 (helper), R-693, R-694; opened and closed R-690, R-692; closed R-657; opened R-695 (teardown finding) and R-696 (the night watch); narrowed R-450, R-463, R-469, R-687, R-688, R-691; R-655 measured. After: 339 rows / 682,159 B.