Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
9.8 KiB
DRILL — 2026-09-25 evening → 26 night: the box converts a PostgreSQL major (docmost first); kept data; the night watch
Brief: "the box converts a PostgreSQL major for its first app (docmost); a reinstall over kept data asks the household…".
Architecture read before any claim: 09-update-architecture.md §3 decisions 13–36, §3b Q5, §6.1a, §6.4 part 10;
07-backup-architecture.md §6; 08-alarm-ladder.md. Evidence: audits/night-2026-09-26/ (part0, A, B, C, D, E, F, G, tools).
Not done, or changed
- Two controller releases, not one — as the brief allowed: v0.273.0 (the conversion; floor raised at 11:32Z so the live catalog could move docmost before the night) and v0.274.0 (kept data; 12:30Z).
- Part E was built by a parallel helper session in its own git worktree (the brief's parts B–D and E touch different files); I reviewed it, found one defect live (R-692) and fixed it before the release, and ran the whole E5 proof myself.
- The live proofs of Part D ran on
0.273.0-rc1, not the released image: the release adds only the recovery line's wording (found live in case (c)) and R-687's two small items. - docmost moved with its own memory limit raised 384M → 512M (decision 39): the bench marked the step
memory_tight, and the move gate refuses a tight step whose limit does not move. The mark stays at 512M (80.4 %) — Node sizes its heap from the limit (R-693). - The bench was a throwaway LXC 9401 on demo-hp, created and destroyed tonight (the first create picked an ARM
template and failed to start; destroyed, name checked first, recreated with amd64). 9202 could not be the bench:
the harness's container names collide with the installed docmost and it ends with
down -v. - Part F3 (adventurelog, R-655): measured, not moved. The answer to "skip or pre-seed download-countries" is in the row; the move needs a choice between two options with different costs.
- My own tool error in Part A4: a bare
docker compose up -din docmost's stack dir (no.envthere — the controller passes the env) started docmost without its secrets; it crash-looped and the box stopped it after 7 restarts (decision 28 working). Started again with the product's Start; the seed read back. - The kept pre-conversion copy was released too early on demo-hp (R-696, found in Part G): the release trusted the unit's refresh time, and the unit's dump was from PostgreSQL 16. Low harm tonight; not fixed (a third release).
Claims in the brief, judged
postgres:18moved its datadir and refuses or warns at/var/lib/postgresql/data— TRUE, and stronger: it REFUSES (exit 1) even an EMPTY volume there (A/A1-images.txt).- The safety dump is
pg_dumpper database with--no-owner --no-privileges, notpg_dumpall— TRUE. For docmost both routes rebuild an identical database (one role, one database); the box usespg_dumpall(decision 38). Restorebrings the old datadir back so the old engine starts — TRUE, by hand (A4) and by the product three times (Part D cases a–c).- The R-487 restore of a removed app picks up the kept drive files — FALSE on 0.272.0: nextcloud came back with no env, no database, its files mounted on the guest's root disk — the unit on the DATA drive was never found (R-690, fixed in v0.274.0, proven live).
- nextcloud needs a rescan for files newer than its backup — TRUE (
occ files:scan --all, 0.6 s); now the template'safter_load:. - Only demo-hp runs docmost — TRUE.
- Also measured: R-657's install LOOP did not recur on 0.272.0 — the install succeeded and the old files were silently unreachable; the choice replaces both shapes.
Part 0 — part0/README.md
Decisions 35, 36 recorded; D3 in STATUS. Last night: demo-felhom's leg did opengist in 20 s at 04:15; its whole-box
backup ran at 07:29, three hours after the leg — nothing waited; demo-hp's leg had nothing and reported
"steps": null (fixed); no whole-box backup was due there. Found: R-689 — demo-hp's restore test picks the golden
template as "the newest settled archive" and fails every 6 h.
Part A — the spike (A/README.md)
Target 18 (decision 37); load from pg_dumpall with zero tolerated errors (decision 38); the check (owners,
encodings, roles, extensions, per-table rows) in 0.5 s; the undo after emptying proven by hand; space: the dump is
0.2 % of the datadir — the box's bound is the volume's size × 1.25 + the 2 GB floor.
Part B — controller v0.273.0 (B/)
internal/stacks/pgconvert.go; 17 tests; 9 red-proofs each seen failing (mark missing → refused; marker re-checked
before emptying; counts compared; PG_VERSION checked; a cut-off dump; a restart during converting undone; the old
copy kept; the release wired; the space check).
Part C — the catalog (C/)
Harness v4 (converts on the bench, moves an 18 mount, writes the mark only when both venues converted); the engine
gate's proof clause with 5 decoys + a red-proof; the postgis family judged. Bench: negative control C3 failed;
docmost run 1 (384M) converted in 25.0 s, proven, memory_tight 90.9 %; run 2 (512M) converted in 11.0 s, proven,
80.4 %; abort (16 images on the 18 datadir) refuses, as expected — the product's route back is the undo, not that.
Part D — the live proof, the floor, the move (D/)
case (9202, drill catalog, 0.273.0-rc1) |
phases | end | PG_VERSION after | seed |
|---|---|---|---|---|
| happy path | safety-dump → pulling → copying → converting 10 s → starting → verifying | done, 42.1 s | 18 | read back |
(a) the load fails (adminpack, gone in 17+) |
… → converting → undoing | undone, 43.6 s | 16 | read back |
| (b) unhealthy on 18 (drill probe :3999) | … converting → starting → verifying (90 s) → undoing | undone, 148.1 s | 16 | read back |
| (c) SIGKILL 1 s after the volume was emptied | … converting → (restart) → undoing | undone, 40 s after the restart | 16 | read back |
Floor 0.273.0 (then 0.274.0) read back from the hub; both demo boxes healthy within ~25 s. docmost moved in the live
catalog at ~14:40 CEST (afd3a60): the engine gate printed ALLOWED … proven on both venues and carries the conversion mark; the Hungarian copy gate caught the RAM line's change on the first push (freeze updated on purpose).
Part E — kept data (E/)
E1 answers above (claims 4, 5). Built: the install choice (409 kept_data_choice), the „Megőrzött adatok" page, the
read-only view, the drive-full naming, after_load:. E5 on 9202 (0.274.0-rc2), both languages, all passed: ask
(use off without a database copy) → start fresh (126 MB renamed to kept/nextcloud/<date>/) → seed, a file, a unit, a
file after it → remove keeping data → use (loaded from the unit on the DATA drive; rescan ran; account + both
files back; a never-written file stays invisible) → start fresh again → Load refused while a leftover occupies the
folder → Delete refused for a wrong name and for an unlisted path → Delete → Load (files moved back, database
loaded, rescan; account + both files back) → a write into the view: Read-only file system. R-692 found and fixed
live (two leftovers named „Filebrowser").
Part F
F1 (R-687 [] + the taken step's log): v0.273.0, 2 red-proofs. F2 (R-688): hub v0.125.0 — the delete dialog lists the
Cloudflare items to remove by hand; live preview read for demo-hp. F3 (R-655): measured, not moved (see the row).
Part G — the night watch
Read-only on both demo boxes (G/). demo-hp converted docmost BY ITSELF in the automatic leg: off-site copy
04:14–04:18 → leg 04:18:25–04:19:20 CEST → CONVERTED 16 → 18 in 9.898s — the check is equal (2 database(s), 50 table(s), 73 row(s)), PG_VERSION 18 → DONE in 54s; hub summary done=1, steps: [docmost 55.1 s]. Content read back
(read-only SELECTs, before 12:41Z / after 02:40Z): users 1 = 1, workspaces 1 = 1, spaces 1 = 1, pages 4 = 4;
PG_VERSION 16 → 18; front door 200; all three containers healthy. The whole-box backup was due since the
evening and held for its window [04:30, 08:30); it started at 04:30:45, eleven minutes after the leg ended — nothing had
to wait (R-687 item 4 still unobserved). demo-felhom: db-dump 02:30, Tier 2 03:30, off-site 04:15, leg nothing to do —
hub reads "steps": [] (R-687's fix, live). 9202's release (B5) worked: 13:30:39Z, one hour after a backup whose
dump (11:57) came after the conversion (11:13). demo-hp's release was premature (R-696): at 02:30:16Z it removed the
16 copy citing "Tier 1 at 02:20:16Z" — the unit's manifest refresh; its docmost-postgres.sql is 02:15:01Z and says
Dumped from database version 16.15.
Teardown
- Machine: 9202 — docmost and nextcloud removed through the product; this session's kept folders deleted through
the Kept-data page (the older romm/paperless leftovers stay as found); back on the LIVE catalog (
controller.yamlfrom the saved copy), controller 0.274.0; one empty folder recreated by the file browser's stale bind removed by hand after the loop was understood (R-695). The drill repo reset to livemain. demo boxes: nothing touched but the two floors (read-only otherwise); docmost on demo-hp now runs PostgreSQL 18 — the intended outcome. - Host: bench LXC 9401 destroyed (name checked), its template removed; demo-hp
pct list9201 + 9202 only. - Hub: global floor 0.274.0 (MinAgent 0.131.0); hub 0.125.0; no customer or appliance created.
Register
Before: 334 rows / 674,532 B. Opened R-689, R-691 (helper), R-693, R-694; opened and closed R-690, R-692; closed R-657; opened R-695 (teardown finding) and R-696 (the night watch); narrowed R-450, R-463, R-469, R-687, R-688, R-691; R-655 measured. After: 339 rows / 682,159 B.