Files
felhom.eu/documentation/audits/DRILL-night-2026-09-26.md
T
2026-09-26 04:42:47 +02:00

130 lines
9.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DRILL — 2026-09-25 evening → 26 night: the box converts a PostgreSQL major (docmost first); kept data; the night watch
Brief: "the box converts a PostgreSQL major for its first app (docmost); a reinstall over kept data asks the household…".
Architecture read before any claim: `09-update-architecture.md` §3 decisions 13–36, §3b Q5, §6.1a, §6.4 part 10;
`07-backup-architecture.md` §6; `08-alarm-ladder.md`. Evidence: `audits/night-2026-09-26/` (part0, A, B, C, D, E, F, G, tools).
## Not done, or changed
- **Two controller releases, not one** — as the brief allowed: **v0.273.0** (the conversion; floor raised at 11:32Z so
the live catalog could move docmost before the night) and **v0.274.0** (kept data; 12:30Z).
- **Part E was built by a parallel helper session** in its own git worktree (the brief's parts B–D and E touch
different files); I reviewed it, found one defect live (R-692) and fixed it before the release, and ran the whole
E5 proof myself.
- **The live proofs of Part D ran on `0.273.0-rc1`**, not the released image: the release adds only the recovery line's
wording (found live in case (c)) and R-687's two small items.
- **docmost moved with its own memory limit raised 384M → 512M** (decision 39): the bench marked the step
`memory_tight`, and the move gate refuses a tight step whose limit does not move. The mark stays at 512M (80.4 %) —
Node sizes its heap from the limit (R-693).
- **The bench was a throwaway LXC 9401 on demo-hp**, created and destroyed tonight (the first create picked an ARM
template and failed to start; destroyed, name checked first, recreated with amd64). 9202 could not be the bench:
the harness's container names collide with the installed docmost and it ends with `down -v`.
- **Part F3 (adventurelog, R-655): measured, not moved.** The answer to "skip or pre-seed download-countries" is in the
row; the move needs a choice between two options with different costs.
- **My own tool error in Part A4:** a bare `docker compose up -d` in docmost's stack dir (no `.env` there — the controller
passes the env) started docmost without its secrets; it crash-looped and **the box stopped it after 7 restarts
(decision 28 working)**. Started again with the product's Start; the seed read back.
- **The kept pre-conversion copy was released too early on demo-hp (R-696, found in Part G):** the release trusted the
unit's refresh time, and the unit's dump was from PostgreSQL 16. Low harm tonight; not fixed (a third release).
## Claims in the brief, judged
1. *`postgres:18` moved its datadir and refuses or warns at `/var/lib/postgresql/data`* — **TRUE, and stronger:** it
REFUSES (exit 1) even an EMPTY volume there (`A/A1-images.txt`).
2. *The safety dump is `pg_dump` per database with `--no-owner --no-privileges`, not `pg_dumpall`* — **TRUE**. For docmost
both routes rebuild an identical database (one role, one database); the box uses `pg_dumpall` (decision 38).
3. *`Restore` brings the old datadir back so the old engine starts* — **TRUE**, by hand (A4) and by the product three
times (Part D cases a–c).
4. *The R-487 restore of a removed app picks up the kept drive files* — **FALSE on 0.272.0**: nextcloud came back with no
env, no database, its files mounted on the guest's root disk — the unit on the DATA drive was never found (R-690,
fixed in v0.274.0, proven live).
5. *nextcloud needs a rescan for files newer than its backup* — **TRUE** (`occ files:scan --all`, 0.6 s); now the
template's `after_load:`.
6. *Only demo-hp runs docmost* — **TRUE**.
- Also measured: R-657's install LOOP did not recur on 0.272.0 — the install succeeded and the old files were silently
unreachable; the choice replaces both shapes.
## Part 0 — `part0/README.md`
Decisions 35, 36 recorded; D3 in STATUS. Last night: demo-felhom's leg did opengist in 20 s at 04:15; its whole-box
backup ran at 07:29, three hours after the leg — nothing waited; demo-hp's leg had nothing and reported
`"steps": null` (fixed); no whole-box backup was due there. **Found: R-689** — demo-hp's restore test picks the golden
template as "the newest settled archive" and fails every 6 h.
## Part A — the spike (`A/README.md`)
Target **18** (decision 37); load from **`pg_dumpall`** with zero tolerated errors (decision 38); the check (owners,
encodings, roles, extensions, per-table rows) in 0.5 s; the undo after emptying proven by hand; space: the dump is
0.2 % of the datadir — the box's bound is the volume's size × 1.25 + the 2 GB floor.
## Part B — controller v0.273.0 (`B/`)
`internal/stacks/pgconvert.go`; 17 tests; **9 red-proofs** each seen failing (mark missing → refused; marker re-checked
before emptying; counts compared; `PG_VERSION` checked; a cut-off dump; a restart during `converting` undone; the old
copy kept; the release wired; the space check).
## Part C — the catalog (`C/`)
Harness v4 (converts on the bench, moves an 18 mount, writes the mark only when both venues converted); the engine
gate's proof clause with 5 decoys + a red-proof; the postgis family judged. Bench: negative control C3 `failed`;
docmost run 1 (384M) converted in 25.0 s, proven, `memory_tight` 90.9 %; run 2 (512M) converted in 11.0 s, proven,
80.4 %; abort (16 images on the 18 datadir) refuses, as expected — the product's route back is the undo, not that.
## Part D — the live proof, the floor, the move (`D/`)
| case (9202, drill catalog, `0.273.0-rc1`) | phases | end | PG_VERSION after | seed |
|---|---|---|---|---|
| happy path | safety-dump → pulling → copying → **converting 10 s** → starting → verifying | **done, 42.1 s** | 18 | read back |
| (a) the load fails (`adminpack`, gone in 17+) | … → converting → undoing | **undone, 43.6 s** | 16 | read back |
| (b) unhealthy on 18 (drill probe :3999) | … converting → starting → verifying (90 s) → undoing | **undone, 148.1 s** | 16 | read back |
| (c) SIGKILL 1 s after the volume was emptied | … converting → (restart) → undoing | **undone, 40 s after the restart** | 16 | read back |
Floor 0.273.0 (then 0.274.0) read back from the hub; both demo boxes healthy within ~25 s. **docmost moved in the live
catalog at ~14:40 CEST** (`afd3a60`): the engine gate printed `ALLOWED … proven on both venues and carries the
conversion mark`; the Hungarian copy gate caught the RAM line's change on the first push (freeze updated on purpose).
## Part E — kept data (`E/`)
E1 answers above (claims 4, 5). Built: the install choice (409 `kept_data_choice`), the „Megőrzött adatok" page, the
read-only view, the drive-full naming, `after_load:`. **E5 on 9202 (`0.274.0-rc2`), both languages, all passed:** ask
(use off without a database copy) → start fresh (126 MB renamed to `kept/nextcloud/<date>/`) → seed, a file, a unit, a
file after it → remove keeping data → **use** (loaded from the unit on the DATA drive; rescan ran; account + both
files back; a never-written file stays invisible) → start fresh again → Load refused while a leftover occupies the
folder → Delete refused for a wrong name and for an unlisted path → Delete → **Load** (files moved back, database
loaded, rescan; account + both files back) → a write into the view: `Read-only file system`. **R-692 found and fixed
live** (two leftovers named „Filebrowser").
## Part F
F1 (R-687 `[]` + the taken step's log): v0.273.0, 2 red-proofs. F2 (R-688): hub v0.125.0 — the delete dialog lists the
Cloudflare items to remove by hand; live preview read for demo-hp. F3 (R-655): measured, not moved (see the row).
## Part G — the night watch
Read-only on both demo boxes (`G/`). **demo-hp converted docmost BY ITSELF in the automatic leg:** off-site copy
04:14–04:18 → leg 04:18:25–04:19:20 CEST → `CONVERTED 16 → 18 in 9.898s — the check is equal (2 database(s), 50
table(s), 73 row(s)), PG_VERSION 18` → `DONE in 54s`; hub summary `done=1`, `steps: [docmost 55.1 s]`. Content read back
(read-only SELECTs, before 12:41Z / after 02:40Z): users 1 = 1, workspaces 1 = 1, spaces 1 = 1, pages 4 = 4;
`PG_VERSION` 16 → 18; front door 200; all three containers healthy. **The whole-box backup** was due since the
evening and held for its window [04:30, 08:30); it started at 04:30:45, eleven minutes after the leg ended — nothing had
to wait (R-687 item 4 still unobserved). demo-felhom: db-dump 02:30, Tier 2 03:30, off-site 04:15, leg nothing to do —
hub reads `"steps": []` (**R-687's fix, live**). **9202's release (B5) worked:** 13:30:39Z, one hour after a backup whose
dump (11:57) came after the conversion (11:13). **demo-hp's release was premature (R-696):** at 02:30:16Z it removed the
16 copy citing "Tier 1 at 02:20:16Z" — the unit's manifest refresh; its `docmost-postgres.sql` is 02:15:01Z and says
`Dumped from database version 16.15`.
## Teardown
- **Machine:** 9202 — docmost and nextcloud removed through the product; this session's kept folders deleted through
the Kept-data page (the older romm/paperless leftovers stay as found); back on the LIVE catalog (`controller.yaml`
from the saved copy), controller 0.274.0; one empty folder recreated by the file browser's stale bind removed by hand
after the loop was understood (R-695). The drill repo reset to live `main`. demo boxes: nothing touched but the two
floors (read-only otherwise); docmost on demo-hp now runs PostgreSQL 18 — the intended outcome.
- **Host:** bench LXC 9401 destroyed (name checked), its template removed; demo-hp `pct list` 9201 + 9202 only.
- **Hub:** global floor 0.274.0 (MinAgent 0.131.0); hub 0.125.0; no customer or appliance created.
## Register
Before: **334 rows / 674,532 B**. Opened R-689, R-691 (helper), R-693, R-694; opened and closed R-690, R-692; closed
R-657; opened R-695 (teardown finding) and R-696 (the night watch); narrowed R-450, R-463, R-469, R-687, R-688, R-691; R-655 measured. After: **339 rows / 682,159 B**.