STATUS + CONTEXT + DRILL record: conversion, kept data, D3; 09 decision 39; R-655 measured
gates / gates (push) Successful in 26s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-25 14:40:23 +02:00
parent 90f1f26d20
commit 305f79c40f
5 changed files with 162 additions and 15 deletions
@@ -441,6 +441,15 @@ R-636's louder repeated alarm.
when they hold no table; `CREATE ROLE <x>;` skipped only for a role that exists — its `ALTER ROLE … PASSWORD`
still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks
(`adminpack`) stopped the load and the box undid it. `audits/night-2026-09-26/A/README.md` A2.
39. **docmost's own memory limit rises 384M → 512M with its PostgreSQL 18 step** — *decided by CC unattended
2026-09-25 — operator may reverse.* *One sentence:* the bench marked the step `memory_tight` (the docmost app, not
the engine) — raise the limit, or leave docmost unmoved? **Options:** (a) keep 384M — the move gate refuses a tight
step whose limit does not move, so docmost stays on 16; (b) raise to 512M in the same commit (the gate's own
remedy, the RomM precedent). **Costs:** (a) the conversion that was built and proven tonight never reaches a box;
(b) +128 MB of reservation per docmost box (`mem_limit` 768M → 896M), and the mark stays (80.4 % at 512M): measured
twice, Node sizes its heap from the limit — 349 MB at 384M, 431 MB at 512M, 0 kills and 0 restarts in both 10-minute
watches (~12 000 requests each). **Why (b):** 91 % of the old limit is too close for a household box, the gate's
rule is followed rather than bypassed, and the mark's weakness for such apps is filed (R-693). Reversible.
---
@@ -0,0 +1,113 @@
# DRILL — 2026-09-25 evening → 26 night: the box converts a PostgreSQL major (docmost first); kept data; the night watch
Brief: "the box converts a PostgreSQL major for its first app (docmost); a reinstall over kept data asks the household…".
Architecture read before any claim: `09-update-architecture.md` §3 decisions 13–36, §3b Q5, §6.1a, §6.4 part 10;
`07-backup-architecture.md` §6; `08-alarm-ladder.md`. Evidence: `audits/night-2026-09-26/` (part0, A, B, C, D, E, F, G, tools).
## Not done, or changed
- **Two controller releases, not one** — as the brief allowed: **v0.273.0** (the conversion; floor raised at 11:32Z so
the live catalog could move docmost before the night) and **v0.274.0** (kept data; 12:30Z).
- **Part E was built by a parallel helper session** in its own git worktree (the brief's parts B–D and E touch
different files); I reviewed it, found one defect live (R-692) and fixed it before the release, and ran the whole
E5 proof myself.
- **The live proofs of Part D ran on `0.273.0-rc1`**, not the released image: the release adds only the recovery line's
wording (found live in case (c)) and R-687's two small items.
- **docmost moved with its own memory limit raised 384M → 512M** (decision 39): the bench marked the step
`memory_tight`, and the move gate refuses a tight step whose limit does not move. The mark stays at 512M (80.4 %) —
Node sizes its heap from the limit (R-693).
- **The bench was a throwaway LXC 9401 on demo-hp**, created and destroyed tonight (the first create picked an ARM
template and failed to start; destroyed, name checked first, recreated with amd64). 9202 could not be the bench:
the harness's container names collide with the installed docmost and it ends with `down -v`.
- **Part F3 (adventurelog, R-655): measured, not moved.** The answer to "skip or pre-seed download-countries" is in the
row; the move needs a choice between two options with different costs.
- **My own tool error in Part A4:** a bare `docker compose up -d` in docmost's stack dir (no `.env` there — the controller
passes the env) started docmost without its secrets; it crash-looped and **the box stopped it after 7 restarts
(decision 28 working)**. Started again with the product's Start; the seed read back.
- **The kept pre-conversion copy's release on 9202** needs the hourly job after a proven backup; its first run is after
this record was written — result in Part G.
## Claims in the brief, judged
1. *`postgres:18` moved its datadir and refuses or warns at `/var/lib/postgresql/data`* — **TRUE, and stronger:** it
REFUSES (exit 1) even an EMPTY volume there (`A/A1-images.txt`).
2. *The safety dump is `pg_dump` per database with `--no-owner --no-privileges`, not `pg_dumpall`* — **TRUE**. For docmost
both routes rebuild an identical database (one role, one database); the box uses `pg_dumpall` (decision 38).
3. *`Restore` brings the old datadir back so the old engine starts* — **TRUE**, by hand (A4) and by the product three
times (Part D cases a–c).
4. *The R-487 restore of a removed app picks up the kept drive files* — **FALSE on 0.272.0**: nextcloud came back with no
env, no database, its files mounted on the guest's root disk — the unit on the DATA drive was never found (R-690,
fixed in v0.274.0, proven live).
5. *nextcloud needs a rescan for files newer than its backup* — **TRUE** (`occ files:scan --all`, 0.6 s); now the
template's `after_load:`.
6. *Only demo-hp runs docmost* — **TRUE**.
- Also measured: R-657's install LOOP did not recur on 0.272.0 — the install succeeded and the old files were silently
unreachable; the choice replaces both shapes.
## Part 0 — `part0/README.md`
Decisions 35, 36 recorded; D3 in STATUS. Last night: demo-felhom's leg did opengist in 20 s at 04:15; its whole-box
backup ran at 07:29, three hours after the leg — nothing waited; demo-hp's leg had nothing and reported
`"steps": null` (fixed); no whole-box backup was due there. **Found: R-689** — demo-hp's restore test picks the golden
template as "the newest settled archive" and fails every 6 h.
## Part A — the spike (`A/README.md`)
Target **18** (decision 37); load from **`pg_dumpall`** with zero tolerated errors (decision 38); the check (owners,
encodings, roles, extensions, per-table rows) in 0.5 s; the undo after emptying proven by hand; space: the dump is
0.2 % of the datadir — the box's bound is the volume's size × 1.25 + the 2 GB floor.
## Part B — controller v0.273.0 (`B/`)
`internal/stacks/pgconvert.go`; 17 tests; **9 red-proofs** each seen failing (mark missing → refused; marker re-checked
before emptying; counts compared; `PG_VERSION` checked; a cut-off dump; a restart during `converting` undone; the old
copy kept; the release wired; the space check).
## Part C — the catalog (`C/`)
Harness v4 (converts on the bench, moves an 18 mount, writes the mark only when both venues converted); the engine
gate's proof clause with 5 decoys + a red-proof; the postgis family judged. Bench: negative control C3 `failed`;
docmost run 1 (384M) converted in 25.0 s, proven, `memory_tight` 90.9 %; run 2 (512M) converted in 11.0 s, proven,
80.4 %; abort (16 images on the 18 datadir) refuses, as expected — the product's route back is the undo, not that.
## Part D — the live proof, the floor, the move (`D/`)
| case (9202, drill catalog, `0.273.0-rc1`) | phases | end | PG_VERSION after | seed |
|---|---|---|---|---|
| happy path | safety-dump → pulling → copying → **converting 10 s** → starting → verifying | **done, 42.1 s** | 18 | read back |
| (a) the load fails (`adminpack`, gone in 17+) | … → converting → undoing | **undone, 43.6 s** | 16 | read back |
| (b) unhealthy on 18 (drill probe :3999) | … converting → starting → verifying (90 s) → undoing | **undone, 148.1 s** | 16 | read back |
| (c) SIGKILL 1 s after the volume was emptied | … converting → (restart) → undoing | **undone, 40 s after the restart** | 16 | read back |
Floor 0.273.0 (then 0.274.0) read back from the hub; both demo boxes healthy within ~25 s. **docmost moved in the live
catalog at ~14:40 CEST** (`afd3a60`): the engine gate printed `ALLOWED … proven on both venues and carries the
conversion mark`; the Hungarian copy gate caught the RAM line's change on the first push (freeze updated on purpose).
## Part E — kept data (`E/`)
E1 answers above (claims 4, 5). Built: the install choice (409 `kept_data_choice`), the „Megőrzött adatok" page, the
read-only view, the drive-full naming, `after_load:`. **E5 on 9202 (`0.274.0-rc2`), both languages, all passed:** ask
(use off without a database copy) → start fresh (126 MB renamed to `kept/nextcloud/<date>/`) → seed, a file, a unit, a
file after it → remove keeping data → **use** (loaded from the unit on the DATA drive; rescan ran; account + both
files back; a never-written file stays invisible) → start fresh again → Load refused while a leftover occupies the
folder → Delete refused for a wrong name and for an unlisted path → Delete → **Load** (files moved back, database
loaded, rescan; account + both files back) → a write into the view: `Read-only file system`. **R-692 found and fixed
live** (two leftovers named „Filebrowser").
## Part F
F1 (R-687 `[]` + the taken step's log): v0.273.0, 2 red-proofs. F2 (R-688): hub v0.125.0 — the delete dialog lists the
Cloudflare items to remove by hand; live preview read for demo-hp. F3 (R-655): measured, not moved (see the row).
## Part G — the night watch
*(filled in during the night)*
## Teardown
*(filled in at the end)*
## Register
Before: **334 rows / 674,532 B**. Opened R-689, R-691 (helper), R-693, R-694; opened and closed R-690, R-692; closed
R-657; narrowed R-450, R-463, R-469, R-687, R-688, R-691; R-655 measured. After: see the final count at the end.
+1 -1
View File
@@ -802,7 +802,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** |
| **R-652** | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** |
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** |
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. **-- MEASURED 2026-09-25 (read from both images on the bench, `audits/night-2026-09-26/F/F3-01-entrypoint.txt`):** v0.13.0's entrypoint has an env switch, `SKIP_WORLD_DATA=1`, that skips `download-countries` (v0.12.1 has none). Both versions use the SAME dataset file, `countries+regions+states-v3.1.json`, kept in the `adventurelog_media` volume, and download only when it is absent — so an UPDATE re-uses the file v0.12.1 left. **The likely crash chain:** v0.12.1 exits only on 137 and ignores any other download failure, so a cut-off file it saved stays; v0.13.0 parses it with `ijson` and fails at every start. **Options for the move (not taken tonight):** (a) `SKIP_WORLD_DATA=1` in the template — no internet dependence at update, but a FRESH install gets no world data; (b) keep the download and have the harness prove the update with the file present (the pre-seed is the volume itself) plus a check that a truncated file is replaced (`--force` is the command's own repair). The healthcheck override fix (`/usr/bin/node`) still rides the image move. | **READY — P2; owner: CC (catalog)** |
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |