diff --git a/CONTEXT.md b/CONTEXT.md index f4404c71..a87b405d 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,19 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-25 (evening) — PostgreSQL conversion (controller v0.273.0), kept data (v0.274.0), hub v0.125.0 + +- **Controller v0.273.0** = `09` §6.4 part 10 for docmost: `stacks/pgconvert.go`, only on a ladder step marked + `engine_conversion`; unmarked PostgreSQL major moves refused (preflight + job). **v0.274.0** = decision 36 (kept data: + install choice, `/kept-data`, read-only file-browser source, `/kept` protected and never backed up) + R-690 + + R-692. Floor 0.274.0 (MinAgent 0.131.0). 9202 runs 0.274.0 on the DRILL catalog (reset to live at teardown). +- **Catalog:** harness v4; the engine gate's proof clause; **docmost `postgres:16` → `18` moved** (`afd3a60`, mount + `/var/lib/postgresql`, docmost limit 512M); nextcloud `after_load`. +- **Decisions taken by CC unattended, operator may reverse (`09` §3):** 37 (target 18, not 17), 38 (`pg_dumpall` from + the old engine, zero tolerated load errors), 39 (docmost 384M → 512M). +- **Hub v0.125.0:** the customer delete lists the Cloudflare items to remove by hand (R-688 half). +- Record: `documentation/audits/DRILL-night-2026-09-26.md`. + ## Rulings 2026-09-25 (evening, operator) - **`09` §3 decision 35 — part 10 is built now, docmost first** (D1 option A). The box converts a PostgreSQL diff --git a/STATUS.md b/STATUS.md index 483517bd..bbf01e73 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,24 +1,36 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-25 (midday). Peti's box is retired. Both demo boxes run controller 0.272.0 and host agent 0.134.0.** +**Updated 2026-09-25 (evening). Both demo boxes run controller 0.274.0 and host agent 0.134.0. Hub 0.125.0.** -**Decisions I took on my own.** None this time. I recorded your two rulings: Peti's box is retired, and a controller restart during the night's updates does not continue them (the apps not reached wait a night). +**Decisions I took on my own** (you may reverse each): +1. **docmost's database goes to PostgreSQL 18, not 17.** Its own makers already use 18. One move instead of two. +2. **The box copies the whole old database before it converts, and stops on any error.** It then puts everything back. +3. **docmost gets 512 MB of memory instead of 384 MB.** Under light load it used 91 % of 384 MB. The move was not allowed without it. **What I did, and it worked.** -- **Peti's box is gone from everything we run.** Its hub customer and all its records went through the hub's own delete, and its off-site folder was removed. The hub keeps only its history (events and the deletion note). -- **His off-site "backup" was never a backup.** His folder held one small key file and no backup at all. None of his 482 reports ever showed an off-site copy. That matches what you said: no user data. -- **Nothing of his was on ep0**, the off-site server. I checked: its backups and its network peers are the same as before, and so are the demo boxes' off-site folders. -- **The protected list is now DooPlex and ep0.** I updated the rules, the runbooks and the architecture pages. Old records keep their text. -- **Controller 0.272.0 is on both demo boxes.** The backup page now says in plain words when a whole-box backup cannot fit. Three small fixes: a restore now also removes the extra copies an update hold kept; a false error line after every undo is gone; the "update available" age is right for a re-tested image. I proved the three fixes on the scratch box. +- **The box now converts a PostgreSQL database to a new main version.** Only for an app whose step was tested. docmost is the first. On the scratch box: the conversion took 42 seconds in total. The data came back. +- **When it goes wrong, the box puts the old version back.** I tested three failures: a load error, an app that does not start, and a controller killed in the middle. Each time the old version came back with its data. +- **docmost moved in the catalog.** Tonight the HP box's automatic update should convert its docmost by itself. I watch it. +- **Kept data is never a dead end now.** When an app is installed over old files, the page asks: "Use my kept data" or "Start fresh". A new "Kept data" page shows every leftover folder. The household can look at it (read only), load it, or delete it (they must type the app's name). I tested all of it on the scratch box, in both languages. +- **The hub's customer delete stops promising to clean Cloudflare.** It now lists what you remove by hand. **What broke, or is not done.** -- **The "does not fit" sentence has not appeared on a real box yet.** No box is short of space now. The tests prove it in both languages. -- **The hub's customer delete promises to remove the Cloudflare tunnel and name, but it does not.** Peti's tokens were deleted with his record. Anything left on Cloudflare's side is not checked. I filed it as a row. -- **Last night's watch did not run.** This session ended in the daytime. +- **The restore of a removed app was broken for data drives.** It came back without its settings or database. Fixed tonight. +- **The Kept data list named two old folders "Filebrowser".** Found on the scratch box. Fixed before release. +- **The other ten PostgreSQL apps still cannot move.** Each needs its own test. +- **The file browser cannot open nextcloud's kept folder.** The files are there and can be loaded or deleted. +- **The HP box's restore test tries to test the golden template file and fails every 6 hours.** Not fixed (agent side). +- **The memory check always calls docmost "tight".** Its memory grows with its limit. Filed. +- **adventurelog stays on its old version.** I found how to skip its internet download. The move needs a choice. -**Rows.** 1 opened, 6 closed. The list went from 339 to 334. +**Rows.** 6 opened, 3 closed. The list went from 334 to 337. + +**Decision for you — D3: may the box ever delete kept data by itself?** +- **A — never; only the household deletes it (I recommend this).** Cost: kept data can fill a drive. The drive-full warning names the kept folders and their sizes, so the household sees what to delete. +- **B — after 90 days, with mails at 30 and 7 days before.** Cost: a small add-on to build; a household that ignores the mails loses the data. +- **If you do nothing:** nothing is ever deleted by itself (A in effect). Nobody is blocked. **What needs you.** -1. **Remove Peti's box from the Claude project instructions** (your own text in the Claude project). I cannot edit those. If you leave it, new sessions still treat his box as protected. -2. **Check Cloudflare for anything left of Peti's domain** (`sajatfelhom.hu`: a tunnel or DNS records). If you do nothing, it stays there unused. -3. **The old Storage Box `PBS-storage-1`** (u629193) is still on the list for you to delete. Neither of our tokens can see it now; it may already be gone. If it still exists, it keeps old test leftovers, including a folder named for Peti. +1. **D3** above. +2. **adventurelog:** skip its world-data download on every start (new installs then have no country list), or keep it and test the update with the file present. If you do nothing, it stays on the old version. +3. **From before:** Peti's box in the Claude project text; Cloudflare leftovers of Peti's domain; the old Storage Box `PBS-storage-1`. If you do nothing, they stay as they are. diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 06a953c6..7a987114 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -441,6 +441,15 @@ R-636's louder repeated alarm. when they hold no table; `CREATE ROLE ;` skipped only for a role that exists — its `ALTER ROLE … PASSWORD` still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks (`adminpack`) stopped the load and the box undid it. `audits/night-2026-09-26/A/README.md` A2. +39. **docmost's own memory limit rises 384M → 512M with its PostgreSQL 18 step** — *decided by CC unattended + 2026-09-25 — operator may reverse.* *One sentence:* the bench marked the step `memory_tight` (the docmost app, not + the engine) — raise the limit, or leave docmost unmoved? **Options:** (a) keep 384M — the move gate refuses a tight + step whose limit does not move, so docmost stays on 16; (b) raise to 512M in the same commit (the gate's own + remedy, the RomM precedent). **Costs:** (a) the conversion that was built and proven tonight never reaches a box; + (b) +128 MB of reservation per docmost box (`mem_limit` 768M → 896M), and the mark stays (80.4 % at 512M): measured + twice, Node sizes its heap from the limit — 349 MB at 384M, 431 MB at 512M, 0 kills and 0 restarts in both 10-minute + watches (~12 000 requests each). **Why (b):** 91 % of the old limit is too close for a household box, the gate's + rule is followed rather than bypassed, and the mark's weakness for such apps is filed (R-693). Reversible. --- diff --git a/documentation/audits/DRILL-night-2026-09-26.md b/documentation/audits/DRILL-night-2026-09-26.md new file mode 100644 index 00000000..590afaf6 --- /dev/null +++ b/documentation/audits/DRILL-night-2026-09-26.md @@ -0,0 +1,113 @@ +# DRILL — 2026-09-25 evening → 26 night: the box converts a PostgreSQL major (docmost first); kept data; the night watch + +Brief: "the box converts a PostgreSQL major for its first app (docmost); a reinstall over kept data asks the household…". +Architecture read before any claim: `09-update-architecture.md` §3 decisions 13–36, §3b Q5, §6.1a, §6.4 part 10; +`07-backup-architecture.md` §6; `08-alarm-ladder.md`. Evidence: `audits/night-2026-09-26/` (part0, A, B, C, D, E, F, G, tools). + +## Not done, or changed + +- **Two controller releases, not one** — as the brief allowed: **v0.273.0** (the conversion; floor raised at 11:32Z so + the live catalog could move docmost before the night) and **v0.274.0** (kept data; 12:30Z). +- **Part E was built by a parallel helper session** in its own git worktree (the brief's parts B–D and E touch + different files); I reviewed it, found one defect live (R-692) and fixed it before the release, and ran the whole + E5 proof myself. +- **The live proofs of Part D ran on `0.273.0-rc1`**, not the released image: the release adds only the recovery line's + wording (found live in case (c)) and R-687's two small items. +- **docmost moved with its own memory limit raised 384M → 512M** (decision 39): the bench marked the step + `memory_tight`, and the move gate refuses a tight step whose limit does not move. The mark stays at 512M (80.4 %) — + Node sizes its heap from the limit (R-693). +- **The bench was a throwaway LXC 9401 on demo-hp**, created and destroyed tonight (the first create picked an ARM + template and failed to start; destroyed, name checked first, recreated with amd64). 9202 could not be the bench: + the harness's container names collide with the installed docmost and it ends with `down -v`. +- **Part F3 (adventurelog, R-655): measured, not moved.** The answer to "skip or pre-seed download-countries" is in the + row; the move needs a choice between two options with different costs. +- **My own tool error in Part A4:** a bare `docker compose up -d` in docmost's stack dir (no `.env` there — the controller + passes the env) started docmost without its secrets; it crash-looped and **the box stopped it after 7 restarts + (decision 28 working)**. Started again with the product's Start; the seed read back. +- **The kept pre-conversion copy's release on 9202** needs the hourly job after a proven backup; its first run is after + this record was written — result in Part G. + +## Claims in the brief, judged + +1. *`postgres:18` moved its datadir and refuses or warns at `/var/lib/postgresql/data`* — **TRUE, and stronger:** it + REFUSES (exit 1) even an EMPTY volume there (`A/A1-images.txt`). +2. *The safety dump is `pg_dump` per database with `--no-owner --no-privileges`, not `pg_dumpall`* — **TRUE**. For docmost + both routes rebuild an identical database (one role, one database); the box uses `pg_dumpall` (decision 38). +3. *`Restore` brings the old datadir back so the old engine starts* — **TRUE**, by hand (A4) and by the product three + times (Part D cases a–c). +4. *The R-487 restore of a removed app picks up the kept drive files* — **FALSE on 0.272.0**: nextcloud came back with no + env, no database, its files mounted on the guest's root disk — the unit on the DATA drive was never found (R-690, + fixed in v0.274.0, proven live). +5. *nextcloud needs a rescan for files newer than its backup* — **TRUE** (`occ files:scan --all`, 0.6 s); now the + template's `after_load:`. +6. *Only demo-hp runs docmost* — **TRUE**. +- Also measured: R-657's install LOOP did not recur on 0.272.0 — the install succeeded and the old files were silently + unreachable; the choice replaces both shapes. + +## Part 0 — `part0/README.md` + +Decisions 35, 36 recorded; D3 in STATUS. Last night: demo-felhom's leg did opengist in 20 s at 04:15; its whole-box +backup ran at 07:29, three hours after the leg — nothing waited; demo-hp's leg had nothing and reported +`"steps": null` (fixed); no whole-box backup was due there. **Found: R-689** — demo-hp's restore test picks the golden +template as "the newest settled archive" and fails every 6 h. + +## Part A — the spike (`A/README.md`) + +Target **18** (decision 37); load from **`pg_dumpall`** with zero tolerated errors (decision 38); the check (owners, +encodings, roles, extensions, per-table rows) in 0.5 s; the undo after emptying proven by hand; space: the dump is +0.2 % of the datadir — the box's bound is the volume's size × 1.25 + the 2 GB floor. + +## Part B — controller v0.273.0 (`B/`) + +`internal/stacks/pgconvert.go`; 17 tests; **9 red-proofs** each seen failing (mark missing → refused; marker re-checked +before emptying; counts compared; `PG_VERSION` checked; a cut-off dump; a restart during `converting` undone; the old +copy kept; the release wired; the space check). + +## Part C — the catalog (`C/`) + +Harness v4 (converts on the bench, moves an 18 mount, writes the mark only when both venues converted); the engine +gate's proof clause with 5 decoys + a red-proof; the postgis family judged. Bench: negative control C3 `failed`; +docmost run 1 (384M) converted in 25.0 s, proven, `memory_tight` 90.9 %; run 2 (512M) converted in 11.0 s, proven, +80.4 %; abort (16 images on the 18 datadir) refuses, as expected — the product's route back is the undo, not that. + +## Part D — the live proof, the floor, the move (`D/`) + +| case (9202, drill catalog, `0.273.0-rc1`) | phases | end | PG_VERSION after | seed | +|---|---|---|---|---| +| happy path | safety-dump → pulling → copying → **converting 10 s** → starting → verifying | **done, 42.1 s** | 18 | read back | +| (a) the load fails (`adminpack`, gone in 17+) | … → converting → undoing | **undone, 43.6 s** | 16 | read back | +| (b) unhealthy on 18 (drill probe :3999) | … converting → starting → verifying (90 s) → undoing | **undone, 148.1 s** | 16 | read back | +| (c) SIGKILL 1 s after the volume was emptied | … converting → (restart) → undoing | **undone, 40 s after the restart** | 16 | read back | + +Floor 0.273.0 (then 0.274.0) read back from the hub; both demo boxes healthy within ~25 s. **docmost moved in the live +catalog at ~14:40 CEST** (`afd3a60`): the engine gate printed `ALLOWED … proven on both venues and carries the +conversion mark`; the Hungarian copy gate caught the RAM line's change on the first push (freeze updated on purpose). + +## Part E — kept data (`E/`) + +E1 answers above (claims 4, 5). Built: the install choice (409 `kept_data_choice`), the „Megőrzött adatok" page, the +read-only view, the drive-full naming, `after_load:`. **E5 on 9202 (`0.274.0-rc2`), both languages, all passed:** ask +(use off without a database copy) → start fresh (126 MB renamed to `kept/nextcloud//`) → seed, a file, a unit, a +file after it → remove keeping data → **use** (loaded from the unit on the DATA drive; rescan ran; account + both +files back; a never-written file stays invisible) → start fresh again → Load refused while a leftover occupies the +folder → Delete refused for a wrong name and for an unlisted path → Delete → **Load** (files moved back, database +loaded, rescan; account + both files back) → a write into the view: `Read-only file system`. **R-692 found and fixed +live** (two leftovers named „Filebrowser"). + +## Part F + +F1 (R-687 `[]` + the taken step's log): v0.273.0, 2 red-proofs. F2 (R-688): hub v0.125.0 — the delete dialog lists the +Cloudflare items to remove by hand; live preview read for demo-hp. F3 (R-655): measured, not moved (see the row). + +## Part G — the night watch + +*(filled in during the night)* + +## Teardown + +*(filled in at the end)* + +## Register + +Before: **334 rows / 674,532 B**. Opened R-689, R-691 (helper), R-693, R-694; opened and closed R-690, R-692; closed +R-657; narrowed R-450, R-463, R-469, R-687, R-688, R-691; R-655 measured. After: see the final count at the end. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 55f08887..f9d015e3 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -802,7 +802,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | | **R-652** | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** | | **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** | -| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** | +| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. **-- MEASURED 2026-09-25 (read from both images on the bench, `audits/night-2026-09-26/F/F3-01-entrypoint.txt`):** v0.13.0's entrypoint has an env switch, `SKIP_WORLD_DATA=1`, that skips `download-countries` (v0.12.1 has none). Both versions use the SAME dataset file, `countries+regions+states-v3.1.json`, kept in the `adventurelog_media` volume, and download only when it is absent — so an UPDATE re-uses the file v0.12.1 left. **The likely crash chain:** v0.12.1 exits only on 137 and ignores any other download failure, so a cut-off file it saved stays; v0.13.0 parses it with `ijson` and fails at every start. **Options for the move (not taken tonight):** (a) `SKIP_WORLD_DATA=1` in the template — no internet dependence at update, but a FRESH install gets no world data; (b) keep the download and have the harness prove the update with the file present (the pre-seed is the volume itself) plus a check that a truncated file is replaced (`--force` is the command's own repair). The healthcheck override fix (`/usr/bin/node`) still rides the image move. | **READY — P2; owner: CC (catalog)** | | **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** | | **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** | | **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |