The update arc's two missing measurements, the lock, and the floor to 0.260.0
gates / gates (push) Successful in 23s

Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.

Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.

Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.

Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.

Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-21 15:00:09 +02:00
parent d19f07ea04
commit c85262111c
18 changed files with 1300 additions and 76 deletions
+8 -1
View File
@@ -714,7 +714,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. | **CLOSED 2026-09-21 — measured on a real version change** |
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. **CORRECTED 2026-09-21 (same day): that sentence left the dangerous half of the question in prose and in no row — see R-610, which carries it and CLOSES it with three measurements.** The history above stands as written; only this pointer is added. | **CLOSED 2026-09-21 — measured on a real version change** |
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. **CLOSED 2026-09-21 — controller v0.260.0.** `stacks.CatalogOrder` (`internal/stacks/updateorder.go`) replaces the three-way comparison with FOUR verdicts — Unknown / Current / Behind / **Ahead** — and **moves out of `web` so the badge and the refusal read ONE verdict**; `web.compareInstalledToTemplate` is now a thin wrapper. An app AHEAD reads „Naprakész" / "Up to date" with `tag-ok` (the same word and class as level — there is nothing for the household to do) and a title saying why (`badge.update.ahead.title`, both bundles). `Manager.UpdatePreflight` refuses with reason `downgrade`, HTTP 409, „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.", logged with both image maps. **The API now renders update refusals through `errText`**, so the new key is not a seam built and never wired. **Ahead is NARROW on purpose:** every differing service must be orderable AND newer, or the verdict falls back to Behind — this gate can BLOCK an update, so it errs towards letting one run. Ordering is `util.Version.Compare` (the house rule: one comparator) behind a tag normaliser — `X.Y`/`X.Y.Z`, optional leading `v`, two-part padded with `.0`, and a trailing suffix that must be IDENTICAL on both sides, so `nextcloud:31.0.14-apache → 31.0.15-apache` orders while `postgres:16-alpine`, `26.05.2-ls310 → -ls311`, `kimai/kimai2:apache-2.57.0`, a date stamp and a digest pin do not. **The suffix rule was found by the fixture, not by design** — the first implementation called every real catalog tag unorderable. **Three red-proofs, each SEEN to fail.** Recorded as `09` §3 decision 10 (decided by CC unattended — operator may reverse). | **CLOSED 2026-09-21 — controller v0.260.0** |
@@ -781,6 +781,13 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** |
| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-607** | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `<data>/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-608** | **[P2-MEDIUM] The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the window proposed for automatic app updates.** FOUND 2026-09-21 by reading the clock, not by a failure. The controller self-updates daily at `self_update.auto_update_time` — **default 04:30** (`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365) — and again from `MaybeAutoUpdate` after ANY hub report once a floor sits above the box, so at any hour. The swap restarts the controller container. `09` §3b **Q1** proposes **02:30–05:00** for automatic app updates. **It contains 04:30.** **MEASURED, and the gap was NARROWER than it first looked — which is why the fix is where it is:** the updater's only busy gate was `backupRunning` (`updater.go` L61, read at L487 dry-run, L512 `TriggerUpdate`, L660 `maybeAutoUpdate`), wired in `main.go` L659 to `backupMgr.IsRunning()`. The guarded update's **`backing-up` phase DOES take the backup single-flight** (`RunAppBackupNow` → `acquireRunning`, `backup/update_guard.go:333`), so that ONE phase was already covered. `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` were not — and the last two are exactly where the new version may already have touched the customer's data. The reverse was absent too: `UpdatePreflight` never asked whether a swap was running. **CLOSED 2026-09-21 — controller v0.261.0.** `stacks.Manager.AnyUpdating()` → `Updater.SetAppUpdatingCheck`, a deliberate sibling of `SetBackupRunningCheck` consulted in the SAME three places; `Updater.IsUpdateRunning` → `Manager.SetSelfUpdatingCheck`, and `UpdatePreflight` refuses `self_updating`. Both wired in `main.go`, the only place holding both objects — **`stacks` never imports `selfupdate`.** Two sentences born as bundle keys. **THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH:** `Stack.Updating` is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it. A latching gate would be a worse failure than the one prevented, and a silent one. Pinned by `TestR608_LockReleasesAfterHold`. Four red-proofs, each seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** |
| **R-609** | **[P3-LOW] An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never".** `UpdateRefusal.Reason` has existed since v0.237.0 (`busy`, `deploying`, `updating`, `held`, `migrating`, `memory`, `disk`, `no_backup`, `downgrade`, and `self_updating` since v0.261.0) and never left the process: the 409 body carried only the translated sentence. **The distinction is not decorative** — `busy`/`updating`/`deploying`/`migrating`/`self_updating` are TRANSIENT and `held`/`downgrade` are TERMINAL until a person acts. A caller that cannot tell them apart either gives up on a passing backup window or presses a terminally-refused button on every pass for ever. `09` §6.2's unattended caller reads exactly this. **CLOSED 2026-09-21 — controller v0.261.0.** The body gains `data: {"reason": "<Reason>"}`, ADDITIVELY; the sentence is unchanged and no page moves. Table-driven test over five reachable paths plus a control that a non-refusal carries none. **FOUND WHILE WRITING THE TEST, NOT BY READING — and it was the reason that matters most:** `actionStack` refuses a HELD app on **its own line, BEFORE `UpdatePreflight`** (`api/router.go` ~L601), so `held` would have been the one reason missing from the wire. That line now carries it too. Red-proof seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** |
| **R-610** | **[P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only.** R-520 (CLOSED 2026-09-21) cut in `pulling`, where **nothing had run**: the pin goes back and that is the easy case. Its own last lines said the `starting` cut was "NOT measured and does not re-open this row", and no row carried it. The dangerous case is the cut AFTER the new version has started and may already have migrated the customer's data — where `RecoverUpdates` (`stacks/update.go:909`) marks the app Updating and RESUMES rather than rolling back. **That behaviour was READ from the source and never observed.** **CLOSED 2026-09-21 — measured THREE times on guest 9202, controller v0.260.0**, with three different apps and two different cut mechanisms: vikunja 2.3.0→2.6.0 and uptime-kuma 2.4.0→2.5.0 by `pct stop` (a real power cut), and wishlist v0.66.0→v0.67.0 by restarting ONLY the controller container (exactly what a self-update does). **All three ended HONEST:** the recovery line appeared, the update resumed, each app came up on the NEW version, and in every case all FOUR version observables agreed — `pinned_images`, `installed_images`, the live compose `image:` line, and `docker inspect` of the running container (digests matched too). No hold, no stuck `Updating`, no surviving journal, no retry loop. The seeded data read back through each app's own front door for the two where a post-cut read-back was taken. **THE DANGEROUS CASE WAS GENUINELY EXERCISED, and the proof is a log line, not an assumption:** vikunja's own log shows `Ran all migrations successfully` and `Vikunja version v2.6.0` at **12:28:26.881 UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema migration had ALREADY been applied to the customer's SQLite database when the power went. Recovery resumed FORWARD, so old-binary-on-migrated-database never happened — **but this branch is one step from it: had the cut landed a second earlier, in `pinning` or `pulling`, the 2.3.0 pin would have been put back onto a 2.6.0 database.** That is not a defect today; it is the reason §4's "no automatic rollback" ruling is right, and it is now evidence rather than argument. **INSTRUMENT LIMIT, stated because it bounds the claim:** `starting` lasts well under a second on this box. Three attempts across two cut mechanisms (`pct stop` returning in 3.0–3.8 s; a controller restart in 1.7 s) ALL landed in `verifying`. No phase was faked. **`RecoverUpdates` handles `starting` and `verifying` in ONE branch, so all three runs exercise the same recovery arm** — the arm under test. A cut that lands inside `starting` itself remains unmeasured and would need an in-process fault injector. Evidence: `audits/update-arc-gaps-2026-09-21/` 04, 05, 07. | **CLOSED 2026-09-21 — measured three times** |
| **R-611** | **[P3-LOW] A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement.** The 2026-09-21 update-arc session's brief contained a Phase 5 spike: one app updated by the box with **nobody pressing anything**, once succeeding and once forced to fail. **It did not run, and nothing said so** — no evidence file, no code, and no sentence in `STATUS.md`, `UPDATE-ARC-STATE-2026-09-21.md`, either `REPORT.md` or `09`. `09` §6.2 was left describing Slice 6 "as it would be built" with no measurement under it, which reads like a considered design rather than an untested one. **Why this is a row and not a grumble:** the missing measurement was recoverable in an afternoon; the missing SENTENCE was not, because the next reader had no way to know it was missing. A skipped phase that is declared costs one line; a skipped phase that is not costs the next session its baseline. **CLOSED 2026-09-21 by the successor session**, which ran it (`audits/update-arc-gaps-2026-09-21/`, scenarios F and G) **and** adopted the standing habit that closes it generally: **the report's FIRST section is "not done", even when empty.** Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason. | **CLOSED 2026-09-21 — run by the successor session; "not done" is now the report's first section** |
| **R-612** | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** | **READY — rank P1-HIGH; owner: CC (catalog + a look at whether a failed first-boot seed can ever be visible)** |
| **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
| **R-614** | **[P3-LOW] A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated.** OBSERVED 2026-09-21 on guest 9202: a newly deployed `uptime-kuma` read `update_phase=done` / „Frissitve" before any update had been run against it — left in the manager's IN-MEMORY stack state by the previous session's update of a since-removed instance of the same name. It cleared on the next controller restart. **Why it matters beyond the cosmetic:** `update_phase` is one of the fields a person (and, after R-609, an unattended caller) reads to decide whether an update happened. A value that outlives the app it described is the same class as the R-166 "absent means unknown" family — a confident answer about something that no longer exists. **Fix shape:** clear the update fields when a stack is removed, beside wherever `Updating`/`updateHeld` are reset; a test that removes and redeploys and asserts the phase is empty. Small. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** |
| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `&#39;`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |