R-672 session: audit README (not done first, wrong claims), STATUS, capability map, register 344->341, topic REPORT
gates / gates (push) Successful in 28s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-24 17:06:09 +02:00
parent 54bff69f88
commit d53bd08442
29 changed files with 395 additions and 24 deletions
+4
View File
@@ -382,4 +382,8 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-666** | **While support is informed, Remove offered to delete the data (decision 27).** v0.269.0: the dialog reads `keep_data_only` and offers only „remove the app, keep my data"; the API refuses data or backup deletion with 409 (hu + en); the no-whole-copy sentence is informal. Live on 9202 (a one-drive hold). | v0.269.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A2/10-held-no-copy-keep-data.*` |
| **R-667** | **A crash loop never reached the alarm, and nothing stopped it (decision 28).** v0.269.0 + hub v0.123.0: ≥ 6 restarts in 10 min (`RestartCount`, not the resettable `restarting_since`) or an OOM storm → the box stops the app, holds it (`unhealthy_stop`), tells household + operator (`app_stopped_unhealthy`); Start = one more try; a repeat in 24 h says support is informed. Live: gokapi trip 1 and 2; chaos rounds 1, 7, 8, 12. **Rule:** Docker's back-off caps a steady loop at ~1 restart/min, so a threshold must be below 10 per 10 min. | v0.269.0 / hub v0.123.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A3/`, `08` §6.2 |
| **R-668** | **The Tier-2 copy chose a registered path that no longer existed, on the app's own disk (P2).** v0.269.0: the same-disk check fails CLOSED. Live: the next copy went to the SSD. Residual: an older same-disk record counts until the next Tier-2 run replaces it. | v0.269.0, 2026-09-24 | `git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-24/A1/01-find-mirror.txt` |
| **R-669** | **After a failed update ended in a restore, the box kept the FAILED step's health check as the pinned version's, and the next undo held the app for nothing (P2).** v0.270.0: the recovery unit captures the pinned version's `.felhom.yml` (`applied-meta`), not the stack dir's file the sync may already have replaced; a restore makes the restored file the applied record. Live on 9202: the sync wrote the bad probe at 14:57:46, the unit captured at 14:57:55 kept the good one; a second-drive restore turned stack 8999 / applied 3000 into 3000 / 3000, the applied record rewritten at the restore. **Rule:** a copy of an app's definition carries the PINNED version's health check, never the catalog's newest. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/30-33*` |
| **R-674** | **The ladder log called a pin at the head "older than the ladder" (P3).** v0.270.0: it says "AT THE HEAD". Unit test + red-proof only — no product path reaches it since R-679 refuses a current app first. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/redproofs/D-r674.txt` |
| **R-679** | **An Update on an app already at the head ran the whole guarded update — dump, pull, restart (P2).** v0.270.0: `409 already_current` before anything moves, hu + en; a re-tested digest of a floating tag still updates. Live on 9202 (privatebin): both languages refused, no backup, no pull. A test comment had called a same-version Update "the repair path" — Restart is. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/10-r679-live.txt` |
| **R-681** | **An install cut off by a controller restart was lost silently (P2).** v0.270.0: an install marker before the compose-up, removed when the install ends; a marker at start → `compose down` (volumes kept), stale pin records cleared, `app_deploy_failed` with the reason, and the apps page says the install was interrupted (hu + en) until the next install; a finished install only loses the marker. Live on 9202: mealie killed 0 s into its pull → reported, page sentence, nothing left running; reinstall cleared the sentence; actualbudget (killed after it finished) left alone. **Rule:** a long-running act the customer started is journaled so a restart finishes or reports it. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/20-26*` |
-4
View File
@@ -806,19 +806,15 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** |
| **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** |
| **R-669** | **[P2-MEDIUM] After a failed update ends in a restore, the box keeps the FAILED step's health check as the pinned version's — and the NEXT failed update's undo judges the correct old version with it and holds the app.** PROVEN LIVE 2026-09-24 on 9202 (v0.269.0): a failing step (probe port 8999) → hold → the second drive's whole restore cleared the hold, but `applied-meta/.felhom.yml` still carried port 8999 (written by `advancePinTo` at 10:39:21; no restore path rewrites it), and the stack's `.felhom.yml` came back from the unit, which the sync had already filled with the failing step's file. The next failed update's undo put the right version and the right data back and then called it unhealthy on port 8999 → **a false hold** (`not_started`). Data safe; the app stopped for nothing. **Fix direction:** a restore that recreates the definition also sets the applied record to the `.felhom.yml` of the RESTORED pin (the ladder step's `steps/<key>.felhom.yml` for those refs, or the catalog's when the catalog head equals the pin), never the unit's copy. Pre-existing since v0.263.2; not a v0.269.0 regression. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P2; owner: CC (controller)** |
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
| **R-672** | **[P1-HIGH] The agent's SCHEDULED restore-test filled the production thin pool and turned a customer guest's disks read-only.** FOUND 2026-09-24 on demo-hp (agent v0.132.0): at 10:29 CEST the restore-test restored 9201's 22 GiB archive as scratch guest 990000 into `local-lvm` — the SAME pool as 9201 — with no free-space check. The pool reached 100 % at 10:35:24 (`out_of_data_space`, `error_if_no_space`); the scratch failed to start; its teardown failed (`lvremove … contains a filesystem in use`) and was "left for Recover" — **which runs only at agent start**, so the pool stayed full for 2.5 h. The agent logged *„a full pool corrupts every guest on it"* every few seconds and took no action. **Consequence, measured:** 9201's controller got `no space left on device` from 10:38, its log stops at 10:39:57, and at 13:12 both its rootfs and `/var/lib/felhom` were READ-ONLY (errors=remount-ro). **Intervention:** an agent restart ran Recover ("destroyed leaked restore-test scratch guest", pool 100 % → 58.9 %). **Still open:** 9201 needs a stop + fsck + start (operator — the session's permission check refused host-level guest operations). **Fix direction:** the restore-test refuses to start unless the target storage has the archive's size plus a margin free and never targets the pool of the guest it tests; a failed teardown retries on a timer, not only at start; the pool-fill WARN pages the operator. `audits/night-2026-09-24/C-02…C-07` **— 2026-09-24 (evening): FIXED IN RELEASE, NOT YET DELIVERED.** Agent **v0.133.0** (tag `9bdb4da`, package verified by download): space preflight before anything is created (UNCOMPRESSED size × 1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage is eligible — demo-hp has none, `nvme-scratch` carries no grant — unknown refuses, reported as a non-pass `skipped`); failed teardown retried every 10 min; thin pool ≥ 90 % requests an immediate report. Hub **v0.124.0** (live): thin pool critical at 90 % of data or metadata, one alarm per pool per 6 h. **The brief's "archive × 1.2 + 5 GiB" would NOT have prevented this incident** (6.9 GB file → 22.6 GB restore). Six red-proofs; live on demo-hp: both margins REFUSE (21.1 GiB restored needs 30.3 GiB, 22.1 GiB free), pool unchanged. The hub HAD alarmed on the day (`storage_fill_critical` at 100 %, operator mail) — late and at the generic bands. **9201 repaired** (stop, `pct fsck` both disks: 17 half-deleted inodes + an orphan block fixed on mp0; second pass clean; start; floor 0.269.1 taken by itself; ONLINE); two Redis AOF tails the pool cut (2,943 B / 6,631 B) truncated with `redis-check-aof --fix` after copies were kept as volumes — the box had meanwhile stopped docmost and romm itself (decision 28, live). **Scheduled restore-test OFF on both demo hosts** (operator ruling) until v0.133.0 is delivered. `audits/r672-2026-09-24/` | **FIXED IN v0.133.0 — awaiting the operator's signed delivery; restore test OFF on demo hosts; owner: operator (sign) / CC** |
| **R-673** | **[P2-MEDIUM] 9201's whole-box backups failed all morning on a stale `snapshot-delete` lock.** demo-hp 2026-09-24: vzdump of 9201 failed at 06:59, 07:17, 07:42 and 08:22 CEST (`CT is locked (snapshot-delete)`), leaving `snap_vm-9201-disk-1_vzdump` (07:42) behind, before R-672's pool fill. The agent's stale-lock scanner ran every 30 s and did not clear it; the lock was gone after the 13:09 agent restart. Cause not read — the session's permission check refused reading the task logs. `audits/night-2026-09-24/C-01-demo-hp-9201-vzdump-errors.txt` **— 2026-09-24 (evening): CAUSE READ, FIX RELEASED.** Not the pool fill (that came at 10:35). The 06:59 and 07:42 backups failed writing the archive to `local` — the host ROOT disk: `zstd: error 70 : … No space left on device` (R-684); the 07:42 failure's cleanup left `snap_vm-9201-disk-1_vzdump` and the `snapshot-delete` lock, and the 08:22 run hit the lock. The agent's own stale-lock recovery cleared it at 13:10, one minute after the agent restart — because it ran ONLY at start. v0.133.0 runs it every 10 min under the one-heavy-operation gate (red-proofed). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **FIXED IN v0.133.0 — awaiting delivery; root cause R-684; owner: CC** |
| **R-674** | **[P3-LOW] The ladder log says an app's pin „matches no update_ladder entry … older than the ladder" when the pin EQUALS the head.** Seen on 9202 2026-09-24 for nextcloud at the head. Misleading to an operator reading why nothing climbed. **Fix:** say "at the head" when the pin equals the newest `to`. | **READY — P3; owner: CC (controller)** |
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` | **OPEN — P3; owner: CC (watch)** |
| **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** |
| **R-678** | **[P3-LOW] After an update step ends `done`, the app's „steps left" and its badge stay STALE until the next scan.** MEASURED 2026-09-24 on 9202 (Part C, first caller run): wishlist read `ladder_steps_left` 1 after its only step, navidrome 2 after both of its steps — for ~50 s, through six presses. A person sees „Frissítés elérhető" for an app that just updated; an automatic caller re-presses. **Fix:** the update's finish refreshes the app's catalog/ladder fields (the same read the scan does). `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P3; owner: CC (controller)** |
| **R-679** | **[P2-MEDIUM] An Update pressed on an app that is already current runs the whole guarded update — backup, pull, restart — and reports `done`.** MEASURED 2026-09-24 on 9202: navidrome at the head was pressed four times by the stale-read caller (R-678); each press made a safety dump and restarted the app (8.5 s of downtime each) and changed nothing. There is no `current` refusal (`UpdatePreflight`, `update.go:366`; `UpdateOrderCurrent` falls through). Harmless by hand, costly for the automatic leg (`09` §6.4.2). **Fix:** the preflight refuses with reason `current` when the pin equals the catalog head and no newer tested digest exists. `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P2; owner: CC (controller)** |
| **R-680** | **[P2-MEDIUM] The box does not remember a failed update step — after an undo it offers the same step again.** MEASURED 2026-09-24 on 9202: vikunja's failing step was undone (104 s) and its badge went straight back to „Frissítés elérhető"; only the test caller's own memory stopped a re-press. Decision 15 („a failed step is never pressed again") therefore holds for a person only by their judgement and not at all for the automatic leg. **Fix (part 7 (b), `09` §6.4.2):** record the failed `to` per app; the leg skips it until the catalog's ladder for that app changes; the page says the step was tried and put back. `audits/night-2026-09-24/C/31-night.json` | **READY — P2; owner: CC (controller, with part 7)** |
| **R-681** | **[P2-MEDIUM] An install interrupted by a controller restart is lost SILENTLY — no event, no page sentence, and its half-written files stay.** MEASURED 2026-09-24 on 9202 (Part E round 2, v0.269.1): n8n's install pressed at 11:48:42Z ("Deploying stack n8n … checking 1 images"); `systemctl restart docker` 20 s later took the controller down with it. After the restart the box logged NOTHING about n8n, the app read `not_deployed`, and the household's page showed it as never installed — while `app.yaml` (with its generated `N8N_ENCRYPTION_KEY`), `applied-compose.yml` and `applied-meta/` stayed in the stack dir. A fresh install through the product afterwards worked (211 s) and was not blocked by the leftovers. Compare the update path, which journals and RESUMES after the same accident (round 3/4). **Fix direction:** journal the install like the update; on boot either resume it or say on the page (and in an event) that it was interrupted, and clean the half-written files. `audits/night-2026-09-24/E/round-02*` | **READY — P2; owner: CC (controller)** |
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
| **R-684** | **[P2-MEDIUM] demo-hp's whole-box backups cannot fit on its backup storage — 9201 has had no new whole-box backup since 2026-09-23.** `local` is the host ROOT disk (40 GB, 84 % used, ~4 GB free); it holds three 9201 archives of 5.8–6.9 GB (retention 3) and a new one needs ~7 GB: the 2026-09-24 runs failed `No space left on device` (06:59, 07:42), leaving the stale lock of R-673. Every night now fails the same way until space is made or the retention/target changes. A Tier-0 box, but the same shape exists on any appliance whose local backup target is its root disk. **Needs a decision** (fewer kept archives, another target, or a larger disk). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **OPEN — P2; owner: operator (the decision) / CC** |