night 2026-09-25 (in progress): part 7 shipped as controller v0.271.0 — decisions 31-33, evidence A/B/C/F, register (R-685..R-687; closed R-672 R-673 R-684 R-680 R-678)
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -386,4 +386,8 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-674** | **The ladder log called a pin at the head "older than the ladder" (P3).** v0.270.0: it says "AT THE HEAD". Unit test + red-proof only — no product path reaches it since R-679 refuses a current app first. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/redproofs/D-r674.txt` |
|
||||
| **R-679** | **An Update on an app already at the head ran the whole guarded update — dump, pull, restart (P2).** v0.270.0: `409 already_current` before anything moves, hu + en; a re-tested digest of a floating tag still updates. Live on 9202 (privatebin): both languages refused, no backup, no pull. A test comment had called a same-version Update "the repair path" — Restart is. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/10-r679-live.txt` |
|
||||
| **R-681** | **An install cut off by a controller restart was lost silently (P2).** v0.270.0: an install marker before the compose-up, removed when the install ends; a marker at start → `compose down` (volumes kept), stale pin records cleared, `app_deploy_failed` with the reason, and the apps page says the install was interrupted (hu + en) until the next install; a finished install only loses the marker. Live on 9202: mealie killed 0 s into its pull → reported, page sentence, nothing left running; reinstall cleared the sentence; actualbudget (killed after it finished) left alone. **Rule:** a long-running act the customer started is journaled so a restart finishes or reports it. | v0.270.0, 2026-09-24 | `git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/D/20-26*` |
|
||||
|
||||
| **R-672** | **The scheduled restore-test filled the production thin pool and turned a customer guest's disks read-only (P1).** Agent v0.133.0: space preflight on the UNCOMPRESSED size, off the tested guest's pool, unknown refuses. **Delivered 2026-09-24 night** to both demo boxes by CC-signed `agent_update` jobs (ruling 1, 2026-09-16); restore test back ON; live: demo-felhom PASS 85 s (pool 3.13 → 5.68 → 3.13 %), demo-hp REFUSED for space (needs 30.3 GiB, has 22.1). Peti's box has not received it. | agent v0.133.0, delivered 2026-09-24 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/r672-2026-09-24/`; `audits/night-2026-09-25/A/` |
|
||||
| **R-673** | **9201's whole-box backups failed on a stale `snapshot-delete` lock (P2).** Agent v0.133.0: the stale-lock sweep every 10 min under the one-heavy-operation gate. Delivered to both demo boxes 2026-09-24 night; demo-hp 9201's next whole-box backup ran clean (no lock, 8.18 GB). | agent v0.133.0, delivered 2026-09-24 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/A/A4-*` |
|
||||
| **R-684** | **demo-hp's whole-box backups could not fit on its root-disk target (P2).** Operator ruling 2026-09-24 evening, option A: retention 3 → 1 (`local_backup_retention`), the two oldest archives pruned by PVE's own prune, and one backup by the product's trigger fitted (8.18 GB; root 89 % → 60 %). The product lesson — warn BEFORE the night — is R-685 (agent v0.134.0 skips with a reason). | ruling 2026-09-24; applied 2026-09-24 night | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/A/A4-hp-backup-space.txt`, `A4-hp-backup-run.txt` |
|
||||
| **R-680** | **The box did not remember a failed update step (P2).** Controller v0.271.0: an undone or held step is recorded in app.yaml (`failed_update_step`, tied to the ladder's print); the automatic leg skips it until the catalog's ladder changes; a person can still press. Live on 9202: vikunja undone night 1, skipped `failed_before` night 2, re-tried after the catalog re-tested it. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/`; `B/redproofs/R680-*` |
|
||||
| **R-678** | **After a step ended `done`, steps-left and the badge stayed stale (P3).** Controller v0.271.0: the update re-reads the app's catalog fields BEFORE it says done (and after an undo). Live on 9202: every automatic step's page read current at the leg's end. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `B/redproofs/R678-*` |
|
||||
|
||||
@@ -808,16 +808,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** |
|
||||
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
|
||||
| **R-672** | **[P1-HIGH] The agent's SCHEDULED restore-test filled the production thin pool and turned a customer guest's disks read-only.** FOUND 2026-09-24 on demo-hp (agent v0.132.0): at 10:29 CEST the restore-test restored 9201's 22 GiB archive as scratch guest 990000 into `local-lvm` — the SAME pool as 9201 — with no free-space check. The pool reached 100 % at 10:35:24 (`out_of_data_space`, `error_if_no_space`); the scratch failed to start; its teardown failed (`lvremove … contains a filesystem in use`) and was "left for Recover" — **which runs only at agent start**, so the pool stayed full for 2.5 h. The agent logged *„a full pool corrupts every guest on it"* every few seconds and took no action. **Consequence, measured:** 9201's controller got `no space left on device` from 10:38, its log stops at 10:39:57, and at 13:12 both its rootfs and `/var/lib/felhom` were READ-ONLY (errors=remount-ro). **Intervention:** an agent restart ran Recover ("destroyed leaked restore-test scratch guest", pool 100 % → 58.9 %). **Still open:** 9201 needs a stop + fsck + start (operator — the session's permission check refused host-level guest operations). **Fix direction:** the restore-test refuses to start unless the target storage has the archive's size plus a margin free and never targets the pool of the guest it tests; a failed teardown retries on a timer, not only at start; the pool-fill WARN pages the operator. `audits/night-2026-09-24/C-02…C-07` **— 2026-09-24 (evening): FIXED IN RELEASE, NOT YET DELIVERED.** Agent **v0.133.0** (tag `9bdb4da`, package verified by download): space preflight before anything is created (UNCOMPRESSED size × 1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage is eligible — demo-hp has none, `nvme-scratch` carries no grant — unknown refuses, reported as a non-pass `skipped`); failed teardown retried every 10 min; thin pool ≥ 90 % requests an immediate report. Hub **v0.124.0** (live): thin pool critical at 90 % of data or metadata, one alarm per pool per 6 h. **The brief's "archive × 1.2 + 5 GiB" would NOT have prevented this incident** (6.9 GB file → 22.6 GB restore). Six red-proofs; live on demo-hp: both margins REFUSE (21.1 GiB restored needs 30.3 GiB, 22.1 GiB free), pool unchanged. The hub HAD alarmed on the day (`storage_fill_critical` at 100 %, operator mail) — late and at the generic bands. **9201 repaired** (stop, `pct fsck` both disks: 17 half-deleted inodes + an orphan block fixed on mp0; second pass clean; start; floor 0.269.1 taken by itself; ONLINE); two Redis AOF tails the pool cut (2,943 B / 6,631 B) truncated with `redis-check-aof --fix` after copies were kept as volumes — the box had meanwhile stopped docmost and romm itself (decision 28, live). **Scheduled restore-test OFF on both demo hosts** (operator ruling) until v0.133.0 is delivered. `audits/r672-2026-09-24/` | **FIXED IN v0.133.0 — awaiting the operator's signed delivery; restore test OFF on demo hosts; owner: operator (sign) / CC** |
|
||||
| **R-673** | **[P2-MEDIUM] 9201's whole-box backups failed all morning on a stale `snapshot-delete` lock.** demo-hp 2026-09-24: vzdump of 9201 failed at 06:59, 07:17, 07:42 and 08:22 CEST (`CT is locked (snapshot-delete)`), leaving `snap_vm-9201-disk-1_vzdump` (07:42) behind, before R-672's pool fill. The agent's stale-lock scanner ran every 30 s and did not clear it; the lock was gone after the 13:09 agent restart. Cause not read — the session's permission check refused reading the task logs. `audits/night-2026-09-24/C-01-demo-hp-9201-vzdump-errors.txt` **— 2026-09-24 (evening): CAUSE READ, FIX RELEASED.** Not the pool fill (that came at 10:35). The 06:59 and 07:42 backups failed writing the archive to `local` — the host ROOT disk: `zstd: error 70 : … No space left on device` (R-684); the 07:42 failure's cleanup left `snap_vm-9201-disk-1_vzdump` and the `snapshot-delete` lock, and the 08:22 run hit the lock. The agent's own stale-lock recovery cleared it at 13:10, one minute after the agent restart — because it ran ONLY at start. v0.133.0 runs it every 10 min under the one-heavy-operation gate (red-proofed). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **FIXED IN v0.133.0 — awaiting delivery; root cause R-684; owner: CC** |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-678** | **[P3-LOW] After an update step ends `done`, the app's „steps left" and its badge stay STALE until the next scan.** MEASURED 2026-09-24 on 9202 (Part C, first caller run): wishlist read `ladder_steps_left` 1 after its only step, navidrome 2 after both of its steps — for ~50 s, through six presses. A person sees „Frissítés elérhető" for an app that just updated; an automatic caller re-presses. **Fix:** the update's finish refreshes the app's catalog/ladder fields (the same read the scan does). `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-680** | **[P2-MEDIUM] The box does not remember a failed update step — after an undo it offers the same step again.** MEASURED 2026-09-24 on 9202: vikunja's failing step was undone (104 s) and its badge went straight back to „Frissítés elérhető"; only the test caller's own memory stopped a re-press. Decision 15 („a failed step is never pressed again") therefore holds for a person only by their judgement and not at all for the automatic leg. **Fix (part 7 (b), `09` §6.4.2):** record the failed `to` per app; the leg skips it until the catalog's ladder for that app changes; the page says the step was tried and put back. `audits/night-2026-09-24/C/31-night.json` | **READY — P2; owner: CC (controller, with part 7)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-684** | **[P2-MEDIUM] demo-hp's whole-box backups cannot fit on its backup storage — 9201 has had no new whole-box backup since 2026-09-23.** `local` is the host ROOT disk (40 GB, 84 % used, ~4 GB free); it holds three 9201 archives of 5.8–6.9 GB (retention 3) and a new one needs ~7 GB: the 2026-09-24 runs failed `No space left on device` (06:59, 07:42), leaving the stale lock of R-673. Every night now fails the same way until space is made or the retention/target changes. A Tier-0 box, but the same shape exists on any appliance whose local backup target is its root disk. **Needs a decision** (fewer kept archives, another target, or a larger disk). `audits/r672-2026-09-24/C5-r673-vzdump-logs.txt` | **OPEN — P2; owner: operator (the decision) / CC** |
|
||||
| **R-685** | **[P2-MEDIUM] A box whose whole-box backup cannot fit must say so BEFORE the night — an operator event and a line on the backup page — never only a nightly failure a log shows.** The product lesson of R-684 (operator ruling 2026-09-24 evening, option A for demo-hp): demo-hp's root-disk target held three 6–7 GB archives with ~4 GB free and failed `No space left on device` every night from 2026-09-23 while nothing but the vzdump log said why. Measured 2026-09-24 night: a 9201 archive is 8.18 GB (22.6 GB uncompressed); PVE prunes AFTER a successful backup, so a target must hold keep-last + 1 archives during the run — lowering the retention alone does not un-stick a full target. **Fix direction:** the agent predicts the archive size (last archive × margin) against the target's free space before a whole-box backup, skips with a reason reported to the hub (operator event), and the controller's backup page shows the sentence. `audits/night-2026-09-25/A/A4-hp-backup-space.txt` | **READY — P2; owner: CC (agent + controller)** |
|
||||
| **R-686** | **[P3-LOW] The automatic update leg is not resumed after a controller restart during the night — the apps it had not reached wait a whole day.** MEASURED 2026-09-24 night on 9202 (v0.271.0, night 3): the controller was killed (kill -9) during romm's step; the step itself was put back correctly (pin back, the page says the update was interrupted), but vikunja — next in line, re-tested and ready — was not pressed that night. **Night 4 (power cut during romm's verify) showed a second effect:** the interrupted step was RESUMED after the boot and ended `done` (data read back), but the leg that pressed it was gone, so `last_auto_update` was never written and romm's page carries no „Automatikus frissítés … — sikeres" line for a step the box did take by itself. By design today: the leg is one call of the `offbox-backup` job, and a Daily job that already fired does not fire again. The controller's own self-update now waits for the leg (decision 32), so the common restart cause is excluded; a crash or a power cut still ends the night's leg. **Fix direction (needs a ruling):** persist "leg started, not finished, window W" and resume at boot while before W+5h — or accept one night's delay. `audits/night-2026-09-25/C/night3-kill/` | **OPEN — P3; owner: CC (operator ruling on resume)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent) — it is Part D's to show on the demo boxes. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user