night 2026-09-24: Part C spike (09 §6.4.2 build brief), demo-hp pool incident evidence, rows R-668..R-680
gates / gates (push) Successful in 41s
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -812,6 +812,19 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-665** | **[P3-LOW] An update pressed soon after a restore was judged with the RESTORED `.felhom.yml`.** OBSERVED 2026-09-24 on 9202 (v0.268.0), not diagnosed: vikunja restored from its unit at 06:19; the drill catalog's `.felhom.yml` named probe port 8999 (a deliberately failing edge); the periodic probe used 8999 until 06:20:41, and the update pressed at 06:21:07 probed 3456 (the unit's own file) and ended `done`. The restore apparently writes the unit's `.felhom.yml` into the stack dir, and the next catalog sync has not yet put the catalog's back. Consequence bounded (the OLD probe, ≤ 15 min), but an update's verdict should not depend on how long ago a restore ran. **Needs:** a measurement of what the restore writes and when the sync overwrites it. Evidence: `audits/ladder-2026-09-24/partA/03-why-g-was-done.txt`. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-666** | **[P3-LOW] A held app whose box has no whole copy is told „ne törölje az alkalmazást" — and its card still offers Eltávolítás.** v0.268.0 (R-659) hides the Mentések button beside the no-whole-copy sentence; the Remove button on the stacks card stays (the household may remove its own app — a promise, not CC's to take away). Measured live on 9202: the page and the sentence disagree. **Needs:** an operator word — hide Remove while support is informed, or soften the sentence. Evidence: `audits/ladder-2026-09-24/partB/01-round11-reproduced.json`. | **WAITING-ON-OPERATOR — P3; owner: operator** |
|
||||
| **R-667** | **[P2-MEDIUM] A crash-looping app whose container reads `running` for a moment between restarts never reaches the crash-loop alarm.** MEASURED 2026-09-24 on 9202 (v0.268.0): `gokapi` (R-644) at **385 restarts**, `restarting_since` = 06:46:30Z at 06:47 — the clock resets whenever a scan catches the container up — while the dead-app heartbeat said *„4 deployed app(s) evaluated, 0 currently down"*. `Stack.CrashLooping` needs `crashLoopAfter` (5 min) of CONTINUOUS `restarting`. Likely also the unattributed second `app_start_failed` at 06:31:41Z (a scan that caught gokapi `exited`): the dropped-event line names no app, so this is inference. **Needs:** crash-loop judged on the container's RestartCount growth over a window, not on an uninterrupted state; a test with a flapping fixture. Evidence: `audits/ladder-2026-09-24/partB/04-states-and-banner.txt`, `02-r660-positive-observables.txt`. | **READY — P2; owner: CC (controller)** |
|
||||
| **R-668** | **[P2-MEDIUM] The Tier-2 copy chose a registered storage path that no longer existed — on the SAME disk as the app.** MEASURED 2026-09-24 on 9202 (v0.268.0): `sameDevice` failed OPEN when it could not `stat` a path, so a removed folder read as "another disk" and a file app's "second drive" copy landed beside its own data. **FIXED v0.269.0:** the check fails CLOSED (an unreadable path is the same disk) — `r668_missing_path_test.go`, red-proofed; LIVE on 9202 the next copy went to the SSD (dev 1793 vs 66304). Residual, recorded not fixed: a Tier-2 record written on the same disk BEFORE the fix still counts as a copy until the next Tier-2 run replaces it. `audits/night-2026-09-24/A1/01-find-mirror.txt`, `redproofs/A1-r668.txt` | **CLOSED 2026-09-24 — v0.269.0, proven live** |
|
||||
| **R-669** | **[P2-MEDIUM] After a failed update ends in a restore, the box keeps the FAILED step's health check as the pinned version's — and the NEXT failed update's undo judges the correct old version with it and holds the app.** PROVEN LIVE 2026-09-24 on 9202 (v0.269.0): a failing step (probe port 8999) → hold → the second drive's whole restore cleared the hold, but `applied-meta/.felhom.yml` still carried port 8999 (written by `advancePinTo` at 10:39:21; no restore path rewrites it), and the stack's `.felhom.yml` came back from the unit, which the sync had already filled with the failing step's file. The next failed update's undo put the right version and the right data back and then called it unhealthy on port 8999 → **a false hold** (`not_started`). Data safe; the app stopped for nothing. **Fix direction:** a restore that recreates the definition also sets the applied record to the `.felhom.yml` of the RESTORED pin (the ladder step's `steps/<key>.felhom.yml` for those refs, or the catalog's when the catalog head equals the pin), never the unit's copy. Pre-existing since v0.263.2; not a v0.269.0 regression. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P2; owner: CC (controller)** |
|
||||
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
|
||||
| **R-672** | **[P1-HIGH] The agent's SCHEDULED restore-test filled the production thin pool and turned a customer guest's disks read-only.** FOUND 2026-09-24 on demo-hp (agent v0.132.0): at 10:29 CEST the restore-test restored 9201's 22 GiB archive as scratch guest 990000 into `local-lvm` — the SAME pool as 9201 — with no free-space check. The pool reached 100 % at 10:35:24 (`out_of_data_space`, `error_if_no_space`); the scratch failed to start; its teardown failed (`lvremove … contains a filesystem in use`) and was "left for Recover" — **which runs only at agent start**, so the pool stayed full for 2.5 h. The agent logged *„a full pool corrupts every guest on it"* every few seconds and took no action. **Consequence, measured:** 9201's controller got `no space left on device` from 10:38, its log stops at 10:39:57, and at 13:12 both its rootfs and `/var/lib/felhom` were READ-ONLY (errors=remount-ro). **Intervention:** an agent restart ran Recover ("destroyed leaked restore-test scratch guest", pool 100 % → 58.9 %). **Still open:** 9201 needs a stop + fsck + start (operator — the session's permission check refused host-level guest operations). **Fix direction:** the restore-test refuses to start unless the target storage has the archive's size plus a margin free and never targets the pool of the guest it tests; a failed teardown retries on a timer, not only at start; the pool-fill WARN pages the operator. `audits/night-2026-09-24/C-02…C-07` | **OPEN — P1; owner: CC (agent) / operator (9201 repair)** |
|
||||
| **R-673** | **[P2-MEDIUM] 9201's whole-box backups failed all morning on a stale `snapshot-delete` lock.** demo-hp 2026-09-24: vzdump of 9201 failed at 06:59, 07:17, 07:42 and 08:22 CEST (`CT is locked (snapshot-delete)`), leaving `snap_vm-9201-disk-1_vzdump` (07:42) behind, before R-672's pool fill. The agent's stale-lock scanner ran every 30 s and did not clear it; the lock was gone after the 13:09 agent restart. Cause not read — the session's permission check refused reading the task logs. `audits/night-2026-09-24/C-01-demo-hp-9201-vzdump-errors.txt` | **OPEN — P2; owner: CC (agent)** |
|
||||
| **R-674** | **[P3-LOW] The ladder log says an app's pin „matches no update_ladder entry … older than the ladder" when the pin EQUALS the head.** Seen on 9202 2026-09-24 for nextcloud at the head. Misleading to an operator reading why nothing climbed. **Fix:** say "at the head" when the pin equals the newest `to`. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-677** | **[P3-LOW] For a floating tag re-tested at a new digest, the Behind badge's age reads the TAG's catalog date („1 napja"), not when the new digest was tested (minutes).** Seen 2026-09-24 on 9202 (Part B). Harmless but confusing. **Fix:** for a digest-only move, age from the ladder entry's `tested_at`. `audits/night-2026-09-24/B/10-floating-tag.json` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-678** | **[P3-LOW] After an update step ends `done`, the app's „steps left" and its badge stay STALE until the next scan.** MEASURED 2026-09-24 on 9202 (Part C, first caller run): wishlist read `ladder_steps_left` 1 after its only step, navidrome 2 after both of its steps — for ~50 s, through six presses. A person sees „Frissítés elérhető" for an app that just updated; an automatic caller re-presses. **Fix:** the update's finish refreshes the app's catalog/ladder fields (the same read the scan does). `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-679** | **[P2-MEDIUM] An Update pressed on an app that is already current runs the whole guarded update — backup, pull, restart — and reports `done`.** MEASURED 2026-09-24 on 9202: navidrome at the head was pressed four times by the stale-read caller (R-678); each press made a safety dump and restarted the app (8.5 s of downtime each) and changed nothing. There is no `current` refusal (`UpdatePreflight`, `update.go:366`; `UpdateOrderCurrent` falls through). Harmless by hand, costly for the automatic leg (`09` §6.4.2). **Fix:** the preflight refuses with reason `current` when the pin equals the catalog head and no newer tested digest exists. `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P2; owner: CC (controller)** |
|
||||
| **R-680** | **[P2-MEDIUM] The box does not remember a failed update step — after an undo it offers the same step again.** MEASURED 2026-09-24 on 9202: vikunja's failing step was undone (104 s) and its badge went straight back to „Frissítés elérhető"; only the test caller's own memory stopped a re-press. Decision 15 („a failed step is never pressed again") therefore holds for a person only by their judgement and not at all for the automatic leg. **Fix (part 7 (b), `09` §6.4.2):** record the failed `to` per app; the leg skips it until the catalog's ladder for that app changes; the page says the step was tried and put back. `audits/night-2026-09-24/C/31-night.json` | **READY — P2; owner: CC (controller, with part 7)** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user