night 2026-09-24: findings doc, decisions 29-30, STATUS/CONTEXT, register 335->344 (7 closed), teardown evidence, floor 0.269.1
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -806,13 +806,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-654** | **[P3-LOW] opengist 1.15 moved every page under `/-/` — a household's `/login` bookmark answers 404 after the update.** MEASURED 2026-09-23 night: 1.13 serves `/login`, `/register`, `/all`; 1.15.2 answers **404** on all three and serves `/-/login`, `/-/register`, `/-/all`; `/` redirects to `/-/all`. The app, its data and its probe (`/healthcheck`) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie `Secure`. **Needs:** a line in opengist's `app_info` if the operator wants households told; nothing in the product. Evidence: `apps/opengist-oldfixture/`, `apps/opengist/`. | **READY — P3; owner: operator (copy decision) / CC (writes it)** |
|
||||
| **R-655** | **[P2-MEDIUM] adventurelog v0.13.0 cannot become healthy in the catalog's template — its update is undone on every box.** MEASURED 2026-09-23 night, both venues. v0.13.0's FRONTEND image adds its own `HEALTHCHECK` (`node -e fetch('http://127.0.0.1:3000/health')`), and `/health` answers **503 `{"ok":false,"backend":"unreachable"}`** unless the backend's `/health/` answers OK to the frontend's own request (read from the image's `django-proxy` chunk: `fetch(${getServerEndpoint()}/health/)`). On the bench the backend served `/api/` 200, the seed read back, the migration ran (6 lines) — and the frontend stayed `unhealthy` for 420 s; on 9202 the guarded Update went `verifying` → **`undoing` → `undone`** (under the drill's 90 s timeout). **MEASURED LATER THE SAME NIGHT — two causes, and the first hypothesis was wrong.** (1) The backend's `/health/` answers `200 {"ok": true}` to the frontend (read on a fresh v0.13.0 install); the unhealthy frontend is **our template's own healthcheck override**, `["CMD", "/nodejs/bin/node", …]` — v0.12.1's distroless image keeps node there, v0.13.0 moved it to `/usr/bin/node`, and Docker's health log reads `exec: "/nodejs/bin/node": no such file or directory` 19 times in a row. (2) With the override removed, a second bench run hit a harder wall: v0.13.0's backend runs `download-countries` at EVERY start, fetching world data from the internet, and a cut-short download (`ijson.common.IncompleteJSONError: Incomplete JSON content`) crash-loops the entrypoint — the backend never became healthy in 15 min. The first bench run's download had succeeded. **So v0.13.0's boot depends on an outside download, and whether an update succeeds depends on it too.** **The catalog did NOT move adventurelog.** **Needs:** the override dropped (or pointed at `/usr/bin/node`) in the SAME commit as the image move; and a measured answer on whether `download-countries` can be skipped or pre-seeded (an env switch, or the data in the volume) before the edge is re-proven on both venues. Evidence: `audits/night-2026-09-23/apps/adventurelog/`. | **READY — P2; owner: CC (catalog)** |
|
||||
| **R-657** | **[P2-MEDIUM] "Remove the app, keep my data", then install it again: nextcloud never installs, and the box only says „unhealthy".** MEASURED 2026-09-23 night on 9202 (v0.267.0): nextcloud was removed through the product keeping its drive folder (the remove with data was refused — R-442's fail-closed guard, as on every drill on 9202 — and the product's keep-data remove taken). An hour later a fresh install of nextcloud on the same box: the template binds `${HDD_PATH}/appdata/nextcloud` to `/var/www/html/data`, the kept folder still holds `admin/`, `appdata_*`, `.ncdata` and a 145 MB `nextcloud.log`, and the image's installer loops **„Login is invalid because files already exist for this user — Retrying install..."**; `occ status` reads `installed: false`. The controller records the deploy as done and the app as `unhealthy`; nothing tells the household that their kept files are what blocks the new install, or what to do. **Why it matters:** keep-data is the choice the product OFFERS a household at remove time — and for nextcloud the kept data makes the app uninstallable. **Needs:** decide the product's promise for a reinstall over kept data, per app class (adopt the data? refuse with a sentence? offer to move it aside?); at minimum a deploy-time refusal or warning when the app's drive folder is not empty. Evidence: `audits/night-2026-09-23/chaos/00-nextcloud-reinstall-over-kept-data.txt`. | **READY — P2; owner: operator (the promise) / CC (the build)** |
|
||||
| **R-661** | **[P2-MEDIUM] The second drive holds a file app WHOLE, and no single action brings the app back from it.** MEASURED FROM SOURCE 2026-09-24 (v0.267.0/v0.268.0): the Tier-2 mirror carries the unit AND the drive files, but „Teljes visszaállítás" there is `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`, whose R-538 guard refuses any app with declared drive files; the file restore (`RestoreTier2Files`) only ADDS missing files and replays no database. So for nextcloud/immich/paperless-ngx/calibre-web the second drive is not a way back after a failed update — which is why R-659's hold (v0.268.0) names only the off-site copy for those apps (`backup.WholeOnTier`). **Needs:** a decision on a Tier-2 "files + database" restore (the off-site full restore's shape), or a ruling that the second drive is files-only for these apps. Evidence: `audits/ladder-2026-09-24/README.md` (the truth table); `backup/update_guard.go` R-659 block. | **WAITING-ON-OPERATOR — P2; owner: operator (the promise) / CC (the build)** |
|
||||
| **R-662** | **[P3-LOW] The unit restore's documented second step („csak az adatbázist és a beállításokat", `accept_missing_files=1`) is dead.** MEASURED FROM SOURCE 2026-09-24: no template sends the field (`grep accept_missing_files internal/web/templates` → 0), and when a request does, the web handler skips its own refusal but calls `RestoreFromRecoveryUnit`, which passes `UnitRestoreOptions{}` — so the backup layer refuses anyway. Nothing ever calls `AcceptMissingFiles: true` outside a test. **Needs:** either wire it end to end (with its own confirm, R-48) or remove the flag and the comment that promises it. Evidence: `web/handlers.go` backupRestoreHandler; `backup/restore_unit.go` L226. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-664** | **[P3-LOW] A ladder step has no `.felhom.yml` of its own.** v0.268.0 pins a step's own COMPOSE file (`steps/<key>.yml`); its health probe, memory check and the recorded `applied-meta/` still come from the catalog's CURRENT `.felhom.yml`. Harmless while a step's probe and limits equal the head's (true for all 8 step files today); wrong the day a step's probe differs. Also: the log line for a step pin says „pin advanced to the catalog's current definition" (it names the step file one line earlier). **Needs:** `steps/<key>.felhom.yml` (or the probe in the ladder entry) and the log line naming the source. Evidence: `audits/ladder-2026-09-24/partD/10-romm-two-steps.log`. | **READY — P3; owner: CC (controller + catalog)** |
|
||||
| **R-665** | **[P3-LOW] An update pressed soon after a restore was judged with the RESTORED `.felhom.yml`.** OBSERVED 2026-09-24 on 9202 (v0.268.0), not diagnosed: vikunja restored from its unit at 06:19; the drill catalog's `.felhom.yml` named probe port 8999 (a deliberately failing edge); the periodic probe used 8999 until 06:20:41, and the update pressed at 06:21:07 probed 3456 (the unit's own file) and ended `done`. The restore apparently writes the unit's `.felhom.yml` into the stack dir, and the next catalog sync has not yet put the catalog's back. Consequence bounded (the OLD probe, ≤ 15 min), but an update's verdict should not depend on how long ago a restore ran. **Needs:** a measurement of what the restore writes and when the sync overwrites it. Evidence: `audits/ladder-2026-09-24/partA/03-why-g-was-done.txt`. | **READY — P3; owner: CC (controller)** |
|
||||
| **R-666** | **[P3-LOW] A held app whose box has no whole copy is told „ne törölje az alkalmazást" — and its card still offers Eltávolítás.** v0.268.0 (R-659) hides the Mentések button beside the no-whole-copy sentence; the Remove button on the stacks card stays (the household may remove its own app — a promise, not CC's to take away). Measured live on 9202: the page and the sentence disagree. **Needs:** an operator word — hide Remove while support is informed, or soften the sentence. Evidence: `audits/ladder-2026-09-24/partB/01-round11-reproduced.json`. | **WAITING-ON-OPERATOR — P3; owner: operator** |
|
||||
| **R-667** | **[P2-MEDIUM] A crash-looping app whose container reads `running` for a moment between restarts never reaches the crash-loop alarm.** MEASURED 2026-09-24 on 9202 (v0.268.0): `gokapi` (R-644) at **385 restarts**, `restarting_since` = 06:46:30Z at 06:47 — the clock resets whenever a scan catches the container up — while the dead-app heartbeat said *„4 deployed app(s) evaluated, 0 currently down"*. `Stack.CrashLooping` needs `crashLoopAfter` (5 min) of CONTINUOUS `restarting`. Likely also the unattributed second `app_start_failed` at 06:31:41Z (a scan that caught gokapi `exited`): the dropped-event line names no app, so this is inference. **Needs:** crash-loop judged on the container's RestartCount growth over a window, not on an uninterrupted state; a test with a flapping fixture. Evidence: `audits/ladder-2026-09-24/partB/04-states-and-banner.txt`, `02-r660-positive-observables.txt`. | **READY — P2; owner: CC (controller)** |
|
||||
| **R-668** | **[P2-MEDIUM] The Tier-2 copy chose a registered storage path that no longer existed — on the SAME disk as the app.** MEASURED 2026-09-24 on 9202 (v0.268.0): `sameDevice` failed OPEN when it could not `stat` a path, so a removed folder read as "another disk" and a file app's "second drive" copy landed beside its own data. **FIXED v0.269.0:** the check fails CLOSED (an unreadable path is the same disk) — `r668_missing_path_test.go`, red-proofed; LIVE on 9202 the next copy went to the SSD (dev 1793 vs 66304). Residual, recorded not fixed: a Tier-2 record written on the same disk BEFORE the fix still counts as a copy until the next Tier-2 run replaces it. `audits/night-2026-09-24/A1/01-find-mirror.txt`, `redproofs/A1-r668.txt` | **CLOSED 2026-09-24 — v0.269.0, proven live** |
|
||||
| **R-669** | **[P2-MEDIUM] After a failed update ends in a restore, the box keeps the FAILED step's health check as the pinned version's — and the NEXT failed update's undo judges the correct old version with it and holds the app.** PROVEN LIVE 2026-09-24 on 9202 (v0.269.0): a failing step (probe port 8999) → hold → the second drive's whole restore cleared the hold, but `applied-meta/.felhom.yml` still carried port 8999 (written by `advancePinTo` at 10:39:21; no restore path rewrites it), and the stack's `.felhom.yml` came back from the unit, which the sync had already filled with the failing step's file. The next failed update's undo put the right version and the right data back and then called it unhealthy on port 8999 → **a false hold** (`not_started`). Data safe; the app stopped for nothing. **Fix direction:** a restore that recreates the definition also sets the applied record to the `.felhom.yml` of the RESTORED pin (the ladder step's `steps/<key>.felhom.yml` for those refs, or the catalog's when the catalog head equals the pin), never the unit's copy. Pre-existing since v0.263.2; not a v0.269.0 regression. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P2; owner: CC (controller)** |
|
||||
| **R-670** | **[P3-LOW] Every undo (and every step-file verify) logs `[ERROR] .felhom.yml backup block rejected … docker-compose.yml unreadable`.** `LoadMetadata` validates the backup block against a compose file that the pre-update-meta directory (undo.go:528, since v0.263.0) and the scratch dir of `loadMetadataFile` (v0.269.0) never hold. The health check it feeds is unaffected; an operator reading ERROR lines after an undo is misled. Seen 10:24:41Z and 10:46:24Z on 9202. **Fix:** a probe-only loader that skips the backup-block validation, or copy the compose beside it. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-671** | **[P3-LOW] The undo copies kept by a hold survive the hold's clearing by a restore.** MEASURED 2026-09-24 on 9202: three `nextcloud_*.pre-update-20260924T103924Z` volumes (~0.9 GiB) were still present after the whole restore cleared that hold at 10:41:52Z, and a second set joined them 6 minutes later. Nothing names them on a page; on a small disk they are the difference between the next update's copy fitting or not. **Fix direction:** the restore that clears an update hold removes that hold's undo copies (they describe the state the restore just replaced), logged. `audits/night-2026-09-24/A2/11-applied-meta-wrong-probe.txt` (volume list) | **READY — P3; owner: CC (controller)** |
|
||||
@@ -826,6 +819,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-679** | **[P2-MEDIUM] An Update pressed on an app that is already current runs the whole guarded update — backup, pull, restart — and reports `done`.** MEASURED 2026-09-24 on 9202: navidrome at the head was pressed four times by the stale-read caller (R-678); each press made a safety dump and restarted the app (8.5 s of downtime each) and changed nothing. There is no `current` refusal (`UpdatePreflight`, `update.go:366`; `UpdateOrderCurrent` falls through). Harmless by hand, costly for the automatic leg (`09` §6.4.2). **Fix:** the preflight refuses with reason `current` when the pin equals the catalog head and no newer tested digest exists. `audits/night-2026-09-24/C/30-night-run1.json` | **READY — P2; owner: CC (controller)** |
|
||||
| **R-680** | **[P2-MEDIUM] The box does not remember a failed update step — after an undo it offers the same step again.** MEASURED 2026-09-24 on 9202: vikunja's failing step was undone (104 s) and its badge went straight back to „Frissítés elérhető"; only the test caller's own memory stopped a re-press. Decision 15 („a failed step is never pressed again") therefore holds for a person only by their judgement and not at all for the automatic leg. **Fix (part 7 (b), `09` §6.4.2):** record the failed `to` per app; the leg skips it until the catalog's ladder for that app changes; the page says the step was tried and put back. `audits/night-2026-09-24/C/31-night.json` | **READY — P2; owner: CC (controller, with part 7)** |
|
||||
| **R-681** | **[P2-MEDIUM] An install interrupted by a controller restart is lost SILENTLY — no event, no page sentence, and its half-written files stay.** MEASURED 2026-09-24 on 9202 (Part E round 2, v0.269.1): n8n's install pressed at 11:48:42Z ("Deploying stack n8n … checking 1 images"); `systemctl restart docker` 20 s later took the controller down with it. After the restart the box logged NOTHING about n8n, the app read `not_deployed`, and the household's page showed it as never installed — while `app.yaml` (with its generated `N8N_ENCRYPTION_KEY`), `applied-compose.yml` and `applied-meta/` stayed in the stack dir. A fresh install through the product afterwards worked (211 s) and was not blocked by the leftovers. Compare the update path, which journals and RESUMES after the same accident (round 3/4). **Fix direction:** journal the install like the update; on boot either resume it or say on the page (and in an event) that it was interrupted, and clean the half-written files. `audits/night-2026-09-24/E/round-02*` | **READY — P2; owner: CC (controller)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user