R-649 closed: a failed install removes what it started (operator ruling); floor 0.266.0
gates / gates (push) Successful in 29s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 18:14:38 +02:00
parent 210394ec6e
commit 899ae18b95
18 changed files with 2324 additions and 5 deletions
+1
View File
@@ -357,4 +357,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-636** | **Six hours of OOM kills sent the same single warning as one hiccup (P2).** v0.265.0 + hub v0.121.0: the kernel `oom_kill` counter (the `OOMKilled` flag is sticky and cannot count); ≥ 20 kills in 30 min of one container run → ONE `app_oom_storm` (error, operator-only, per-app cooldown). | **CLOSED 2026-09-23 — PROVEN-LIVE on the controller (`45-*`, `46-*`: RomM at 320M, storm at 21 kills, still one at 49); hub side by unit tests** | as above |
| **R-647** | **Three leftovers of the update mail (P3).** v0.265.0: a held update's error is the key `update.error.held`, rendered per reader on both pages and the API; `copy_holds` travels as its key; the two log wordings fixed. | **CLOSED 2026-09-23 — PROVEN-LIVE for (1) (`43-*`, `44-*`); (2)(3) by red-proofed tests** | as above |
| **R-648** | **The drill's „Mentés most" was whole-box (P3).** No per-app backup endpoint exists; the harness (`audits/cleanup-2026-09-23/walk.py` `backup_now`) now presses nothing and the guarded update's own `backing-up` phase backs up the throwaway app alone. | **CLOSED 2026-09-23 — PROVEN-LIVE (`43-*`: phase `backing-up` for vikunja only)** | as above |
| **R-649** | **A failed install could leave containers under „not deployed" (P2).** Operator ruling 2026-09-23 (option a): the controller cleans up. Closed in **v0.266.0** (`964ae75`): `runComposeDeploy`'s failure branch runs `compose down` (volumes kept) before the record reads not-deployed. | **CLOSED 2026-09-23 — PROVEN-LIVE (`audits/r649-2026-09-23/`: outline with a never-healthy redis — 0 containers left, 3 volumes kept)** | `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
-1
View File
@@ -806,7 +806,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** |
| **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** |
| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** |
| **R-649** | **[P2-MEDIUM] OPERATOR QUESTION: when `docker compose up` fails for the deploy's OWN reasons and leaves containers behind, what should the record say?** Found while diagnosing R-634 (2026-09-23). The backup race is fixed (v0.265.0), but `runComposeDeploy`'s failure branch (`stacks/deploy.go:421-435`) still writes `Deployed=false` without asking whether containers exist — e.g. a dependency's healthcheck that times out after its containers started. Since v0.262.0 the household can REMOVE such an app (`halfStateEvidence`), so nothing is stranded; but the page says „not installed" over containers that may be running. **Options:** (a) the failed deploy runs `compose down` on what it started, so „failed" means nothing runs — cost: a slow-but-healthy start is thrown away, and the household must press Deploy again; (b) keep the containers and record the deploy as FAILED-WITH-CONTAINERS, a state the page shows with a Remove button and the compose error in both languages — cost: a new state every surface must learn; (c) leave it (today). **Recommendation: (a)** — it keeps the R-634 promise (never containers under „not deployed") with the least new surface, and a deploy that failed is re-pressed anyway. Not taken unattended: it changes what a failed deploy does to what the household started (rule 1). | **OPEN — P2; owner: operator (decision), then CC** |
| **R-650** | **[P3-LOW] A controller unit test that falls through to the REAL docker acts on the build host — DooPlex.** MEASURED 2026-09-23: my own first draft of `TestR634_VolumeLegNeverStopsADeployingApp` drove the real `DumpAppVolumesSafe` for its control case, and `docker run … tar` CREATED an empty volume `outline_outline_data` on DooPlex (removed by name, verified empty and unused; nothing else touched). The same draft of the stacks test ran a real `docker compose down` from a temp dir named `nextcloud` (no such project existed). Both tests now use seams (`dumpVolumesSafe`; `composeCmd` + an empty PATH). **The class is unguarded:** any test that reaches `exec.Command("docker", …)` on DooPlex acts on production Docker. Fix shape: a test-binary guard in `composeExec`/`execCommand` that refuses a real docker exec when running under `go test` unless an explicit env opt-in is set, plus a decoy test that proves it refuses. | **OPEN — P3; owner: CC (controller)** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.