STATUS, report, register (R-691 closed live, R-704..R-706), Part E + D2 evidence
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -806,16 +806,18 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-676** | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-682** | **[P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held).** MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household `502 Bad Gateway`; after the restart `chaoscrash` read deployed, stopped, `unhealthy_stop`, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. **Fix direction:** the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. `audits/night-2026-09-24/E/round-09*.json`, `E/round-09b-remove-again.txt` | **READY — P3; owner: CC (controller)** |
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). | **WATCHING — P3; owner: CC (live off-site proof)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). **-- 2026-09-28 afternoon: LIVE-PROVEN on demo-hp 9201 (controller 0.278.0), endpoint level.** A throwaway nextcloud, seeded through its own front door, joined the off-site copy; the off-site run-now pushed snapshot `6cb379a8`. (a) The full off-site restore (prepare → download → reconstitute): 3 volumes + the database replayed, the seed read back, a marker user written after the snapshot read ABSENT, same versions. (b) Remove keeping the data + deleting the local copies → reinstall: the choice and the kept list both named "távoli mentés, 2026-09-28 15:40" / "the off-site copy, 2026-09-28 15:40"; "use my kept data" downloaded the unit alone and loaded 3/3 volumes + 1/1 database in 55 s; the seed and the kept files read back. App removed with its data; demo-hp's app list equals the list before. `audits/kept-offsite-2026-09-28/E/`. | **CLOSED — controller v0.277.0, live 2026-09-28** |
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` **-- 2026-09-28 14:13 CEST:** the first scheduled cycle after the trim was refused again ("needs 31.0 GiB free, has 23.9 GiB"). **Answered: each refusal DOES reach the hub** — `[WARN] host demo-hp-bb76ea restore-test FAILED …` at every host-report (every 15 min). | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
|
||||
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. | **OPEN — P1; owner: operator (fix shape), CC implements** |
|
||||
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. **-- 2026-09-28 later: FIXED in the catalog (`9555e73`, alone in its commit): memory 768M → 1536M.** Measured on 9202 through the drill catalog: at 2048M a 12-minute watch read `anon` steady ~780 MiB, 0 kills; at 1536M, sampled every 2 s from the container's birth, `anon` peaked at **817 MiB (53 %)** during start, `memory.peak` 1075 MiB, 0 kills, healthy. 1024M would sit at the 80 % `memory_tight` line. The seed route works at the new limit (`Calcom` fixture, catalog `b35fc7f`). `…/box/R703-0*.txt` | **CLOSED — catalog `9555e73`, 2026-09-28** |
|
||||
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. | **OPEN — P3; owner: CC** |
|
||||
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. **-- 2026-09-28 later: SECOND and worse instance, then FIXED in controller v0.278.0.** demo-hp's fresh nextcloud (installed 10:13) carried an UPDATE hold from a nextcloud of 2026-09-13 (set before v0.242.0 made removals clear update holds; nothing ever swept it). At the manual off-site run (15:17) the backup leg logged `Skipping volume dump for nextcloud — the app is HELD stopped`, captured no unit, and pushed a snapshot that `carried NO database dump and NO volume tar` — a freshly installed app silently NOT backed up. **Fix:** a removal also clears the crash-loop stop (`settings.ClearUpdateHold`), and a new install (plain or "use my kept data") drops a leftover update/crash-loop hold of an app that is not installed (`Router.dropLeftoverHold`); restore holds (R-379) untouched. Red-proofed RP4–RP6 (`audits/kept-offsite-2026-09-28/redproofs/`). Floor 0.278.0. **STILL OPEN: live proof of the install-time drop** (a box with a leftover hold on an uninstalled app). | **WATCHING — P2; owner: CC (install-time drop, live)** |
|
||||
| **R-705** | **[P3-LOW] There is no way to run the night's chain now — only its pieces.** Asked by the operator 2026-09-28 (to finish a proof in the day). What exists (read from source, controller v0.278.0): the backup page's off-site run-now (`POST /backup/offbox/run`) runs the dump leg first (the R-44 pre-phase: DB dumps, volume dumps with brief app stops, unit capture) and then the push — used live on demo-hp 2026-09-28 15:17, 3m57s; the debug API has `backup/dbdump`, `backup/crossdrive` (Tier 2), `backup/integrity`, `backup/offsite-proof`. **Missing:** the automatic update leg (`RunUpdateLeg`, chained only to the scheduled off-site job) and the whole-guest backup (the agent's, on its own 24 h / 7 d cadence) have no manual trigger; the only way to run the chain in order is to move the backup window (`POST /backups/window`), which takes W..W+2h at least. **Needs:** a debug action "run tonight's chain now" (dump → Tier 2 → off-site → update leg, in order, one at a time), and an agent-side "whole-guest backup now" for demo boxes. Not built. | **OPEN — P3; owner: CC** |
|
||||
| **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** Measured 2026-09-28 on demo-hp: after a full off-site restore of nextcloud (which leaves the downloaded copy in `backups/offsite-restore/nextcloud`, ~1 GB, by design, for the household to inspect), `POST /api/stacks/nextcloud/remove` with `remove_backups: true` removed the unit and listed `backup_paths_removed` WITHOUT the verification copy; it stayed until the restore page's own delete (`POST /backup/offbox/verify-copy/delete`, 302 `scratch_deleted`). A household that removes an app to free space keeps 1 GB it cannot see on the app list. **Fix direction:** the removal with backups also deletes the app's verification copy (the same `DeleteOffsiteRestoreCopy`). Evidence: `audits/kept-offsite-2026-09-28/E/E9-teardown.txt`. | **OPEN — P3; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user