audits: demo-hp's first scheduled restore test on the NVMe passed (R-701)
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -812,7 +812,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` **-- 2026-09-28 14:13 CEST:** the first scheduled cycle after the trim was refused again ("needs 31.0 GiB free, has 23.9 GiB"). **Answered: each refusal DOES reach the hub** — `[WARN] host demo-hp-bb76ea restore-test FAILED …` at every host-report (every 15 min). **-- 2026-09-28 evening: CLOSED by option (b) (operator, `09` §3 decision 44).** demo-hp `restore_storage` → `nvme-scratch`; Proxmox first refused it (403 `Datastore.AllocateSpace` — the agent had no grant there, the 03 doc said so); with the operator's word the agent's `FelhomAgentStore` role was granted on `/storage/nvme-scratch` (user + token), and one restore test passed: restored + booted + verified + torn down in 8m46s, NVMe 50.3 → 74.1 GiB used at peak → back, `local-lvm` 58.46 % throughout. `audits/logins-nvme-2026-09-28/C/`. | **CLOSED — option (b), 2026-09-28** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` **-- 2026-09-28 14:13 CEST:** the first scheduled cycle after the trim was refused again ("needs 31.0 GiB free, has 23.9 GiB"). **Answered: each refusal DOES reach the hub** — `[WARN] host demo-hp-bb76ea restore-test FAILED …` at every host-report (every 15 min). **-- 2026-09-28 evening: CLOSED by option (b) (operator, `09` §3 decision 44).** demo-hp `restore_storage` → `nvme-scratch`; Proxmox first refused it (403 `Datastore.AllocateSpace` — the agent had no grant there, the 03 doc said so); with the operator's word the agent's `FelhomAgentStore` role was granted on `/storage/nvme-scratch` (user + token), and one restore test passed: restored + booted + verified + torn down in 8m46s, NVMe 50.3 → 74.1 GiB used at peak → back, `local-lvm` 58.46 % throughout. **The first SCHEDULED cycle (22:25) passed too:** preflight on nvme-scratch, restored, `scheduled restore-test passed` in 516 s (`C/C5-scheduled-cycle.txt`). `audits/logins-nvme-2026-09-28/C/`. | **CLOSED — option (b), 2026-09-28** |
|
||||
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. **-- 2026-09-28 evening: claper FIXED (catalog `9dc8a05`, controller v0.279.0 `after_install`)** — on a fresh install the seeded password is replaced by a generated one shown on the app page; measured on 9202: the default no longer authenticates, the generated one does, a wrong one does not, a restore keeps it. **The widened question (every app, operator ruling = decision 45) continues in R-707.** | **CLOSED — claper fixed 2026-09-28; the class → R-707** |
|
||||
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. **-- 2026-09-28 later: FIXED in the catalog (`9555e73`, alone in its commit): memory 768M → 1536M.** Measured on 9202 through the drill catalog: at 2048M a 12-minute watch read `anon` steady ~780 MiB, 0 kills; at 1536M, sampled every 2 s from the container's birth, `anon` peaked at **817 MiB (53 %)** during start, `memory.peak` 1075 MiB, 0 kills, healthy. 1024M would sit at the 80 % `memory_tight` line. The seed route works at the new limit (`Calcom` fixture, catalog `b35fc7f`). `…/box/R703-0*.txt` | **CLOSED — catalog `9555e73`, 2026-09-28** |
|
||||
| **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (`unhealthy_stop`, `crash_loop`, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused `409 held` with the crash-loop sentence („…újra és újra összeomlott…"), and `GET /api/stacks/calcom` carried the old `hold_reason` while `state=running`. Start lifted it (`the unhealthy stop is LIFTED by Start`). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. **Fix direction:** the remove clears the app's box-set holds, as `DeleteAppBackupPrefs` clears its backup preferences (R-474). Evidence: `audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt`, `…/box/calcom/move.txt`. **-- 2026-09-28 later: SECOND and worse instance, then FIXED in controller v0.278.0.** demo-hp's fresh nextcloud (installed 10:13) carried an UPDATE hold from a nextcloud of 2026-09-13 (set before v0.242.0 made removals clear update holds; nothing ever swept it). At the manual off-site run (15:17) the backup leg logged `Skipping volume dump for nextcloud — the app is HELD stopped`, captured no unit, and pushed a snapshot that `carried NO database dump and NO volume tar` — a freshly installed app silently NOT backed up. **Fix:** a removal also clears the crash-loop stop (`settings.ClearUpdateHold`), and a new install (plain or "use my kept data") drops a leftover update/crash-loop hold of an app that is not installed (`Router.dropLeftoverHold`); restore holds (R-379) untouched. Red-proofed RP4–RP6 (`audits/kept-offsite-2026-09-28/redproofs/`). Floor 0.278.0. **STILL OPEN: live proof of the install-time drop** (a box with a leftover hold on an uninstalled app). | **WATCHING — P2; owner: CC (install-time drop, live)** |
|
||||
|
||||
Reference in New Issue
Block a user