docs: R-691 (2) built (07 §6.5), R-701 (a) measured, R-702 claper default admin, R-703 calcom OOM; evidence
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -808,11 +808,13 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
|
||||
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
|
||||
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** | **OPEN — P3; owner: CC — NARROWED to (2)** |
|
||||
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). | **WATCHING — P3; owner: CC (live off-site proof)** |
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. | **OPEN — P3; owner: operator (which option), CC measures** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
|
||||
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. | **OPEN — P1; owner: operator (fix shape), CC implements** |
|
||||
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. | **OPEN — P2; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user