From 5ba1702fcf81f57b2ea76456906cb5e64aa116fa Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 7 Oct 2026 09:30:55 +0200 Subject: [PATCH] Read-back done, releases delivered (hub 0.141.0, agent 0.150.0, controller 0.302.0, catalog 872039d), R-892 proven on the Tester 1 box; 7 rows closed (137 -> 130) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 6 ++ REPORT.md | 94 +++++-------------- STATUS.md | 18 +++- .../night-burndown-2026-10-06/NIGHT-LOG.md | 10 ++ .../s3/box/C1-repoint-drill.txt | 2 + .../s3/box/T0-leftovers.txt | 24 +++++ .../s3/box/T1-repoint-live.txt | 4 + .../s3/box/T2-drill-reset.txt | 1 + .../s3/box/T3-app-list-after.txt | 10 ++ .../vaultwarden/box-verdict-vaultwarden.json | 18 ++++ .../s3/box/vaultwarden/step.txt | 39 ++++++++ .../s3/tools/repoint.py | 37 ++++++++ .../s3/tools/vwstep.py | 87 +++++++++++++++++ .../delivery/bundle-0.150.0-readback.txt | 15 +++ ...er-0.302.0-and-agent-0.150.0-delivered.txt | 4 + .../delivery/floors-0.302.0.txt | 7 ++ .../delivery/hub-0.141.0-deploy.txt | 3 + .../delivery/sign-agent-update-0.150.0.txt | 13 +++ .../delivery/sign-bundle-0.150.0.txt | 10 ++ .../delivery/vouch-agent-0.150.0.txt | 4 + documentation/backlog/CLOSED-ITEMS.md | 14 +++ documentation/backlog/OPEN-ITEMS.md | 17 +--- 22 files changed, 350 insertions(+), 87 deletions(-) create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/C1-repoint-drill.txt create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/T0-leftovers.txt create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/T1-repoint-live.txt create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/T2-drill-reset.txt create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/T3-app-list-after.txt create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/box-verdict-vaultwarden.json create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/step.txt create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/tools/repoint.py create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/tools/vwstep.py create mode 100644 documentation/audits/readback-2026-10-07/delivery/bundle-0.150.0-readback.txt create mode 100644 documentation/audits/readback-2026-10-07/delivery/controller-0.302.0-and-agent-0.150.0-delivered.txt create mode 100644 documentation/audits/readback-2026-10-07/delivery/floors-0.302.0.txt create mode 100644 documentation/audits/readback-2026-10-07/delivery/hub-0.141.0-deploy.txt create mode 100644 documentation/audits/readback-2026-10-07/delivery/sign-agent-update-0.150.0.txt create mode 100644 documentation/audits/readback-2026-10-07/delivery/sign-bundle-0.150.0.txt create mode 100644 documentation/audits/readback-2026-10-07/delivery/vouch-agent-0.150.0.txt diff --git a/CONTEXT.md b/CONTEXT.md index 114d09f6..674b11fc 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,12 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-07 (morning) — read-back B + D, the releases, the Tester 1 proof (`09` §3 160–161).** R-518 read back (night +> stop 91 s on demo-hp, one press 80 s; the page text holds; the two-tier night ~2026-10-08 is left). Released + delivered +> to demo-hp, demo-felhom, Tester 1: hub 0.141.0, agent 0.150.0 (+ bundle, probe 68/68), controller 0.302.0 (floors with +> declared MinAgent 0.131.0; golden 0.301.0, waiver to 2026-10-13), catalog `872039d`. R-892 proven on the Tester 1 box +> (vaultwarden 14.4 s). Rule 11 (helpers get the full fences). Register 137 → 130. + > **2026-10-06 (night) — the second burn-down night (in progress; morning note `audits/night-burndown-2026-10-06/`).** > **Decided by CC unattended — operator may reverse:** `09` §3 decision 159 (R-805: an empty bind after the persistence > exercise is a note in the reasons, the verdict unchanged). Agent main carries R-894 (`74b5eae`, the newest backup per diff --git a/REPORT.md b/REPORT.md index 0986a360..a5c907f8 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,81 +1,33 @@ -# REPORT — night read-back brief, the night-free parts (2026-10-06 evening; Parts B and D follow after 08:30 on 2026-10-07) +# REPORT — the read-back, the releases and the Tester 1 proof (2026-10-07 morning) -The brief said „start after 08:30 on 2026-10-07". The operator then said „start A, C, E, F now". Parts B (the night -read-back) and D (one press, outside 02:00–08:30) wait for tomorrow. **Nothing was delivered to a box tonight**, so the -night of 2026-10-06→07 runs controller v0.301.0 and agent v0.149.0, and Part B reads a clean result. The hub was -released (it is not a box). +Operator instruction 2026-10-07 08:40: Parts B and D of `TASK-night-readback-offsite-gap-tester1-2026-10-07.md`, then +release hub 0.141.0 → agent 0.150.0 → controller → the held catalog branch, deliver to demo-hp, demo-felhom and the +Tester 1 box, then the Tester 1 update test with the operator present; rule 11 in the shared rule file. Recorded as +`09` §3 decisions 160–161. **The task file is not on DooPlex** (searched the disk and git); Parts B and D were taken from +the evening report and the R-518 row. | Part | Result | |---|---| -| **A** — the rulings | **done** — `09` §3 decisions 157 (R-528 option A + the Docker-approval check) and 158 (R-892: VM 341 on the HP box); R-892's "no Proxmox host known" corrected | -| **B** — the night read-back (R-518) | **waits for 08:30 2026-10-07** | -| **C** — demo-hp's off-site copy | **diagnosed: the premise was wrong — no 10-day gap.** Last copy 2026-10-01 20:15Z (ep0's own listing, verify ok); every box within 7 days. Two real findings: an off-site tier read DUE after an agent restart while its storage was unreachable (R-894, filed), and the alarm's operator mail failed and was never retried (**fixed**, hub v0.140.0) | -| **D** — one press, measured | **waits for after 08:30 2026-10-07** | -| **E** — the Tester 1 box in the update test (R-892) | **route built and identity matched; the live proof is BLOCKED** — DooPlex's key is not authorized on VM 341, and the permission check refused fetching its vaulted password (the brief's rule: a refusal stops that item) | -| **F** — the Docker-approval memory-kill check (R-528) | **built**: hub v0.140.0 LIVE (the approval waits for a passing check); the agent's wrapper merged, unreleased (ships as v0.150.0 after Part B); proven by hand on demo-hp's guest | -| small check — wger's 100 % peak | file cache, not a kill (kill counter 0, restarts 0, anon 50.3 %); the memory hint stays | +| **B** — the night read-back (R-518) | demo-hp (9 apps) stop **~91 s** (was 5 min 47 s); demo-felhom (1 app) ~11 s; local tier only, off-site not due; two channels each. `audits/readback-2026-10-07/RESULT-B-D.md` | +| **D** — one press, measured | demo-hp: **80 s** from the press to the last app (per app 39–79 s); the copy finished 4 min later with the apps up; only the local tier ran; the page's „kb. 1–1,5 perc" holds — no text change | +| hub 0.141.0 | built, manifest `011a481a`, Synced/Healthy at HEAD, image 0.141.0, `starting` log, healthz + System 200 at 06:57:42Z (40 s after the sync) | +| agent 0.150.0 | released (sha `a23d1c90…`, bundle `88456b38…`, tag `v0.150.0`), vouched (golden 0.301.0 and MinAgent 0.131.0 unchanged); signed `agent_update` → all three on 0.150.0 by 07:07Z; signed `agent_config_update` → `BUNDLE DONE written=1 same=24 self-check=ok`, probe 68/68 on all three by 07:21Z | +| controller 0.302.0 | MinAgent 0.131.0; floors 0.302.0 with declared MinAgent for demo-hp, demo-felhom, tester-1 → all three `0.302.0 (healthy)` by 07:10Z. Global floor and Tester 2 not touched. Golden waiver issued to 2026-10-13 (the weekly bake's date) | +| catalog | `night-held-2026-10-06` merged (`872039d`); synced on all three boxes (read in their templates) | +| **Tester 1 proof (R-892)** | vaultwarden 1.36.0 → 1.37.4 through the guarded Update: **proven, 14.4 s**, seed read back before and after; removed with its volume; app list equal; pointer restored byte-identical; drill reset. `audits/night-burndown-2026-10-06/s3/box/` | +| rule 11 | „Every helper prompt carries the brief's fences in full" — in all five copies (one md5) | | Rows before | Rows after | Opened | Closed | |---|---|---|---| -| **137** | **138** | **1** (R-894) | **0** (so far) | +| **137** | **130** | **0** | **7** (R-892, R-894, R-542, R-777, R-612, R-805, R-806) | -## Part C — what happened on demo-hp, with times +**Fixed without a row:** none. **One slip:** the catalog merge was pushed in the same command as its unit-test run, +before reading the result (standing rule 1). Read right after: 156 tests, the single known error (`test_pg_conversion` +is a script, not a test module — the same on `main` before the merge); `catalog_gates.py --fast` was green before the push. -- **My last report was wrong**: it said „demo-hp had no successful off-site run in 10 days". I read a log view cut by - `tail -40` that started on 2026-10-04. The agent's full log and ep0's own listing both show a completed off-site copy - on **2026-10-01 20:15Z** (15.2 GB, verify `ok`). With the 7-day cadence the next one is due ~2026-10-08. -- ep0, read only (`C/C2-ep0-snapshot-listing.txt`): demo-hp 2026-10-01T20:15Z, demo-felhom 2026-10-06T04:21Z, tester-1 - 2026-10-04T19:57Z, Tester-2 2026-10-04T16:31Z — every box within 7 days. -- **2026-10-05 04:25Z (06:25 local), demo-hp:** the off-site storage answered *Can't connect to 10.77.0.1:8007*; the - agent had restarted at 02:57Z (04:57 local), 1.5 hours before; its per-tier backup record is in memory only - (`internal/backup/store.go`, R-348), so the due-check fell back to an EMPTY record and read the tier DUE although it - was not (`internal/localapi/server.go` `newestArchiveOn` → unknown → in-memory). The controller asked; vzdump failed - (*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks; whether it did - that night is **not known** (the controller's log was lost to a later restart). → **R-894**. -- **Who was told:** the hub recorded `whole_guest_backup_failed` (error) at 04:27Z; the household was not mailed (by - design: operator only); **the operator's mail FAILED** (Resend: *context deadline exceeded*) and was never retried. - It is the only failed operator mail of 692 since 2026-02-16 (`C/C4-hub-failed-mails.txt`). -- **Fixed** (hub v0.140.0): a failed operator mail is tried again after 1, 5 and 15 minutes; each try is a row; giving - up is an ERROR line. Red-proof `C/red-operator-mail-retry.txt` („the failed mail was never sent again (tries=1)"). -- Not read: the household's backups page on demo-hp (no session tonight). +**Instruction-file edits:** `.claude/rules/unprompted-work.md` §3 item 11 added in all five copies (operator ruling, +decision 160). -## Part E — the Tester 1 box - -- Identity (`E1-identity-match.txt`): VM 341 `night1004-tester1` on demo-hp, at 192.168.0.154 (found by its MAC); - its certificate `CN=felhom.enkicsifelhom.hu`; the agent's own hub report for `tester-1-d70be4` says `host.node=felhom`; - its guest at 192.168.0.101 answers `felhom.enkicsifelhom.hu` with the Felhom login (200) and demo-hp's domain 404. -- Built (catalog `d63ea35`): `box_walk.py` `TARGETS` (9202, 9201, tester-1 via `-J demo-hp`), `BOX_ADMIN_SEED_GUESTS` - adds `("tester-1", "9201")`; tests BoxWalkTargets (red-proved). `operations/nodes.md` has the box. -- **Blocked:** `ssh -J demo-hp root@192.168.0.154` → `Permission denied (publickey,password)`. The VM has no guest agent; - editing its disk needs a VM stop (a reboot — not allowed). I tried to read the hub's code for revealing the vaulted - console password; **the session's permission check refused it**, and I did not try another way. R-892 now asks the - operator. - -## Part F — the memory-kill check - -- Hub v0.140.0 (LIVE 17:34Z): `DockerStatus` needs, per ring-0 box, a passing `oom_check` with the set; failed, errored or - missing blocks. Red-proof `F/red-hub-docker-approval.txt`; the System page hides the button and says why. -- Agent (`acccb66`, unreleased): after a Docker step the wrapper runs a throwaway container from the running - controller's image (`--pull never`, `--network none`, 64 MB cap, one 200 MB block); pass = `OOMKilled=true` AND the - `oom` event; always removed. 12 red-proofs in `F/`. The helper also found 11 wrapper tests that never ran (a - `unittest.main()` mid-file) — moved; all pass. -- **By hand on demo-hp's guest** (`F/F1`, `F/F2`): `OOMKilled=true`, exit 137, 21 → 21 containers, none left. The `oom` - event was MISSING when the events window ended in the same second as the run, and present (`create attach start - oom die`) with the window ending a second later — the wrapper waits 2 s and reads to epoch + 1. -- Until agent v0.150.0 reaches the ring-0 boxes, no Docker set can be approved (none is pending). - -## Instruction-file edits - -None tonight. - -## CI - -felhom.eu `e0bdd52` → 1448 success (and `eed1dbd` 1446, `01c4a5d` 1447); catalog `d63ea35` → 1445; agent `7e82f32` → 1449; -this commit checked after its push. - -## What is left for 2026-10-07 after 08:30 - -Part B (read both demo boxes' night under v0.301.0), Part D (one press on a demo box, measured; the page text), the -agent v0.150.0 release + bundle and its delivery, the controller release for Part D's text, and Part E's live proof if -the operator authorizes the key. Teardown tonight: none needed (no box provisioned; the hub DB copies deleted; the -by-hand check containers removed — 0 left). +**Teardown:** Tester 1 — vaultwarden removed through the product (no container, no volume), the saved +`controller.yaml.pre-r892` shredded, pointer restored identical; demo-hp — the one press made a normal local copy (kept +by retention), nothing else; hub — nothing provisioned. The night report stays in `REPORT-night-burndown-2026-10-06.md`. diff --git a/STATUS.md b/STATUS.md index ea802786..0db843fc 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,10 +2,20 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-07 05:50: hub 0.140.0; demo-hp, demo-felhom and Tester 1 run controller 0.301.0, agent 0.149.0 — -nothing was delivered tonight. The open-items list is at 137. Night: `documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md`. -Day: Parts B and D of the read-back brief after 08:30, then the releases (hub 0.141.0, agent 0.150.0, controller, the -held catalog branch).** +**Updated 2026-10-07 09:35: hub 0.141.0; demo-hp, demo-felhom and Tester 1 run controller 0.302.0, agent 0.150.0. The +open-items list is at 130. Report: `REPORT.md`.** + +## Morning (2026-10-07): read back, released, delivered — and the Tester 1 test passed + +- **The shorter backup stop works.** Last night the apps on demo-hp were down about 1.5 minutes (it was almost + 6 minutes before). One press of „Mentés most" stopped them for about 1 minute. The page tells the truth. +- **Released and delivered to all three boxes:** hub 0.141.0, agent 0.150.0 (with its root files), controller 0.302.0, + and the night's catalog fixes. Tester 2 was not touched (it is off). +- **The update test ran on the Tester 1 box with you present:** vaultwarden updated in 14 seconds, the test account + survived, and the box is back as it was. +- **New rule:** every helper gets the full fence list, word for word. + +**Needs you (none urgent):** the six one-page designs from the night (each has one question). If nothing: they wait. ## Night (2026-10-06 → 07): the second burn-down night — fixes on main, nothing delivered diff --git a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md index 14184c37..512588b9 100644 --- a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md @@ -56,3 +56,13 @@ no reboot. | R-872 | **closed** — the 05:00 run judged Tester 2 on the longer lines (dump missed=1, backup missed=0 correctly) and the mail reached the operator inbox (second channel) | 10 | (this batch) | | R-892 | 07:10 (operator: key added, „continue"): SSH works, app list read; **the walk is blocked again** — no dashboard password for this box on DooPlex. Nothing changed on the box | 6 | (this batch) | | R-892 | 07:41 dashboard password reset through the box's own reset (operator's word; new password only in a 0600 file); the walk was then **refused by the permission check** before any change — box unchanged | 20 | (this batch) | + +## Morning 2026-10-07 (operator present) + +| Act | Result | Commit | +|---|---|---| +| Parts B + D | R-518 read back: night stop 91 s / 11 s; one press 80 s; page text holds | `audits/readback-2026-10-07/RESULT-B-D.md` | +| releases | hub 0.141.0, agent 0.150.0 (+ bundle), controller 0.302.0, catalog merge `872039d` — all delivered to the three boxes | `readback-2026-10-07/delivery/` | +| R-892 | **closed** — vaultwarden 1.36.0 → 1.37.4 on the Tester 1 box, proven in 14.4 s; box back as before | `s3/box/` | +| rows closed on delivery | R-894, R-542, R-777, R-612, R-805, R-806 | — | +| rule 11 | every helper prompt carries the brief's fences in full (five copies) | decision 160 | diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/C1-repoint-drill.txt b/documentation/audits/night-burndown-2026-10-06/s3/box/C1-repoint-drill.txt new file mode 100644 index 00000000..05806162 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/C1-repoint-drill.txt @@ -0,0 +1,2 @@ +24: repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/T0-leftovers.txt b/documentation/audits/night-burndown-2026-10-06/s3/box/T0-leftovers.txt new file mode 100644 index 00000000..d398214f --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/T0-leftovers.txt @@ -0,0 +1,24 @@ +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +total 20 +drwxr-xr-x 2 root root 4096 Oct 7 07:26 . +drwxr-xr-x 63 root root 4096 Oct 4 19:41 .. +-rw-r--r-- 1 root root 7661 Oct 6 11:51 .felhom.yml +-rw-r--r-- 1 root root 3447 Oct 7 07:25 docker-compose.yml diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/T1-repoint-live.txt b/documentation/audits/night-burndown-2026-10-06/s3/box/T1-repoint-live.txt new file mode 100644 index 00000000..387c5dc5 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/T1-repoint-live.txt @@ -0,0 +1,4 @@ +24: repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git +1 +RESTORED-IDENTICAL + diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/T2-drill-reset.txt b/documentation/audits/night-burndown-2026-10-06/s3/box/T2-drill-reset.txt new file mode 100644 index 00000000..dac48395 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/T2-drill-reset.txt @@ -0,0 +1 @@ +drill=872039d live=872039d diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/T3-app-list-after.txt b/documentation/audits/night-burndown-2026-10-06/s3/box/T3-app-list-after.txt new file mode 100644 index 00000000..a40d152e --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/T3-app-list-after.txt @@ -0,0 +1,10 @@ +bookstack Up 5 hours (healthy) +bookstack-db Up 5 hours (healthy) +cloudflared Up 47 hours (healthy) +felhom-controller Up 45 seconds (healthy) +filebrowser Up 47 hours (healthy) +paperless-postgres Up 5 hours (healthy) +paperless-redis Up 5 hours (healthy) +paperless-webserver Up 5 hours (healthy) +privatebin Up 5 hours (healthy) +traefik Up 47 hours diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/box-verdict-vaultwarden.json b/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/box-verdict-vaultwarden.json new file mode 100644 index 00000000..980e4380 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/box-verdict-vaultwarden.json @@ -0,0 +1,18 @@ +{ + "app": "vaultwarden", + "venue": "the Tester 1 box (VM 341, guest 9201; drill catalog), the product's guarded Update; seeded through the admin invite inside the box (R-892)", + "from": { + "vaultwarden": "vaultwarden/server:1.36.0-alpine" + }, + "to": { + "vaultwarden": "vaultwarden/server:1.37.4-alpine" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "duration_s": 14.4, + "final_phase": "done", + "measured_at": "2026-10-07T07:25:12Z", + "evidence": "felhom.eu/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/step.txt" +} \ No newline at end of file diff --git a/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/step.txt b/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/step.txt new file mode 100644 index 00000000..8ef8ee58 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/step.txt @@ -0,0 +1,39 @@ +drill (old pin): 61c5082 DRILL vaultwarden: back to vaultwarden/server:1.36.0-alpine for the Tester 1 box proof (R-892) + vaultwarden: self-registration http=400 (closed by design, R-512 — 400 expected) + vaultwarden: test-box admin sign-in and invite (inside the box) -> ('200', 'yes', '200') + vaultwarden: invited registration http=200 + vaultwarden: token for the seeded account http=200 ok=True +C1 seed reads back BEFORE: True +before: pinned={'vaultwarden': 'vaultwarden/server:1.36.0-alpine'} +drill: 93eaacb DRILL vaultwarden: vaultwarden/server:1.36.0-alpine -> vaultwarden/server:1.37.4-alpine (Tester 1 box proof, R-892) +badge before: {'hu': [{'title': 'Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.', 'text': 'Frissítés elérhető — 1 napja'}], 'en': [{'title': 'A newer version of this app is available. Select the Update button to start it.', 'text': 'Update available — 1 day ago'}]} + phase +0.0s checking | err=None + phase +2.1s backing-up | err=None + phase +4.1s pulling | err=None + phase +8.2s copying | err=None + phase +9.3s verifying | err=None + phase +14.4s done | err=None + vaultwarden: token for the seeded account http=200 ok=True +2026/10/07 07:25:13 [INFO] [stacks] update vaultwarden: accepted — guarded update started +2026/10/07 07:25:13 [INFO] [stacks] update vaultwarden: phase checking +2026/10/07 07:25:13 [INFO] [stacks] update vaultwarden: ladder — the last step (1 of 1) — the catalog's current definition +2026/10/07 07:25:15 [INFO] [stacks] update vaultwarden: no usable copy on any tier — younger than 24h0m0s and not older than this install's deploy (2026-10-07T07:24:33Z) (found: none) — backing up first +2026/10/07 07:25:15 [INFO] [stacks] update vaultwarden: phase backing-up +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: precondition met after the backup — Tier 2 (second drive) copy from 2026-10-07T07:25:15Z +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: phase safety-dump +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: safety dump done (0 file(s)) [] +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: the undo copy will hold 1 named volume(s), 0.3 MiB +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: phase pinning +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/vaultwarden/docker-compose.yml (vaultwarden=vaultwarden/server:1.37.4-alpine) +2026/10/07 07:25:16 [INFO] [stacks] update vaultwarden: phase pulling +2026/10/07 07:25:20 [INFO] [stacks] update vaultwarden: phase copying +2026/10/07 07:25:21 [INFO] [stacks] update vaultwarden: phase copying +2026/10/07 07:25:21 [INFO] [stacks] update vaultwarden: copied vaultwarden_vaultwarden_data → vaultwarden_vaultwarden_data.pre-update-20261007T072521Z in 387ms +2026/10/07 07:25:21 [INFO] [stacks] update vaultwarden: phase starting +2026/10/07 07:25:22 [INFO] [stacks] update vaultwarden: phase verifying +2026/10/07 07:25:27 [INFO] [stacks] update vaultwarden: healthy after 5s (the app's health check passed) +2026/10/07 07:25:27 [INFO] [stacks] update vaultwarden: DONE in 14s + +badge after: {'hu': [{'title': 'Ez az alkalmazás a legfrissebb elérhető változatot futtatja.', 'text': 'Naprakész'}], 'en': [{'title': 'This app is running the newest version available.', 'text': 'Up to date'}]} +RESULT final_phase=done after={'vaultwarden': 'vaultwarden/server:1.37.4-alpine'} seed_after=True verdict=proven (14.4 s) +remove -> 200 diff --git a/documentation/audits/night-burndown-2026-10-06/s3/tools/repoint.py b/documentation/audits/night-burndown-2026-10-06/s3/tools/repoint.py new file mode 100644 index 00000000..d152e4f3 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/tools/repoint.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Point the box (TARGET) at the drill catalog, or put the saved controller.yaml back. `09` §6.5.""" +import io, os, re, sys +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +SAVE = f"{VOL}/controller.yaml.pre-r892" +DRILL = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" +def creds(): + for l in io.open(os.path.expanduser("~/.git-credentials")).read().split("\n"): + m = re.match(r"https://(admin):([^@]+)@gitea\.dooplex\.hu", l) + if m: return m.group(1), m.group(2) + sys.exit("no admin credential") +if sys.argv[1] == "drill": + u, t = creds() + out = w.guest(f"""set -e +test -f {SAVE} || cp -p {VOL}/controller.yaml {SAVE} +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml"; s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml +""") + print(out.replace(t, "")) +elif sys.argv[1] == "restore": + print(w.guest(f"""set -e +cp -p {SAVE} {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml; grep -c 'token: ""' {VOL}/controller.yaml || true +cmp {SAVE} {VOL}/controller.yaml && echo RESTORED-IDENTICAL""")) diff --git a/documentation/audits/night-burndown-2026-10-06/s3/tools/vwstep.py b/documentation/audits/night-burndown-2026-10-06/s3/tools/vwstep.py new file mode 100644 index 00000000..2afb6ac5 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/tools/vwstep.py @@ -0,0 +1,87 @@ +"""vwstep.py — R-892: vaultwarden's step on the Tester 1 box (TARGET=tester-1, drill catalog), ONE process, through the product. + + 1 install vaultwarden fresh at the live pin (this run installs it — the admin seed refuses otherwise); + 2 seed through the household's door, the invite through the admin page INSIDE the box + (FELHOM_BOX_ADMIN_SEED=1, upgrade_fixtures_box.box_admin_seed_allowed); read it back (C1); + 3 a DRILL-only commit moves the image and adds a ladder entry; sync, rescan; + 4 the product's guarded Update; the seed read back; box verdict JSON; + 5 remove through the product. +Evidence: ../box/vaultwarden/step.txt + box-verdict-vaultwarden.json. The live entry is written ONLY by +`upgrade-test.py --write-ladder` from both verdicts.""" +import json, os, re, subprocess, sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +import upgrade_fixtures_box as fixtures + +APP, SUB, SVC = "vaultwarden", "vault", "vaultwarden" +to = sys.argv[1] +HERE = os.path.dirname(os.path.abspath(__file__)) +EVD = os.path.join(HERE, "..", "box", APP); os.makedirs(EVD, exist_ok=True) +log = open(f"{EVD}/step.txt", "a", buffering=1) +D = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" + + +def say(*a): + w.say(*a); log.write(" ".join(map(str, a)) + "\n") + + +fx = fixtures.FIXTURES[APP] +w.login() +if w.stack(APP).get("deployed"): + sys.exit(say("vaultwarden is already installed on this box — STOP (not ours to remove)") or 1) +FROM = sys.argv[2] # the old pin to install first, e.g. vaultwarden/server:1.36.0-alpine +subprocess.run(["git", "-C", D, "pull", "-q", "--rebase", "origin", "main"], check=True) +_c = f"{D}/templates/{APP}/docker-compose.yml"; _s = open(_c).read() +_cur = re.search(r"^\s+image:\s*(\S+)", _s, re.M).group(1) +if _cur != FROM: + open(_c, "w").write(_s.replace("image: " + _cur, "image: " + FROM, 1)) + subprocess.run(["git", "-C", D, "commit", "-q", "-am", f"DRILL {APP}: back to {FROM} for the Tester 1 box proof (R-892)"], check=True) + subprocess.run(["git", "-C", D, "push", "-q", "origin", "main"], check=True, capture_output=True) +say("drill (old pin):", subprocess.run(["git", "-C", D, "log", "--oneline", "-1"], capture_output=True, text=True).stdout.strip()) +w.sync_rescan(APP, FROM) +if not w.deploy(APP, SUB): + sys.exit(say("RESULT the install did not complete") or 1) +tok = fx.seed(w, SUB, say) +if tok is None or not fx.verify(w, SUB, tok, say): + say(f"RESULT C1 failed: {getattr(fx, 'tried', '')}") + w.remove(APP); sys.exit(1) +say("C1 seed reads back BEFORE: True") +before = (w.stack(APP).get("app_config") or {}).get("pinned_images") +say(f"before: pinned={before}") +subprocess.run(["git", "-C", D, "pull", "-q", "--rebase", "origin", "main"], check=True) +comp, fy = f"{D}/templates/{APP}/docker-compose.yml", f"{D}/templates/{APP}/.felhom.yml" +s = open(comp).read() +frm = re.search(r"^\s+image:\s*(\S+)", s, re.M).group(1) +open(comp, "w").write(s.replace("image: " + frm, "image: " + to, 1)) +entry = {"from": {SVC: frm}, "to": {SVC: to}, "verdict": "proven", "tested_at": "DRILL", "harness_version": 5, + "evidence": "DRILL (box proof in progress)", "marks": {"files_may_change": False, "needs_person": None, "memory_tight": False}} +f = open(fy).read() +if to in f and frm in f: + pass +else: + f = (f.rstrip("\n") + "\n - " + json.dumps(entry) + "\n") if "update_ladder:" in f else (f.rstrip("\n") + "\nupdate_ladder:\n - " + json.dumps(entry) + "\n") +open(fy, "w").write(f) +subprocess.run(["git", "-C", D, "commit", "-q", "-am", f"DRILL {APP}: {frm} -> {to} (Tester 1 box proof, R-892)"], check=True) +subprocess.run(["git", "-C", D, "push", "-q", "origin", "main"], check=True, capture_output=True) +say("drill:", subprocess.run(["git", "-C", D, "log", "--oneline", "-1"], capture_output=True, text=True).stdout.strip()) +w.sync_rescan(APP, to) +say(f"badge before: {w.badges(APP)}") +since = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip() +res = w.press_update(APP, poll=1, cap_s=1800) +for p in res.get("phases", []): + log.write(f" phase +{p['t']}s {p['phase']} | err={p['error']}\n") +time.sleep(10) +read = fx.verify(w, SUB, tok, say) +lines = w.guest(f"docker logs --since {since} felhom-controller 2>&1 | grep -E 'update {APP}' | grep -v DEBUG | cut -c1-400") +log.write(lines + "\n") +st = w.stack(APP); after = (st.get("app_config") or {}).get("pinned_images") +verdict = {"app": APP, "venue": "the Tester 1 box (VM 341, guest 9201; drill catalog), the product's guarded Update; seeded through the admin invite inside the box (R-892)", + "from": before, "to": after, + "verdict": "proven" if (res.get("final_phase") == "done" and read and (after or {}).get(SVC) == to) else "failed", + "seed_read_before": True, "seed_read_after": read, "healthy_after": st.get("state") == "running", + "duration_s": res.get("duration_s"), "final_phase": res.get("final_phase"), "measured_at": since, + "evidence": "felhom.eu/documentation/audits/night-burndown-2026-10-06/s3/box/vaultwarden/step.txt"} +json.dump(verdict, open(f"{EVD}/box-verdict-{APP}.json", "w"), indent=2) +say(f"badge after: {w.badges(APP)}") +say(f"RESULT final_phase={res.get('final_phase')} after={after} seed_after={read} verdict={verdict['verdict']} ({res.get('duration_s')} s)") +say(f"remove -> {w.remove(APP)}") diff --git a/documentation/audits/readback-2026-10-07/delivery/bundle-0.150.0-readback.txt b/documentation/audits/readback-2026-10-07/delivery/bundle-0.150.0-readback.txt new file mode 100644 index 00000000..9027b67b --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/bundle-0.150.0-readback.txt @@ -0,0 +1,15 @@ +== hp +Oct 07 09:21:47 demo-hp felhom-agent[4029318]: time=2026-10-07T09:21:47.848+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)" +Oct 07 09:21:47 demo-hp felhom-agent[4029318]: time=2026-10-07T09:21:47.848+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.150.0 written=1 same=24 self-check=ok signers-created=False" +Oct 07 09:21:48 demo-hp felhom-agent[4029318]: time=2026-10-07T09:21:48.765+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded="" +Oct 07 09:21:48 demo-hp felhom-agent[4029318]: time=2026-10-07T09:21:48.765+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=23d9f79da880c635 op=agent_config_update +== felhom-pve +Oct 07 09:21:41 demo-felhom felhom-agent[2298873]: time=2026-10-07T09:21:41.733+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)" +Oct 07 09:21:41 demo-felhom felhom-agent[2298873]: time=2026-10-07T09:21:41.733+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.150.0 written=1 same=24 self-check=ok signers-created=False" +Oct 07 09:21:42 demo-felhom felhom-agent[2298873]: time=2026-10-07T09:21:42.308+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded="" +Oct 07 09:21:42 demo-felhom felhom-agent[2298873]: time=2026-10-07T09:21:42.308+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=883aca6d5a4d973d op=agent_config_update +== root@192.168.0.154 +Oct 07 09:21:46 felhom felhom-agent[3741283]: time=2026-10-07T09:21:46.282+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)" +Oct 07 09:21:46 felhom felhom-agent[3741283]: time=2026-10-07T09:21:46.282+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.150.0 written=1 same=24 self-check=ok signers-created=False" +Oct 07 09:21:47 felhom felhom-agent[3741283]: time=2026-10-07T09:21:47.109+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded="" +Oct 07 09:21:47 felhom felhom-agent[3741283]: time=2026-10-07T09:21:47.109+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=2aad600292ed8add op=agent_config_update diff --git a/documentation/audits/readback-2026-10-07/delivery/controller-0.302.0-and-agent-0.150.0-delivered.txt b/documentation/audits/readback-2026-10-07/delivery/controller-0.302.0-and-agent-0.150.0-delivered.txt new file mode 100644 index 00000000..9e818042 --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/controller-0.302.0-and-agent-0.150.0-delivered.txt @@ -0,0 +1,4 @@ +2026-10-07T07:22:13Z hp guest 9201: gitea.dooplex.hu/admin/felhom-controller:0.302.0 Up 12 minutes (healthy); agent felhom-agent 0.150.0 +2026-10-07T07:22:15Z felhom-pve guest 9201: gitea.dooplex.hu/admin/felhom-controller:0.302.0 Up 12 minutes (healthy); agent felhom-agent 0.150.0 +2026-10-07T07:22:16Z root@192.168.0.154 guest 9201: gitea.dooplex.hu/admin/felhom-controller:0.302.0 Up 12 minutes (healthy); agent felhom-agent 0.150.0 +Tester 2: not touched (off) diff --git a/documentation/audits/readback-2026-10-07/delivery/floors-0.302.0.txt b/documentation/audits/readback-2026-10-07/delivery/floors-0.302.0.txt new file mode 100644 index 00000000..d23cad6d --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/floors-0.302.0.txt @@ -0,0 +1,7 @@ +== floors 2026-10-07T07:09:19Z: 0.302.0 with min_agent 0.131.0 (golden stays 0.301.0); global floor not touched; Tester 2 not touched +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/07 09:09:19 [INFO] Customer demo-hp controller-version floor override set to "0.302.0" (declared MinAgent "0.131.0") +2026/10/07 09:09:20 [INFO] Customer demo-felhom controller-version floor override set to "0.302.0" (declared MinAgent "0.131.0") +2026/10/07 09:09:20 [INFO] Customer tester-1 controller-version floor override set to "0.302.0" (declared MinAgent "0.131.0") diff --git a/documentation/audits/readback-2026-10-07/delivery/hub-0.141.0-deploy.txt b/documentation/audits/readback-2026-10-07/delivery/hub-0.141.0-deploy.txt new file mode 100644 index 00000000..3b400db7 --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/hub-0.141.0-deploy.txt @@ -0,0 +1,3 @@ +2026-10-07T06:57:48Z +hub 0.141.0: sync=Synced health=Healthy rev=011a481a325ab6cec3f7cfadf7d39858837067ef image=gitea.dooplex.hu/admin/felhom-hub:0.141.0 +2026/10/07 08:57:09 [INFO] felhom-hub 0.141.0 starting; sync 06:57:02Z, pod ready 06:57:35Z, healthz 200 + /system 200 at 06:57:42Z diff --git a/documentation/audits/readback-2026-10-07/delivery/sign-agent-update-0.150.0.txt b/documentation/audits/readback-2026-10-07/delivery/sign-agent-update-0.150.0.txt new file mode 100644 index 00000000..d87112f3 --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/sign-agent-update-0.150.0.txt @@ -0,0 +1,13 @@ +== agent_update 0.150.0 (sha a23d1c90…) signed with felhom-op-1, ttl 45m, 2026-10-07T07:00:25Z +-- demo-hp-bb76ea +signed: op=agent_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=e32907d89af4721e514c265fd9e070f5 expires=2026-10-07T07:45:25Z +{"op_blob_b64":"eyJleHBpcmVzX2F0IjoiMjAyNi0xMC0wN1QwNzo0NToyNVoiLCJpc3N1ZWRfYXQiOiIyMDI2LTEwLTA3VDA3OjAwOjI1WiIsImtleV9pZCI6ImZlbGhvbS1vcC0xIiwibm9uY2UiOiJlMzI5MDdkODlhZjQ3MjFlNTE0YzI2NWZkOWUwNzBmNSIsIm9wIjoiYWdlbnRfdXBkYXRlIiwicGFyYW1zIjp7InNoYTI1NiI6ImEyM2QxYzkwODViYzdmZDRmYzQ4ZmUwMzI3ZjY1MDRhYTgzZTYzMzE1MTBmYjRhM2QyM2UwNDJkZGRiMjZmOWMiLCJ2ZXJzaW9uIjoiMC4xNTAuMCJ9LCJ0YXJnZXQiOnsiZ3Vlc3RfaWQiOiIiLCJob3N0X2lkIjoiZGVtby1ocC1iYjc2ZWEifX0=","sig_armored":"-----BEGIN SSH SIGNATURE-----\nU1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgvzPSoI2ADfHbHEAHRCPmujzBoW\nMZnxNr+xYHre2UvD4AAAAMZmVsaG9tLW9wLXYxAAAAAAAAAAZzaGE1MTIAAABTAAAAC3Nz\naC1lZDI1NTE5AAAAQPaXiF/S1nw49tYDUIdJmvknW/ZvWbaP8ZjBsXB40tN7TYCuDLrsCH\n4ic3xaDV4BZa09MVpYE3YnEevC3by60QM=\n-----END SSH SIGNATURE-----\n"} +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=0c838937b7174329c68effdd91d8aaaf expires=2026-10-07T07:45:25Z +{"op_blob_b64":"eyJleHBpcmVzX2F0IjoiMjAyNi0xMC0wN1QwNzo0NToyNVoiLCJpc3N1ZWRfYXQiOiIyMDI2LTEwLTA3VDA3OjAwOjI1WiIsImtleV9pZCI6ImZlbGhvbS1vcC0xIiwibm9uY2UiOiIwYzgzODkzN2I3MTc0MzI5YzY4ZWZmZGQ5MWQ4YWFhZiIsIm9wIjoiYWdlbnRfdXBkYXRlIiwicGFyYW1zIjp7InNoYTI1NiI6ImEyM2QxYzkwODViYzdmZDRmYzQ4ZmUwMzI3ZjY1MDRhYTgzZTYzMzE1MTBmYjRhM2QyM2UwNDJkZGRiMjZmOWMiLCJ2ZXJzaW9uIjoiMC4xNTAuMCJ9LCJ0YXJnZXQiOnsiZ3Vlc3RfaWQiOiIiLCJob3N0X2lkIjoiZGVtby1mZWxob20tODM2M2I1In19","sig_armored":"-----BEGIN SSH SIGNATURE-----\nU1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgvzPSoI2ADfHbHEAHRCPmujzBoW\nMZnxNr+xYHre2UvD4AAAAMZmVsaG9tLW9wLXYxAAAAAAAAAAZzaGE1MTIAAABTAAAAC3Nz\naC1lZDI1NTE5AAAAQApxXKRWy1eB8LFZgOSZNXf/1qTNV2abKQ4bS/JgTW7v4B/s+0Ythq\n6PSEuRlK1Mh/fJPua0Cq0PIXQfkfXa0Qo=\n-----END SSH SIGNATURE-----\n"} +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +signed: op=agent_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=7dfc6a5ca403c26cdf9801e478348b3f expires=2026-10-07T07:45:25Z +{"op_blob_b64":"eyJleHBpcmVzX2F0IjoiMjAyNi0xMC0wN1QwNzo0NToyNVoiLCJpc3N1ZWRfYXQiOiIyMDI2LTEwLTA3VDA3OjAwOjI1WiIsImtleV9pZCI6ImZlbGhvbS1vcC0xIiwibm9uY2UiOiI3ZGZjNmE1Y2E0MDNjMjZjZGY5ODAxZTQ3ODM0OGIzZiIsIm9wIjoiYWdlbnRfdXBkYXRlIiwicGFyYW1zIjp7InNoYTI1NiI6ImEyM2QxYzkwODViYzdmZDRmYzQ4ZmUwMzI3ZjY1MDRhYTgzZTYzMzE1MTBmYjRhM2QyM2UwNDJkZGRiMjZmOWMiLCJ2ZXJzaW9uIjoiMC4xNTAuMCJ9LCJ0YXJnZXQiOnsiZ3Vlc3RfaWQiOiIiLCJob3N0X2lkIjoidGVzdGVyLTEtZDcwYmU0In19","sig_armored":"-----BEGIN SSH SIGNATURE-----\nU1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgvzPSoI2ADfHbHEAHRCPmujzBoW\nMZnxNr+xYHre2UvD4AAAAMZmVsaG9tLW9wLXYxAAAAAAAAAAZzaGE1MTIAAABTAAAAC3Nz\naC1lZDI1NTE5AAAAQM2jDDmhZ80fYh3+ye6eeaAtZkb+8UfJaBOEru4/bnUjND1bD0P6oH\n5Ngnh4N2vjj8BWDZ+JLaNvQ/RHKZRaqAU=\n-----END SSH SIGNATURE-----\n"} +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/readback-2026-10-07/delivery/sign-bundle-0.150.0.txt b/documentation/audits/readback-2026-10-07/delivery/sign-bundle-0.150.0.txt new file mode 100644 index 00000000..388a7508 --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/sign-bundle-0.150.0.txt @@ -0,0 +1,10 @@ +== agent_config_update 0.150.0 (bundle 88456b38…) signed with felhom-op-1, ttl 45m, 2026-10-07T07:09:11Z +-- demo-hp-bb76ea +signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=172c387e5c67a55df5c757e40a3e2141 expires=2026-10-07T07:54:11Z +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=4152bdbf80d35e572bbd5fe526ea273f expires=2026-10-07T07:54:11Z +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +signed: op=agent_config_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=f55e84c49d56ba88dca3711e933b2527 expires=2026-10-07T07:54:11Z +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/readback-2026-10-07/delivery/vouch-agent-0.150.0.txt b/documentation/audits/readback-2026-10-07/delivery/vouch-agent-0.150.0.txt new file mode 100644 index 00000000..e090cbc7 --- /dev/null +++ b/documentation/audits/readback-2026-10-07/delivery/vouch-agent-0.150.0.txt @@ -0,0 +1,4 @@ +== vouch 2026-10-07T06:59:41Z: agent 0.150.0, golden 0.301.0 (unchanged), min_agent 0.131.0 (unchanged) +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +2026/10/07 09:00:00 [INFO] Artifact manifest set: agent=0.150.0 golden=0.301.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="88456b386d9b1027bd22861cac8c23df004bf9fd9f67644d6595bfca8c94498e" diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 067b4b64..13b88065 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,20 @@ --- +## 2026-10-07 (morning) — the read-back, the releases, the Tester 1 proof + +The full text of every row below: `git show 011a481a32:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-612** | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** (P3) | CLOSED 2026-10-07 — DELIVERED (catalog 872039d) | Wishlist's healthcheck needs the seed (bench old rc 0 / new rc 1); `cat/R-612-red.txt`. | +| **R-894** | **After an agent restart, an UNREADABLE off-site storage makes the off-site tier look DUE, so the box asks for a copy that cannot be made.** (P3) | CLOSED 2026-10-07 — DELIVERED (agent v0.150.0) | Agent 0.150.0 on demo-hp, demo-felhom and Tester 1 (signed `agent_update` 07:06Z, bundle 07:21Z, capability probe 68/68). Built `74b5eae`, four red-proofs `audits/night-burndown-2026-10-06/s4/`; `07` §6.1. Not yet seen live: a restart followed by an unreadable storage (it needs an outage). | +| **R-542** | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** (P3) | CLOSED 2026-10-07 — DELIVERED (controller v0.302.0) | `/api/disks/candidates` drops drives backing a registered path; two red-proofs `audits/night-burndown-2026-10-06/ctrl/R-542-red.txt`; 0.302.0 healthy on the three boxes. | +| **R-777** | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** (P2) | CLOSED 2026-10-07 — DELIVERED (catalog 872039d) | Measured on 9202 (remote-access-OFF user signed in from the internet: 200 → 403 after; LAN 200); Jellyfin KnownProxies + Emby LocalNetworkSubnets seeded on a fresh volume or an empty list; synced to all three boxes (KnownProxies present in their jellyfin template). `audits/night-burndown-2026-10-06/r777/RESULT.md` | +| **R-892** | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** (P4) | CLOSED 2026-10-07 — PROVEN LIVE | The update test ran on the Tester 1 box (TARGET=tester-1, operator present): vaultwarden 1.36.0-alpine → 1.37.4-alpine through the product's guarded Update in 14.4 s (checking → backing-up → pulling → copying → verifying → done), the seeded account read back before and after, badge „Naprakész" after; app removed with its volume; app list equal to before; catalog pointer restored byte-identical; drill reset to live. Dashboard password reset through the box's own reset on the operator's word (kept in a 0600 file on DooPlex). `audits/night-burndown-2026-10-06/s3/box/` | +| **R-805** | **[P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount.** (P4) | CLOSED 2026-10-07 — DELIVERED (catalog 872039d) | An empty bind is a note in the reasons, verdict unchanged (`09` §3 decision 159, CC unattended — operator may reverse); `cat/R-805-red.txt`. | +| **R-806** | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** (P4) | CLOSED 2026-10-07 — DELIVERED (catalog 872039d) | gramps-web runs 2 gunicorn workers (8 filled 1 GB); synced to the three boxes; `cat/R-806-red.txt`. | + ## 2026-10-06 (night) — the second burn-down night The full text of every row below: `git show e866a66b56:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 48c739bd..17a24f82 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -122,12 +122,11 @@ stopping line that lies. | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | -## Apps & catalog — 10 rows (P3 4, P4 6) +## Apps & catalog — 9 rows (P3 3, P4 6) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** **2026-10-06 night: not started — the row needs the operator's word on the Hungarian format (decimal comma, date style) before any code.** | — | — | CC | -| **R-612** | Apps & catalog | P3 | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** **— FIXED 2026-09-23 night (catalog `a5a729a`): 512M.** Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (`The seed command has been executed`), and wishlist then moved v0.66.0 → v0.67.1 with its test record. **Still open:** a failed first-boot seed is invisible to the box — nothing reads the seed's exit. | **READY — P3, narrowed to making a failed seed visible; owner: CC (catalog)** **Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains.** **2026-10-06 night: fixed on catalog branch `night-held-2026-10-06`** (`d4150d7`) — the healthcheck needs the role rows and a group or user (node:sqlite, read-only); bench: seedless DB old check rc 0, new rc 1; seeded both rc 0 (`audits/night-burndown-2026-10-06/cat/R-612-red.txt`). Reaches main after the read-back. | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** **2026-10-06: NARROWED — the files are served.** Catalog `cf1ed43`: `DJANGO_DEBUG=False` + a `wger-files` nginx serving `/static` and `/media`, a definition step proven on the bench (harness v5) and on 9202 (photo 404 → 200, CSS 404 → 200). **What remains is R-755's merged half:** wger still runs Django's `runserver`; upstream's gunicorn runs 3 workers, which do not fit wger's 384 MB — a memory decision and a new proof. Known cost of the fix: the collected static files are 283 MB in a named volume, in every backup of wger. Ready to show? Its files and sign-up are fixed and proven; the server question is open — `audits/design-build-2026-10-06/`F/. **2026-10-06 night: the server question MEASURED on bench 9401** (wger 2.7, template 384M, `wger-files` in front; `audits/night-burndown-2026-10-06/r762/RESULT.md`): runserver anon 201 MB (52 %), gunicorn 1 worker 165 MB (43 %), 2 workers 157 MB (41 % — `--preload` shares pages); oom_kill 0 and restarts 0 in all three; login page and its CSS 200 every time; 20 GETs, 10 at once: avg 394 / 241 / 124 ms. The worker count is gunicorn's `WEB_CONCURRENCY` (the image passes no `-w`). Pick the numbers support: `WGER_USE_GUNICORN=True` + `WEB_CONCURRENCY=2`. Not measured: a signed-in page. Not built: the template change. | — | Decide the gunicorn worker count against the memory limit, prove it on both venues; then the operator decides whether wger is shown. If nothing: wger stays hidden | CC | | **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: not a catalog fix** — the setgid chain is set by the controller's FileBrowser setup; a controller row. | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | @@ -146,7 +145,7 @@ stopping line that lies. | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | -## Backup & restore — 34 rows (P2 8, P3 12, P4 14) +## Backup & restore — 33 rows (P2 8, P3 11, P4 14) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -157,7 +156,6 @@ stopping line that lies. | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** **2026-10-06 night: the wrong verdict was already gone** (agent v0.138.0, R-727 skips another key's archives); demo-hp's August archives are pruned by ep0's keep-last 2 (inferred, ep0 fenced). **NEW, the real gap:** the hub retained an escrow only on a restic-password change, so a reinstall's new backup key K overwrote the only copy of the old one — **fixed on felhom.eu main (hub, unreleased; red-proved `audits/night-burndown-2026-10-06/r366/`)**. Design `audits/night-burndown-2026-10-06/design-R-366.md` (pick B: retain + tell the operator; C, key continuity, is the operator's). Not closed: the operator signal (archives made with another key) is the design's second slice. | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC | | **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | -| **R-894** | Backup & restore | P3 | **After an agent restart, an UNREADABLE off-site storage makes the off-site tier look DUE, so the box asks for a copy that cannot be made.** MEASURED 2026-10-05 on demo-hp, read 2026-10-06 (`audits/readback-2026-10-07/C/`): the last off-site copy was 2026-10-01 20:15Z (ep0's own listing, verify `ok`), so the 7-day tier was NOT due; the agent had restarted at 04:57 local; at 06:25 `GET …/storage/felhom-pbs/content` answered 500 *Can't connect to 10.77.0.1:8007*; `newestArchiveOn` returned `unknown` and fell back to the in-memory record (`internal/localapi/server.go`), which a restart empties (`internal/backup/store.go` — memory only, R-348); so the tier read DUE, the controller requested it, and vzdump failed (*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks (`07` §6.4); whether it did that night is NOT KNOWN (the controller's log was lost to a later restart). The hub got `whole_guest_backup_failed` (error); its operator mail then failed (fixed in the hub this session: a failed operator mail is retried). The code's rule is deliberate: an unreadable storage must not suppress a backup. | **OPEN — filed 2026-10-06** **2026-10-06 night: fixed on agent main `74b5eae` (the row's first option — the newest success per tier on disk, read only when the storage cannot be read; `07` §6.1); ships with agent v0.150.0, closes when delivered.** | a design: keep the newest success per tier on disk (as `RestoreTestState` does) so the fallback is the last known copy, not "never"; or report an unreachable storage so no app is stopped for it | Decide the fallback; build it in the agent with a test that restarts the agent and then cannot read the storage; measure once | CC | | **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-822.md`) — the past-dated residual can shrink real history to the last 7 days in ~2 windows, unalarmed, because the hub's window check trusts the box's own counts (`hub/internal/offsitekeys/service.go:284,343`); the same box-trusted count already sits under decision 68. Pick: accept + close (operator's word); the hub's own snapshot count is the new row R-895. | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | | **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | @@ -185,23 +183,21 @@ stopping line that lies. | **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator | | **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** **2026-10-05 (burn-down night): NEEDS A LIVE MEASUREMENT** — a timed catch-up on 9202 for an app with a multi-GB volume, then the skip-or-warn choice. | — | measure a large volume first | CC | -## Storage & devices — 7 rows (P3 5, P4 2) +## Storage & devices — 6 rows (P3 4, P4 2) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-298** | Storage & devices | P3 | **The `/storage` page's unregistered list is filtered by `role==='user-data'`, so a drive that is also the backup target can never be registered from it.** `storage.html:363` routes anything not `user-data` into the read-only protected group with NO actions. On the rebuilt `demo-hp` the NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`/etc/pve/storage.cfg`: `dir: felhom-backup` → `/mnt/nvme-1tb`), so it renders locked. **This is the SECOND reason that page was empty** during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action | **READY (S) — NEW 2026-08-10** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Registering a drive the agent classes as the backup target conflicts with the agent's eject/decommission rule (403 for that role). First: is a drive holding both app data and backups a supported layout? | R-280 | Split the role gate: `user-data` keeps destructive actions; any mounted role may be REGISTERED | CC | | **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** **2026-10-06 night: the wire half BUILT on main (unreleased):** the agent sends SMART 187 `reported_uncorrect`, 188 `command_timeout` (raw, vendor-packed on some drives — carried as reported), 199 `udma_crc_errors` (pointer + omitempty: unknown is absent, never 0); the controller decodes them and the hub's host-report mirror models them (G-1 wire gate green; it convicted the agent half alone). **Carried only** — no verdict, banner, mail or alarm reads them (`TestR330_CountersChangeNoVerdictYet`); using them changes what a household is told and is the next slice, with persisting samples. Red-proofs `audits/night-burndown-2026-10-06/r330/`. Ships with agent v0.150.0 and the next controller and hub releases. | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC | | **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC | -| **R-542** | Storage & devices | P3 | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** **2026-10-06 night: fixed on controller main `f65ace0`** — `/api/disks/candidates` drops every drive backing a registered storage path from initialize and attach (joined through the guest mount table; an unreadable table empties initialize); `TestR542_*`, red-proofs `audits/night-burndown-2026-10-06/ctrl/R-542-red.txt`. Ships with the next controller release; closes when delivered. | — | — | CC | | **R-756** | Storage & devices | P3 | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). **Checked from source 2026-10-05 (burn-down round 2):** Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names .../scratch_hdd/use | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS A LIVE READING.** The refusal shows that app's recorded HDD_PATH is a sub-folder of the drive (the R-839 shape); loosening the check would weaken the boot start gate. Next: that app's HDD_PATH and `findmnt` inside 9202. | — | — | CC | | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -## Security & access — 15 rows (P2 2, P3 11, P4 2) +## Security & access — 14 rows (P2 1, P3 11, P4 2) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** **2026-10-06 night: MEASURED on 9202 through the simulated tunnel, and fixed on the held catalog branch `night-held-2026-10-06` (`e5a5984`).** A remote-access-OFF user signed in from the internet on both apps (200): Jellyfin saw traefik `172.18.0.5`, Emby saw cloudflared `172.16.253.2`. Fix: Jellyfin `KnownProxies 172.16.0.0/12` in `network.xml`; Emby `LocalNetworkSubnets 10.0.0.0/8, 192.168.0.0/16` in `system.xml`; both seeded at start, only on a fresh volume or into an empty list. After: tunnel 403 (Jellyfin logs the visitor `198.51.100.66`), forged XFF 403, LAN household 200. **The row's „no setting fixes Emby" was wrong** (read in source, never measured) — no operator route needed. Cost: a 172.16–31.x home LAN counts as remote. Closes when merged to main and synced. `audits/night-burndown-2026-10-06/r777/RESULT.md` | — | — | CC + operator | | **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | @@ -277,13 +273,12 @@ stopping line that lies. | **R-89** | Business & legal | P4 | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 22 rows (P3 2, P4 20) +## Process & tooling — 19 rows (P3 2, P4 17) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** **2026-10-06 night: NARROWED** — the harness records each container's `swap_peak` and the venue's swap (catalog branch `night-held-2026-10-06`, `3f4611c`; SwapRecorded red-proved; bench SwapTotal 0 kB, 9202 524288 kB). LEFT: the golden's swap in its bake evidence, the box walk's record, and whether proofs run with swap off. | — | — | CC | | **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. **NIGHT WATCH 2026-10-06/07 (second burn-down night): 0 jobs lost** — every push checked by its commit (felhom.eu, controller, agent, catalog main and the held branch), each completed `success`, one filtered call per check. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | -| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** **2026-10-06 (night): the route is BUILT in the box walk** (catalog `d63ea35`: `TARGET=tester-1`, `ssh -J demo-hp root@192.168.0.154`, guest 9201, `felhom.enkicsifelhom.hu`; `BOX_ADMIN_SEED_GUESTS` has it) and the identity is matched (the agent's report `host.node=felhom` = VM 341's certificate; the guest answers its domain 200, demo-hp's 404). **Blocked:** DooPlex's key is not authorized on VM 341 (`Permission denied (publickey,password)`); the VM has no guest agent and a disk edit needs a VM stop (a reboot, not allowed); fetching its vaulted password from the hub was refused by the session's permission check. `audits/readback-2026-10-07/` **2026-10-06 20:27 (night brief §3): still blocked** — the operator ran `ssh-copy-id` at 20:20, but it copied the key of the operator's own SSH agent; this shell has no agent, and DooPlex's two keys (`id_rsa`, `id_ed25519`) are still refused, direct and through demo-hp (`audits/night-burndown-2026-10-06/s3/`). The fix: run `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` from DooPlex as kisfenyo. **2026-10-07 07:10: SSH works now** (the operator added the key; `hostname` → `felhom`; the guest app list read, `audits/night-burndown-2026-10-06/s3/s3-tester1-readonly.txt`). **Still blocked:** the walk signs in to the box's dashboard, and DooPlex holds no dashboard password for this box (the claim of 2026-10-04 kept none; the four old `.ctlpw` files predate the install). Needs: the operator puts the box's dashboard password in a 0600 file for `SC=… TARGET=tester-1`. **2026-10-07 07:20: SSH and the dashboard password are in place** (operator key; on the operator's word the dashboard password was reset through the box's own reset — code mailed to tester1@, read in the mailbox — and is kept only in `/home/kisfenyo/.felhom-tester1/.ctlpw`, 0600). **The walk did not run:** the session's permission check refused it before any change — nothing installed, the catalog pointer untouched, the app list unchanged (`audits/night-burndown-2026-10-06/s3/`). Needs: the operator allows the walk, or a day session runs it (`SC=/home/kisfenyo/.felhom-tester1 TARGET=tester-1`). | DooPlex's key on VM 341 | Operator: authorize DooPlex's public key on VM 341 (one line in `/root/.ssh/authorized_keys`, from its console), or allow CC to read the vaulted password; then prove one step there | operator | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | @@ -299,8 +294,6 @@ stopping line that lies. | **R-739** | Process & tooling | P4 | **[P3-LOW] The test bench cannot run wanderer at all: its web server calls the database at the PUBLIC name `https://.`, which the bench has no name or TLS for.** MEASURED 2026-09-30 (more-night-apps brief, Part B): on bench 9401 the catalog template (v0.20.0) came up `wanderer-db` healthy, `wanderer-search` healthy, `wanderer` **unhealthy** for 12 min, every page 500 („Error 0: Something went wrong"); `PUBLIC_POCKETBASE_URL=https://hike-db.gate.invalid`. So the harness can only ever answer `inconclusive` at FROM for wanderer, never about an update. Its step `v0.20.0 → v0.21.0` (web + db) and meilisearch `v1.36 → v1.54` stay untested; **the brief's question — does the search index survive meilisearch's move or is it rebuilt — is NOT measured.** **Needs:** a bench venue that gives the stack the DB name (an `extra_hosts` + plain-http override in the bench's render only, stated in the verdict), or wanderer proven on the box venue alone with the operator's word. `audits/more-night-apps-2026-09-30/B/wanderer-probe.txt` **-- 2026-09-30 late: the bench CAN run wanderer now** — `upgrade-test.py` `BENCH_ENV_OVERRIDES` points `PUBLIC_POCKETBASE_URL` at the database container on the bench only (the only address the web image reads, measured in v0.20.0 and v0.21.0), recorded in every verdict: web 200, sign-up (`PUT /api/v1/user`) 200, login 200. **The meilisearch question, answered:** v1.54.2 REFUSES v1.36's database („Your database version (1.36.0) is incompatible”); with `MEILI_UPGRADE_DB=true` it upgrades in place and all three indexes (actors, lists, trails) are there — so wanderer's step needs that switch in the template first. Not done: a fixture (creating a list answered PocketBase's „Failed to create record” — likely a rule for unverified users), whether a household's record survives in the index, the step itself. `audits/night-rulings-2026-09-30/` | **NARROWED — the bench runs it; the step needs a fixture and `MEILI_UPGRADE_DB`; owner: CC** **Re-ranked 2026-10-03: P3→P4: test bench coverage; the template switch and fixture are tooling work.** | — | — | CC | | **R-759** | Process & tooling | P4 | **[P3-LOW] wger's onboarding record (the checklist pilot) keeps rows open that no other row owns.** 2026-10-01, `app-catalog-felhom.eu/onboarding/wger.md`: **2.5** no backup → remove → restore → read back of wger exists — the box has no per-app backup press outside an Update (R-648) and wger has no newer step to carry one; **3.7** changing the password and adding a family member not measured, and the template has no `add_people` text; **6.3** no forced-fail undo for wger; **8.2** the app page not read on 9202 this session; **9.1** the runtime volume-persistence gate not re-run (last CLEAN 2026-08-02). The other open rows have their own: 1.5 (R-755), 1.7/2.8 (R-762), 3.4 (R-763), 7.1 (R-764). wger is exempt from the onboarding gate (published before the checklist), so nothing blocks; this row is what keeps the record honest. **Needs:** the five measured on 9202 — 2.5 and 6.3 ride wger's next ladder step (the update's backing-up phase is the per-app backup). `audits/new-app-checklist-2026-10-01/` | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: record-keeping for a hidden app; nothing blocks.** | — | — | CC | | **R-786** | Process & tooling | P4 | **[P3-LOW] SparkyFitness's onboarding record has six open rows** (`app-catalog-felhom.eu/onboarding/sparkyfitness.md`): 0.5 runtime internet (food search providers), 0.7 the phone app's sign-in route through traefik, 1.6 the env names the server reads, 1.7 the entrypoint read, 5.4 a second memory watch at another limit, 8.3 no logo/screenshots on felhom.eu (404). Everything else measured this session (bench + 9202). **Needs:** each row measured, or n/a with a reason. | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: onboarding record completeness; no household meets it directly.** | — | — | CC | -| **R-805** | Process & tooling | P4 | **[P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): Grimmory's `/app/data` bind held 0 files and read CLEAN (its state is in MariaDB); komga, paperless-ngx, radarr and sonarr also carry empty binds that are not judged. **Needs:** decide whether an empty bind after the exercise is UNDETERMINED like a named volume, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate rule; no household meets it.** **2026-10-06 night: fixed on catalog branch `night-held-2026-10-06`** (`816c557`) — an empty bind is a NOTE in the reasons, the verdict unchanged (`09` §3 decision 159, decided by CC unattended — operator may reverse); `TestEmptyBind` red-proved (`cat/R-805-red.txt`). Closes when the branch reaches main. | — | — | CC | -| **R-806** | Process & tooling | P4 | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): both reached data only through their fixture seed. **Needs:** read the `loadbalancer.server.scheme` label (upgrade_boxport already does); find why gramps-web's first GET got no answer. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate defect; no household meets it.** **NARROWED 2026-10-05 (app-catalog `4828dc7`): the harness GET now uses the traefik scheme label (https + -k for crafty-controller). LEFT: gramps-web :5000 not answering — a live look on a box.** **2026-10-06 night: the :5000 silence is measured and fixed on catalog branch `night-held-2026-10-06`** (`1859903`): the image's default 8 gunicorn workers filled the 1024M limit (anon at limit, oom_kill 3, 4/5 GETs timed out); `GUNICORN_NUM_WORKERS=2` → first answer 10 s, ~2 ms, anon 342 MB, 0 kills (`cat/R-806-red.txt`). Closes when delivered. | — | — | CC | | **R-807** | Process & tooling | P4 | **[P3-LOW] 13 apps stay UNDETERMINED because their seeds write only to the database, so their upload / media / cache / redis volumes stay empty.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): claper, crafty-controller, dawarich (redis), docmost, gramps-web, immich (ML cache), outline (+redis), sparkyfitness, tandoor, vikunja, wger, wishlist, zipline — plus plex (no seed route: a plex.tv claim token) and wanderer (no fixture; unhealthy under the gate). None is BROKEN: nothing was written outside a preserved folder. **Needs:** per app, a seed that uploads one file (or the volume named n/a with a reason: a redis cache, an ML model cache). | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test coverage; nothing was found broken.** | — | — | CC | | **R-896** | Process & tooling | P4 | **The catalog's `volume-persistence` gate runs its canary and the apps on the LOCAL Docker, so a full gate run on DooPlex acts on the production host — and there its self-test fails for every app.** SEEN 2026-10-06 night: a helper's full `catalog_gates.py` run on DooPlex built `felhom-volgate-canary:1` and ran the canary pair locally (`scripts/check-volume-persistence.py` canary self-test, ~line 830); it reported the self-test failing for every app, also on `main` (checked with gokapi). Not re-run (it would act on DooPlex again). Same class as R-650 (tests that reach real Docker act on DooPlex). | **OPEN — filed 2026-10-06 night** | — | Point the gate at the bench guest (as the upgrade harness does) or refuse to run on a host without a marker; then find why the self-test fails | CC |