From fe998b8847445d12d86e9763bc2a1d2e4fc114dd Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 6 Oct 2026 08:50:08 +0200 Subject: [PATCH] morning after: R-889 delivered (controller v0.299.0, df's percent read back), R-776/R-613 proven on 9202 and live, R-585/R-621 delivered; ten operator questions on one page in STATUS; report (164 -> 150; 0 opened, 14 closed) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 8 ++ REPORT-morning-after-2026-10-06.md | 85 +++++++++++++++++ STATUS.md | 33 +++++-- .../delivery/controller-delivery.txt | 3 + .../delivery/vouch-golden-floors.txt | 12 +++ .../live-9202/README.md | 28 ++++++ .../live-9202/kimai.txt | 16 ++++ .../live-9202/nextcloud.txt | 33 +++++++ .../live-9202/teardown.txt | 91 +++++++++++++++++++ .../live-9202/tools/burst.sh | 10 ++ .../live-9202/tools/km.py | 34 +++++++ .../live-9202/tools/kmlogin.sh | 14 +++ .../live-9202/tools/nc.py | 40 ++++++++ .../live-9202/tools/nc2.py | 21 +++++ .../live-9202/tools/nc3.py | 33 +++++++ .../live-9202/tools/repoint.py | 37 ++++++++ .../live-9202/tools/rm.py | 19 ++++ .../live-9202/tools/tun.sh | 12 +++ .../live-9202/tools/vk.py | 40 ++++++++ .../live-9202/tools/zl.py | 36 ++++++++ .../live-9202/vikunja.txt | 20 ++++ .../live-9202/zipline.txt | 15 +++ .../r889-readback.txt | 8 ++ .../r889-red-proof.txt | 3 + documentation/backlog/CLOSED-ITEMS.md | 12 +++ documentation/backlog/OPEN-ITEMS.md | 7 +- 26 files changed, 658 insertions(+), 12 deletions(-) create mode 100644 REPORT-morning-after-2026-10-06.md create mode 100644 documentation/audits/morning-after-2026-10-06/delivery/controller-delivery.txt create mode 100644 documentation/audits/morning-after-2026-10-06/delivery/vouch-golden-floors.txt create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/README.md create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/kimai.txt create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/nextcloud.txt create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/teardown.txt create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/burst.sh create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/km.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/kmlogin.sh create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/nc.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/nc2.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/nc3.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/repoint.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/rm.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/tun.sh create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/vk.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/tools/zl.py create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/vikunja.txt create mode 100644 documentation/audits/morning-after-2026-10-06/live-9202/zipline.txt create mode 100644 documentation/audits/morning-after-2026-10-06/r889-readback.txt create mode 100644 documentation/audits/morning-after-2026-10-06/r889-red-proof.txt diff --git a/CONTEXT.md b/CONTEXT.md index 5bdd0fcd..61b41695 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,14 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-06 (morning) — the morning after (rulings `09` §3 137–138).** Register 164 → 150 (0 opened, 14 closed). +> Installer 1.32.0 PUBLISHED (tag `installer-v1.32.0` = `32a1520833`, pins `dbede797`; public sha = tag). R-889: every +> disk percent is `df`'s (`system.DFUsedPercent`, controller `bee2c2d`) — controller v0.299.0 (`ff4a99a`, MinAgent 0.131.0) +> + golden 0.299.0 (`9948aa16…`, pinned Docker set), vouched with agent 0.148.0, floors 0.299.0 for the three boxes; +> read back demo-hp SSD 21 % → 23 % = df. Catalog `3896cb0`: R-776 (kimai/zipline/vikunja/nextcloud trust the docker +> networks) + R-613 (nextcloud probe `"installed":true`), each proven on 9202 with a control. Scratch 9202's controller +> now 0.299.0 (set by hand). Ten operator questions: `STATUS.md`. Report: `REPORT-morning-after-2026-10-06.md`. + > **2026-10-06 (night) — the burn-down night (releases).** Register 199 → 164 (1 opened: R-889 disk percent vs `df`; > 36 closed). Releases in window 1 (00:30–00:58): hub v0.138.0 (manifest `edde9e13`; R-879 seal — 16 legacy values sealed, > raw DB copy 0 plaintext; roll-back = `felhom-hub -unseal-box-secrets` in the pod first), agent v0.148.0 (tag = `861d32a`, diff --git a/REPORT-morning-after-2026-10-06.md b/REPORT-morning-after-2026-10-06.md new file mode 100644 index 00000000..7f2cd1df --- /dev/null +++ b/REPORT-morning-after-2026-10-06.md @@ -0,0 +1,85 @@ +# REPORT — the morning after: installer 1.32.0 published, the real disk percentage (R-889), the four held catalog fixes proven, ten questions on one page — 2026-10-06 + +| Part | Result | +|---|---| +| **A** — publish installer 1.32.0 | **done** — tag `installer-v1.32.0` (= `32a1520833`), both `webpage.yaml` pins moved (`dbede797`), synced; the public URL serves `SCRIPT_VERSION="1.32.0"`, sha256 `70a4ea9f…` = the tag's file; 9 installer rows closed | +| **B** — the real disk percentage (R-889) | **done** — cause measured (the root reserve was in the denominator); one function `DFUsedPercent` for every disk percent; controller **v0.299.0** + golden 0.299.0 delivered to demo-hp, demo-felhom, Tester 1; read back: demo-hp SSD 21 % → 23 % = `df` | +| **C** — the four waiting catalog fixes | **done** — all proven on scratch 9202 with a control each, pushed to the live catalog (`3896cb0`); R-776 and R-613 closed | +| **D** — the ten questions on one page | **done** — `STATUS.md` „Ten questions for you", each with two options, the cost, what happens if nothing, and a pick | + +| Rows before | Rows after | Opened | Closed | +|---|---|---|---| +| **164** | **150** | **0** | **14** | + +Counted by `register_shape_gate.py`'s method. Closed: R-275, R-276, R-881, R-306, R-130, R-180, R-179, R-310, R-274 (installer), +R-889, R-776, R-613, R-585, R-621. + +## Baselines and rulings + +Verified 07:49: felhom.eu `32a1520833` · hub v0.138.0 · agent v0.148.0 · controller v0.298.0 (+ unreleased) · register 164. +The two rulings were recorded first as `09` §3 decisions 137 (publish 1.32.0) and 138 (R-889: `df`'s number, alarm levels kept). + +## Part A — the installer + +Tag → push (pre-push gates OK) → both pins → commit `dbede797` → ArgoCD sync → `felhom-webpage` rolled out. The first +read-back hit a short 502 during the pod swap (the downloaded file was an nginx error page); the second, with `curl -f`, +read 1.32.0 and the tag's exact bytes (`audits/morning-after-2026-10-06/installer-publish.txt`). CI 1404 (tag) and 1405 +(pins) green. + +## Part B — R-889 + +**Measured, not guessed** (demo-hp 9201, `stat -f` + `df -P` on `/mnt/sys_drive`): blocks 18016108, free 14150899, avail +13229302 → the old `(blocks − free) / blocks` = 21.5 %, `df` = 23 %. The gap is the 5 % reserved for root, which `df` +leaves out of the denominator. Fix: `system.DFUsedPercent(used, avail)` = used / (used + avail), used by `GetDiskUsage`, +`readDiskUsage`, the recovery-unit headroom projection and the deploy page's free percent (`bee2c2d`). The tile's „(/)" +label measured the docker data volume, so it now says „Rendszer" / „System" (7 parity fixtures changed by exactly those +bytes). Tests: `TestDFUsedPercent_MatchesDFOnRealNumbers` (the demo-hp numbers and the R-516 shape) and +`TestR889_EveryDiskPercentIsTheDFOne` (a source scan); red-proved by putting the old line back. Full suite rc 0, gates rc 0. + +Release v0.299.0 (`ff4a99a`, MinAgent 0.131.0; also carries R-585, R-621, R-516 from the night), image built, **golden +0.299.0 baked with the pinned Docker set from the start** (sha `9948aa16…`, round trip equal, token leak 0 with a working +control, VM back on `virgin`), vouched (agent 0.148.0, min_agent 0.131.0), floors 0.299.0 for the three boxes only. +Delivered by 06:17:33Z. **Read back** (the hub page, from each box's report, vs `df` in the guest — two channels): demo-hp +SSD **23 %** (df 23 %; was 21 %), NVMe 6 % (df 7 %; df rounds up), demo-felhom 1 % (df 1 %). Tester 1 reports 0.299.0; its +`df` was not read (no direct access). CI 1406 green. + +## Part C — the four catalog fixes on 9202 + +Method (`audits/morning-after-2026-10-06/live-9202/README.md`): 9202's controller set by hand to 0.299.0; the drill repo +reset to live `d955df1` + the five held commits; 9202 pointed at it per `09` §6.5; every request through the simulated +tunnel (a curl container at cloudflared's address), the stranger forging a new leftmost `X-Forwarded-For` each try; each +app also run as a **control** with its new line removed from the stack's compose and restarted through the product. + +| App | With the fix | Control | +|---|---|---| +| vikunja | stranger 10 × 403 then 429; household **200** | household **429** | +| zipline | stranger 7 × 400 then 429; household **200** | household **429** | +| kimai | stranger 5 × wrong then THROTTLED; household **ok** | household **THROTTLED** | +| nextcloud | brute-force attempts: visitor **6**, forged/tunnel/traefik 0 | visitor **0** | +| nextcloud probe | `status.php` `{"installed":true,…}`, controller `running` | — | + +**Not established:** where nextcloud's control tries were counted (its throttle lives in Redis; traefik's two addresses +read 0). Controls held: demo-hp's catalog cache and the live `main` stayed `d955df1` during the test. **The stall of last +night was not repeated:** this time the lead ran the proof itself, so there was no helper to stall. Teardown in the +README; the catalog pointer was restored byte-identical (`cmp`). + +**Found on the way, fixed or noted:** the walk tool's default data path (`/userdata/`) is a folder inside a +drive, which R-839's new refusal rejects — passed the drive itself instead (the tool still carries the old default; it is +only reached on 9202). One `docker ps` of mine was filtered by a „perl" locale filter that also dropped the `paperless-*` +lines; paperless had run since 00:30:09Z, before the work — said in the README so nobody reads it as a start caused here. + +## Part D + +`STATUS.md`, „Ten questions for you": R-444, R-99, R-618, R-645, R-856, R-747, R-734, R-624, R-502, R-774. A reply of +„all as picked" is enough. + +## CI, last commit of every repo + +felhom.eu: this commit (its run is checked by its commit after the push; the previous one, `e6cfe6d`, is 1407 success); felhom-controller `ff4a99a` → 1406 success; +app-catalog `3896cb0` → 1408 success; felhom-agent `37e98f4` (unchanged) → 1399 success. + +## Teardown + +Drill VM: build guest destroyed, secrets shredded, off, `virgin`. 9202: four throwaway apps removed through the product, +their own folders removed by name, the catalog pointer restored byte-identical, tools and password files deleted; its +controller stays on 0.299.0. demo-hp host: nothing changed. Hub: the vouch and three floors. Scratch secrets shredded. diff --git a/STATUS.md b/STATUS.md index 28ceff1e..45665a69 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,11 +1,32 @@ # STATUS — what works, what's broken, what's next -**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was -sent to it.** +**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-06 06:00 (the burn-down night): every box of ours healthy on hub 0.138.0, agent 0.148.0, controller -0.298.0. The open-items list is at 164 (was 199). Morning note: `documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md`; -report: `REPORT-burndown3-2026-10-06.md`.** +**Updated 2026-10-06 09:00: every box of ours healthy on hub 0.138.0, agent 0.148.0, controller 0.299.0; installer 1.32.0 +is public. The open-items list is at 150. Report: `REPORT-morning-after-2026-10-06.md`.** + +## This morning (2026-10-06): your two answers done; the list at 150 + +- **Installer 1.32.0 is published.** The public script reads 1.32.0 and is byte-identical to its tag. +- **The box shows the real disk use** (the number `df` gives). demo-hp's SSD went from 21 % to 23 %, the same as `df`. + Alarm levels are unchanged, so a full disk alarms a little earlier. Released as controller 0.299.0 to all three boxes. +- **The four waiting app fixes passed on the scratch box and are live** (visitor addresses for kimai, zipline, vikunja, + nextcloud; nextcloud's health check). Each was tested against a control. + +## Ten questions for you — answer „all as picked", or name the ones you want differently + +| # | Question | Option A | Option B | If you decide nothing | My pick | +|---|---|---|---|---|---| +| 1 | **R-444** — Should the boxes trim their customer guests' disks once a week, so deleted data stops filling the disk pool? | Yes: one new admin permission (`pct fstrim`), weekly, outside the night, measured once on demo-hp first | No: leave it; the pool keeps space nobody uses | Nothing changes; a full pool can one day stop every guest on that box | **A** | +| 2 | **R-99** — Should broken leftovers of aborted backups be deleted on the backup server? | Yes, by hand, by a runbook, when one is seen | No: they are detected, harmless, and do not affect what is kept | They stay; one small leftover per aborted upload | **B** | +| 3 | **R-618** — May an app update count Docker's own „healthy" as proof that the new version works? | No: keep our own check only | Yes: Docker's healthy can end a wait early | Nothing changes (A) | **A** | +| 4 | **R-645** — When you lift a held update by hand, how do we stop the night backup from saving the broken version over the good copy? | The night backup skips an app whose saved version is not what it runs | Lifting a hold by hand is refused; you use „Undo" instead | The good copy can be overwritten within seconds of a hand lift | **A** | +| 5 | **R-856** — After a crash restart, should app mails wait longer (about 15 minutes), so the household gets one message, not three? | Yes: a longer quiet time after a crash boot (two repositories change) | No: close it; one crash can send several true mails | Several mails after a crash, as now | **A** | +| 6 | **R-747** — Should mealie's app page tell the household that five wrong logins lock the account for 1–2 hours? | Yes: one sentence on the page | No | A locked household does not know why or for how long | **A** | +| 7 | **R-734** — Should the update test ignore small marker files an app rewrites at every start (immich)? | Yes: a per-app list of such files, each with a reason | No: keep marking it; immich updates then need a fresh full copy first | immich's night update is skipped where no fresh copy exists | **A** | +| 8 | **R-624** — May the update test hold an app's admin password to create test data in apps that refuse strangers (vaultwarden, zipline)? | Yes, on the test bench only | No: write down that these apps are tested by hand | No change; those two apps stay hand-tested | **B** | +| 9 | **R-502** — May a slow check that starts Docker containers run on DooPlex (the ISO's first-boot test)? | Yes, in full runs only (never on every push), with a clear „not checked" when Docker is missing | No: it stays a hand step of each ISO release | It stays a hand step; it can be forgotten | **B** | +| 10 | **R-774** — Should Karakeep's page say that its phone app sends crash reports to its makers? | Yes: one sentence on the page | No | The household is not told | **A** | ## Tonight (2026-10-06): the list at 164 @@ -14,7 +35,7 @@ hub 0.138.0 (secrets in the hub database are sealed; checked on a copy: none rea files), controller 0.298.0 and a new install image 0.298.0. The app catalog was updated. Six small decisions were taken by CC unattended (`09` §3 decisions 131–136); each can be reversed. -**Needs you (none urgent; if you do nothing, each stays as it is):** +**Needs you (none urgent; if you do nothing, each stays as it is):** *(items 1 and 2 answered 2026-10-06 07:45 and done; item 3 is the table above)* 1. **Publish installer 1.32.0?** Nine uninstall/pre-flight fixes wait on `main`. If nothing: new installs keep the old installer. Pick: yes. 2. **R-889 — the disk percentage** reads ~5 points low (not `df`'s formula), so fill alarms come late. Pick: use `df`'s diff --git a/documentation/audits/morning-after-2026-10-06/delivery/controller-delivery.txt b/documentation/audits/morning-after-2026-10-06/delivery/controller-delivery.txt new file mode 100644 index 00000000..d407352a --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/delivery/controller-delivery.txt @@ -0,0 +1,3 @@ +== controller delivery 2026-10-06T06:47:49Z +demo-hp + demo-felhom: gitea.dooplex.hu/admin/felhom-controller:0.299.0 (healthy) at 06:17:33Z (docker ps) +tester-1: hub customer page 'Controller version 0.299.0', 'Controller elindult (0.299.0)' diff --git a/documentation/audits/morning-after-2026-10-06/delivery/vouch-golden-floors.txt b/documentation/audits/morning-after-2026-10-06/delivery/vouch-golden-floors.txt new file mode 100644 index 00000000..250b2bd4 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/delivery/vouch-golden-floors.txt @@ -0,0 +1,12 @@ +== before: hub customer page demo-hp 'SSD 21% 14.7 / 68.7 GB' (df in the guest: /mnt/sys_drive 23%) +== vouch 2026-10-06T06:16:27Z: agent 0.148.0, golden 0.299.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +== floors 2026-10-06T06:16:56Z: min_controller_version=0.299.0 min_agent=0.131.0 +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/06 08:16:56 [INFO] Artifact manifest set: agent=0.148.0 golden=0.299.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="a6fa4f589d184b58c9911303bd087e300be1e75b3647e4302c9594df6989c4de" +2026/10/06 08:16:56 [INFO] Customer demo-hp controller-version floor override set to "0.299.0" (declared MinAgent "0.131.0") +2026/10/06 08:16:56 [INFO] Customer demo-felhom controller-version floor override set to "0.299.0" (declared MinAgent "0.131.0") +2026/10/06 08:16:56 [INFO] Customer tester-1 controller-version floor override set to "0.299.0" (declared MinAgent "0.131.0") diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/README.md b/documentation/audits/morning-after-2026-10-06/live-9202/README.md new file mode 100644 index 00000000..7f8e887b --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/README.md @@ -0,0 +1,28 @@ +# Scratch 9202 — the four held catalog fixes proven (2026-10-06 morning) + +Venue: scratch guest 9202 on demo-hp, controller set BY HAND to 0.299.0 (was 0.296.0; allowed only there), pointed at the +drill catalog `admin/app-catalog-drill` = live catalog `d955df1` + the five held commits (`c5674ce`), by `tools/repoint.py` +(`09` §6.5: saved copy `controller.yaml.pre-morning1006`, cache removed, restart). Controls: demo-hp's 9201 cache and the +live catalog's `main` both stayed `d955df1`. Every request went through the SIMULATED tunnel: a curl container at +172.16.253.2 (cloudflared's address, the only one traefik trusts for forwarded headers) sending +`X-Forwarded-For: , ` — the stranger changed the forged leftmost entry on every try. + +| App (row) | With the fix | Control: the line removed from the stack's compose, restart through the product | +|---|---|---| +| nextcloud (R-776) | 6 wrong WebDAV logins → `occ security:bruteforce:attempts`: **visitor 6**, every forged address 0, the tunnel 0, traefik 0 | visitor **0** (traefik 0 too: this nextcloud keeps its throttle in Redis, so where the control's tries were counted was not found) | +| nextcloud (R-613) | `status.php` = `{"installed":true,"maintenance":false,…}`; the controller's state `running` with the new probe | — | +| vikunja (R-776) | stranger: 10 × 403 then 429; **household 200** | stranger the same; **household 429** | +| zipline (R-776) | stranger: 7 × 400 then 429; **household 200** | stranger the same; **household 429** | +| kimai (R-776) | stranger: 5 × wrong then THROTTLED; **household ok** | stranger the same; **household THROTTLED** | + +So with the fix each app throttles the VISITOR, a stranger cannot escape by forging the leftmost address, and the +household is not locked out by a stranger; without it the household is. + +**Teardown:** each app removed through the product (keep drive data, drop backups), its own folders under the scratch +drive removed by name (`teardown.txt`); `controller.yaml` restored byte-identical (`cmp`); the tools and password files +deleted in the guest. 9202's controller stays on 0.299.0 (newer than before; `/etc/felhom-controller-image.pre-morning1006` +holds the old line). The drill repo keeps the tested branch. Host demo-hp and the hub: nothing changed. + +**Said plainly:** a log filter used during this work drops lines containing „perl" (locale noise) — it also dropped the +`paperless-*` containers from one `docker ps`, which made it look as if paperless started during the work. It did not: +it has run since 2026-10-06 00:30:09Z (`docker inspect` StartedAt), before any of this. diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/kimai.txt b/documentation/audits/morning-after-2026-10-06/live-9202/kimai.txt new file mode 100644 index 00000000..c8737708 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/kimai.txt @@ -0,0 +1,16 @@ +##### kimai on 9202 — controller gitea.dooplex.hu/admin/felhom-controller:0.299.0 +deploy -> True +env: ['TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12'] +control first: the household logs in at once -> ok +## WITH the fix (TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12) — 2026-10-06T06:41:42Z + stranger 198.51.100.66, 7 wrong for admin@example.com (a new forged leftmost each): ['wrong', 'wrong', 'wrong', 'wrong', 'wrong', 'THROTTLED', 'THROTTLED'] + household 203.0.113.10, the right password, at once: ok +## CONTROL: the line removed 0 + restart -> 200 + env: ['TRUSTED_PROXIES=nginx,localhost,127.0.0.1'] +## CONTROL — the template before R-776 — 2026-10-06T06:42:52Z + stranger 198.51.100.66, 7 wrong for admin@example.com (a new forged leftmost each): ['wrong', 'wrong', 'wrong', 'wrong', 'wrong', 'THROTTLED', 'THROTTLED'] + household 203.0.113.10, the right password, at once: THROTTLED +## put the template's compose back: 1 + restart -> 200 + env back: ['TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12'] diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/nextcloud.txt b/documentation/audits/morning-after-2026-10-06/live-9202/nextcloud.txt new file mode 100644 index 00000000..aeb287c4 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/nextcloud.txt @@ -0,0 +1,33 @@ +##### nextcloud on 9202 — controller gitea.dooplex.hu/admin/felhom-controller:0.299.0 — 2026-10-06T06:21:57Z +deploy -> True +the container env: 172.16.0.0/12 +R-613 status.php bytes (inside the container): {"installed":true,"maintenance":false,"needsDbUpgrade":false,"version":"34.0.4.1","versionstring":"34.0.4","edition":"","productname":"Nextcloud","extendedSupport":false} +R-613 the controller's state for the app (the probe expects '"installed":true'): running | health: +trusted_proxies in config.php: 172.16.0.0/12 +stranger 198.51.100.66, 6 wrong basic-auth tries with a new forged leftmost each: [['401'], ['401'], ['401'], ['401'], ['401'], ['401']] +occ bruteforce:attempts 198.51.100.66 (the real visitor): - bypass-listed: false - attempts: 6 - delay: 6400 +occ bruteforce:attempts 10.5.0.1 (forged / the tunnel): - bypass-listed: false - attempts: 0 - delay: 0 +occ bruteforce:attempts 10.5.0.6 (forged / the tunnel): - bypass-listed: false - attempts: 0 - delay: 0 +occ bruteforce:attempts 172.16.253.2 (forged / the tunnel): - bypass-listed: false - attempts: 0 - delay: 0 +occ bruteforce:attempts 172.16.253.3 (traefik): - bypass-listed: false - attempts: 0 - delay: 0 +occ bruteforce:attempts 172.18.0.5 (traefik): - bypass-listed: false - attempts: 0 - delay: 0 +occ bruteforce:attempts 203.0.113.10 (another visitor, never failed): - bypass-listed: false - attempts: 0 - delay: 0 +## CONTROL — trusted_proxies removed by hand (the template before R-776) — 2026-10-06T06:24:58Z +delete: System config value trusted_proxies deleted | now: 172.16.0.0/12 +stranger 198.51.100.77, 3 wrong tries: [['401'], ['401'], ['401']] +attempts 198.51.100.77 (visitor): - bypass-listed: false - attempts: 3 - delay: 800 +attempts 172.16.253.3 (traefik): - bypass-listed: false - attempts: 0 - delay: 0 +attempts 172.18.0.5 (traefik): - bypass-listed: false - attempts: 0 - delay: 0 +restore: System config value trusted_proxies => 0 set to string 172.16.0.0/12 | now: 172.16.0.0/12 +## CONTROL 2 — TRUSTED_PROXIES removed from the stack's compose (= the template before R-776) — 2026-10-06T06:26:13Z +0 +System config value trusted_proxies deleted +restart through the product -> 200 +env now: (unset) | trusted_proxies: +stranger 198.51.100.88, 3 wrong tries: [['401'], ['401'], ['401']] +attempts 198.51.100.88 (visitor): - bypass-listed: false - attempts: 0 - delay: 0 +attempts 172.16.253.3 (traefik): - bypass-listed: false - attempts: 0 - delay: 0 +attempts 172.18.0.5 (traefik): - bypass-listed: false - attempts: 0 - delay: 0 +## put the template's compose back: 1 +restart through the product -> 200 +env back: 172.16.0.0/12 diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/teardown.txt b/documentation/audits/morning-after-2026-10-06/live-9202/teardown.txt new file mode 100644 index 00000000..a9c61a37 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/teardown.txt @@ -0,0 +1,91 @@ +nextcloud: stop -> 200 +nextcloud: remove (keep drive data, drop backups) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_ht +nextcloud: tidy own folders by name: removed /mnt/felhom-drives/scratch_hdd/appdata/nextcloud +removed /mnt/felhom-drives/scratch_hdd/userdata/nextcloud +nextcloud: after: deployed=False leftovers='/opt/docker/stacks/nextcloud' +vikunja: stop -> 200 +vikunja: remove (keep drive data, drop backups) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_path +vikunja: tidy own folders by name: +vikunja: after: deployed=False leftovers='/opt/docker/stacks/vikunja' +zipline: stop -> 200 +zipline: remove (keep drive data, drop backups) -> 200 {'ok': True, 'data': {'removed': 'zipline', 'volumes_removed': ['zipline_zipline_postgres_data', 'zipline_zipline_public +zipline: tidy own folders by name: +zipline: after: deployed=False leftovers='/opt/docker/stacks/zipline' +kimai: stop -> 200 +kimai: remove (keep drive data, drop backups) -> 200 {'ok': True, 'data': {'removed': 'kimai', 'volumes_removed': ['kimai_kimai_db_data', 'kimai_kimai_var'], 'hdd_paths_remo +kimai: tidy own folders by name: +kimai: after: deployed=False leftovers='/opt/docker/stacks/kimai' +26: repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git +3 +RESTORED-IDENTICAL + +0 +filebrowser +felhom-controller +paperless-webserver +paperless-postgres +paperless-redis +traefik +actualbudget +adventurelog +audiobookshelf +bentopdf +bookstack +calcom +calibre-web +chaoscrash +chaosoom +chaosoomb +claper +code-server +crafty-controller +dawarich +docmost +emby +filebrowser +ghost +gitea +glance +gokapi +grafana +gramps-web +grimmory +home-assistant +homebox +homepage +immich +jellyfin +karakeep +kimai +komga +mealie +metube +n8n +navidrome +nextcloud +onlyoffice +opengist +outline +paperless-ngx +papra +plant-it +plex +privatebin +radarr +radicale +rallly +recipe-importer +romm +seerr +sonarr +sparkyfitness +tandoor +termix +traefik +uptime-kuma +vaultwarden +vikunja +wanderer +wger +wishlist +zipline diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/burst.sh b/documentation/audits/morning-after-2026-10-06/live-9202/tools/burst.sh new file mode 100644 index 00000000..1d8eb46d --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/burst.sh @@ -0,0 +1,10 @@ +#!/bin/sh +# burst.sh — n wrong POSTs from stranger 198.51.100.66 (a new forged +# leftmost each), then ONE right POST from household 203.0.113.10, all through the simulated tunnel, back to back. +H=$1; P=$2; N=$3; BAD=$4; GOOD=$5; out="" +i=1; while [ $i -le $N ]; do + c=$(curl -sk -o /dev/null -w '%{http_code}' -X POST "https://traefik$P" -H "Host: $H" -H "X-Forwarded-For: 10.5.0.$i, 198.51.100.66" \ + -H "CF-Connecting-IP: 198.51.100.66" -H "Content-Type: application/json" --data "$BAD"); out="$out $c"; i=$((i+1)); done +g=$(curl -sk -o /dev/null -w '%{http_code}' -X POST "https://traefik$P" -H "Host: $H" -H "X-Forwarded-For: 203.0.113.10" \ + -H "CF-Connecting-IP: 203.0.113.10" -H "Content-Type: application/json" --data @"$GOOD") +echo "stranger:$out | household: $g" diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/km.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/km.py new file mode 100644 index 00000000..e4195b85 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/km.py @@ -0,0 +1,34 @@ +import sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +APP, SUB, EMAIL = "kimai", "time", "admin@example.com" +EVF = open(f"{w.EV}/kimai.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +def login(v, f, pw): + return w.guest(f"docker run --rm --network felhom-tunnel --ip 172.16.253.2 -v /root/kmlogin.sh:/s.sh:ro -v /root/.kmpw:/pw:ro " + f"curlimages/curl:8.11.1 sh /s.sh {v} {f} {EMAIL} {pw}").strip().splitlines()[-1] +def rounds(label): + say(f"## {label} — {time.strftime('%FT%TZ', time.gmtime())}") + say(" stranger 198.51.100.66, 7 wrong for", EMAIL, "(a new forged leftmost each):", [login("198.51.100.66", f"10.5.0.{i}", "-wrong") for i in range(1, 8)]) + say(" household 203.0.113.10, the right password, at once:", login("203.0.113.10", "203.0.113.10", "/pw")) +w.login() +say(f"##### kimai on 9202 — controller {w.guest('docker inspect felhom-controller --format {{.Config.Image}}').strip()}") +say("deploy ->", w.deploy(APP, SUB)) +pw = (w.GENERATED.get(APP) or {}).get("ADMIN_PASSWORD") +if not pw: sys.exit("no generated ADMIN_PASSWORD") +w.guest(f"umask 077; printf '%s' '{pw}' > /root/.kmpw; chmod 644 /root/.kmpw") +w.wait_app(SUB, "/en/login", tries=60) +env = lambda: [l for l in w.guest("docker inspect kimai --format '{{range .Config.Env}}{{println .}}{{end}}'").splitlines() if "TRUSTED_PROXIES" in l] +say("env:", env()) +say("control first: the household logs in at once ->", login("203.0.113.10", "203.0.113.10", "/pw")) +rounds("WITH the fix (TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12)") +D = "/opt/docker/stacks/kimai" +say("## CONTROL: the line removed", w.guest(f"cp -p {D}/docker-compose.yml /root/km.with; sed -i '/- TRUSTED_PROXIES=/d' {D}/docker-compose.yml; grep -c TRUSTED_PROXIES {D}/docker-compose.yml").strip()) +c, d = w.ctl("POST", f"/api/stacks/{APP}/restart"); say(" restart ->", c); time.sleep(15); w.wait_app(SUB, "/en/login", tries=60) +say(" env:", env()) +rounds("CONTROL — the template before R-776") +say("## put the template's compose back:", w.guest(f"cp -p /root/km.with {D}/docker-compose.yml; grep -c TRUSTED_PROXIES {D}/docker-compose.yml").strip()) +c, d = w.ctl("POST", f"/api/stacks/{APP}/restart"); say(" restart ->", c); time.sleep(15) +say(" env back:", env()) +w.guest("rm -f /root/.kmpw") diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/kmlogin.sh b/documentation/audits/morning-after-2026-10-06/live-9202/tools/kmlogin.sh new file mode 100644 index 00000000..553dba6c --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/kmlogin.sh @@ -0,0 +1,14 @@ +#!/bin/sh +# kmlogin.sh — ONE Kimai form login through the SIMULATED tunnel. +# Prints: ok | wrong | THROTTLED | ? +V=$1; F=$2; E=$3; P=$4; H=time.enkisfelhom.hu; J=/tmp/jar.$$ +x="X-Forwarded-For: $F, $V"; c="CF-Connecting-IP: $V" +curl -sk -c $J -b $J "https://traefik/en/login" -H "Host: $H" -H "$x" -H "$c" -o /tmp/p.$$ +T=$(grep -o 'name="_csrf_token" value="[^"]*"' /tmp/p.$$ | head -1 | sed 's/.*value="//;s/"$//') +if [ "$P" = "-wrong" ]; then PW="wrong-$$"; else PW=$(cat "$P"); fi +L=$(curl -sk -c $J -b $J -o /dev/null -w '%{http_code} %{redirect_url}' "https://traefik/en/login_check" -H "Host: $H" -H "$x" -H "$c" \ + --data-urlencode "_csrf_token=$T" --data-urlencode "_username=$E" --data-urlencode "_password=$PW") +case "$L" in *"/login"*) curl -sk -c $J -b $J "https://traefik/en/login" -H "Host: $H" -H "$x" -H "$c" -o /tmp/q.$$ + if grep -qiE 'too many|try again in' /tmp/q.$$; then echo THROTTLED; else echo wrong; fi;; + "302 "*) echo ok;; *) echo "?$L";; esac +rm -f $J /tmp/p.$$ /tmp/q.$$ diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc.py new file mode 100644 index 00000000..87b38b97 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc.py @@ -0,0 +1,40 @@ +import sys, time, json +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +EVF = open(f"{w.EV}/nextcloud.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +APP, SUB = "nextcloud", "cloud" +H = f"{SUB}.{w.DOMAIN}" +def tun(v, f, m, p, b="", ct="", x=""): + q = lambda s: "'" + s.replace("'", "'\\''") + "'" + return w.guest(f"docker run --rm --network felhom-tunnel --ip 172.16.253.2 -v /root/tun.sh:/t.sh:ro curlimages/curl:8.11.1 " + f"sh /t.sh {H} {v} {f} {m} {q(p)} {q(b)} {q(ct)} {q(x)}").strip().splitlines()[-1:] +w.login() +say(f"##### nextcloud on 9202 — controller {w.guest('docker inspect felhom-controller --format {{.Config.Image}}').strip()} — {time.strftime('%FT%TZ', time.gmtime())}") +say("deploy ->", w.deploy(APP, SUB, {"HDD_PATH": "/mnt/felhom-drives/scratch_hdd"})) +say("the container env:", w.guest("docker exec nextcloud printenv TRUSTED_PROXIES").strip()) +# R-613: wait until the app is up, then read status.php's exact bytes and the controller's own verdict +for i in range(90): + st = w.guest("docker exec nextcloud curl -s http://localhost/status.php").strip() + if '"installed":true' in st: break + time.sleep(10) +say("R-613 status.php bytes (inside the container):", st[:200]) +for i in range(30): + s = w.stack(APP); state = s.get("state") + if state == "running": break + time.sleep(10) +say("R-613 the controller's state for the app (the probe expects '\"installed\":true'):", state, "| health:", (s.get("health") or s.get("health_status") or "")) +say("trusted_proxies in config.php:", w.guest("docker exec -u www-data nextcloud php occ config:system:get trusted_proxies 2>&1 | tr '\\n' ' '").strip()) +# R-776 3.6: a stranger fails WebDAV basic auth 6 times through the simulated tunnel, each with a NEW forged leftmost +V, LAN = "198.51.100.66", "203.0.113.10" +codes = [tun(V, f"10.5.0.{i}", "PROPFIND", "/remote.php/dav/files/nobody/", "", "", "Authorization: Basic bm9ib2R5Ondyb25n") for i in range(1, 7)] +say(f"stranger {V}, 6 wrong basic-auth tries with a new forged leftmost each:", codes) +time.sleep(3) +occ = lambda ip: w.guest(f"docker exec -u www-data nextcloud php occ security:bruteforce:attempts {ip} 2>&1 | tr '\\n' ' '").strip() +say(f"occ bruteforce:attempts {V} (the real visitor):", occ(V)) +for f in ("10.5.0.1", "10.5.0.6", "172.16.253.2"): + say(f"occ bruteforce:attempts {f} (forged / the tunnel):", occ(f)) +tip = w.guest("docker inspect traefik --format '{{range .NetworkSettings.Networks}}{{.IPAddress}} {{end}}'").split() +for ip in tip: say(f"occ bruteforce:attempts {ip} (traefik):", occ(ip)) +say(f"occ bruteforce:attempts {LAN} (another visitor, never failed):", occ(LAN)) diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc2.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc2.py new file mode 100644 index 00000000..63e7a0d2 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc2.py @@ -0,0 +1,21 @@ +import sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +EVF = open(f"{w.EV}/nextcloud.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +H = "cloud.enkisfelhom.hu" +def tun(v, f, m, p, x=""): + q = lambda s: "'" + s.replace("'", "'\\''") + "'" + return w.guest(f"docker run --rm --network felhom-tunnel --ip 172.16.253.2 -v /root/tun.sh:/t.sh:ro curlimages/curl:8.11.1 " + f"sh /t.sh {H} {v} {f} {m} {q(p)} '' '' {q(x)}").strip().splitlines()[-1:] +occ = lambda a: w.guest(f"docker exec -u www-data nextcloud php occ {a} 2>&1 | tr '\\n' ' '").strip() +say(f"## CONTROL — trusted_proxies removed by hand (the template before R-776) — {time.strftime('%FT%TZ', time.gmtime())}") +say("delete:", occ("config:system:delete trusted_proxies"), "| now:", occ("config:system:get trusted_proxies")) +V2 = "198.51.100.77" +say(f"stranger {V2}, 3 wrong tries:", [tun(V2, f"10.6.0.{i}", "PROPFIND", "/remote.php/dav/files/nobody/", "Authorization: Basic bm9ib2R5Ondyb25n") for i in range(1, 4)]) +time.sleep(3) +say(f"attempts {V2} (visitor):", occ(f"security:bruteforce:attempts {V2}")) +for ip in w.guest("docker inspect traefik --format '{{range .NetworkSettings.Networks}}{{.IPAddress}} {{end}}'").split(): + say(f"attempts {ip} (traefik):", occ(f"security:bruteforce:attempts {ip}")) +say("restore:", occ("config:system:set trusted_proxies 0 --value=172.16.0.0/12"), "| now:", occ("config:system:get trusted_proxies")) diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc3.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc3.py new file mode 100644 index 00000000..f0b07cdb --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/nc3.py @@ -0,0 +1,33 @@ +import sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +EVF = open(f"{w.EV}/nextcloud.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +H = "cloud.enkisfelhom.hu" +def tun(v, f, m, p, x=""): + q = lambda s: "'" + s.replace("'", "'\\''") + "'" + return w.guest(f"docker run --rm --network felhom-tunnel --ip 172.16.253.2 -v /root/tun.sh:/t.sh:ro curlimages/curl:8.11.1 " + f"sh /t.sh {H} {v} {f} {m} {q(p)} '' '' {q(x)}").strip().splitlines()[-1:] +occ = lambda a: w.guest(f"docker exec -u www-data nextcloud php occ {a} 2>&1 | tr '\\n' ' '").strip() +w.login() +D = "/opt/docker/stacks/nextcloud" +say(f"## CONTROL 2 — TRUSTED_PROXIES removed from the stack's compose (= the template before R-776) — {time.strftime('%FT%TZ', time.gmtime())}") +say(w.guest(f"cd {D} && cp -p docker-compose.yml /root/nc-compose.with && sed -i '/- TRUSTED_PROXIES=/d' docker-compose.yml && grep -c TRUSTED_PROXIES docker-compose.yml; docker exec -u www-data nextcloud php occ config:system:delete trusted_proxies").strip()) +code, d = w.ctl("POST", "/api/stacks/nextcloud/restart"); say("restart through the product ->", code) +for _ in range(40): + if '"installed":true' in w.guest("docker exec nextcloud curl -s http://localhost/status.php"): break + time.sleep(5) +say("env now:", w.guest("docker exec nextcloud printenv TRUSTED_PROXIES || echo '(unset)'").strip(), "| trusted_proxies:", occ("config:system:get trusted_proxies")) +V2 = "198.51.100.88" +say(f"stranger {V2}, 3 wrong tries:", [tun(V2, f"10.7.0.{i}", "PROPFIND", "/remote.php/dav/files/nobody/", "Authorization: Basic bm9ib2R5Ondyb25n") for i in range(1, 4)]) +time.sleep(3) +say(f"attempts {V2} (visitor):", occ(f"security:bruteforce:attempts {V2}")) +for ip in w.guest("docker inspect traefik --format '{{range .NetworkSettings.Networks}}{{.IPAddress}} {{end}}'").split(): + say(f"attempts {ip} (traefik):", occ(f"security:bruteforce:attempts {ip}")) +say("## put the template's compose back:", w.guest(f"cp -p /root/nc-compose.with {D}/docker-compose.yml && grep -c TRUSTED_PROXIES {D}/docker-compose.yml").strip()) +code, d = w.ctl("POST", "/api/stacks/nextcloud/restart"); say("restart through the product ->", code) +for _ in range(40): + if '"installed":true' in w.guest("docker exec nextcloud curl -s http://localhost/status.php"): break + time.sleep(5) +say("env back:", w.guest("docker exec nextcloud printenv TRUSTED_PROXIES").strip()) diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/repoint.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/repoint.py new file mode 100644 index 00000000..8db39c74 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/repoint.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Point 9202 at the drill catalog, or put the saved controller.yaml back. `09` §6.5.""" +import io, os, re, sys +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +SAVE = f"{VOL}/controller.yaml.pre-morning1006" +DRILL = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" +def creds(): + for l in io.open(os.path.expanduser("~/.git-credentials")).read().split("\n"): + m = re.match(r"https://(admin):([^@]+)@gitea\.dooplex\.hu", l) + if m: return m.group(1), m.group(2) + sys.exit("no admin credential") +if sys.argv[1] == "drill": + u, t = creds() + out = w.guest(f"""set -e +test -f {SAVE} || cp -p {VOL}/controller.yaml {SAVE} +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml"; s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml +""") + print(out.replace(t, "")) +elif sys.argv[1] == "restore": + print(w.guest(f"""set -e +cp -p {SAVE} {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml; grep -c 'token: ""' {VOL}/controller.yaml || true +cmp {SAVE} {VOL}/controller.yaml && echo RESTORED-IDENTICAL""")) diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/rm.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/rm.py new file mode 100644 index 00000000..42bb8fcc --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/rm.py @@ -0,0 +1,19 @@ +import sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +name = sys.argv[1]; dirs = sys.argv[2:] +EVF = open(f"{w.EV}/teardown.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +w.login() +c, d = w.ctl("POST", f"/api/stacks/{name}/stop"); say(f"{name}: stop -> {c}") +for _ in range(24): + time.sleep(5) + if w.stack(name).get("state") != "running": break +c, d = w.ctl("POST", f"/api/stacks/{name}/remove", {"remove_hdd_data": False, "remove_backups": True}); say(f"{name}: remove (keep drive data, drop backups) -> {c} {str(d)[:120]}") +time.sleep(5) +for p in dirs: + assert p.startswith("/mnt/felhom-drives/scratch_hdd/") and p.count("/") >= 5 and p.rstrip("/").split("/")[-1] == name, p +say(f"{name}: tidy own folders by name:", w.guest("for p in " + " ".join(dirs) + "; do [ -d $p ] && rm -rf $p && echo removed $p; done; true").strip()) +st = w.stack(name) +say(f"{name}: after: deployed={st.get('deployed')} leftovers={w.guest(f'ls -d /opt/docker/stacks/{name} 2>/dev/null; docker ps -a --filter label=com.docker.compose.project={name} --format {{{{.Names}}}}; docker volume ls -q --filter label=com.docker.compose.project={name}').strip()!r}") diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/tun.sh b/documentation/audits/morning-after-2026-10-06/live-9202/tools/tun.sh new file mode 100644 index 00000000..ad8588f7 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/tun.sh @@ -0,0 +1,12 @@ +#!/bin/sh +# tun.sh [body] [content-type] [extra-header] +# ONE request through the SIMULATED tunnel (run in a curl container at 172.16.253.2, cloudflared's address). +# Prints the HTTP status code only. +H=$1; V=$2; F=$3; M=$4; P=$5; B=$6; CT=${7:-application/json}; X=$8 +if [ -n "$B" ]; then + curl -sk -o /dev/null -w '%{http_code}' -X "$M" "https://traefik$P" -H "Host: $H" -H "X-Forwarded-For: $F, $V" \ + -H "CF-Connecting-IP: $V" -H "Content-Type: $CT" ${X:+-H "$X"} --data "$B" +else + curl -sk -o /dev/null -w '%{http_code}' -X "$M" "https://traefik$P" -H "Host: $H" -H "X-Forwarded-For: $F, $V" \ + -H "CF-Connecting-IP: $V" ${X:+-H "$X"} +fi diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/vk.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/vk.py new file mode 100644 index 00000000..b7a1d7fd --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/vk.py @@ -0,0 +1,40 @@ +import sys, time, json, secrets +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +APP, SUB = "vikunja", "tasks"; H = f"{SUB}.{w.DOMAIN}" +EVF = open(f"{w.EV}/vikunja.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +def tun(v, f, body): + q = lambda s: "'" + s.replace("'", "'\\''") + "'" + return w.guest(f"docker run --rm --network felhom-tunnel --ip 172.16.253.2 -v /root/tun.sh:/t.sh:ro curlimages/curl:8.11.1 " + f"sh /t.sh {H} {v} {f} POST /api/v1/login {q(body)} application/json").strip().splitlines()[-1] +pw = "Drill-" + secrets.token_hex(10) +open(f"{w.SC}/.vkpw", "w").write(pw) +def rounds(label): + say(f"## {label} — {time.strftime('%FT%TZ', time.gmtime())}") + bad = json.dumps({"username": "household", "password": "wrong"}) + r = [tun("198.51.100.66", f"10.5.0.{i}", bad) for i in range(1, 15)] + say(" stranger 198.51.100.66, 14 wrong logins, a new forged leftmost each:", r) + good = json.dumps({"username": "household", "password": pw}) + say(" household 203.0.113.10, the right password, at once:", tun("203.0.113.10", "203.0.113.10", good)) +w.login() +say(f"##### vikunja on 9202 — controller {w.guest('docker inspect felhom-controller --format {{.Config.Image}}').strip()}") +say("deploy ->", w.deploy(APP, SUB)) +w.wait_app(SUB, "/api/v1/info", tries=40) +rc, code, body = w.app_curl(SUB, "/api/v1/register", "-H", "Content-Type: application/json", method="POST", + data=json.dumps({"username": "household", "email": "household@example.com", "password": pw})) +say("register the household's account (before the gate opens) ->", code) +c, d = w.ctl("POST", f"/apps/{APP}/setup-gate/open", {}); say("setup-gate/open ->", c) +time.sleep(20); w.wait_app(SUB, "/api/v1/info", tries=40) +say("env:", w.guest("docker exec vikunja printenv VIKUNJA_SERVICE_IPEXTRACTIONMETHOD").strip()) +rounds("WITH the fix (VIKUNJA_SERVICE_IPEXTRACTIONMETHOD=xff)") +D = "/opt/docker/stacks/vikunja" +say("## CONTROL: the line removed from the stack's compose, restart through the product", w.guest(f"cp -p {D}/docker-compose.yml /root/vk.with; sed -i '/IPEXTRACTIONMETHOD/d' {D}/docker-compose.yml; grep -c IPEXTRACTION {D}/docker-compose.yml").strip()) +say(" waiting 70 s for the minute window"); time.sleep(70) +c, d = w.ctl("POST", f"/api/stacks/{APP}/restart"); say(" restart ->", c); time.sleep(10); w.wait_app(SUB, "/api/v1/info", tries=40) +say(" env:", w.guest("docker exec vikunja printenv VIKUNJA_SERVICE_IPEXTRACTIONMETHOD || echo '(unset)'").strip()) +rounds("CONTROL — the template before R-776 (direct)") +say("## put the template's compose back:", w.guest(f"cp -p /root/vk.with {D}/docker-compose.yml; grep -c IPEXTRACTION {D}/docker-compose.yml").strip()) +c, d = w.ctl("POST", f"/api/stacks/{APP}/restart"); say(" restart ->", c); time.sleep(10) +say(" env back:", w.guest("docker exec vikunja printenv VIKUNJA_SERVICE_IPEXTRACTIONMETHOD").strip()) diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/tools/zl.py b/documentation/audits/morning-after-2026-10-06/live-9202/tools/zl.py new file mode 100644 index 00000000..0fb6a1bc --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/tools/zl.py @@ -0,0 +1,36 @@ +import sys, time, json, secrets +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +APP, SUB = "zipline", "img"; H = f"{SUB}.{w.DOMAIN}" +EVF = open(f"{w.EV}/zipline.txt", "a", buffering=1) +def say(*a): + w.say(*a); EVF.write(" ".join(map(str, a)) + "\n") +pw = "Drill-" + secrets.token_hex(10) +w.guest(f"umask 077; printf '%s' '{json.dumps({'username': 'household', 'password': pw})}' > /root/.zlgood; chmod 644 /root/.zlgood") +def burst(label): + say(f"## {label} — {time.strftime('%FT%TZ', time.gmtime())}") + bad = json.dumps({"username": "household", "password": "wrong"}) + say(" ", w.guest(f"docker run --rm --network felhom-tunnel --ip 172.16.253.2 -v /root/burst.sh:/b.sh:ro -v /root/.zlgood:/g:ro " + f"curlimages/curl:8.11.1 sh /b.sh {H} /api/auth/login 10 '{bad}' /g").strip().splitlines()[-1]) +w.login() +say(f"##### zipline on 9202 — controller {w.guest('docker inspect felhom-controller --format {{.Config.Image}}').strip()}") +say("deploy ->", w.deploy(APP, SUB)) +w.wait_app(SUB, "/api/healthcheck", tries=60) +rc, code, body = w.app_curl(SUB, "/api/setup", "-H", "Content-Type: application/json", method="POST", + data=json.dumps({"username": "household", "password": pw})) +say("first account through the front door (before the gate opens) ->", code, body[:80].replace(pw, "")) +c, d = w.ctl("POST", f"/apps/{APP}/setup-gate/open", {}); say("setup-gate/open ->", c) +time.sleep(15); w.wait_app(SUB, "/api/healthcheck", tries=40) +env = lambda: [l for l in w.guest("docker inspect zipline --format '{{range .Config.Env}}{{println .}}{{end}}'").splitlines() if "TRUST" in l] +say("env:", env()) +burst("WITH the fix (CORE_TRUST_PROXY=true + CORE_TRUSTED_PROXIES=172.16.0.0/12)") +D = "/opt/docker/stacks/zipline" +say("## CONTROL: both lines removed from the stack's compose", w.guest(f"cp -p {D}/docker-compose.yml /root/zl.with; sed -i '/CORE_TRUST_PROXY=/d;/CORE_TRUSTED_PROXIES=/d' {D}/docker-compose.yml; grep -c CORE_TRUST {D}/docker-compose.yml").strip()) +time.sleep(12) +c, d = w.ctl("POST", f"/api/stacks/{APP}/restart"); say(" restart ->", c); time.sleep(10); w.wait_app(SUB, "/api/healthcheck", tries=40) +say(" env:", env()) +burst("CONTROL — the template before R-776") +say("## put the template's compose back:", w.guest(f"cp -p /root/zl.with {D}/docker-compose.yml; grep -c CORE_TRUST {D}/docker-compose.yml").strip()) +c, d = w.ctl("POST", f"/api/stacks/{APP}/restart"); say(" restart ->", c); time.sleep(10) +say(" env back:", env()) +w.guest("rm -f /root/.zlgood") diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/vikunja.txt b/documentation/audits/morning-after-2026-10-06/live-9202/vikunja.txt new file mode 100644 index 00000000..283381b0 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/vikunja.txt @@ -0,0 +1,20 @@ +##### vikunja on 9202 — controller gitea.dooplex.hu/admin/felhom-controller:0.299.0 +deploy -> True +register the household's account (before the gate opens) -> 200 +setup-gate/open -> 200 +env: OCI runtime exec failed: exec failed: unable to start container process: exec: "printenv": executable file not found in $PATH +## WITH the fix (VIKUNJA_SERVICE_IPEXTRACTIONMETHOD=xff) — 2026-10-06T06:30:16Z + stranger 198.51.100.66, 14 wrong logins, a new forged leftmost each: ['403', '403', '403', '403', '403', '403', '403', '403', '403', '403', '429', '429', '429', '429'] + household 203.0.113.10, the right password, at once: 200 +## CONTROL: the line removed from the stack's compose, restart through the product 0 + waiting 70 s for the minute window + restart -> 200 + env: OCI runtime exec failed: exec failed: unable to start container process: exec: "printenv": executable file not found in $PATH +(unset) +## CONTROL — the template before R-776 (direct) — 2026-10-06T06:32:36Z + stranger 198.51.100.66, 14 wrong logins, a new forged leftmost each: ['403', '403', '403', '403', '403', '403', '403', '403', '403', '403', '429', '429', '429', '429'] + household 203.0.113.10, the right password, at once: 429 +## put the template's compose back: 1 + restart -> 200 + env back: OCI runtime exec failed: exec failed: unable to start container process: exec: "printenv": executable file not found in $PATH +VIKUNJA_SERVICE_IPEXTRACTIONMETHOD=xff diff --git a/documentation/audits/morning-after-2026-10-06/live-9202/zipline.txt b/documentation/audits/morning-after-2026-10-06/live-9202/zipline.txt new file mode 100644 index 00000000..7e881031 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/live-9202/zipline.txt @@ -0,0 +1,15 @@ +##### zipline on 9202 — controller gitea.dooplex.hu/admin/felhom-controller:0.299.0 +deploy -> True +first account through the front door (before the gate opens) -> 200 {"firstSetup":false,"user":{"id":"pun6ifh0vehx2edctxo1w55g","username":"househol +setup-gate/open -> 200 +env: ['CORE_TRUSTED_PROXIES=172.16.0.0/12', 'CORE_TRUST_PROXY=true'] +## WITH the fix (CORE_TRUST_PROXY=true + CORE_TRUSTED_PROXIES=172.16.0.0/12) — 2026-10-06T06:36:47Z + stranger: 400 400 400 400 400 400 400 429 429 429 | household: 200 +## CONTROL: both lines removed from the stack's compose 0 + restart -> 200 + env: [] +## CONTROL — the template before R-776 — 2026-10-06T06:37:31Z + stranger: 400 400 400 400 400 400 400 429 429 429 | household: 429 +## put the template's compose back: 2 + restart -> 200 + env back: ['CORE_TRUST_PROXY=true', 'CORE_TRUSTED_PROXIES=172.16.0.0/12'] diff --git a/documentation/audits/morning-after-2026-10-06/r889-readback.txt b/documentation/audits/morning-after-2026-10-06/r889-readback.txt new file mode 100644 index 00000000..2eec611a --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/r889-readback.txt @@ -0,0 +1,8 @@ +== R-889 read-back 2026-10-06T06:46:57Z: hub customer page (from the box's report) vs df in the guest (a different channel) +-- demo-hp hub: SSD 23% 14.7 / 68.7 GB NVME 1TB 6% 54.6 / 937.8 GB + df: /mnt/sys_drive=23% /mnt/felhom-drives/hdd_1=7% +-- demo-felhom hub: SSD 1% 1.6 / 245.0 GB + df: /mnt/sys_drive=1% +Storage SSD 6% 4.1 / 68.7 GB Adatlemez 0% 0.2 / 48.9 GB Bac +Controller version 0.299.0 +Controller elindult (0.299.0) diff --git a/documentation/audits/morning-after-2026-10-06/r889-red-proof.txt b/documentation/audits/morning-after-2026-10-06/r889-red-proof.txt new file mode 100644 index 00000000..dcab4b61 --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/r889-red-proof.txt @@ -0,0 +1,3 @@ +--- FAIL: TestR889_EveryDiskPercentIsTheDFOne (0.05s) + dfpercent_test.go:66: a disk percent divides by the whole filesystem (not df's used/(used+avail)) — use DFUsedPercent: + ../../internal/system/mounts_linux.go:85: info.UsedPercent = float64(used) / float64(total) * 100 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 65220127..2289a4fe 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,18 @@ --- +## 2026-10-06 (morning) — R-889 delivered, the held catalog fixes proven and pushed + +The full text of every row below: `git show e6cfe6d6:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** (P3) | CLOSED 2026-10-06 — FIXED, PROVEN ON 9202 AND PUSHED: nextcloud's probe needs a finished install | catalog `c5674ce` (live `3896cb0`): `body_contains: '"installed":true'`. Measured on 9202: `status.php` = `{"installed":true,"maintenance":false,…}` and the controller reads the app `running` with the new probe (`live-9202/nextcloud.txt`). The sweep found no other wizard that reads healthy unnoticed. | +| **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** (P4) | CLOSED 2026-10-06 — FIXED AND DELIVERED: a held update keeps the app's log and shows it | controller `0b93e1a` (v0.298.0: logs page + API) + `0f2eab7` (v0.299.0: the hold panel links to it, `09` decision 136). Delivered 2026-10-06 06:17Z: controller v0.299.0 (golden 0.299.0, floors for demo-hp, demo-felhom, tester-1; `audits/morning-after-2026-10-06/delivery/`). | +| **R-889** | **The disk percent the box shows and alarms on is `used / total`, not what `df` shows (`used / (used + available)`), so on a full system disk it reads ~5 points LOW — and the fill alarms (85 % / 95 %) fire that much late.** (P3) | CLOSED 2026-10-06 — FIXED AND DELIVERED (operator ruling, `09` §3 decision 138): every disk percent is df's | controller `bee2c2d` (v0.299.0): `system.DFUsedPercent` (used / (used + available)) serves the four places that computed one; tests on numbers measured on demo-hp 9201 + a source scan, red-proved (`audits/morning-after-2026-10-06/r889-red-proof.txt`). Cause MEASURED: the 5 % reserved for root was in the denominator (demo-hp `/mnt/sys_drive`: `stat -f` blocks 18016108, free 14150899, avail 13229302 → old 21.5 %, df 23 %). **Read back after delivery (hub page vs `df` in the guest):** demo-hp SSD 21 % → **23 %** (df 23 %), NVMe 6 % (df 7 %, df rounds up), demo-felhom SSD 1 % (df 1 %) (`r889-readback.txt`). The tile now says „Rendszer” (it measured the docker data volume, not /). Alarm levels unchanged — a full disk alarms a little earlier. Delivered 2026-10-06 06:17Z: controller v0.299.0 (golden 0.299.0, floors for demo-hp, demo-felhom, tester-1; `audits/morning-after-2026-10-06/delivery/`). | +| **R-776** | **[P3-LOW] Right-walking catalog apps need ONE setting to see each visitor since v0.286 (R-753); without it they keep the tunnel's one address (a stranger can still trip their per-address limits for everyone).** (P3) | CLOSED 2026-10-06 — FIXED, PROVEN ON 9202 AND PUSHED: four apps see each visitor behind the tunnel | catalog `462e66f` kimai, `ccc2be5` zipline, `5f7bf77` vikunja, `1d3af8e` nextcloud (live catalog `3896cb0`, 06:46Z). Measured on 9202 through the simulated tunnel, each with a control (the line removed, restart through the product): vikunja household 200 vs 429; zipline 200 vs 429; kimai ok vs THROTTLED; nextcloud attempts counted for the visitor 6 vs 0 — and a stranger changing the forged leftmost address was throttled every time (`audits/morning-after-2026-10-06/live-9202/`). | +| **R-585** | **[P3-LOW] Six event producers still send Hungarian only, so an English household can see one Hungarian line in some mails.** (P3) | CLOSED 2026-10-06 — FIXED AND DELIVERED: every customer-facing producer follows the household's language | controller `ca89e70` (v0.298.0) + `ff1758a` (v0.299.0); `local_api_endpoint_drift` is operator-only and English by design (the row's list was wrong about it). Delivered 2026-10-06 06:17Z: controller v0.299.0 (golden 0.299.0, floors for demo-hp, demo-felhom, tester-1; `audits/morning-after-2026-10-06/delivery/`). | + ## 2026-10-06 (morning) — installer 1.32.0 published The full text of every row below: `git show 32a15208:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 8c8f2eab..d9f1779d 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -117,7 +117,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-250** | Install & onboarding | P3 | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | -| **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** **2026-10-05 (burn-down night): items 1 and 4 FIXED on controller `main`** (`d7efa7b`): the menu says „Hibakeresés” (English „Debug”); the storage pages use the te-form (formal-form ceiling 17 → 14; 116 parity fixtures changed by exactly those bytes). Item 2 was already fixed. **Left:** items 5–6 (apps' own screens) and the rest. **2026-10-06: items 1 and 4 DELIVERED** (controller v0.298.0). **2026-10-06 (burn-down night, later): items 7–10 and 12 FIXED on controller `main`** (`e4774e6`, `2e5ef7e`): the disconnect time in local time; one banner per unplugged drive; „nézd meg”; the disk/memory/CPU/temperature banners in the household's language (`health.*` keys; the wire text to the hub unchanged and pinned); the last eleven counted formal forms are te-form (ceiling 14 → 0; 119 parity fixtures changed by exactly those bytes). **Left:** items 5–6 (apps' own screens); about 20 formal forms the gate does not count yet („írja be”, „adja meg”, „hozzon létre”, …) on the debug, storage, network-storage and security pages. Item 11 moved to its own row **R-889**. **2026-10-06 02:40: the bundle has NO formal form left** (controller `4c3c203`, unreleased): 49 sentences in the te-form; the gate's stem list widened (it counted 61 on the previous bundle, 0 now) with decoys and one listed third-person exception. **Left (narrowed):** Felhom-owned formal forms in Go and template literals the gate cannot read yet — the network-storage attach errors, escrow and share handlers, the single-copy notice, a settings refusal, the SMART mail, a fill-watch „Kérjük”, and the setup wizard pages; each moves into the bundle in the te-form. Items 5–6 are the apps' own screens (not ours). | — | — | CC | +| **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** **2026-10-05 (burn-down night): items 1 and 4 FIXED on controller `main`** (`d7efa7b`): the menu says „Hibakeresés” (English „Debug”); the storage pages use the te-form (formal-form ceiling 17 → 14; 116 parity fixtures changed by exactly those bytes). Item 2 was already fixed. **Left:** items 5–6 (apps' own screens) and the rest. **2026-10-06: items 1 and 4 DELIVERED** (controller v0.298.0). **2026-10-06 (burn-down night, later): items 7–10 and 12 FIXED on controller `main`** (`e4774e6`, `2e5ef7e`): the disconnect time in local time; one banner per unplugged drive; „nézd meg”; the disk/memory/CPU/temperature banners in the household's language (`health.*` keys; the wire text to the hub unchanged and pinned); the last eleven counted formal forms are te-form (ceiling 14 → 0; 119 parity fixtures changed by exactly those bytes). **Left:** items 5–6 (apps' own screens); about 20 formal forms the gate does not count yet („írja be”, „adja meg”, „hozzon létre”, …) on the debug, storage, network-storage and security pages. Item 11 moved to its own row **R-889**. **2026-10-06 02:40: the bundle has NO formal form left** (controller `4c3c203`, unreleased): 49 sentences in the te-form; the gate's stem list widened (it counted 61 on the previous bundle, 0 now) with decoys and one listed third-person exception. **Left (narrowed):** Felhom-owned formal forms in Go and template literals the gate cannot read yet — the network-storage attach errors, escrow and share handlers, the single-copy notice, a settings refusal, the SMART mail, a fill-watch „Kérjük”, and the setup wizard pages; each moves into the bundle in the te-form. Items 5–6 are the apps' own screens (not ours). **2026-10-06: items 7–10, 12 and the bundle's 49 formal forms DELIVERED** in controller v0.299.0. | — | — | CC | | **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | @@ -128,7 +128,6 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-612** | Apps & catalog | P3 | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** **— FIXED 2026-09-23 night (catalog `a5a729a`): 512M.** Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (`The seed command has been executed`), and wishlist then moved v0.66.0 → v0.67.1 with its test record. **Still open:** a failed first-boot seed is invisible to the box — nothing reads the seed's exit. | **READY — P3, narrowed to making a failed seed visible; owner: CC (catalog)** **Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains.** | — | — | CC | -| **R-613** | Apps & catalog | P3 | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. **— UPDATE NIGHT 2026-09-21:** the update night could not seed `uptime-kuma` for the same reason and left it out rather than faking it. **— FIXED 2026-09-23 night for uptime-kuma (catalog `a5a729a`):** `UPTIME_KUMA_DB_TYPE=sqlite` — the database choice is made, the wizard never appears, the real server starts (`/api/entry-page` → `entryPage`, `/metrics` 401), the database lands in `/app/data` (the backed-up volume). Red-proofed on 9202 through the product: before, the box read `running` over `setup-database`; after, the real server. (A first "after" run used a stale template — R-607's lag — and is kept.) **Still open:** the sweep for other templates whose probe passes on a setup wizard. | **READY — P3, narrowed to the sweep; owner: CC (catalog)** **Re-ranked 2026-10-03: P2->P3: uptime-kuma fixed; only a sweep of other templates remains.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** | — | — | CC | | **R-774** | Apps & catalog | P3 | **[P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers.** READ/MEASURED 2026-10-01: Karakeep's `smtp_mapping` (plaintext :2526, as Cal.com) was proven only with mail OFF (a fresh install boots) — 9202 has no hub, so the relay cannot be exercised there; the official mobile app ships Sentry crash reporting with a hard-coded DSN (`apps/mobile/app/_layout.tsx`, FIT.md). **Needs:** one password-reset mail from Karakeep on a hub-enabled box (demo-hp 9201, a throwaway install); and a sentence on the page about the phone app (copy, freeze). `audits/new-apps-2026-10-01/bench/karakeep-mail-off-boot.txt` | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | @@ -151,7 +150,6 @@ stopping line that lies. | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | | **R-618** | App updates | P4 | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. **CLOSED 2026-09-22.** All three fixed in one commit (`app-catalog-felhom.eu@793c4fb`): tandoor `8080 -> 80`, wger `80 -> 8000`, zipline `/api/health -> /api/healthcheck`. No `image:` line moved, so no `catalog_since` moved. **RED-PROOFED LIVE ON 9202 THROUGH THE PRODUCT, BOTH DIRECTIONS.** Before the fix, at the LIVE pin, all three read **`Nem egészséges` / `Not healthy`** on their own app page while docker reported every container healthy and each front door served a real page through the household's own route — tandoor 200 `Login / Sign In`, zipline 200 `Zipline`, wger 200 `wger Workout Manager`. The fix was applied through the REAL sync (`POST /api/sync` answered *frissítve: tandoor, wger, zipline*) and all three read **`Fut` / `Running`** at the next poll, with no redeploy and no restart. **AND THE EDGE THAT FAILED WAS RE-WALKED AND PASSED:** tandoor `2.6.13 -> 2.6.15` via the drill catalog ended **`done` at +41.1 s** with the seed read back through tandoor's own front door and both containers running with zero restarts — where the identical edge on 2026-09-21 entered `verifying` at +58.4 s and ended `failed` at **+361.9 s** with the app stopped. **Same app, same versions, same button; the only change is one port number.** tandoor's verdict moved `failed -> proven` and it is now on the live catalog. **THE GATE SHIPPED WITH IT:** `scripts/check-probe-matches-compose.py`, a `--fast` row in `catalog_gates.py`, comparing the probe against the SAME service's own compose healthcheck. The rule was read out of `healthprobe.go` rather than guessed: a wrong PORT refuses for every check type; a wrong PATH refuses only for `type: api` WITH an `expect` block and WARNS otherwise, which is why `home-assistant` is warned about and not convicted. Four red-proofs and five decoys, plus five more with PyYAML shadowed out (R-630's sibling problem — see below). **Residual, filed separately:** R-630 (paperless-ngx's probe can never run at all) and R-631 (five templates the static rule cannot judge). | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the controller-side idea — let `verifying` accept docker's own `healthy` before it stops a working app — was raised and never decided) — **CLOSED 2026-09-22 — three probes fixed, red-proofed live both ways, gate shipped with decoys; tandoor re-walked `failed -> proven`** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** The three wrong probes and the gate shipped (2026-09-22); what is left is whether `verifying` may accept docker's own `healthy` — that widens the update guard (R-635: a probe can be green on a broken app). Next: the ruling. | — | — | CC + operator | -| **R-621** | App updates | P4 | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. **FIXED in v0.262.0.** `failAndHold` now writes each service's log into `/hold-logs//compose-logs.txt` (`compose logs --no-color --tail 400`) **before** the `down` that destroys them, best-effort by design: a hold must never fail because its evidence could not be written. This is what `adventurelog` cost — nine migrations ran, the app never bound its port, and the only log that could have said why was gone before anyone looked. | **VERIFY** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fix shape also said to show the captured hold log on the app page's hold panel; the row does not say that was done) — **CLOSED 2026-09-22 — v0.262.0: the hold keeps the app's own log before stopping it** **2026-10-05 (burn-down night): FIXED on controller `main`** (`0b93e1a` — a held app's logs page and API show the log the hold kept; test + red-proof). Ships with the next controller release; close after delivery. Left: the hold panel itself does not show it. **2026-10-06: DELIVERED** in controller v0.298.0 (held app's logs page and API show the kept log). Left: the hold panel itself. **2026-10-06 (burn-down night, later): the hold panel FIXED on controller `main`** (`0f2eab7`): it says the app's own log from before the stop was kept and links to it (`09` §3 decision 136: a link, not 400 raw lines inline). **Ships with the next controller release — close after delivery.** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | ## Backup & restore — 34 rows (P2 8, P3 11, P4 15) @@ -197,7 +195,6 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-889** | Storage & devices | P3 | **The disk percent the box shows and alarms on is `used / total`, not what `df` shows (`used / (used + available)`), so on a full system disk it reads ~5 points LOW — and the fill alarms (85 % / 95 %) fire that much late.** Split out of R-516 item (11) on 2026-10-06 (burn-down night): the dashboard tile read „Rendszer (/) 61.8 GB / 68.7 GB (90%)" while `df` said 95 % (reserved blocks ignored), and „(/)" labels the data volume. Read in source: `controller/internal/system/mounts_linux.go` `GetDiskUsage` (`used = total - Bfree`, `UsedPercent = used/total`); the same number feeds `fillwatch` (`WarnUsedPercent` 85, `CritUsedPercent` 95) and the hub report. **Not a small fix:** moving to df's formula raises every box's figure at once and moves when the fill alarms fire (a box near 90 % today would alarm after the release). | **OPEN — filed 2026-10-06** | — | Decide: df's formula with the thresholds as they are (alarms come earlier — the honest reading), or the thresholds re-based; then fix the label and add a test with reserved blocks | CC | | **R-298** | Storage & devices | P3 | **The `/storage` page's unregistered list is filtered by `role==='user-data'`, so a drive that is also the backup target can never be registered from it.** `storage.html:363` routes anything not `user-data` into the read-only protected group with NO actions. On the rebuilt `demo-hp` the NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`/etc/pve/storage.cfg`: `dir: felhom-backup` → `/mnt/nvme-1tb`), so it renders locked. **This is the SECOND reason that page was empty** during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action | **READY (S) — NEW 2026-08-10** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Registering a drive the agent classes as the backup target conflicts with the agent's eject/decommission rule (403 for that role). First: is a drive holding both app data and backups a supported layout? | R-280 | Split the role gate: `user-data` keeps destructive actions; any mounted role may be REGISTERED | CC | | **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC | | **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC | @@ -222,7 +219,6 @@ stopping line that lies. | **R-747** | Security & access | P3 | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** | — | — | CC | | **R-763** | Security & access | P3 | **[P2-MEDIUM] On wger a stranger can make an account after the household's setup, and every anonymous visit to the dashboard creates a guest account.** MEASURED 2026-10-01 on 9202 (live template `82fff32`), found by checklist row 3.4: after the admin existed, a stranger with no dashboard session `POST /en/user/registration` → 302, signed in with it → 302, read the API → 200; two anonymous `GET /en/dashboard` raised the user count from 2 to 4 (wger's middleware `create_temporary_user`, `utils/middleware.py:69`). Settings read inside the app: `ALLOW_REGISTRATION True`, `ALLOW_GUEST_USERS True` (the image defaults; the template sets neither). `GET /en/user/demo-entries` as a stranger answered 500. wger is FIRST-ADMIN class 3 (a known default login, fixed by `after_install`), so it never got decision 47's sign-up lock — it was not in R-711's list. Every crawler visit adds a user row to the household's database. **Needs:** per decision 47, close it after the first admin: `ALLOW_REGISTRATION=False` and `ALLOW_GUEST_USERS=False` (env switches the settings read — measure that the admin can still add family members, row 3.7), proven on 9202 as a stranger. `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** **2026-10-06: PUSHED** to the live catalog (`c265b37`; wger stays hidden). | — | — | CC | | **R-775** | Security & access | P3 | **[P2-MEDIUM] Grimmory: a stranger's 5 wrong sign-ins lock EVERY visitor out of the web login for 15 minutes — so Grimmory was not published.** MEASURED 2026-10-01 on 9202 (drill catalog, v3.4.1, through traefik): after 5 wrong tries for `admin` every further sign-in answered 429 — the household's right password AND a different name — and stayed 429 for 10+ minutes of retries. Read in the jar: `AuthRateLimitService` — Caffeine `expireAfterWrite(ofMinutes(15))`, `MAX_ATTEMPTS 5`, keys `login:ip:` and `login:user:`; Spring `forward-headers-strategy: native` takes the address from X-Forwarded-For, and behind the tunnel every visitor is the tunnel container's address (R-753) — the wger shape (R-752), with no setting to change it. Everything else in the checklist passed (bench + box step v3.4.1 → v3.5.0, gate by its own probe, OPDS through traefik); two smaller findings for the publishing session: on a reinstall over the first install's kept books, a new upload was saved to the drive but not added to the library (`box/grimmory/reinstall-c1.txt`, not investigated); and the remove + restore round trip (2.5) cannot be shown on 9202 for a drive app — its backup lives on the scratch drive, which is not a registered drive (R-756). The template waits in `audits/new-apps-2026-10-01/wip/grimmory/`. **Needs (operator):** (A) publish with a sentence on the page that wrong guesses by others can lock the login for 15 minutes (MEASURED: during the lock an e-reader's OPDS feed still answered 200 with its own login, wrong 401 — `box/grimmory/opds-under-lock.txt`), or (B) wait until the box passes each visitor's real address (R-753). `audits/new-apps-2026-10-01/box/grimmory/throttle.txt` **-- 2026-10-01 (evening):** option B's precondition SHIPPED (controller v0.286.1, R-753): Grimmory's Tomcat RemoteIpValve walks from the right and counts `172.16.0.0/12` as a proxy (READ in source, Spring Boot 4.1.1 — not yet measured with Grimmory's own lock), so `login:ip:` becomes per visitor; `login:user:` still lets a stranger lock the public name `admin` 15 min. A third route was spiked and passed: Grimmory behind the permanent family gate with its e-reader paths excepted (R-780). Recommendation: publish behind the family gate if R-780 is built; otherwise B with a measured 3.6. **UPDATE 2026-10-02 — NARROWED, Grimmory PUBLISHED behind the family gate:** a stranger cannot reach Grimmory's web sign-in at all (6 tries through the simulated tunnel: the gate's 401, then the household signs in 200 — `audits/family-gate-2026-10-02/A/items.txt`), and the e-reader exceptions keep Grimmory's own login. The 2.5 round trip now WORKS on 9202 (`audits/family-gate-2026-10-02/B/box/life.txt`). **What is left:** a family member past the gate can still lock a NAME (the admin's) for 15 minutes with 5 wrong tries — hard-coded in Grimmory; and the reinstall-over-kept-books finding (`new-apps-2026-10-01/box/grimmory/reinstall-c1.txt`) is still not investigated. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P2→P3: Grimmory is now behind the family gate; only a family member can still lock a name.** | — | — | CC | -| **R-776** | Security & access | P3 | **[P3-LOW] Right-walking catalog apps need ONE setting to see each visitor since v0.286 (R-753); without it they keep the tunnel's one address (a stranger can still trip their per-address limits for everyone).** READ in source (`audits/visitors-2026-10-01/A/sweep/`): kimai `TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12`; zipline `CORE_TRUST_PROXY=true` + `CORE_TRUSTED_PROXIES=172.16.0.0/12`; vikunja `VIKUNJA_SERVICE_IPEXTRACTIONMETHOD=xff`; nextcloud `TRUSTED_PROXIES=172.16.0.0/12`; n8n `N8N_PROXY_HOPS=2` (optional). BookStack `APP_PROXIES` measured this session (see its catalog commit). **Needs:** per app, the setting + checklist 3.6 re-measured on 9202 through the simulated tunnel (the BookStack shape, `audits/visitors-2026-10-01/tools/bookstack_36.py`). Count readers (calibre-web, tandoor, wger) stay as they are — a fixed count is wrong for one of the two paths. | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-831** | Security & access | P3 | **The Hetzner storage API token (`HETZNER_TOKEN`, the storage project's token in `Secret/storagebox`) was printed into the 2026-10-03 session transcript** — CC read the gitignored `manifests/storagebox.secret.yaml` and its redaction pattern missed the quoted value. It can create, reset and delete Storage Box sub-accounts. Not rotated by the operator's choice (decision 73). **Rotation, whenever chosen (3 steps):** create a new token in the storage project in the Hetzner console → patch `Secret/storagebox` key `HETZNER_TOKEN` in `felhom-system` and `kubectl rollout restart deployment/hub` → delete the old token in the console. Rule for sessions: never print a file that holds secrets — read the one field needed. | **WAITING-ON-OPERATOR — rotation is his call** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | | **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | @@ -255,7 +251,6 @@ stopping line that lies. | **R-388** | Monitoring & notifications | P3 | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor | | **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Covering one-app deletions needs per-app snapshot counts from the controller and a per-tag threshold measured on real counts — two repos, a mechanism nobody has measured. | — | — | CC | | **R-521** | Monitoring & notifications | P3 | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`800b32c` — controller half: apps a lost drive stopped no longer mail app_start_failed one by one; test + red-proof). Ships with the next controller release; close after delivery. Left: the hub's cooldown outliving a recovery (F6/F7) and whether a household gets a mail for a lost drive (operator). **2026-10-06: controller half DELIVERED** (v0.298.0). | — | — | CC + operator | -| **R-585** | Monitoring & notifications | P3 | **[P3-LOW] Six event producers still send Hungarian only, so an English household can see one Hungarian line in some mails.** FOUND 2026-09-18 by localisation slice 3 Part B (R-558, controller v0.256.0), which converted 19 of the customer-facing producers and left these: `backup_failed`, `db_dump_failed`, `backup_integrity_ok`, `backup_integrity_failed`, `offbox_enlarge_blocked` and `local_api_endpoint_drift`. **Why they were left:** each receives its sentence already FINISHED from another package, so the key and its arguments no longer exist by the time the notifier sees it — converting them means changing their callers, not the notifier. The 15 operator-tier types are deliberately excluded and are NOT part of this row: the operator reads Hungarian. **Why it matters more than it looks: `offbox_enlarge_blocked` has no `customerMessages` entry on the hub**, so its raw sentence IS the household's mail rather than an extra line under a translated headline — for that one type an English household gets a wholly Hungarian mail, not a mostly-English one. **Fix shape:** push the key and its arguments down from each caller (the shape slice 2 release B already used for errors, `util.MsgError`), then add each to the `convertedProducers` table in `internal/notify/message_customer_test.go`, which is the list both language tests walk. | **READY - rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): PARTIAL on controller `main`** (`ca89e70`): `db_dump_failed`, the off-box backup failure and `offbox_enlarge_blocked` (which had no hub sentence — an English household got a wholly Hungarian mail) follow the household's language. **Left:** `backup_integrity_ok`/`_failed`, the app-stop backup failure, `local_api_endpoint_drift`. **2026-10-06: the three converted producers DELIVERED** (controller v0.298.0). **2026-10-06 (burn-down night, later): the rest FIXED on controller `main`** (`ff1758a`): `backup_integrity_ok`/`_failed` follow the household's language (Hungarian bytes pinned); the app-stop backup failure was sending the operator's ENGLISH sentence to households — it is now built per language (English byte-identical to the log line, pinned). `local_api_endpoint_drift` needed nothing: it is operator-only and English by design (the row's list was wrong about it). **All producers done; ships with the next controller release — close after delivery.** | — | — | CC | | **R-724** | Monitoring & notifications | P3 | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** | — | — | CC | | **R-177** | Monitoring & notifications | P4 | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the same missing operator door as R-314 and R-279. | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator |