09 decisions 180-184 (the 09:04/09:07 rulings); R-904 (Cloudflare, deferred) opened; R-900 closed (FAQ published); R-901/R-762/R-304 ruled; 131 -> 131
gates / gates (push) Successful in 4m0s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-08 09:47:44 +02:00
parent b2dce901b2
commit 97c260c56b
3 changed files with 26 additions and 4 deletions
+8
View File
@@ -26,6 +26,14 @@
---
## 2026-10-08 (afternoon) — the morning's answers built
The full text of every row below: `git show b2dce901b2:documentation/backlog/OPEN-ITEMS.md`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-900** | **The website's FAQ said the household's data is not with a third party, while copies and traffic go to processors.** (P2) | CLOSED 2026-10-08 — PUBLISHED: the GDPR answer on `gyik.html` and `en/faq.html` (visible text and structured data) now names the box, the encrypted copies at Hetzner in the EU, and Cloudflare's view of remote-access traffic; text approved by the operator in chat (`09` §3 decision 180) | website commit `b2dce901`; live read-back 2026-10-08: the new sentence 2× on each page (visible + JSON-LD), the old „nem harmadik félnél" 0×; rest of the site searched (ASCII fragments, positive control) — no other page makes the claim. Left for R-813: the contact form's consent line „harmadik félnek nem adjuk ki", which the consent draft already replaces. |
## 2026-10-07 — the operator's answers on the night designs
The full text of every row below: `git show 5ba1702fcf:documentation/backlog/OPEN-ITEMS.md`.
+4 -4
View File
@@ -128,7 +128,7 @@ stopping line that lies.
|---|---|---|---|---|---|---|---|
| **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** **2026-10-06 night: not started — the row needs the operator's word on the Hungarian format (decimal comma, date style) before any code.** | — | — | CC |
| **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC |
| **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. **-- 2026-10-08 (day):** the gunicorn half PROVEN ON THE BENCH (9401, `wger/server:2.7`, limit 384M): `WGER_USE_GUNICORN=True` + `WEB_CONCURRENCY=2` (the image runs gunicorn only on that switch and passes no `-w`; the container log shows 2 × „Booting worker"); anon peak 166–170 MiB = 43–44 %, oom_kill 0, restarts 0 over a 606 s watch (11,584 requests; harness verdict `proven`); login 200, all three CSS links 200 via `wger-files`, a photo uploaded 201 and read back 200; an unknown `/static/` file 404 (control). **Not done: 9202** — wger is hidden, the live catalog refuses a hidden deploy, and the drill-catalog write that un-hides it there was refused by the permission check (stop, per the brief). **Nothing pushed**; the tested definition is in the evidence. **No ladder step applies:** the change moves no image (from == to), the ladder writer refuses a same-digest re-test, and `09` §5.4 renders the catalog template when the images equal the pinned ones, so the compose change reaches an installed wger at its next `up -d`. Still true: the collected style files (283 MB) go into every wger backup. `audits/day-2026-10-08/r762/`. | **READY — narrowed 2026-10-08: the bench proof is done; the 9202 proof and the push wait for the operator's word on the drill-catalog write.** wger stays `lifecycle: hidden`. | operator (drill-catalog write) | Operator allows the drill-catalog write (fast-forward the drill to live main, un-hide wger there, point 9202 at it) → prove on 9202 (login + CSS + photo 200, memory, 10-min watch) → push the two env lines (no ladder entry) → close | CC |
| **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. **-- 2026-10-08 (day):** the gunicorn half PROVEN ON THE BENCH (9401, `wger/server:2.7`, limit 384M): `WGER_USE_GUNICORN=True` + `WEB_CONCURRENCY=2` (the image runs gunicorn only on that switch and passes no `-w`; the container log shows 2 × „Booting worker"); anon peak 166–170 MiB = 43–44 %, oom_kill 0, restarts 0 over a 606 s watch (11,584 requests; harness verdict `proven`); login 200, all three CSS links 200 via `wger-files`, a photo uploaded 201 and read back 200; an unknown `/static/` file 404 (control). **Not done: 9202** — wger is hidden, the live catalog refuses a hidden deploy, and the drill-catalog write that un-hides it there was refused by the permission check (stop, per the brief). **Nothing pushed**; the tested definition is in the evidence. **No ladder step applies:** the change moves no image (from == to), the ladder writer refuses a same-digest re-test, and `09` §5.4 renders the catalog template when the images equal the pinned ones, so the compose change reaches an installed wger at its next `up -d`. Still true: the collected style files (283 MB) go into every wger backup. `audits/day-2026-10-08/r762/`. | **READY — narrowed 2026-10-08: the bench proof is done; the 9202 proof is ALLOWED since the operator's ruling 2026-10-08 09:04 (`09` §3 decision 182: CC may write the shared test catalog).** wger stays `lifecycle: hidden`. | — | Operator allows the drill-catalog write (fast-forward the drill to live main, un-hide wger there, point 9202 at it) → prove on 9202 (login + CSS + photo 200, memory, 10-min watch) → push the two env lines (no ladder entry) → close | CC |
| **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: not a catalog fix** — the setgid chain is set by the controller's FileBrowser setup; a controller row. | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC |
| **R-577** | Apps & catalog | P4 | **[P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them.** FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (`launcher_shared`, `launcher_share_password`) deliberately do NOT, and `TestGuestSharePagesHaveNoGlobe` pins that so it stays a decision rather than an oversight. **Why it is the operator's and not CC's:** a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The `felhom_lang` cookie already built would fit them exactly (display-only, their own browser, never the household's setting). **Fix shape, if the operator says yes:** add `{{template "lang_globe" .}}` to both shells with the anonymous form, and one render case per page per language. | **READY - rank P3-LOW; owner: operator (the decision), CC (the change)** **Re-ranked 2026-10-03: P3->P4: feature decision for the operator; Hungarian default works today.** | — | — | operator |
| **R-707** | Apps & catalog | P4 | **[P2] 37 apps still start with a login a stranger can take (`09` §3 decision 45).** Audit of all 53 apps: `app-catalog-felhom.eu/FIRST-ADMIN.md` (class, fix route, status, measured or read). Open: **3 hard-coded defaults** — calibre-web (`admin / admin123`, measured working on demo-hp and 9202; its own `cps.py -s` route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (`changeme@example.com / MyPassword`), wger (`admin / adminadmin`); **34 open first-run screens** (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). **Stale notes:** romm's `default_creds` `admin / admin` answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. **Measured on demo-hp 2026-09-28 (read-only):** bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via `after_install:`, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). **2026-09-29 (controller v0.280.0, catalog `d0e7e2e`):** every class-3 app fixed — mealie, wger, calibre-web by `after_install` (calibre-web with the new `password:24:special`), proven on 9202 fresh installs (`audits/login-gate-2026-09-29/D/`); the setup gate (decision 46, spike PASSED) built and live on immich, n8n, audiobookshelf (probes measured) and uptime-kuma (button) (`…/C/`); romm's and zipline's stale notes removed. **Left: 30 class-4 apps** — gate each (probe measured on 9202 where one exists — 11 upstream candidates listed in `…/B/B-VERDICT.md` §3; the button otherwise). **2026-09-29 afternoon (controller v0.281.0, catalog `6faf432`):** 28 more class-4 apps gated — 32 of 34 — each proven on 9202 (`audits/gate-rollout-2026-09-29/`B): stranger → gate page / 401, household reached the first-setup screen, the gate opened (9 by a measured probe, the rest by the press), the app answered after. seerr, outline, rallly: gated, their opening needs a media server / e-mail (not proven). **Left:** wanderer (R-714); plant-it is not installable. | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — seerr, outline and rallly are gated, but the gate OPENING is not proven (needs a media server / e-mail)) — **CLOSED — 2026-09-29 (the rest → R-714)** | — | — | CC |
@@ -150,7 +150,7 @@ stopping line that lies.
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-08 — owner Viktor.** (a) DONE with the operator's yes in chat: `notify_failure` now mails admin@felhom.eu through Resend; proven by one test mail that reached the inbox (`audits/day-2026-10-08/r232/`; no backup was started). (b) partly: the hub database leaves DooPlex nightly to ep0 (R-173); everything else stays on the box. (c)–(h) unchanged. **READY** for the rest | — | — | operator |
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **WAITING-ON-OPERATOR — narrowed 2026-10-08.** Most of the „told the code is wrong" defect closed with R-311 (2026-08-13: the agent tries the retained packages, 422 „correct, earlier package"). The rest fixed on main today: the agent said „wrong code" also when it had NOT tried every earlier package (withheld rows, the caps, a malformed package, an unreadable list) — now 424 `older_unchecked`, and the screen says „we do not know whether your code is wrong … contact support" in Hungarian and English (agent `91b9405`, controller `75b3b39`; ships tomorrow). The route to really serve an old copy is R-312's decision (operator-only); the one-page design and two questions: `audits/day-2026-10-08/design-R-304.md`. | R-198, R-199, R-224, R-241 | Operator answers the two questions in the design (the promise wording; the operator mail, option C) | operator + CC |
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — ruled 2026-10-08 09:04 (`09` §3 decision 183): the operator gets a mail when a household's code opens, or may open, an older escrow package (design option C).** The honesty fix is on main (agent `91b9405`, controller `75b3b39`, ships tomorrow); design `audits/day-2026-10-08/design-R-304.md`. The promise-wording question (design Q1) was not answered: the wording stays as it is. | R-198, R-199, R-224, R-241 | Build the operator mail (controller event + hub allowlist and operator-only), ships with the next releases | CC |
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC |
| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC |
| **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** **2026-10-07 07:58: kept open for later (`09` §3 decision 166).** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC |
@@ -209,6 +209,7 @@ stopping line that lies.
| **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator |
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
| **R-904** | Security & access | P4 | **Cloudflare can read every household's app traffic; replacing it with our own relay is a later item.** Facts (reviewer discussion 2026-10-08; `01`): app traffic and the dashboard reach the box through the Cloudflare Tunnel (`01` §5 trust table, rows end-user ↔ apps and customer ↔ controller UI; §7), and the tunnel's public end is Cloudflare's edge, where TLS ends — so Cloudflare can technically read that traffic (the FAQ says so since 2026-10-08, R-900; „TLS ends at the edge" is not written in `01` — add it there). What Cloudflare gives today, free: inbound reach with no router setup, the CGNAT answer (`01` §4, §7); certificates (`01` §7, the free tier covers one level below a zone); the geo-WAF the hub enforces (`01` §5 last row, §7); flood protection (not in `01`). The alternative named: our own EU relay over WireGuard with TLS passthrough by SNI, certificates on the box, the geo-block on the relay. Its costs: one more machine the operator keeps up, and a single point of reach for every box; weaker flood protection; about a week of work after a spike. Operator ruling 2026-10-08 09:07 (`09` §3 decision 184): a later item. | **DEFERRED — after the first customers (operator ruling 2026-10-08 09:07)** | — | A spike after the first customers (the relay's reach, cost and flood behaviour, measured) | operator |
## Box system & updates — 6 rows (P2 1, P3 5)
@@ -267,8 +268,7 @@ stopping line that lies.
| **R-789** | Business & legal | P2 | **[P2-MEDIUM] Tandoor's licence is AGPL-3.0 WITH the Commons Clause: it forbids selling "a product or service whose value derives, entirely or substantially, from the functionality of the Software" — fees for hosting or support included.** READ 2026-10-02 at 2.6.15 (`audits/licences-2026-10-02/TABLE.md`). Felhom charges for installing and caring for the household's apps; whether that value comes "substantially" from Tandoor is the question. **Needs (operator):** keep (the fee is for the box, not Tandoor), hide for new installs, or ask the authors. Nothing changed meanwhile. **RULED 2026-10-02 afternoon (`09` §3 decision 66):** treated like SparkyFitness — stays offered; the operator asks the authors for written permission; without it before the first paying customer Tandoor is hidden (`lifecycle: hidden`). STATUS "Before the first paying customer". | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (the request; trigger: first paying customer)** | — | — | operator |
| **R-802** | Business & legal | P2 | **[P2-MEDIUM] A lawyer reviews the non-OSI licence list before the first paying customer.** Operator ruling 2026-10-02 (`09` §3 decision 66): Tandoor (R-789), SparkyFitness (R-784), Emby, n8n, Plex (kept — R-790..R-792), the EE/BUSL parts (R-793), redis 7.4 (R-794); the table is `audits/licences-2026-10-02/TABLE.md`. STATUS "Before the first paying customer". | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (trigger: first paying customer)** | — | — | operator |
| **R-813** | Business & legal | P2 | **[P2] The website collects personal data but publishes no privacy notice, no terms and no imprint.** CHECKED 2026-10-03 (read-only): `website/` holds nine Hungarian pages and one English page; none is an ÁSZF, an adatkezelési tájékoztató or an impresszum, and no page links to one (ASCII-fragment search `aszf`, `adatkezel`, `impresszum`, `impressum`, `privacy` over `website/`; positive control: the same search finds `adatkezel` in the contact form). The contact form makes the visitor tick a data-processing consent (`website/kapcsolat.html:123-128`) whose text names no controller, no retention and no rights, and links nowhere. The papers around it — contract, data-processing agreement, billing — are the intention **R-809** in `ROADMAP.md`. **2026-10-08 (website refresh): the English twins are PUBLIC since today** (`felhom.eu/en/…`, operator choice B) — **the legal pages are needed in English too**, or an English line saying the legal texts are Hungarian, linked from every English page. The English contact form translates the consent text as it is. The refresh CUT the FAQ's „önálló modell" (no row backs it, brief Part A) and LEFT the GDPR answer („nem harmadik félnél") for R-900; the English FAQ carries the same answer. | **WAITING-ON-OPERATOR — the operator writes or commissions the texts; CC drafts on request; owner: operator** **2026-10-08: first drafts written (not published, pending lawyer review R-802): `documentation/legal/` — `DRAFT-aszf.md`, `DRAFT-adatkezelesi-tajekoztato.md`, `DRAFT-impresszum.md`, `DRAFT-kapcsolat-hozzajarulas.md` (proposed consent text + link). Every fact the operator supplies is a `[[…]]` placeholder; the lawyer list and the guesses are at the end of each draft. Owner stays operator.** | — | — | operator |
| **R-900** | Business & legal | P2 | **The website's FAQ says the household's data is not with a third party, while the system sends copies and traffic to processors.** FOUND 2026-10-08 while drafting the privacy notice (R-813): `website/gyik.html` (GDPR answer, „az adataid a saját infrastruktúrádon vannak — nem harmadik félnél") — but the encrypted off-site copies go to Hetzner (Storage Box, ep0), app traffic passes Cloudflare's tunnel (TLS ends at its edge), box reports and on-request log tails go to the hub, the hub's mails go through Resend. The drafting session also read an offer of an „önálló modell" (remote access removed) in the FAQ; after the same day's website refresh that phrase no longer appears (ASCII and accented search, 0 hits) — re-check the refreshed text against `07` §2 (operator root access is part of the product). A promise the product makes is the operator's to reword. Lawyer list and processor table: `documentation/legal/DRAFT-adatkezelesi-tajekoztato.md` §11. | **WAITING-ON-OPERATOR — owner: operator** | R-813 | Reword the FAQ answers to match the privacy notice (or decide the product changes); not changed by CC — a promise to customers | operator |
| **R-901** | Business & legal | P2 | **Two kinds of a household's data outlive the deletion of the customer, and no document says when they go.** FOUND 2026-10-08 (R-813 drafting), read in source: the hub keeps `events` and `notification_log` on purpose after a customer is deleted (`hub/internal/store/customer_delete.go:24-26`, „the audit trail outlives every lifecycle tier") and nothing prunes `notification_log`; DooPlex's copy of ep0 (`ep0-copy`) pulls with `remove-vanished false` and prunes only to keep-weekly 8 (`runbooks/ep0-datastore-copy.md`), so a deleted customer's last 8 weekly encrypted whole-guest copies stay on DooPlex with no end date. A privacy notice cannot promise deletion until this is decided. | **WAITING-ON-OPERATOR — owner: operator** | R-813 | Decide the retention after deletion (audit rows: how long; ep0-copy: remove a deleted customer's namespace, by hand or by job) and write it into the privacy notice | operator |
| **R-901** | Business & legal | P2 | **Two kinds of a household's data outlive the deletion of the customer, and no document says when they go.** FOUND 2026-10-08 (R-813 drafting), read in source: the hub keeps `events` and `notification_log` on purpose after a customer is deleted (`hub/internal/store/customer_delete.go:24-26`, „the audit trail outlives every lifecycle tier") and nothing prunes `notification_log`; DooPlex's copy of ep0 (`ep0-copy`) pulls with `remove-vanished false` and prunes only to keep-weekly 8 (`runbooks/ep0-datastore-copy.md`), so a deleted customer's last 8 weekly encrypted whole-guest copies stay on DooPlex with no end date. A privacy notice cannot promise deletion until this is decided. | **READY — ruled 2026-10-08 09:04 (`09` §3 decision 181): audit rows (`events`, `notification_log`) of a deleted customer kept 1 year, then deleted; the customer's `ep0-copy` namespace removed within 30 days; both go into the privacy-notice draft.** Close only when both are live. | R-813 | Build the hub's daily deletion (ships with the next hub release) and the DooPlex `ep0-copy` removal job (written; run needs the operator's word); write both times into the privacy-notice draft | CC |
| **R-89** | Business & legal | P4 | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC |
| **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC |