Backlog triage Part A: four roadmap items R-808..R-811 (OS updates, legal/business, independence spike, dashboard 2FA); findings R-812 (no OS security updates) + R-813 (no legal pages); 00 gap rows + new §H; CONTEXT records the 2026-10-03 request
gates / gates (push) Successful in 28s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-03 09:06:39 +02:00
parent 93052884b2
commit 9e2786c907
4 changed files with 35 additions and 1 deletions
+9
View File
@@ -16,6 +16,15 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **Operator request 2026-10-03 — recorded before the work (backlog triage).** (1) Four roadmap items: box OS security
> updates (R-808, finding R-812), legal pages and business papers (R-809, finding R-813), independence — "if the household
> leaves Felhom, or Felhom stops" (R-810, spike), a second login step for the dashboard (R-811). (2) Finished rows leave
> `OPEN-ITEMS.md` for `CLOSED-ITEMS.md` **in the same commit that closes them**, and a gate refuses a finished row left in
> the open register. (3) Every open row carries ONE category and ONE severity, and new rows are filed into their category.
> (4) `ROADMAP.md` is cleaned. (5) The operator then decides what is worked on next. **Reviewer's defaults, operator may
> reverse:** the severity scale (P1 now · P2 before the first paying customer · P3 during the first customers · P4 later)
> and the eleven categories, both written at the top of `OPEN-ITEMS.md`.
> **2026-10-02 (afternoon) — the persistence re-sweep (R-801/R-788 CLOSED), controller v0.288.0 (R-800), golden 0.288.0.**
> Rulings `09` §3 66 (licences: Tandoor like SparkyFitness; Emby/Plex/n8n kept; lawyer review before the first paying
> customer, R-802) and 67 (userdata stays on every remove; dialog + result name it — `07` §6.5). The gate now sends
@@ -194,6 +194,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **The box tells visitors apart — a stranger's wrong passwords lock only the stranger** (dashboard, setup gate, apps that read the visitor from the right) | controller v0.286.1 (`felhom-tunnel`, traefik trust of cloudflared only, `clientaddr.go`), catalog router resets | **PARTIAL** | `audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt` (real tunnel: the dashboard counter keyed on the visitor `37.191.56.193`, locked after 5 while forging a new X-Forwarded-For each try); `A/L1-9202-live.txt` (simulated tunnel: the household from another address in at once); `A/bookstack-3.6.txt` | The second OUTSIDE address on the real tunnel was refused by Cloudflare before the box → R-779; right-walking app settings → R-776; Emby/Jellyfin LAN rights → R-777 |
| **The family gate — a permanent door with each family member's own login in front of an app** (Grimmory with its e-reader exceptions, MeTube with none) | controller v0.287.0 (`internal/family`, `stacks/family_gate.go`, the Család card), catalog `family_gate:` + `check-family-gate.py`, golden 0.287.0 | **PROVEN** on 9202 through the product: a stranger reached nothing (36 requests, LAN + simulated tunnel), members with their own logins, reset/remove/logout end access at the next request, a stranger's guesses lock only the stranger, each exception keeps the app's own login and a look-alike stays gated, the controller down → error never the app; survives an update and a remove + restore. Demo boxes on 0.287.0. Not yet: a household's phone through the real internet. | `audits/family-gate-2026-10-02/A/items.txt`, `B/box/` | R-780 closed; R-775 (narrowed), R-796, R-797 |
| Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 |
| A second login step for the dashboard (a TOTP code or a passkey) | — | **MISSING** | — | One password, one bcrypt hash (`controller/internal/web/auth.go:37-44`) → R-811 (added 2026-10-03) |
| The household can leave Felhom, or outlive it — the box runs without the hub, the household owns its domain, tunnel and off-site account, and can export everything | — | **MISSING** (as a written answer) | the LOST-hub half only: `_recovery-inventory-2026-07-28.md` §D2.4, `07` §8 row 11b | Leaving and hand-over are answered nowhere → R-810 (spike, added 2026-10-03) |
| WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` |
| The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe |
| The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing |
@@ -230,5 +232,17 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03) |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
## H. Business & legal (outside the platform — added 2026-10-03)
> Not platform capabilities, but what must exist before a household pays. Listed here so the coupling rule
> (every roadmap item names the row it flips) has a row to point at.
| Scenario | Components | Status | Evidence | Gap / roadmap |
|---|---|---|---|---|
| Website legal pages: ÁSZF, adatkezelési tájékoztató, impresszum | website | **MISSING** | `website/kapcsolat.html:123-128` asks for data-processing consent and links to no notice | → R-813 (finding), R-809 |
| Customer contract and data-processing agreement | — | **MISSING** | — | → R-809 |
| Billing and invoicing | — | **MISSING** | — | → R-809 |
+2
View File
@@ -919,6 +919,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-805** | **[P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): Grimmory's `/app/data` bind held 0 files and read CLEAN (its state is in MariaDB); komga, paperless-ngx, radarr and sonarr also carry empty binds that are not judged. **Needs:** decide whether an empty bind after the exercise is UNDETERMINED like a named volume, with a decoy. | **READY — rank P3-LOW; owner: CC** |
| **R-806** | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): both reached data only through their fixture seed. **Needs:** read the `loadbalancer.server.scheme` label (upgrade_boxport already does); find why gramps-web's first GET got no answer. | **READY — rank P3-LOW; owner: CC** |
| **R-807** | **[P3-LOW] 13 apps stay UNDETERMINED because their seeds write only to the database, so their upload / media / cache / redis volumes stay empty.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): claper, crafty-controller, dawarich (redis), docmost, gramps-web, immich (ML cache), outline (+redis), sparkyfitness, tandoor, vikunja, wger, wishlist, zipline — plus plex (no seed route: a plex.tv claim token) and wanderer (no fixture; unhealthy under the gate). None is BROKEN: nothing was written outside a preserved folder. **Needs:** per app, a seed that uploads one file (or the volume named n/a with a reason: a redis cache, an ML model cache). | **READY — rank P3-LOW; owner: CC** |
| **R-812** | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **READY — design first (R-808); owner: CC (spike, design) + operator (go)** |
| **R-813** | **[P2] The website collects personal data but publishes no privacy notice, no terms and no imprint.** CHECKED 2026-10-03 (read-only): `website/` holds nine Hungarian pages and one English page; none is an ÁSZF, an adatkezelési tájékoztató or an impresszum, and no page links to one (ASCII-fragment search `aszf`, `adatkezel`, `impresszum`, `impressum`, `privacy` over `website/`; positive control: the same search finds `adatkezel` in the contact form). The contact form makes the visitor tick a data-processing consent (`website/kapcsolat.html:123-128`) whose text names no controller, no retention and no rights, and links nowhere. The papers around it — contract, data-processing agreement, billing — are the intention **R-809** in `ROADMAP.md`. | **WAITING-ON-OPERATOR — the operator writes or commissions the texts; CC drafts on request; owner: operator** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
+9
View File
@@ -32,6 +32,15 @@
---
## New 2026-10-03 — the operator's four items (intentions; severity per `OPEN-ITEMS.md`'s scale)
| ID | Item | Size | Status | Notes / map rows flipped |
|----|------|------|--------|--------------------------|
| R-808 | **[P2] Box system security updates.** Goal: every box receives operating-system security patches on a schedule, and a failed update is undone. Why: a box lives in a home for years and today keeps the packages it was installed with — nothing updates the Proxmox host, the guest's Debian or its Docker engine (the finding is **R-812**). Scope: (1) regular security patches, host and guest; (2) reboots and their timing — inside the night window, never across a backup; (3) host kernel updates (a kept fallback boot entry); (4) Docker engine updates in the guest (an engine restart stops every app — quiesce like a backup); (5) the Proxmox MAJOR upgrade path (PVE 9 → 10) as its own later step, drilled first; (6) how the household and the operator are told (an event, a dashboard line); (7) how a failed update is undone (host: the vzdump/PBS copy plus the boot entry; guest: a snapshot before the run). Spike first: what the agent may run under its sudoers fence, and what a half-applied `apt` run leaves behind. | L | idea — 2026-10-03 (operator request) | Flips `00` §G *"Box operating-system security updates"* (MISSING, added 2026-10-03). Finding half: **R-812** in `OPEN-ITEMS.md`. |
| R-809 | **[P2] Legal pages and business papers before the first paying customer.** Goal: Felhom may legally take money from a household. Pieces: the website's ÁSZF, adatkezelési tájékoztató and impresszum (none exist — finding **R-813**); the customer contract; a data-processing agreement (Felhom monitors boxes and holds encrypted off-site backups, so it processes household data); billing and invoicing; the lawyer's licence review already on STATUS's list (**R-802**). Connects to the old **R-11** rulings (contact channel — RULED 2026-07-21; the tester agreement — never written). Owner: **operator**; CC drafts a text on request. | M | idea — 2026-10-03 (operator request) | Flips `00` §H rows *"Website legal pages"*, *"Customer contract and data-processing agreement"*, *"Billing and invoicing"* (MISSING, added 2026-10-03). Finding half: **R-813**. |
| R-810 | **[P3] Independence — "if the household leaves Felhom, or Felhom stops".** Goal: a written answer the data-sovereignty pitch can point at. Questions: does the box keep working without the hub (partly answered — `architecture/_recovery-inventory-2026-07-28.md` §D2.4 and `07` §8 row 11b cover a LOST hub: a day is invisible, a week loses alarms, resets and convergence); who owns the domain (the customer, `01` §7), the Cloudflare tunnel and the off-site storage account; how a household exports everything; what a hand-over to the household or another provider looks like. Status **spike**: nothing is known to be wrong; no document answers the leave/hand-over half. | M | idea — spike, 2026-10-03 (operator request) | Flips `00` §E *"The household can leave Felhom, or outlive it"* (MISSING, added 2026-10-03). |
| R-811 | **[P3] A second login step for the dashboard.** Goal: a stolen or guessed password alone does not open the dashboard, which controls the whole box and is on the internet. Today: one password, one bcrypt hash (`felhom-controller/controller/internal/web/auth.go:37-44`). Options: a TOTP code or a passkey; recovery when the phone is lost modelled on the existing reset-code and recovery-code designs. Relates to **R-15** (member accounts) and the family gate (`09` §3 decisions 63–65): the same door should carry both. | M | idea — 2026-10-03 (operator request) | Flips `00` §E *"A second login step for the dashboard"* (MISSING, added 2026-10-03). |
## P1 — closed-alpha blockers
| ID | Item | Size | Status | Notes / map rows flipped |