diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 664a8a94..8ea25443 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -303,7 +303,7 @@ stopping line that lies. | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | -## Box system & updates — 16 rows (P2 3, P3 10, P4 3) +## Box system & updates — 17 rows (P2 3, P3 11, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -311,6 +311,7 @@ stopping line that lies. | **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC | | **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **READY — design first (R-808); owner: CC (spike, design) + operator (go)** | — | — | CC + operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | +| **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | | **R-121** | Box system & updates | P3 | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-190** | Box system & updates | P3 | **A storage ACL that demonstrably WORKED in the morning was gone by mid-morning, and nothing recorded its removal.** On demo-felhom, a `vzdump` by `felhom-agent@pve!agent` with `--storage felhom-backup` completed **OK at 04:44:50 CEST 2026-08-03** (task log read in full). From **09:24:56** the same path returned `HTTP 403 … missing privilege Datastore.Allocate at /storage/felhom-backup`, six times through the day, until the grant was re-applied by hand at 18:54. By ~14:50 `pveum acl list` showed **no row at all** for that path | **NARROWED** — **MITIGATION SHIPPED 2026-08-04** (agent **v0.124.0 → v0.124.1**) — **MECHANISM STILL OPEN** | — | **Why this is not just R-185 restated:** R-185's mechanism (the installer's Scenario-F arm resolves a pre-existing target without granting) explains a box that NEVER had the grant. This box HAD it and lost it, inside five hours, with the machine up throughout. **Ruled out, each by measurement:** a host reinstall (`uptime` = 12 days); any `pveum`/ACL/`user.cfg` activity in syslog between 04:00 and 10:00 (none); any ACL entry in `/cluster/log` (none). **Correlated, not established:** `host_leaf_changed` at 09:15 and `controller_started` at 09:19 — guest 9201 was reprovisioned nine minutes before the first 403. PVE removes ACLs at `/vms/` when a guest is destroyed (`AccessControl::remove_vm_access`, the F-LEAK mechanism); whether any path can take a `/storage/` row with it has NOT been established and is the first thing to check. **Why it matters more than the grant did:** a permission that can vanish silently makes every ACL-based guarantee on these hosts provisional, and the agent's new store-grant probe (v0.123.0) now detects the STATE but says nothing about the TRANSITION. **Worth pairing with:** whether the probe should report a grant it once had and no longer has as a distinct, louder signal than one it never had **THE ROW NOW REFLECTS THE MITIGATION, NOT THE CAUSE — stated plainly because the two are different things.** The box is resilient; the loss is still unexplained. **Mitigation:** when the store-grant probe finds the grant absent, the agent runs the EXISTING root wrapper (`felhom-backup-target-apply grant `) and **re-reads once** to confirm — the pbsdr R-22 self-grant shape, including its restraint. **No new privileged surface:** the sudoers vector `grant *` already covers any storage id (confirmed in `configs/felhom-agent.sudoers`, not assumed), and the verb already grants BOTH user and token. The verb existed, was permitted, and had only ever been called at storage CREATION — the *built but never wired* shape in a verb rather than a seam, this project's seventh instance. Bounded at one attempt per tier per hour (a storage can be unreadable for reasons an ACL cannot fix; re-granting every cycle is a repair loop wearing a fix's clothes). **THE RECORD IS THE HALF THIS ROW IS ABOUT, and v0.124.0 got it wrong in production while every unit test passed.** It reported degraded for *one cycle* — meaning the probe call that repaired. But `probeAll` is invoked INDEPENDENTLY by the self-check log and by the collector building a host-report: on the box the repairing call was the log's (`09:39:34`, journal shows the repair and `degraded=1`) and the report three seconds later found the grant present and sent **`ok`**. The agent's journal had the record, the hub had nothing, and the operator would have learned nothing — the exact silence this row exists for, re-created inside its own mitigation. **v0.124.1** replaces it with a latch on TIME (20 min > the 900 s report interval), so at least one report must carry it. **PROVEN LIVE, twice, on demo-felhom** (grant deleted by hand, both rows): agent logs `store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190) target=felhom-backup privilege=Datastore.AllocateSpace action="felhom-backup-target-apply grant felhom-backup" confirmed_by=re-read`; the ACL rows return; and on v0.124.1 the **host-report at 08:00:30Z carried `status=degraded`** with the explanation and the hub raised `agent_capability_degraded` and **e-mailed the operator** at 08:00:40. Nothing new was built to carry it — the hub's existing ok→degraded→ok edge is the channel, and the text rides `Feature` because that is the field the hub interpolates into the e-mail (`Reason` does not travel). **PART 3 — the single bounded pass at the mechanism, with the negatives named.** The lead §8.6 nominated is real as a CLASS and is documented in our own installer: *"`pveum user token remove` purges the token's ACL, so re-applying post-rotate is mandatory"*. **It does NOT fit this box.** A rotation purges ALL of the token's ACLs and mints a NEW secret; demo-felhom's token still authenticates with the same secret (`--selftest` OK), it retained its other three storage grants throughout, and only `felhom-backup` was refused. No installer run is evidenced (no 2026-08-03 install log; host uptime 12 days at the time). Previously ruled out and unchanged: a host reinstall, any `pveum`/ACL/`user.cfg` activity in syslog 04:00–10:00, any cluster-log ACL entry. **Ruled out on THIS box; NOT ruled out fleet-wide** — any installer run still purges and re-grants only the hardcoded `PVE_STORAGES` set, though installer 1.24.0's reuse-arm fix now re-grants the backup target on that path. **A NEW OBSERVATION FROM THE LIVE RUNS, relevant to the timeline:** PVE **caches permissions** — after deleting both ACL rows the probe still read the privilege as present for ~40 s in one run and ~16 min in another. Detection is only as prompt as that cache, and a cache expiry could equally explain why a box kept working for hours after a grant was removed. → **R-194**. **The alert pair CLOSED on its own at 10:30:40** (`degraded → ok`, `agent_capability_recovered`) once the 20-minute latch expired — one lost grant, one e-mail, one recovery, nothing further. | CC | diff --git a/documentation/backlog/ROADMAP-HISTORY.md b/documentation/backlog/ROADMAP-HISTORY.md index 02bad00d..6ded6caa 100644 --- a/documentation/backlog/ROADMAP-HISTORY.md +++ b/documentation/backlog/ROADMAP-HISTORY.md @@ -114,3 +114,42 @@ | **R-44** | **[P2-HIGH] A manual offsite push ships an unrefreshed DB dump — "backed up now" is false for the DB half.** Shipped in v0.148.0. | **SHIPPED controller v0.148.0 (2026-07-19)** | full text: `git show fddfe00ce268:documentation/backlog/ROADMAP.md` | | **R-47** | **[P2-HIGH] The DB replay races the application's own schema repair.** Shipped in 0.153.0, 0.90.0, v0.153.0. | **SHIPPED — controller v0.153.0, 2026-07-20** | full text: `git show fddfe00ce268:documentation/backlog/ROADMAP.md` | | **R-49** | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** | **MOVED 2026-08-22 -> OPEN-ITEMS.md** (was: idea) | full text: `git show fddfe00ce268:documentation/backlog/ROADMAP.md` | + +## 2026-10-03 — the triage: shipped, killed, closed and moved items out of `ROADMAP.md` + +| ID | Item | Final state | Full text | +|---|---|---|---| +| **R-116** | **The drive-absent alarm and its recovery are a mismatched pair — generic on the way out, specific on the way back** | **CLOSED — the register row is closed (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-113** | **The drive-absent gate cannot fire on device loss — E-2b's alarm is wired to an unreachable condition** | **CLOSED — the register row is closed (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-112** | **E-2's degraded banner and offer have no UI consumer — correct endpoint, invisible to the customer** | **CLOSED — the register row is closed (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that vanished** | **CLOSED — the register row is closed (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-7b** | **Share backup EXECUTION** | ****SHIPPED (controller v0.145.0, 2026-07-18)**** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-202** | **The orphan card promises recoverability unconditionally** | **CLOSED 2026-10-03 — register row verified finished and moved (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-201** | **The wipe-and-recover drill** | **CLOSED — both halves passed 2026-08-04/07 (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **D5** | **Move app secrets into the LOCAL recovery unit** | ****SHIPPED + PROVEN-LIVE** — controller v0.188.0, 2026-07-30** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **E-2** | **Drive-role machinery around the moved vzdump target.** | **CLOSED — PARTIALLY PROVEN (Session C, 2026-07-29; `CLOSED-ITEMS.md`, the id-less rows)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **G-1** | **Gate C5 — the cross-repo tag-reachability check.** | ****BUILT AND CLOSED 2026-08-08** — `scripts/wire_contract_gate.py`, `--fast`, registered in `repo_gates.py`** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **G-6** | **C2 — NOT mechanically gateable.** | **KILLED — recorded as a no: "names a route" is a judgement, not a predicate** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **G-7** | **C6 — NOT gateable, and the measurement is the finding.** | **KILLED — recorded as a no, with evidence: `deadcode` re-found neither known instance** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-25b** | **RULED: customer DELETE becomes a guided full-teardown cascade.** | ****SHIPPED hub v0.69.0 (2026-07-21)**** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-87** | **The restic (app-data offsite) tier is NEVER restore-tested** | ****SHIPPED 2026-08-31 — controller v0.231.0, RE-SCOPED by its own spike.** Not built as written: the spike measured that an unattended scratch restore would have caught ONE of five drill-found restore defects, so it proves the snapshot CONTAINS a recoverable …** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-48** | **[P2-HIGH] Restore controls are separable only by layout — and the difference between them is whether the data comes back.** | **SHIPPED — controller v0.154.0 (2026-07-21): the offsite restore wizard, one entry per app (`controller/internal/web/restore_wizard.go:12`); verified 2026-10-03** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-29** | **The design-v2 green gates are not enforced anywhere — one has been RED for 16 releases.** | **CLOSED — the register row is closed (`CLOSED-ITEMS.md`)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-127** | **`data_key: true` is unreliable (4+ encryption keys unflagged, contradicting the catalog's own labels), and O4 can regenerate a DB password that no longer matches the restored data directory** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: READY — NEW 2026-07-30)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-133** | **The vaulted break-glass console credential is PLAINTEXT AT REST, so every hub DB backup is a fleet-wide console-credential dump.** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: READY — NEW 2026-07-31)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-126** | **A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS.** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: READY — 2026-07-30)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-95** | **The restic offsite tier's credential CAN DELETE — R-89's "parallel question", now ANSWERED** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: idea — established read-only 2026-07-27)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-162** | **`docker diff` is the gate's only witness and its failure mode is quiet** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: WATCHING — 2026-08-02)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-164** | **C2's chain — the DB volume tar cannot be dropped until a sound dump predicate exists** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: BLOCKED — on the predicate (2026-08-02))** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-91** | **The old 13 GB datastore copy is still on ep0's root disk** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: WATCHING — gated on demo-felhom's first post-migration PBS backup)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-92** | **The hub's PBS-DR gauge is 0.1 GB-granular, so small deltas are unverifiable** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: idea — 2026-07-27)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-93** | **`drill-r50` is both a blocked customer and the only drift fixture** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: idea — 2026-07-27)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-169** | **CI can only report, because there is no gate in the road** | **MOVED → `OPEN-ITEMS.md` — a FINDING; its register row is the record (moved out of this file 2026-10-03; status here was: idea — minted 2026-08-02, **WAITING-ON-OPERATOR**)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** | **COLLAPSED 2026-10-03** — parts 1–10 shipped, part 11 deferred by ruling; the open rows are listed in `ROADMAP.md` §UPDATE-ARC. Status at collapse: **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** **2026-09-13 — SLICE 4 COLLAPSED: SHIPPED** (controller v0.237.0/v0.238.0/v0.238.1, R-448 CLOSED, proven live — … | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| pre-invite: golden 0.146.0 rebuild | **golden 0.146.0 rebuild** | **DONE — baked + published 2026-07-18** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| pre-invite: golden ≥0.147.x carries all four infra images | **golden ≥0.147.x carries all four infra images** | **DONE — `build-golden.sh` v2.1.0 derives the pre-pull list (2026-07-19)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| pre-invite: freemail.hu test-send (R-4) | **freemail.hu test-send (R-4)** | **DONE — R-4 closed 2026-07-21 (all three halves)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| pre-invite: C6 — a customer performs a restore unassisted | **C6 — a customer performs a restore unassisted** | **SUPERSEDED — the unaided recovery journey was proven 2026-08-07 (`00` §A); the capability row "a customer performs a restore via UI alone" stays as R-3's target** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| pre-invite: R-11 rulings (contact channel, tester agreement, alert thresholds) | **R-11 rulings (contact channel, tester agreement, alert thresholds)** | **PARTLY RULED — the contact channel ruled 2026-07-21; the tester agreement → R-809 (2026-10-03)** | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| stray text | **Two paragraphs that sat between ROADMAP rows** (a boot-gate log excerpt about immich after the R-50b row; three paragraphs on a quiesce-cycle amplifier after the R-87 row) | **KEPT IN HISTORY** — their rows are history; nothing open depends on them | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | +| **R-50b** | **A root-owned privileged host artifact is delivered unversioned from `main`** | **MOVED → `OPEN-ITEMS.md`** (Box system & updates, P3) — a FINDING, found when `one_register_gate.py` stopped skipping suffix ids (2026-10-03) | full text: `git show 9f77865:documentation/backlog/ROADMAP.md` | diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index b4daad4d..70dd14c3 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -27,145 +27,98 @@ > **`scripts/one_register_gate.py` enforces it:** a row here that is neither an idea nor done, and > has no counterpart in the register, fails the push. > -> **Priorities:** P1 = closed-alpha blocker · P2 = close during alpha · P3 = post-alpha. -> Existing loose notes in this folder (`FOLLOWUP-*`, `FIX-M*`) are absorbed as references below. +> **Severity — ONE scale for this file and `OPEN-ITEMS.md` (2026-10-03; operator may reverse):** **P1** now (a +> household can lose or leak data, a box can stop or be taken over, or a promise is false — today) · **P2** before the +> first paying customer · **P3** during the first customers · **P4** later / nice to have. Each item's tag leads its +> Item cell; an older tag later in the text (`[P2-HIGH]`) is history, and a re-rank says why in one dated line. +> +> **Cleaned 2026-10-03.** Shipped, killed, ruled and moved items went to `ROADMAP-HISTORY.md` (compressed; full text +> `git show 9f77865:documentation/backlog/ROADMAP.md`). Rows that are FINDINGS live in `OPEN-ITEMS.md` only. The loose +> notes of this folder have their verdicts in `README.md`. + --- -## New 2026-10-03 — the operator's four items (intentions; severity per `OPEN-ITEMS.md`'s scale) +## Intentions — by severity (2026-10-03) + +> One table, P2 first. **There is no P1 intention**: nothing on this page is a harm happening today — those are +> findings, and they live in `OPEN-ITEMS.md`. Inside a severity, the order is the operator's request first, then +> by how directly a household meets it. + +### P2 — before the first paying customer | ID | Item | Size | Status | Notes / map rows flipped | |----|------|------|--------|--------------------------| | R-808 | **[P2] Box system security updates.** Goal: every box receives operating-system security patches on a schedule, and a failed update is undone. Why: a box lives in a home for years and today keeps the packages it was installed with — nothing updates the Proxmox host, the guest's Debian or its Docker engine (the finding is **R-812**). Scope: (1) regular security patches, host and guest; (2) reboots and their timing — inside the night window, never across a backup; (3) host kernel updates (a kept fallback boot entry); (4) Docker engine updates in the guest (an engine restart stops every app — quiesce like a backup); (5) the Proxmox MAJOR upgrade path (PVE 9 → 10) as its own later step, drilled first; (6) how the household and the operator are told (an event, a dashboard line); (7) how a failed update is undone (host: the vzdump/PBS copy plus the boot entry; guest: a snapshot before the run). Spike first: what the agent may run under its sudoers fence, and what a half-applied `apt` run leaves behind. | L | idea — 2026-10-03 (operator request) | Flips `00` §G *"Box operating-system security updates"* (MISSING, added 2026-10-03). Finding half: **R-812** in `OPEN-ITEMS.md`. | | R-809 | **[P2] Legal pages and business papers before the first paying customer.** Goal: Felhom may legally take money from a household. Pieces: the website's ÁSZF, adatkezelési tájékoztató and impresszum (none exist — finding **R-813**); the customer contract; a data-processing agreement (Felhom monitors boxes and holds encrypted off-site backups, so it processes household data); billing and invoicing; the lawyer's licence review already on STATUS's list (**R-802**). Connects to the old **R-11** rulings (contact channel — RULED 2026-07-21; the tester agreement — never written). Owner: **operator**; CC drafts a text on request. | M | idea — 2026-10-03 (operator request) | Flips `00` §H rows *"Website legal pages"*, *"Customer contract and data-processing agreement"*, *"Billing and invoicing"* (MISSING, added 2026-10-03). Finding half: **R-813**. | + +### P3 — during the first customers + +| ID | Item | Size | Status | Notes / map rows flipped | +|----|------|------|--------|--------------------------| | R-810 | **[P3] Independence — "if the household leaves Felhom, or Felhom stops".** Goal: a written answer the data-sovereignty pitch can point at. Questions: does the box keep working without the hub (partly answered — `architecture/_recovery-inventory-2026-07-28.md` §D2.4 and `07` §8 row 11b cover a LOST hub: a day is invisible, a week loses alarms, resets and convergence); who owns the domain (the customer, `01` §7), the Cloudflare tunnel and the off-site storage account; how a household exports everything; what a hand-over to the household or another provider looks like. Status **spike**: nothing is known to be wrong; no document answers the leave/hand-over half. | M | idea — spike, 2026-10-03 (operator request) | Flips `00` §E *"The household can leave Felhom, or outlive it"* (MISSING, added 2026-10-03). | | R-811 | **[P3] A second login step for the dashboard.** Goal: a stolen or guessed password alone does not open the dashboard, which controls the whole box and is on the internet. Today: one password, one bcrypt hash (`felhom-controller/controller/internal/web/auth.go:37-44`). Options: a TOTP code or a passkey; recovery when the phone is lost modelled on the existing reset-code and recovery-code designs. Relates to **R-15** (member accounts) and the family gate (`09` §3 decisions 63–65): the same door should carry both. | M | idea — 2026-10-03 (operator request) | Flips `00` §E *"A second login step for the dashboard"* (MISSING, added 2026-10-03). | +| R-19 | **[P3]** Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE **Flips (2026-10-03):** `00` §E "LAN access when internet is down" IMPLEMENTED → PROVEN-LIVE. | +| R-34 | **[P3]** **Backup data lifecycle management.** An "inactive backups" section on „Távoli mentés": apps that have snapshots but no active backup — **disabled OR uninstalled** — listed with name / size / last snapshot / restorable, plus an explicit **double-confirmed per-app delete** via `restic forget --tag` + nightly prune. | M | idea | **RULING: the offsite toggle NEVER offers deletion — policy and destruction stay decoupled.** Turning backups off must never be a data-destroying act, and deletion must never hide behind a toggle. Origin: 2026-07-18 rehearsal. Pairs with R-32 (that one is the operator's view of dead bytes; this one is the customer's) **Flips (2026-10-03):** `00` §C "Offsite (restic → Hetzner Storage Box)" — adds the inactive-backups leg; a new §C row when built. | +| R-46 | **[P3]** **[P2] Verification copies need a customer-visible browse surface and an expiry.** v0.147.0 made them *visible* (listed with path/size/date, individually deletable) — but the customer still cannot LOOK INSIDE a verification restore to confirm the file they wanted is really there, which is the entire point of a verification restore, and nothing ever removes them. | S–M | idea | Origin: 2026-07-19 feedback slice 4a, registered as the explicit follow-up to it. Two gaps, deliberately designed together because they are the same object: (a) **the invisible-result gap** — a read-only browse of `backups/offsite-restore/` (the FileBrowser infra stack already exists and already serves scoped roots, so this may be a mount rather than new code); (b) **the disk-lifecycle gap** — auto-expiry after N days with the count/size surfaced before it fires, so a drive is never quietly filled by verification restores nobody remembers taking. Pairs with R-43: a browse surface is also how a customer would discover that a DB-indexed app's files came back but the app still cannot see them **Flips (2026-10-03):** `00` §C "Offsite restore: local-preferred scratch …" — the browse-and-expiry leg. **Re-ranked 2026-10-03:** [P2] → P3: a verification restore works today; what is missing is a way to look inside it and its expiry — a household meets it, with a workaround (FileBrowser). **Finding-shaped** (it states a fact about the shipped product): check it against today's product before building — an R-424 instance. | +| R-58 | **[P3]** **[P2] Assisted disk-picker install mode — the installer should let the operator CHOOSE the target disk instead of requiring the serial up front.** Today an install is either unattended (the answer file pins one `ID_SERIAL_SHORT`, which you can only know by first booting the machine) or match-nothing safety (aborts by design). That forces a two-boot dance for every new box: boot the safety ISO to read the serial, rebuild the ISO armed, boot again. | S–M | **idea — operator ruling 2026-07-21** | **Operator's argument, verbatim:** *"the installer should list the available storage devices (excluding the installation media) and let us select one, and continue."* **Shape:** a THIRD ISO mode alongside the two that exist — unattended-serial and match-nothing-safety. It enumerates candidate disks with **size / model / serial**, excludes the installation media itself, takes a selection plus a confirm, and proceeds. **Unattended+serial REMAINS the appliance/factory mode** — it is the right shape when the machine is provisioned in bulk and nobody is standing there; the picker is for the case where somebody is. **Slice 1 (cheap, same code surface, do this first):** improve the abort screen. On filter-no-match the installer currently just fails safe and says nothing useful — it should print the candidate table (size/model/serial) plus the one-line hint naming which serial to put in the profile. That alone collapses the two-boot dance from "boot, guess, go read docs, rebuild" to "boot, copy the serial off the screen, rebuild", and it is the same enumeration code the full picker needs. **Why it matters beyond convenience:** it is the BYO / reinstall flow — a customer's existing hardware, or a rebuild of a box whose disk layout nobody recorded, is exactly where the serial is unknown and a wrong guess is destructive. The current fail-safe is correct but mute. Origin: TASK-G, arming the HP install ISO — the serial had to be read off the board by hand between two boots **Flips (2026-10-03):** `00` §A "Bare-metal Felhom ISO" — adds a third, assisted install mode. **Re-ranked 2026-10-03:** [P2] → P3: the operator installs every box today and the answer-file route works. | +| R-69 | **[P3]** **F14-full: an operator push channel that actually interrupts (ntfy / Telegram / similar), beyond mail-client priority flags.** F14-light (v0.71.0 headers + Gmail filter) nudges a mail client; a 15:29 node_down should reach the operator's pocket in seconds regardless of inbox hygiene. Needs: channel choice (self-hosted ntfy on k3s vs Telegram bot), dispatcher fan-out seam, per-severity routing, quiet hours. | M | idea | Origin: `AUDIT-power-outage-recovery-2026-07-22.md` F14. Deliberately NOT built in the v0.71.0 train (scope-forked per the task spec) **Flips (2026-10-03):** `00` §F "Operator alerting (Healthchecks → monitoring@felhom.eu)" — a channel that interrupts. | -## P1 — closed-alpha blockers +### P4 — later / nice to have | ID | Item | Size | Status | Notes / map rows flipped | |----|------|------|--------|--------------------------| -| R-116 | **The drive-absent alarm and its recovery are a mismatched pair — generic on the way out, specific on the way back** | S | idea — **PROVEN LIVE 2026-07-29** | Absent fires `storage_disconnected`; return fires `backup_target_restored`. `backup_target_absent` never fires at all (count 0 across a full Session-C run), so an operator gets an alarm they cannot match to its recovery — exactly what `notifyDriveReturned`'s own comment forbids. Root cause: `notifyDriveAbsent` (`intermediary.go:635-646`) branches on `isTarget[a.Path]` with `a.Path` the GUEST path, and `driveTargetByPath` (`:602-616`) builds it as `out[GuestPath] = d.BackupTarget` — but **the drive is TWO `/disks` rows and the flag and the guest path sit on different ones**: the `felhom-backup` storage row has `BackupTarget: true` (`felhom-agent/internal/localapi/disks.go:211`) and gets a guest path only while classified user-data, while the registry union row has the guest path and **never assigns `BackupTarget`** (`disks.go:265-267`). Absent ⇒ the flagged row loses its guest path ⇒ the union row writes `false` ⇒ generic. On return the rows rejoin ⇒ specific. v0.184.1 fixed the KEYING, not this. **Only reachable because R-113 made the gate fire at all.** Fix likely agent-side; decide the repo first. Blocks E-2's C5. Evidence: `audits/SESSION-C-2026-07-29.md` §5 | -| R-113 | **The drive-absent gate cannot fire on device loss — E-2b's alarm is wired to an unreachable condition** | M | idea — **PROVEN LIVE 2026-07-29** | `planDriveGates` (`felhom-controller/internal/web/intermediary.go:216-262`) treats a path as present by OR-ing in `d.BoundUnderParent`, which the agent derives from `GuestSeesMount()` — *"is this path a mount target in the guest's `/proc//mountinfo`"* (`internal/localapi/disks.go:210`). The raw drive mount is a **device-bound systemd unit** and dies with the device; **the agent's own bind under the shared parent is not device-bound and its mountinfo entry outlives the device**, so the gate reads it as present and `notifyDriveAbsent` is never called. Live on a fresh box: target drive hot-detached, agent said `enrolled drive absent by UUID` every 20 s for 4½ min, controller logged **0** `[gate]` lines, hub received **zero** events — neither `backup_target_absent` nor the generic `storage_disconnected`. Not a virtualisation artefact (device-bound-mount vs manual-bind is the same on metal); caveat: SCSI hot-detach, physical unplug not staged. **Sixth instance of seam-built-but-never-wired — E-2b wired the seam to a condition that cannot occur.** Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.2 | -| R-112 | **E-2's degraded banner and offer have no UI consumer — correct endpoint, invisible to the customer** | S | idea — **PROVEN LIVE 2026-07-29** | `GET /api/storage/backup-target` returns byte-exact Hungarian copy (verified on a live box), and nothing in the product asks for it: `grep 'backup-target'` across every `*.html`/`*.js`/`*.css` → **0 hits**; no template references `OfferPath`/`Degraded`/the copy; `resolveBackupTargetState` and `degradedMessageFor` are consumed **only** by the JSON handler, with **no page handler injecting the state**. Decisive contrast: the templates fetch **18 distinct `/api/storage/*` endpoints** — `backup-target` and `backup-target/assign` are the only two with zero references. The handler's own comment calls itself *"the dashboard's source for the degraded banner and the offer"*. v0.185.1 shipped as *"the offer endpoints were mounted where nothing routed to them"* and fixed the **mount**, stopping one layer short of the **render**; its test pins dispatch, not reachability. **Fifth instance of the class. Fix R-114 first** — wiring this alone starts showing customers a wrong message. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.1 | -| R-114 | **On target-drive loss the customer is told the wrong story and offered the drive that vanished** | S | idea — **PROVEN LIVE 2026-07-29** | With the assigned target absent the endpoint returned `degraded:true, target:"felhom-backup"` **plus** the *"ugyanazon a lemezen van, mint a rendszer"* message — false, the target is a missing drive, not the system disk — **and** an `offer_path` pointing at the drive that just disappeared. `resolveBackupTargetState` falls through to the generic degraded branch whenever no disk satisfies `d.BackupTarget && d.MountPath != ""`, never distinguishing *never configured* from *configured and now missing*. Shares R-113's root cause (two disagreeing presence signals), different code path and fix. **Invisible today only because of R-112.** Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.3 | -| R-3 | Friend-alpha onboarding runbook (generalized from `pilot/RUNBOOK-peti-return-2026-07-13`): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step | M | idea | Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of `PETI-tester-agreement.md`. **Next from-scratch rehearsal to include customer DELETE + re-create** — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). **S6b (manual stale-host delete before re-enroll) is OBSOLETE** — re-enrollment upserts the existing host row cleanly (`store.UpsertHost` ON CONFLICT DO UPDATE; `handleAdminCreateHost` no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. **NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE):** one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). **Remaining consistency gap:** the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b | +| R-3 | **[P4]** Friend-alpha onboarding runbook (generalized from `pilot/RUNBOOK-peti-return-2026-07-13`): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step | M | idea | Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of `PETI-tester-agreement.md`. **Next from-scratch rehearsal to include customer DELETE + re-create** — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). **S6b (manual stale-host delete before re-enroll) is OBSOLETE** — re-enrollment upserts the existing host row cleanly (`store.UpsertHost` ON CONFLICT DO UPDATE; `handleAdminCreateHost` no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. **NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE):** one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). **Remaining consistency gap:** the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b **Flips (2026-10-03):** `00` §C "A customer (not the operator) performs a restore via UI alone" (MISSING). **May be superseded** by `runbooks/day0-install.md` and the first-hour drills — check before building. | +| R-6 | **[P4]** **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. **Flips (2026-10-03):** `00` §E "Media to TV via DLNA" (MISSING) — the spike half, done; R-8 is the build. | +| R-8 | **[P4]** DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** **Flips (2026-10-03):** `00` §E "Media to TV via DLNA" (MISSING → IMPLEMENTED). | +| R-9 | **[P4]** Uninstaller trio (from 07-15 Peti session): cluster-aware `felhom_guests` guard (node-local `pct list` deletes cluster-wide pveum objects); saferemove detection + time estimate + opt-in `--quick-remove` (never mutate `storage.cfg`); smarter `restore_storage` default for BYO clusters (shared storage, not local-lvm) | M | idea | Second item's rejected alternative (temp-disable-and-restore) stays rejected — crash window silently downgrades cluster wipe policy **Flips (2026-10-03):** `00` §A "Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling" — the cluster/BYO half. | +| R-27c | **[P4]** **Customer self-bind, slice 2 — console-passphrase bind.** Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. | M | idea | **Security constraints from the session ruling, all load-bearing:** passphrase **issued at customer creation**; the global-lookup endpoint must be **spray-hardened** — per-appliance **and** per-IP caps, constant-time comparison, a **single generic failure** (no oracle), alerting on abuse; an **accent-free wordlist** (console keymaps are not Hungarian); the **web capability-link path is RETAINED**; **claim-by-email is RETAINED** as the delivery-channel proof. **Also under this item:** the self-bind email gains the **public universal-ISO download link + two-line instructions** (the DIY case). **Secret-bearing per-customer ISOs are ruled OUT.** Sibling of R-27b (second-box flow) — different axis, both build on the same `/bind/` page **Flips (2026-10-03):** `00` §A "Customer binds their own appliance (self-service)" — a second bind route. | +| R-56 | **[P4]** **[P3] Apps do not say how technical they are, so a beginner can be ambushed by a config-heavy one.** The catalog presents every app as equally approachable — one Telepítés button, the same Hungarian copy — but they are not. Glance needs a hand-written `glance.yml` before it does anything; some apps need a reverse-proxy or API concept to configure; others genuinely are install-and-use. A tester who picks the wrong first app concludes the PRODUCT is broken, not that they picked an advanced app. | S | **idea (filed 2026-07-21)** | Origin: TASK-E Part 3 — **filed, deliberately not implemented**. Shape: a `difficulty:` field in `.felhom.yml` (`kezdő` / `haladó` / `technikás`) surfaced as a catalog-card badge and repeated on the deploy screen. Cheap and incremental: one optional metadata field plus a badge, classifiable app-by-app with no migration — an app with no `difficulty:` simply shows no badge. **This is the constructive half of the glance ruling**: glance STAYS in the catalog (operator ruling 2026-07-21 — it is a legitimate app, not a broken one; its missing seeded `glance.yml` is a known pre-existing finding), and the honest fix is to LABEL it rather than hide it. Pairs with R-41: that gate proves an app CAN still deploy; this field tells a customer whether THEY should be the one deploying it. **Badge plumbing is ALREADY BUILT (controller v0.158.0)** — `web.MetaBadge` + the `meta_badge` template partial + the `lifecycleBadge` funcmap entry were written generic for exactly this: a `difficultyBadge` funcmap function returning the same `*MetaBadge`, plus a `difficulty:` field on `stacks.Metadata`, is the whole remaining job. No new markup, no new CSS. Re-sized accordingly **Flips (2026-10-03):** `00` §B "Deploy an app from the catalog" — a difficulty label on the card. | +| R-62 | **[P4]** **[P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape.** | XS | **idea (operator, 2026-07-22)** | Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; **wording only** — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). **Flips (2026-10-03):** `00` §G "Customer/host management" — the delete dialog. | +| R-64 | **[P4]** **„Felhom↔Felhom media pairing blessed" — the two-box SMB pairing (one box shares, the other mounts it as NAS storage) becomes a supported, documented flow.** | XS–S | idea (2026-07-22) | Origin: the operator ran the pairing drill on the live demo pair and it WORKS — the drill itself is the pending evidence leg (a written run-through with the R-66 surfaces in play). R-66 shipped the enabling visibility: the serving box's address is now on its own Beállítások → Rendszer „Hálózat" card, and the add form names the NetBIOS trap. Blessing = a short customer-facing recipe (`documentation/controller/network-storage-nas.md` naming-caveat paragraph is the seed) + one supported-path sentence in the capability map. Flips: would add a "Felhom↔Felhom media pairing" capability row (currently unlisted). Pairs with R-65 (same two-box topology, entirely different transport + guarantees) **Flips (2026-10-03):** `00` §E "Files from Windows Explorer / Mac Finder (SMB server)" — box-to-box pairing as a supported flow. | +| R-26 | **[P4]** **Guided old-history recovery via a retained superseded escrow + the recovery code.** Enabled by hub v0.60.0 (Part B) which now RETAINS superseded escrow blobs (`host_escrow_superseded`, `ListSupersededEscrow`). Build the flow that, given the customer's recovery code, unwraps a retained old blob → recovers the old repo passphrase → mounts/reads the moved-aside `.orphaned-` repo for restore. | M | idea (enabled by v0.60.0) | Turns "history recoverable in principle" into a real customer-drivable path; pairs with the controller v0.142.0 orphaned-repo move-aside. Origin `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Flips (2026-10-03):** `00` §C "Offsite restore …" — old history via a retained escrow. **Parked by a decision:** R-312 (decided 2026-08-13) keeps any route to a set-aside store operator-only until a real customer asks (`07` §11). | +| R-27b | **[P4]** **Customer self-bind, second-box flow (controller side).** For a customer who ALREADY has a bound box and installs another, the controller shows a dismissable "bind another box" prompt (and a bind-later entry under settings) that walks to the hub `/bind/` page — so a returning customer isn't emailed a fresh operator-sent link for every box. Mechanism sketched in the hub v0.66.0 REPORT; NOT built (R-27 slice 1 deliberately did not touch the controller). | M | idea (minted by hub v0.66.0) | Origin: hub v0.66.0 slice-1 ship (first-box only). Reuses the same `/bind/` public page + tokenized-link machinery; adds a controller-side entry point + the operator "mint a link for an existing customer" affordance **Flips (2026-10-03):** `00` §A "Customer binds their own appliance (self-service)" — the second-box flow. | +| R-12 | **[P4]** Cluster mode: agent-follows-guest, bind-mount reconciliation on HA migration | XL | idea | Scoped 07-15; interim = HA-group pin to one node. Driven by Peti's two-node cluster **Flips (2026-10-03):** `00` §G — a new row "a guest follows its node in a cluster" when built. | +| R-14 | **[P4]** Headscale/WireGuard spike: Minecraft/gaming port connectivity (CGNAT-proof, sovereign DERP fallback) | M | idea | **Flips (2026-10-03):** `00` §E — a new row "game-server ports reachable behind CGNAT" when built. | +| R-15 | **[P4]** Multi-user dashboard accounts (household members, roles) | L | idea | Single password is a stated alpha limitation (R-11). **Launcher coupling — REVISED (controller v0.165.0):** the "share the launcher outside the household" need is now met WITHOUT member accounts — the **Indítópult megosztása** capability-URL guest link (`/s/`, information-only, no account) shipped in v0.165.0. What remains for this arc is member-specific: **per-member tile visibility** (each member sees only their apps) and the launcher-as-member-landing-page — both live inside this SSO/members arc; the guest-link ruling explicitly SUPERSEDES the earlier "members are how you share the launcher" framing **Flips (2026-10-03):** `00` §E "Multiple household users / per-person accounts" (MISSING). Relates to **R-811** (a second login step): the same door should carry both. | +| R-72 | **[P4]** Curate `brand_color` for the top catalog apps | XS | idea | Parked follow-up to the v0.163.0 launcher. `.felhom.yml` `brand_color` (`#rgb`/`#rrggbb`) overrides the deterministic slug-hash tile color; no catalog app sets it yet. Pick brand-accurate colors for the most-installed apps so their launcher tiles match their real brand. Catalog-only change (`app-catalog-felhom.eu`), `brand_color` is already `omitempty` and consumed by the controller **Flips (2026-10-03):** `00` §E "Indítópult (app launcher)" — brand colours. | +| R-74 | **[P4]** **Island control plane on a CLUSTER (Peti's 2 nodes)** — bring R-50's island bridge to a multi-node PVE cluster. | M | idea (Phase C of R-50, parked) | R-50 shipped the island for the ONE-host fleet (demo-hp, demo-felhom). A cluster needs **bridge parity on every node**: either per-node identical `/etc/network/interfaces` `vmbr9` stanzas (simplest, drift-prone) or — preferred at ≥2 nodes — a Proxmox **SDN zone/vnet** defined cluster-wide (one definition, auto-applied per node). The guest island IP is per-guest + node-independent; the **agent-follows-guest** rule holds (each node's agent binds its own `vmbr9` `169.254.253.1`). Migration order per the spike: drill-proven → demo (done) → **Peti (this row)**. Its own supervised runbook, coordinated with Peti (a live customer). Completes the capability-map "site/network change" row for clustered installs. Source: `audits/SPIKE-island-bridge-2026-07-25.md` (cluster-parity finding) + `RUNBOOK-island-migration.md` (single-host procedure to generalise) **Flips (2026-10-03):** `00` §G — a new row "the island control plane on a multi-node cluster" when built. | +| R-41 | **[P4]** **[SLICE 1 SHIPPED 2026-07-21] The catalog has no standing "does every template still deploy?" check.** Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found **5 that had NEVER been deployable**: papra (missing required `AUTH_SECRET`), zipline (v4 renamed `CORE_DATABASE_URL` → `DATABASE_URL`), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the `v` tag prefix + new required env), glance (needs a seeded `glance.yml` the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus **7 broken healthchecks** and 2 apps whose images no longer resolve at all (plant-it, wanderer). | M | idea | Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in `scripts/check-image-pins.py` — a mechanical gate run on every change. Cheap first slice: a **resolvability gate** (`docker manifest inspect` every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. **Silent rot is the real risk** — an app can die upstream and nobody learns until a customer clicks Telepítés | **SLICE 1 SHIPPED 2026-07-21 — `app-catalog-felhom.eu/scripts/check-image-resolvable.py`** (+ 14 fixture tests, no network, resolver injected). Resolves every unique pin with `docker manifest inspect`, ONE image at a time; exit 0 / 1 (the registry says GONE) / 2 (inconclusive). **Two traps encoded, both hit live while building it:** (a) `docker manifest inspect` prints `toomanyrequests: …` and **still exits 0** — the same exits-0-on-failure shape as the ISO tooling's `validate-answer`, so stderr is inspected even on rc=0; (b) the inverse and more dangerous one — the first full sweep called **24 of 65 pins dead, including `postgres:16-alpine` and `redis:7-alpine`**, purely because Docker Hub throttled it partway through. Ambiguity therefore resolves to INCONCLUSIVE and never to an accusation: a gate that cries wolf gets ignored, and then it protects nothing. **The full sweep is still OWED** — DooPlex is not logged in to Docker Hub, so the 52-app table needs one re-run after `docker login`. Wired into `CLAUDE.md` + `REUSE.md` as a start-of-campaign / pre-publish-train step. It immediately paid for itself: it is what turned plant-it and wanderer from 'images do not resolve' into two DIFFERENT diagnoses (see the 2026-07-21 catalog entry). Full slice — the periodic deploy-all sweep on the demo box — remains open **Flips (2026-10-03):** `00` §B "Catalog sync … validation choke point" — a standing deploy-all check. | +| R-45 | **[P4]** **[P2] Unified async-job feedback.** Every long operation invents its own progress surface, or none. Tonight produced three more one-off cards (v0.147.x: samba bring-up, offsite progress, restore result) on top of two existing patterns (deploy 3-step panel; storage-init/netstorage status poll). They agree on nothing: some use `{ok,data}` envelopes and some raw JSON, some poll 1 s / 1.5 s / 3 s, some are in-memory-only and lie after a restart, and each re-implements single-flight + snapshot + phase→Hungarian mapping. | M | idea | Origin: 2026-07-19 feedback slice 1 (controller v0.147.0). The cases to generalise from are all in-tree: `web/storage_init_job.go` (the best shape — acquire/release/set/snapshot), `web/netstorage_job.go`, `web/samba_ensure_job.go`, `backup/opstatus.go`, `backup/offbox_progress.go`. Shape: one job registry + one poll endpoint + one client-side renderer, phases declared per job. **Two lessons tonight that any framework must encode:** (1) a terminal state must be **probed, not inferred** — `compose up -d` exits 0 on a crash-loop; (2) a progress source that reports nothing is normal, not broken — restic reports 0 bytes for a whole incremental run, and a bar that sits at 0% is worse than no bar. Also fixes the restart hole: in-memory job state currently vanishes and the card silently disagrees with reality **2026-07-20 — the first bill for NOT having this arrived, and it was customer-facing.** The samba card's poll (`web/samba_ensure_job.go` + `sharing.html`) mixed a job EDGE and a service LEVEL on one JSON field, and `/sharing` reload-looped at ~1.2 s for every customer with sharing enabled until controller v0.151.0 (`audits/DIAG-sharing-2026-07-20.md`, S-1/S-4). v0.151.0 fixed THAT card's contract only — the framework is still this item. **Third lesson for it to encode, beside the two already listed:** a phase a client answers with a one-shot action must be an EDGE the registry SERVES ONCE, and must never be synthesised from a level; if it can be re-read, it will be re-acted on. **Flips (2026-10-03):** `00` §B "Deploy an app from the catalog (… health-aware progress)" — one job surface for every long operation. **Re-ranked 2026-10-03:** [P2] → P4: polish; every long operation already shows a progress surface of its own. | +| R-65 | **[P4]** **Buddy-box backup replication, cross-household — two Felhom boxes in different homes replicate backups to each other.** | L | idea (post-alpha, spike-first, 2026-07-22) | The natural big sibling of R-64: two households each hosting the other's encrypted backup tier. Explicitly **spike-first** — the transport is NOT SMB (R-64's live-share protocol is wrong for backup replication across the internet: no auth story between households, no resumability, cleartext LAN assumptions); candidates to spike: restic rest-server / rclone / syncthing over the existing WG/tailnet plumbing, encryption keyed so the buddy can never read the payload. Sits on top of the offsite tier's FILL/OVERSUB thresholds thinking (R-5 aggregate). Flips: would add a "cross-household buddy replication" capability row (currently unlisted). Pairs with R-64 (same topology, different transport + guarantees) **Flips (2026-10-03):** `00` §C — a new row "backups replicated to a second household" when built. | -## P2 — during alpha +## UPDATE-ARC — what is still open (collapsed 2026-10-03) -> **Sub-rank `P2-HIGH` = close before the first REMOTE tester.** These are the 2026-07-18 N100 -> rehearsal's findings (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`). They are not P1 — the -> rehearsal proved the product flow works — but each one either misleads the operator, misleads the -> customer, or hides a failure, and all of that gets materially worse the moment the box is somewhere -> you cannot walk over to. +The app-update arc (`audits/SPIKE-app-update-2026-09-01.md`; its living design is `architecture/09-update-architecture.md`). +**Parts 1–10 SHIPPED** (slices 1, 1b, 2, 3, 4 and 5 in controller v0.233.0 – v0.238.1, then the tested-steps, +ladder, hold, undo and night-update parts through controller v0.271.0 — `09` §6.4); **part 11 deferred by ruling** +(`09` §3 decision 18: the fleet sweep waits until the fleet grows). The full previous entry is in +`ROADMAP-HISTORY.md`. **The open work is in the register, by id:** R-450 (slice 6 — automatic within a major, never +across one), R-451 (slice 7 — the fleet sweep, DEFERRED), R-469 (lift the engine-major rule per app), R-440, R-446, +R-458, R-462, R-607, R-610, R-618, R-621, R-622, R-683, R-738, R-785 — all under **App updates** in `OPEN-ITEMS.md`. +Flips: `00` §B "Automatic updates at night — the update leg" and "Whether the box UPGRADES an app by itself". -| ID | Item | Size | Status | Notes | -|----|------|------|--------|-------| -| **UPDATE-ARC** | **The app-update arc — seven slices, from "nobody knows what any box runs" to "an update is a decision the box can take safely."** Opened after `audits/SPIKE-app-update-2026-09-01.md` measured what an update actually does. | L | **slices 1, 1b, 2, 3 & 5 SHIPPED (controller v0.233.0 / v0.234.0 / v0.235.0; harness `app-catalog/scripts/upgrade-test.py`); slices 4, 6, 7 open** **2026-09-13 — SLICE 4 COLLAPSED: SHIPPED** (controller v0.237.0/v0.238.0/v0.238.1, R-448 CLOSED, proven live — `audits/slice4-2026-09-13/`). Remaining: slice 6 (R-450), slice 7 (R-451); R-469 unblocked. | **The reasoning has a home and it is the point of the exercise: `architecture/09-update-architecture.md`** — a LIVING document, updated by every slice in the same session, created because its ABSENCE was a finding (R-438: the mechanism was chosen deliberately and written down nowhere). **Capability-map rows this flips:** the App-lifecycle row (`00-capability-map.md`, "start/stop/restart/update/logs/remove/redeploy") and its 2026-09-01 sibling ("what restart and update do to a deployed app whose compose file the catalog already moved") — both recorded that the ACTIONS work and said nothing about VERSIONS. Slices 1 and 2 add the version half. **The findings live in the register, per the ONE REGISTER ruling:** R-438, R-440, R-441, R-443 (existing) and R-446..R-452 (this session) — one row per remaining slice, each with a rank and an owner, so the arc is visible in the register and not only in a task file. **Slice 3 SHIPPED 2026-09-06 (controller v0.235.0) on the operator's ruling — Option 1, *freeze the version, keep the fixes flowing*.** It was BLOCKED until then, deliberately: changing what the syncer does to a deployed app reverses a decision, not a bug. R-447, R-441 and R-438 are CLOSED by it; **Slice 5 SHIPPED 2026-09-06 as a spike** — `scripts/upgrade-test.py`, 7 edges over 3 apps, proven by a RED negative control, closing R-449 and opening R-459/R-460/R-462. **Its headline changes an assumption this arc was carrying: whether an update can be undone is a property of the individual APP, not of updates** — docmost refuses, privatebin does not — so any design assuming one answer for all 53 is designing against a checked-and-false fact. **R-459 was measured 2026-09-06 and narrowed to a decision, not a defect:** the skipped MariaDB conversion is stable but never self-resolving, and correcting it costs 7 s and does **not** cost the ability to abort — so the trade that row was expected to produce does not exist. Opened R-463 (the PostgreSQL analogue, which fails in the OPPOSITE direction — it refuses to start, and 8 templates sit on `postgres:16-alpine`) and R-464 (an entrypoint line that says `upgrade not required` on an unsupported downgrade). **R-448 (slice 4, the guarded update) is now the head of the arc** and is where the backup precondition goes — slice 3 did NOT make the Update button safer and must not be read as having done so. | -| R-6 | **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. | -| R-7b | **Share backup EXECUTION** — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) | M | **SHIPPED (controller v0.145.0, 2026-07-18)** | **Viktor's ruling: Model B′ — a SIBLING shares source.** New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving **every per-app engine path byte-identical** — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → `RunSharesTier2` (legs grouped by SOURCE drive → `backups/secondary/_shares//`, payload at `_payload/`, layout marker LAST). Tier 3 → `runOffboxSharesLeg`: ONE extra `restic backup --tag felhom-offbox --tag _shares` placed after the app loop and BEFORE retention, so `forget --group-by host,tags` covers the new group with no flag change; a quota-blocked push degrades to the **manifest only, never to nothing**. Restore → „Megosztások" on `/backups/restore`: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then `ReconcileSamba`, then the credential. The **payload** (`_shares-manifest.json` + a best-effort secret-bearing `passdb.tar`) is what makes DR return files + configuration + password rather than loose bytes. **Fold-in: samba joins the liveness set** — `EffectiveProtected` adds the CONTAINER `felhom-samba` exactly while sharing is on. **FULLY PROVEN-LIVE on demo (2026-07-18), all four legs.** (1) tier-2: real `/api/backup/tier2` trigger → `_shares` tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot **`e0b9d723`** (tags `felhom-offbox,_shares`) with the payload dir + both share folders; a second run via the „Távoli mentés" button → **`4e2b15ec`**, containing `_shares-manifest.json` (418 B) AND `passdb.tar` (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the `dokumentumok` DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → `1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true`; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, `smb.conf` re-rendered, `filmek` untouched. (4) liveness: samba stopped → `health_critical` pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). **Correction:** an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (`offbox_target`) instead of the real one (`offbox`); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (`nbNameRe` accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: `controller/sharing.md`; ship report `felhom-controller/REPORT.md`. | -| R-8 | DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** | -| R-9 | Uninstaller trio (from 07-15 Peti session): cluster-aware `felhom_guests` guard (node-local `pct list` deletes cluster-wide pveum objects); saferemove detection + time estimate + opt-in `--quick-remove` (never mutate `storage.cfg`); smarter `restore_storage` default for BYO clusters (shared storage, not local-lvm) | M | idea | Second item's rejected alternative (temp-disable-and-restore) stays rejected — crash window silently downgrades cluster wipe policy | -| R-202 | **The orphan card promises recoverability unconditionally**, which after R-198 is true going forward and false for anything already orphaned | S | **CLOSED 2026-10-03 — the register row moved to `CLOSED-ITEMS.md` (triage)** — **gate hit 2026-08-04 — card untouched, sentence still live** | Blocked on knowing which escrow generation an orphaned repo belongs to (R-199/R-201). A single ACK boolean can say a retained recoverable blob EXISTS but not that one COVERS this repo; a conditional promise that can still be false is worse on that surface than a hedged one | -| R-201 | **The wipe-and-recover drill** | L | **CLOSED 2026-10-03 — the register row moved to `CLOSED-ITEMS.md` (triage)** — **PREPARED, HALTED BEFORE THE WIPE (2026-08-04)** | Nothing irreversible done. Established live for the first time: a rebuilt box's off-site run REFUSES (the orphan card, not a silent fresh history — closes R-193 Q3), the orphan reset works move-aside-never-delete, and demo-hp's pre-rebuild key is permanently gone (superseded four hours before v0.93.0). Resume needs R-203 | -| R-19 | Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE | -| R-34 | **Backup data lifecycle management.** An "inactive backups" section on „Távoli mentés": apps that have snapshots but no active backup — **disabled OR uninstalled** — listed with name / size / last snapshot / restorable, plus an explicit **double-confirmed per-app delete** via `restic forget --tag` + nightly prune. | M | idea | **RULING: the offsite toggle NEVER offers deletion — policy and destruction stay decoupled.** Turning backups off must never be a data-destroying act, and deletion must never hide behind a toggle. Origin: 2026-07-18 rehearsal. Pairs with R-32 (that one is the operator's view of dead bytes; this one is the customer's) | -| R-27c | **Customer self-bind, slice 2 — console-passphrase bind.** Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. | M | idea | **Security constraints from the session ruling, all load-bearing:** passphrase **issued at customer creation**; the global-lookup endpoint must be **spray-hardened** — per-appliance **and** per-IP caps, constant-time comparison, a **single generic failure** (no oracle), alerting on abuse; an **accent-free wordlist** (console keymaps are not Hungarian); the **web capability-link path is RETAINED**; **claim-by-email is RETAINED** as the delivery-channel proof. **Also under this item:** the self-bind email gains the **public universal-ISO download link + two-line instructions** (the DIY case). **Secret-bearing per-customer ISOs are ruled OUT.** Sibling of R-27b (second-box flow) — different axis, both build on the same `/bind/` page | -| R-50b | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\|reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | S–M | **(a) SHIPPED 2026-07-21; (b)/(c) open** | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. | +## Gating candidates — Campaign 12, Part 4 (2026-08-08) — all **P4** -> `14:29:28 [gate] boot 1784525102-11906045: live bind confirmed — recreating drive-backed app calibre-web (state=stopped) onto /mnt/felhom-drives/hdd_1` -> `14:29:29 [gate] boot 1784525102-11906045: 1 drive-backed app(s) left stopped — zero containers means the customer stopped them on purpose` - -**immich is absent from the recreate list and came back STOPPED (0 containers)** — on 2026-07-21 before this fix, the identical fixture brought it back RUNNING. calibre-web recreated, bookstack back, `[bootrecon] no boot-orphaned apps` (consistent — nothing was left orphaned for it to adopt), **ZERO alerts**, whole convergence ~15 s from reboot to steady state. The `left stopped` INFO line fired in production for the first time, so the honoured path is observable rather than silent | -| R-56 | **[P3] Apps do not say how technical they are, so a beginner can be ambushed by a config-heavy one.** The catalog presents every app as equally approachable — one Telepítés button, the same Hungarian copy — but they are not. Glance needs a hand-written `glance.yml` before it does anything; some apps need a reverse-proxy or API concept to configure; others genuinely are install-and-use. A tester who picks the wrong first app concludes the PRODUCT is broken, not that they picked an advanced app. | S | **idea (filed 2026-07-21)** | Origin: TASK-E Part 3 — **filed, deliberately not implemented**. Shape: a `difficulty:` field in `.felhom.yml` (`kezdő` / `haladó` / `technikás`) surfaced as a catalog-card badge and repeated on the deploy screen. Cheap and incremental: one optional metadata field plus a badge, classifiable app-by-app with no migration — an app with no `difficulty:` simply shows no badge. **This is the constructive half of the glance ruling**: glance STAYS in the catalog (operator ruling 2026-07-21 — it is a legitimate app, not a broken one; its missing seeded `glance.yml` is a known pre-existing finding), and the honest fix is to LABEL it rather than hide it. Pairs with R-41: that gate proves an app CAN still deploy; this field tells a customer whether THEY should be the one deploying it. **Badge plumbing is ALREADY BUILT (controller v0.158.0)** — `web.MetaBadge` + the `meta_badge` template partial + the `lifecycleBadge` funcmap entry were written generic for exactly this: a `difficultyBadge` funcmap function returning the same `*MetaBadge`, plus a `difficulty:` field on `stacks.Metadata`, is the whole remaining job. No new markup, no new CSS. Re-sized accordingly | -| R-58 | **[P2] Assisted disk-picker install mode — the installer should let the operator CHOOSE the target disk instead of requiring the serial up front.** Today an install is either unattended (the answer file pins one `ID_SERIAL_SHORT`, which you can only know by first booting the machine) or match-nothing safety (aborts by design). That forces a two-boot dance for every new box: boot the safety ISO to read the serial, rebuild the ISO armed, boot again. | S–M | **idea — operator ruling 2026-07-21** | **Operator's argument, verbatim:** *"the installer should list the available storage devices (excluding the installation media) and let us select one, and continue."* **Shape:** a THIRD ISO mode alongside the two that exist — unattended-serial and match-nothing-safety. It enumerates candidate disks with **size / model / serial**, excludes the installation media itself, takes a selection plus a confirm, and proceeds. **Unattended+serial REMAINS the appliance/factory mode** — it is the right shape when the machine is provisioned in bulk and nobody is standing there; the picker is for the case where somebody is. **Slice 1 (cheap, same code surface, do this first):** improve the abort screen. On filter-no-match the installer currently just fails safe and says nothing useful — it should print the candidate table (size/model/serial) plus the one-line hint naming which serial to put in the profile. That alone collapses the two-boot dance from "boot, guess, go read docs, rebuild" to "boot, copy the serial off the screen, rebuild", and it is the same enumeration code the full picker needs. **Why it matters beyond convenience:** it is the BYO / reinstall flow — a customer's existing hardware, or a rebuild of a box whose disk layout nobody recorded, is exactly where the serial is unknown and a wrong guess is destructive. The current fail-safe is correct but mute. Origin: TASK-G, arming the HP install ISO — the serial had to be read off the board by hand between two boots | -| R-62 | **[P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape.** | XS | **idea (operator, 2026-07-22)** | Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; **wording only** — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). | -| R-64 | **„Felhom↔Felhom media pairing blessed" — the two-box SMB pairing (one box shares, the other mounts it as NAS storage) becomes a supported, documented flow.** | XS–S | idea (2026-07-22) | Origin: the operator ran the pairing drill on the live demo pair and it WORKS — the drill itself is the pending evidence leg (a written run-through with the R-66 surfaces in play). R-66 shipped the enabling visibility: the serving box's address is now on its own Beállítások → Rendszer „Hálózat" card, and the add form names the NetBIOS trap. Blessing = a short customer-facing recipe (`documentation/controller/network-storage-nas.md` naming-caveat paragraph is the seed) + one supported-path sentence in the capability map. Flips: would add a "Felhom↔Felhom media pairing" capability row (currently unlisted). Pairs with R-65 (same two-box topology, entirely different transport + guarantees) | -| R-69 | **F14-full: an operator push channel that actually interrupts (ntfy / Telegram / similar), beyond mail-client priority flags.** F14-light (v0.71.0 headers + Gmail filter) nudges a mail client; a 15:29 node_down should reach the operator's pocket in seconds regardless of inbox hygiene. Needs: channel choice (self-hosted ntfy on k3s vs Telegram bot), dispatcher fan-out seam, per-severity routing, quiet hours. | M | idea | Origin: `AUDIT-power-outage-recovery-2026-07-22.md` F14. Deliberately NOT built in the v0.71.0 train (scope-forked per the task spec) | - -### Recovery-model gaps (2026-07-28, `07-backup-architecture.md` §10.2) - -> Minted when `07-backup-architecture.md` was rewritten as the recovery model. Every one of these is -> a divergence between that model and the system as it is, and each is cited there. **They are filed -> at P2 as the neutral default, not ranked** — ranking them needs the per-scenario RTO/RPO targets -> that `07` §11-C records as never having been stated. Flips: `00-capability-map.md` §C rows, which -> now cite the matrix rather than restating the route. - -| ID | Item | Size | Status | Notes / map rows flipped | -|----|------|------|--------|--------------------------| -| R-127 | **`data_key: true` is unreliable (4+ encryption keys unflagged, contradicting the catalog's own labels), and O4 can regenerate a DB password that no longer matches the restored data directory** | S/M | READY — NEW 2026-07-30 | Found by **D5's Part 0**, and the reason D5's boundary became `type: secret` rather than `data_key`. Leg (a): flag the missing keys (catalog-only) + pin flag-vs-label agreement; the residual risk after D5 is that the **fail-closed gate** keys on `data_key`, so an unflagged key missing from both sources lets the restore proceed onto undecryptable data. Leg (b): a regenerated DB password is silently wrong — `POSTGRES_PASSWORD` is ignored once PGDATA is non-empty, so the app cannot authenticate while the replay still succeeds over the local trust socket (proven live on `postgres:16-alpine`). v0.188.0 corrected the false "stored data is unaffected" WARN but added no guard. Flips: `07` §7.4 | -| R-133 | **The vaulted break-glass console credential is PLAINTEXT AT REST, so every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim (`hub/internal/store/host_recovery.go` — *"a hub-held secret, operator-retrievable (NOT zero-knowledge like escrow)"*), so anything that copies the SQLite DB — a Longhorn snapshot, a PBS backup of the hub PVC, a hand-taken copy during a diagnosis — carries root console access to every Felhom host in one file, with no second factor and no key to withhold | M | READY — NEW 2026-07-31 | **The DEFERRED LEG of hub v0.84.0**, filed as its own ID because v0.84.0 changed only WHO can retrieve the secret, never how it is stored — the at-rest shape predates it and is untouched by it. **v0.84.0 made it more worth doing, not more broken:** putting retrieval behind the hub session means the hub login password alone now unlocks console root fleet-wide, so the DB and the login are the whole of the protection. Shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held OUTSIDE the DB** (k8s Secret / out-of-band file, the way `manifests/` already keeps the bearer out of git), so a DB copy is opaque the way escrow blobs already are — the contrast is the argument, since the hub already proves it can hold a secret it cannot itself read. Constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the entire point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Flips: `architecture/00-capability-map.md` **"Break-glass management-plane recovery"** — the row that today reads IMPLEMENTED with a plaintext-at-rest caveat | -| D5 | ~~**Move app secrets into the LOCAL recovery unit** so Tier-1/Tier-2 restore stop needing the guest and stop needing R~~ | M | **SHIPPED + PROVEN-LIVE** — controller v0.188.0, 2026-07-30 | **The arc's architectural centrepiece. Tier-1/2 no longer depend on the whole-guest tier — a customer needs the DRIVE AND NOTHING ELSE.** Part 0 tested this row's own premise and **rejected** it: data-keys-only is both insufficient and unsafe, because `data_key` is unreliable (→ **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty, so a regenerated value leaves the app unable to reach its own restored rows while the dump replay still reports success — proven on `postgres:16-alpine`). **Operator ruling: `type: secret` travels (45 fields), `type: password` never (7) plus a code register (`vaultwarden/ADMIN_TOKEN`); plaintext, because withholding the internet-reachable class is what licenses it — the two are coupled.** `stacks.PortableSecretEnvVars` is the single boundary; the register is code, not a catalog flag (R-97a). **Precedence: the UNIT WINS** (its secrets match the data being restored, not merely the newest), pinned both directions. Fail-closed data-key gate UNCHANGED. Manifest schema 2; schema-1 units still restore. Proven live on a scratch drill guest: AdventureLog restored with the guest `app.yaml` moved aside (`secrets recovered=2/2`, 27.6 s) and **the app read the seeded row over TCP with its own credential**; Grafana's admin password withheld with **0 hits** across the backup namespace. 4 red-proofs each verified to land. `audits/D5-drive-alone-restore-2026-07-30.md` Flips `07` §3/§7.1/§7.3/§7.4/§8/§10.1 + a new capability-map row | -| R-126 | **A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | S | READY — 2026-07-30 | Split out of R-108 on its closure. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (`07` §7.3 records the reasoning). Fix = filter network paths from the export destination list, or force the bundle password when the destination is a share. Flips: `07` §5 | -| E-2 | **Drive-role machinery around the moved vzdump target.** The 2026-07-28 runbook proved the architecture change by hand on both demo boxes; this is the machinery: a **backup-target role** on `StoragePath` beside `Schedulable`/`IsDefault`/`Kind`; **assignment in the storage wizard** (suggest by attribute, refuse the absurd, never decide by transport or `removable` — on the reference hardware demo-felhom's target IS a USB HDD and BOTH drives report `removable=0`); **unassigned drives do nothing automatically**; **stickiness** (never silently retarget); `felhom-host-install.sh` creating the target with `--is_mountpoint 1` **and** issuing the `FelhomAgentStore` ACL; **absent-target policy**; **retention/space accounting** on a drive the customer shares; the honest **single-drive label**; remaining fleet migration | M | READY — 2026-07-28 | Full scope + rationale in `runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md` §7. Two traps already paid for live: the storage `path` must BE the mountpoint or the agent reports the target `disconnected` forever (`internal/storage/observe.go:321`), and the per-storage `FelhomAgentStore` grant is mandatory or every backup 403s. Absent-drive behaviour today is **fail-loudly, no silent retarget** (`is_mountpoint 1` proven live) — which is NOT the intended fall-back-and-alarm design. Flips: matrix row 4 | - -## Gating candidates — Campaign 12, Part 4 (2026-08-08) - -**Ranked by what a gate would be worth, using Campaign 12's own instance counts as the evidence.** -Nothing here was built; §4 of `audits/CAMPAIGN-12-class-sweep-2026-08-08.md` carries the reasoning and -the measurements. The recurring lesson these rank against: **a pattern found three times is not closed -by looking a fourth time.** +**Ranked by what a gate would be worth**, using Campaign 12's own instance counts as the evidence +(`audits/CAMPAIGN-12-class-sweep-2026-08-08.md` §4). G-1 was built (`scripts/wire_contract_gate.py`); G-6 and G-7 +were recorded as a NO with evidence — all three are in `ROADMAP-HISTORY.md`. **Flips:** none — gates change no +capability; they keep the map's claims true. | Rank | Item | Size | Status | Notes | |----|------|------|--------|-------| -| ~~**G-1**~~ | ~~**Gate C5 — the cross-repo tag-reachability check.**~~ | S | **BUILT AND CLOSED 2026-08-08** — `scripts/wire_contract_gate.py`, `--fast`, registered in `repo_gates.py` | Shipped as ranked. **Built BEFORE the fixes and seen failing on 40 fields** (`documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md`) — the order was the method, because `deadcode` had been rejected for C6 the night before precisely for failing that test. 210 tags checked across 3 declared wires, 51 skipped (generic / opaque / allowlisted, each with a reason). Carries a `--selftest` that plants an unreachable tag on a real root and asserts conviction, and publishes its blind spots in both its docstring and its output. **Two instrument defects the control caught before it was trusted:** a substring false negative (`grep -F healed_at` matched `privsep_healed_at`), and treating `dr_recipe` as wholly opaque when its top-level section keys ARE decoded through an allow-list that already cost `offsite_restic` (R-122) — it is now opaque only BELOW depth 1. **Estimate held:** the size guess was right and the `--fast` judgement was right. **Not covered, and stated in the gate itself:** the hub's desired-state (served as raw stored JSON, no typed emitter) and the agent's local API (no single root). → R-260 CLOSED, R-247 CLOSED, leftover appetite R-264 | | **G-2** | **Gate C3 — a success verdict may not be set where an incompleteness signal is in scope.** Assert that every literal success-status assignment either has no gap/skip/missing signal available at that point, or consults it. | S | candidate | The whole population in the controller is **9 sites** — Campaign 12 read all of them, which is why this class is the one where "no others exist" is supportable. Small enough to gate by enumeration rather than by inference. **Known miss:** verdicts expressed as booleans, enum constants, or the absence of an error — and the controller does use those elsewhere. Instances: R-240 (open), R-258 (new). | | **G-3** | **Gate C4 — every rendered count/size/percentage needs a `*Known` companion.** | M | candidate — **UNBLOCKED 2026-08-08: the convention decision it was waiting for has been made** | The decision owed was *"how does this codebase say 'we could not look'"*, and it is now ruled (`CONTEXT.md` **S-39**, shipped in controller v0.210.0 / R-259): **an explicit `…Known bool` companion beside the figures, checked in the template before anything is rendered** — the shape `Offbox.StatsKnown` already used, whose own comment carries the reasoning (*"a 0%-wide bar over an unread store is a picture of emptiness, and a picture is a claim"*). Pointers and separate error fields remain legitimate Go and both still exist here; the ruling is that **new** three-state figures use the companion, because a codebase with three dialects cannot be gated by a name-based check. **Existing call sites were deliberately NOT converted** — that conversion is the bulk of this item's M and is what remains before a gate can be turned on without a wall of false positives. **Next step is therefore a survey, not a gate:** count the rendered figures that lack a companion, decide which are genuinely three-state, convert those, then gate. Instances so far: R-225 (fixed, the pattern's origin), R-259 (fixed, the ruling). | | **G-4** | **Complete C1's runtime body assertion — 4 of 27 pages today.** | L | candidate — the expensive one, and honestly so | `secret_in_markup_gate.py` covers all 36 templates on the NAME-based check and is **blind to a secret under a neutral page-data key** — verified 2026-08-08 by replaying the three pre-fix templates through it: it convicts 2 of 3 and not the third. The runtime assertion catches all three; extending it means constructing each remaining page's data in a test, which is a per-page cost and is the real reason it has not been done. **Do NOT adopt the Go-side mirror Campaign 12 wrote as a gate on its own** — 27 candidates, 0 findings is bad signal-to-noise in front of every push. → R-255 | | **G-5** | **Gate C7, narrowly — uniqueness claims only.** A comment saying "X is the ONLY writer/place/caller of Y" is mechanically falsifiable; assert it. | S | candidate — narrow by construction | Covers ~60 of the 2652 production invariant comments. The other ~97% of the vocabulary (`never`, `always`, `must not`, `guarantees`) is not mechanical and a gate must not pretend otherwise. Instance: R-263. | -| **G-6** | **C2 — NOT mechanically gateable.** | — | **recorded as a no** | "Names a route" is a judgement, not a predicate. The most a check could do is enforce a *convention* (e.g. every customer-visible refusal string ends in an imperative clause), which would be gamed rather than followed. Better served by the UI-copy review the `felhom-ui-design` skill already governs. Instances: R-256, R-257. | -| **G-7** | **C6 — NOT gateable, and the measurement is the finding.** | — | **recorded as a no, with evidence** | The class *looks* the most mechanical of the seven and is the one where the off-the-shelf tool measurably fails: `golang.org/x/tools/cmd/deadcode` re-found **neither** known instance, and a **planted probe** showed why — it reports an unreachable exported FUNCTION but not an unreachable exported METHOD on a widely-used type, and both known instances are methods. A bespoke gate would be conservative by construction and would spend most of its output on **inert dead accessors** (4 of the 5 residual candidates were inert). **Recommendation: gate C5 instead** — it catches a strict subset of the same "the answer was available and discarded" family with none of the ambiguity. → R-261 | | **G-8** | **R-242's untouched half — catch a SKIPPED VOUCH.** | S then M | candidate — **the one with a live recurrence** | Measured during Campaign 12's own bake: `golden_currency_gate.py` **flipped red→green the moment the evidence DIRECTORY existed**, before the round-trip download finished and with no vouch near it. **(a) Cheapest, now:** have the bake session re-read `/configuration` after the operator's Save and write the observed `golden_version` into the evidence README as a machine-readable line; the gate then requires that line rather than the directory. Detects the FORGETTING, which is the actual failure mode. **(b) Loudest, when the hub is next touched:** a hub-side daily check comparing the vouched golden against the newest controller the fleet reports — it fires within a day and catches a silent ROLLBACK too, which nothing in git can ever see. **(c) A checklist item is not a fix** — R-242 already was a rule without a mechanism and it recurred the next day. → R-242 | -## P3 — post-alpha +## Before the first paying customer -| ID | Item | Size | Status | Notes | -|----|------|------|--------|-------| -| R-26 | **Guided old-history recovery via a retained superseded escrow + the recovery code.** Enabled by hub v0.60.0 (Part B) which now RETAINS superseded escrow blobs (`host_escrow_superseded`, `ListSupersededEscrow`). Build the flow that, given the customer's recovery code, unwraps a retained old blob → recovers the old repo passphrase → mounts/reads the moved-aside `.orphaned-` repo for restore. | M | idea (enabled by v0.60.0) | Turns "history recoverable in principle" into a real customer-drivable path; pairs with the controller v0.142.0 orphaned-repo move-aside. Origin `DIAGNOSE-offbox-repo-orphaned-2026-07-17` | -| R-27b | **Customer self-bind, second-box flow (controller side).** For a customer who ALREADY has a bound box and installs another, the controller shows a dismissable "bind another box" prompt (and a bind-later entry under settings) that walks to the hub `/bind/` page — so a returning customer isn't emailed a fresh operator-sent link for every box. Mechanism sketched in the hub v0.66.0 REPORT; NOT built (R-27 slice 1 deliberately did not touch the controller). | M | idea (minted by hub v0.66.0) | Origin: hub v0.66.0 slice-1 ship (first-box only). Reuses the same `/bind/` public page + tokenized-link machinery; adds a controller-side entry point + the operator "mint a link for an existing customer" affordance | -| R-25b | **RULED: customer DELETE becomes a guided full-teardown cascade.** The middle-tier Customer RESET (hub v0.61.0) runs the full external teardown (Hetzner sub-account/box + PBS namespace/groups/token) and refuses while any host row exists. The Danger-zone DELETE still (a) leaves host rows and (b) does NOT run that teardown. | M (was S) | **SHIPPED hub v0.69.0 (2026-07-21)** | **operator ruling 2026-07-21**: **DELETE subsumes the whole cascade, behind explicit consent.** Three separate acknowledgements, each its own checkbox — (1) the host(s) will be deleted, (2) the customer will be RESET including external teardown and offsite data destruction, (3) the customer record and escrow will be purged — plus a **typed customer-name confirmation** before the button arms. Internal order is **host-delete → RESET → delete**, which preserves every existing invariant rather than relaxing any: RESET keeps its no-hosts precondition (hosts are already gone by then), and escrow keeps its demote-then-purge custody rule (host delete DEMOTES to retained custody, the final delete PURGES — the one true purge point). **Re-sized S → M: this is a multi-step destructive wizard with three acks and a typed confirmation, not a checkbox.** Implementation is explicitly NOT part of TASK-E; the row carries the ruling and awaits its own spec. **It no longer blocks R-3** — the model is decided, so the friend-alpha runbook can be written against it. **IMPLEMENTED per the ruling (TASK-I, hub v0.69.0):** `POST /configs/{id}/delete` now runs `hosts → RESET → purge`; three acks + typed customer-id + a stale-preview check + the ONLINE-host refusal, all gates before any write (zero side effects on refusal); custody purged exactly ONCE in leg 3 (leg 2 runs with `purgeEscrow=false`); ruling-3 preserved BY CONSTRUCTION and asserted from inside leg 2; failed legs retain the journal and the dialog offers Resume. Standalone RESET byte-identical. 5 red-proofs. Offboarding guidance: `runbooks/RUNBOOK-onboarding-draft-v4.md` §G. **v0.70.0 follow-up (same day, found validating against the live hub):** a completed delete still left the customer on the Customers list and still ALERTING, because `GetCustomers()` is report-derived and no tier ever deleted a report — new **residue** leg (reports/telemetry/log-tails/notif-prefs + the credential-bearing `appliance_registrations`/`selfbind_tokens`), and **ghost customers are now deletable** (404 = nothing here, not no-config-row). **v0.70.1 (2026-07-22): the ghost delete was implemented but UNREACHABLE** — the Danger-zone card (and the `customerDeleteOpen` script) sat inside `{{if .HasConfig}}`, so a ghost rendered no Delete button at all (the fourth inert-seam defect; handler tests POST directly and proved nothing about reachability). Render gate split: RESET stays HasConfig-gated, Danger zone gates on `Deletable` (the exact negation of the preview's 404 predicate), Block/Unblock stay config-only; render tests per branch + 2 red-proofs. **Operator live leg: the demo-vm-felhom ghost delete click — PENDING** (doubles as the v0.70.0+v0.70.1 live validation; expect `residue=ok customer_delete=ok` with `skipped_no_config` Hetzner/descriptor legs, staleness emails stop) | -| R-12 | Cluster mode: agent-follows-guest, bind-mount reconciliation on HA migration | XL | idea | Scoped 07-15; interim = HA-group pin to one node. Driven by Peti's two-node cluster | -| R-14 | Headscale/WireGuard spike: Minecraft/gaming port connectivity (CGNAT-proof, sovereign DERP fallback) | M | idea | | -| R-15 | Multi-user dashboard accounts (household members, roles) | L | idea | Single password is a stated alpha limitation (R-11). **Launcher coupling — REVISED (controller v0.165.0):** the "share the launcher outside the household" need is now met WITHOUT member accounts — the **Indítópult megosztása** capability-URL guest link (`/s/`, information-only, no account) shipped in v0.165.0. What remains for this arc is member-specific: **per-member tile visibility** (each member sees only their apps) and the launcher-as-member-landing-page — both live inside this SSO/members arc; the guest-link ruling explicitly SUPERSEDES the earlier "members are how you share the launcher" framing | -| R-72 | Curate `brand_color` for the top catalog apps | XS | idea | Parked follow-up to the v0.163.0 launcher. `.felhom.yml` `brand_color` (`#rgb`/`#rrggbb`) overrides the deterministic slug-hash tile color; no catalog app sets it yet. Pick brand-accurate colors for the most-installed apps so their launcher tiles match their real brand. Catalog-only change (`app-catalog-felhom.eu`), `brand_color` is already `omitempty` and consumed by the controller | -| R-74 | **Island control plane on a CLUSTER (Peti's 2 nodes)** — bring R-50's island bridge to a multi-node PVE cluster. | M | idea (Phase C of R-50, parked) | R-50 shipped the island for the ONE-host fleet (demo-hp, demo-felhom). A cluster needs **bridge parity on every node**: either per-node identical `/etc/network/interfaces` `vmbr9` stanzas (simplest, drift-prone) or — preferred at ≥2 nodes — a Proxmox **SDN zone/vnet** defined cluster-wide (one definition, auto-applied per node). The guest island IP is per-guest + node-independent; the **agent-follows-guest** rule holds (each node's agent binds its own `vmbr9` `169.254.253.1`). Migration order per the spike: drill-proven → demo (done) → **Peti (this row)**. Its own supervised runbook, coordinated with Peti (a live customer). Completes the capability-map "site/network change" row for clustered installs. Source: `audits/SPIKE-island-bridge-2026-07-25.md` (cluster-parity finding) + `RUNBOOK-island-migration.md` (single-host procedure to generalise) | -| R-87 | ~~**The restic (app-data offsite) tier is NEVER restore-tested**~~ | M | **SHIPPED 2026-08-31 — controller v0.231.0, RE-SCOPED by its own spike.** Not built as written: the spike measured that an unattended scratch restore would have caught ONE of five drill-found restore defects, so it proves the snapshot CONTAINS a recoverable unit rather than testing the restore code. See `audits/SPIKE-restic-restore-test-2026-08-31.md`. | **R-85 covers whole-guest vzdump tiers only** (`local`, `felhom-pbs`); the agent has no restic surface at all. restic is the CONTROLLER's app-data offsite backup to the Hetzner Storage Box, a separate mechanism — so the tier that is arguably most important to a customer is the one nothing verifies. It is the only tier that survives losing the box **and** carries their actual app data: the whole-guest snapshot deliberately excludes the bind-mounted data drives (`/mnt/felhom-drives`). Restore code exists and has been exercised BY HAND (the immich destroy-and-recover drill, PROVEN-LIVE), but nothing tests it unattended — **exactly the state PBS was in before R-85: it works when someone tries it, and nobody would know if it stopped.** Needs its own design: a restic restore-test is controller-side, has no scratch-guest analogue, and would verify into a scratch dir rather than a booted guest, so R-85's machinery does not transfer. | +**One list, not two:** root `STATUS.md`, section "Before the first paying customer". The old "Pre-invite checklist" +(2026-07-18, the remote-tester era) is in `ROADMAP-HISTORY.md`; its open legal half — the **R-11** rulings +(the tester agreement was never written) — is now part of **R-809**. -Each attempt runs the **full quiesce cycle**, so every customer app stack is STOPPED and RESTARTED for a backup that cannot succeed. Measured on demo-felhom: `07:07:58 quiescing 4 stack(s): [bookstack calibre-web docmost immich]` → `07:08:17 unquiescing (backup failed)` → `07:08:45 failed` — **~19 s of app downtime per cycle (~50 s per full cycle), every 5 minutes.** +## Loose notes in this folder -**The amplifier, and the part worth designing against:** the agent answers `Due: true, Reason: "no successful backup recorded yet", AgeSecs: nil`, and that **nil age does double duty**. `scheduledRunAllowed` (`quiesce.go:466-480`) returns `true` whenever `lastAgeSecs == nil` — *"no recorded backup yet — never withhold the first one"* — so the same nil that makes every poll due **also bypasses the time-of-day gate** `[W+2h, W+6h)`. On the live box the gate was `[04:30, 08:30)` and the cycles ran at 09:02–09:12 Budapest, i.e. **outside the backup window entirely**. So the fault stops customer apps every 5 minutes *at any hour, including business hours* — the one protection specifically built to prevent that is switched off by the same missing value. A safety valve written for a genuine first-ever backup is being triggered by an unreachable storage read, which is not the same thing at all. - -Self-resolves the moment the target answers (the storage read succeeds, sees the archive, tier stops being due) — which is why it can hide indefinitely: it needs an offsite outage to appear at all. **PHASE-0 ROOT CAUSE, established at source 2026-07-27 — it is AGENT-side, case (a).** The storage read **errored** (`could not read the backup storage for the due-check … err=…` at 09:02:57/09:07:58/09:12:57 CEST), so this was never an empty-success. The failure is a **type boundary**: `newestArchiveOn` (`localapi/server.go:1095-1111`) documents *"Errors and unsupported services degrade to unknown, never to 'no backup'"* — but its `(time.Time, bool)` signature **cannot represent unknown**, so an error and a genuinely-empty storage both collapse to `(zero, false)`, and `handleBackupDue` (`server.go:934-941`) then emits a POSITIVE claim: `Due: true, Reason: "no successful backup recorded yet", AgeSecs: nil`. The fail-safe that *does* exist — `targetStoragePresent`'s "a storage-view error must never be read as 'not there'" (`server.go:1131-1151`) — answers a different question (does the storage exist) and behaved correctly. **Decisive for scoping: the errored path and the genuine-never path are BYTE-IDENTICAL on the wire** — same `Due`, same `Reason` string, same nil `AgeSecs` — so the controller has nothing to discriminate on and Part 2 CANNOT be done controller-side. **Two further P0 findings:** the agent restarted **4× on 2026-07-27** (07:36:39, 07:54:06, 08:50:16, 11:31:52 CEST) — all deliberate (`NRestarts=0`, `Restart=on-failure`, `Result=success`), zero self-update — so the trigger is armed by ordinary operator/config work far more often than "only when ep0 is down"; and **the loop alerted NOBODY** — zero `backup_failed` events despite the hub allowlist carrying that type, because **`internal/quiesce` does not import `internal/notify` at all**. Its only trace was `07:13:27 info app_start_failed "Telepített alkalmazás nem fut: BookStack"` — a customer-tier, Hungarian, info-severity SYMPTOM of the third cycle catching BookStack mid-restart. **The whole-guest backup tier R-82 built has no failure signal to the hub → its own item.** **Shape:** distinguish *storage unreachable* from *storage readable and empty*. Unreachable is UNKNOWN — defer the due-verdict rather than resolving it either way, exactly as R-81 made the hub do with a missing report. Only a target that is reachable AND has no archive is genuinely due. **Fix the window bypass in the same slice:** `AgeSecs == nil` must stop meaning "run now regardless of the hour". Either the agent distinguishes *never backed up* from *cannot tell* in what it reports, or `scheduledRunAllowed` gates on the former only — otherwise any future nil-age path re-opens the same hole. Note this does NOT weaken R-84's fail-safe intent: a tier whose storage is merely slow or briefly unreadable should still err toward backing up — it is specifically the **cold-store + unreachable** pair that must defer, because there the fallback has no information at all, only an empty default that looks like a fact. | -| R-95 | **The restic offsite tier's credential CAN DELETE — R-89's "parallel question", now ANSWERED** | M | idea — established read-only 2026-07-27 | **The exposure closed on the weekly PBS tier is fully open on the daily restic tier**, which holds the customer's actual documents and photos and is the only tier that survives losing the box. Established without mutating anything: **(1) Identity** — a per-customer *subaccount* on `storage-box-pool-1` (box 611714, bx11, `u629488`): `u629488-sub1` home `felhom-demo-felhom`, `sub2` peti-felhom, `sub3` demo-hp, each labelled `felhom-customer`. Auth is an **SSH key stored ON THE BOX** (`…/felhom-controller-data/_data/data/offbox/ssh_key`, 0600, beside `repo_password` + a pinned `known_hosts`) — customer-side, not hub-side, so a compromised guest holds it. **(2) Read-write: YES** — the API reports **`readonly=False` on all three subaccounts**, and it is not merely latent: the controller runs `restic forget --group-by host,tags --keep-daily 7 --keep-weekly … --prune` **from the box** (`backup/offbox.go:984`, also `:1070`). Delete rights are exercised on every run. **(3) Append-only: NO, and not expressible** — the repo is built as `sftp:` (`offbox.go:482`); restic's append-only mode requires the **REST server** backend, which plain SFTP cannot provide. **(4) A zero-code mitigation exists and is unused:** the box type carries `snapshot_limit=10` and the API reports `snapshot_plan=null` with **0 snapshots** and `size_snapshots=0`. Hetzner Storage Box snapshots are taken **server-side, outside the SFTP namespace** — an SFTP subaccount cannot delete them — so they are a genuine immutability layer at no extra cost and with no code change. **Rule once for both tiers, per R-89.** Options, cheapest first: enable a snapshot plan (operator click, immediate); split backup-write from prune so pruning runs somewhere the box cannot reach; or move the repo to restic's REST server with `--append-only`. Flips the capability-map row for offsite immutability | -| R-162 | **`docker diff` is the gate's only witness and its failure mode is quiet** | XS | WATCHING — 2026-08-02 | A limitation, not a defect. The gate's power is `docker diff` excluding mounted paths; on a driver where it is unsupported or lies, the gate degrades to mount-occupancy + writability **and would not say so**. It fails closed (the canary self-test stops reporting BROKEN and the gate then refuses to report), but the message blames the prober rather than the driver. Revisit only if a non-overlay driver ships | -| R-164 | **C2's chain — the DB volume tar cannot be dropped until a sound dump predicate exists** | S | BLOCKED — on the predicate (2026-08-02) | The unit holds a volume tar **and** a SQL dump and the restore uses both: the dump is authoritative and replayed after the tar so it WINS (F17), with only the DB service up (R-47) — `restore_unit.go:262-266`. Dropping the DB tar would halve DB-app units and close R-127(b)'s initdb-skip trap. **The obvious gate is dead, measured:** `ValidateDump`'s empty-`accounts` warning was **correct** (the DB truly had 0 rows; seeding one stopped the warning and put the row in the dump) — but **a fresh appliance legitimately has zero accounts**, so gating on it blocks every new customer's first backup. Order: sound predicate (dump vs **live** per-table counts) → warn→gate → tar-drop. Pairs with **R-127** | -| R-91 | **The old 13 GB datastore copy is still on ep0's root disk** | XS | WATCHING — gated on demo-felhom's first post-migration PBS backup | The datastore moved to a Hetzner Cloud Volume on 2026-07-27 (`/dev/sdb`, 100 GiB, attached 06:29:40 UTC, now `/mnt/pbs-datastore`, 13 G used of 98 G). The pre-migration copy survives at **`/srv/pbs-felhom`, 13 G**, on `/` (38 G total, 16 G used, 21 G free). **Do not delete yet:** demo-hp has landed two post-migration snapshots (07-27 08:25:47Z, 09:37:29Z) but **demo-felhom's newest is 2026-07-26T12:21:48Z — before the migration**, so the new volume has not yet proven a write for that namespace. Delete once it has. **Doc drift to fix in the same commit:** `CONTEXT.md:1018` still records the datastore at `/srv/pbs-felhom` | -| R-92 | **The hub's PBS-DR gauge is 0.1 GB-granular, so small deltas are unverifiable** | XS | idea — 2026-07-27 | The PBS-DR box card rounds to 0.1 GB, which is coarser than the changes an operator wants to confirm after a prune or a GC — a successful prune of a small namespace moves the number by less than one displayed digit, so the UI cannot distinguish "it worked" from "nothing happened". Cosmetic today; it becomes load-bearing the moment retention (R-89) is customer-visible and someone needs to see that a policy change took effect | -| R-93 | **`drill-r50` is both a blocked customer and the only drift fixture** | XS | idea — 2026-07-27 | The drill customer is blocked in the hub (so it stops alarming) yet it is also the only record exercising the endpoint-drift path R-77 added. Blocking hides it from `GetActiveCustomerIDs`, so the fixture it provides is silently inert — a monitor with no live subject reads exactly like a monitor that passes. Decide: retire it and build a synthetic fixture, or unblock it and silence per-customer instead (the operator has a per-alert silencing feature planned). Related to the R-50 drill VM, now shut down | -| R-29 | **The design-v2 green gates are not enforced anywhere — one has been RED for 16 releases.** `controller/scripts/docker_run_volume_path_gate.py` has failed continuously since **2026-07-14 (v0.129.0)** and nobody noticed until R-7b's close-out ran it by hand at v0.145.0. Two separable parts. **(a) The finding itself is benign and the fix is 3 lines.** The flagged call is `internal/appexport/estimate.go:179` `docker run --rm -v :/vol:ro alpine du` — a **NAMED-VOLUME** mount, i.e. daemon-side with no host path, which is the *safe* shape and byte-for-byte the same pattern as three entries already on the gate's ALLOWLIST (`export.go` `volName+":/vol"`, `backup.go` `volName+":/vol:ro"`, `restore.go` `volName+":/vol"`). It is NOT the v0.124.0 path-strand class the gate exists to catch — the author of the v0.129.0 F-A fix explicitly avoided that class (see the function's own comment) and simply never added the allowlist entry. So the fix is an ALLOWLIST addition WITH ITS WHY, **not** a docker-cp rewrite; anyone who 'fixes' this by rewriting the call has misread the gate. **(b) The systemic half is the real item:** the gates run only when a human remembers to run them, so a gate can sit red across 16 releases while every REPORT says 'green'. This is the SECOND instance of the class — cf. the v0.123.0 note *'Windows green gate silently red (read-only fsync)'*. Decide where they run (pre-push hook, `build.sh` step, or a CI job) and make a red gate block the train the way the Go green gate does. | S (a) / M (b) | idea | Origin: R-7b close-out, `felhom-controller` REPORT §4(f) — CC correctly left it alone as out-of-scope and pre-existing, and verified by stashing that it fails identically on the unmodified tree. Flips no capability-map row (engineering hygiene, no customer-visible behaviour). Affected gates to audit for the same rot: controller `template_id_gate` / `emoji_gate` / `native_confirm_gate` / `offbox_rename_gate` / `mojibake_gate` / `app_row_dedup_gate` / `docker_run_volume_path_gate`, hub `hub_confirm_gate`, manifests `manifest_bearer_gate`, website `site_gates`. **Do not bundle (a) into an unrelated feature commit** — it is a one-line behavioural claim about a mount's safety and deserves its own reviewed diff. **2026-07-18 rehearsal note:** the run's finding list independently re-raised "assign the pre-existing `docker_run_volume_path_gate` failure its ID so red stops normalizing" — **that is this item; no second ID was minted.** **2026-07-29 — audit list extended, and a THIRD independent re-raise absorbed under the same rule (again no new ID):** add `scripts/hostinstall_gates.py`, which **postdates this item** (it comes from drill F-1, 2026-07-12) and is therefore not a design-v2 gate — but it is the identical failure shape and is tracked as **R-94 leg (b)**. It is **RED as of 2026-07-29**: `hub Setup-tab hostInstallVersion=1.19.0 != SCRIPT_VERSION=1.22.0`, exit 1, with its nine other assertions green. `scripts/hub_confirm_gate.py`, already on the list above, was **verified orphan on the same date**. Both confirmed by repo-wide grep across all file types plus sibling repos, `~/.claude` settings/skills/hooks, `.git/hooks` (no non-sample hooks exist), a Makefile/justfile/Taskfile find (only `hub/Makefile`, zero `gate` occurrences) and a CI-directory find (**`felhom.eu` has no CI configuration at all**) — all 19 hits are docstrings, code comments or prose; **zero are invocations.** Only `site_gates.py` is mandated (`CLAUDE.md:153`); `manifest_bearer_gate.py` is named in `runbooks/secrets.md:76`. **Now also filed in `OPEN-ITEMS.md`** — this item predates the 2026-07-27 register rebuild and was never carried across, so an open item about work not getting done was itself missing from the page that decides what gets done. **2026-07-30 — THE FIRST ENTRY ON THE OTHER SIDE OF THE LEDGER, recorded so the contrast is not lost:** the **R-120 golden-staleness gate** (hub v0.82.0, `hub/internal/web/configs.go` `handleSetArtifacts`) **IS enforced.** It is not a script in `scripts/` that someone must remember; it sits inside the only UI path that writes `SetArtifactManifest`, so it runs on every vouch whether or not anyone chose to run it, and it **refuses** (operator ruling, 2026-07-30) rather than warning — because this row's whole finding is that a non-blocking check reads as coverage it is not providing. It compares the submitted golden against the newest controller any box has reported (`store.NewestReportedControllerVersion`) and is pinned by four tests driven through the production handler over `httptest`, not an injected seam, plus a red-proof: deleting the block makes the stale golden vouchable again. **Note the near-miss worth keeping:** the first draft read `guests.controller_version`, a column that exists in the schema and that **nothing writes** — it would have been an inert gate, i.e. this row's exact failure shape, caught by grepping for a writer before trusting the column. **The three orphans above are unchanged and still orphaned** — this entry proves the pattern is available, not that the backlog moved **UPDATE 2026-08-02 — leg (a) CLOSED** (`felhom-controller` `c432f70`, its own reviewed diff); **leg (b) HALF-SHIPPED**: every repo now has ONE entry point wired to `.githooks/pre-push --fast`, each mandated in its `CLAUDE.md`. The census that drove it: 13 gates, and every gate a `CLAUDE.md` names was green while two of the four unnamed ones were red. Stays open for the automatic half → **R-168** **CLOSED 2026-08-02** on the demonstrated alarm, not on a green run: both halves are live (local hook refuses; CI notices a bypass and emails). Remaining is a working-style choice → R-169 | -| R-169 | **CI can only report, because there is no gate in the road** | S | idea — minted 2026-08-02, **WAITING-ON-OPERATOR** | Making CI blocking needs branch protection on `main` plus a pull-request workflow instead of direct-to-`main` pushes — both change how the operator works, so neither was done. Current arrangement is two nets: the pre-push hook refuses locally, R-168's runner notices a `--no-verify` bypass and emails. Decide only if that window ever costs something. Detail: `OPEN-ITEMS.md` R-169 | - -| R-41 | **[SLICE 1 SHIPPED 2026-07-21] The catalog has no standing "does every template still deploy?" check.** Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found **5 that had NEVER been deployable**: papra (missing required `AUTH_SECRET`), zipline (v4 renamed `CORE_DATABASE_URL` → `DATABASE_URL`), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the `v` tag prefix + new required env), glance (needs a seeded `glance.yml` the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus **7 broken healthchecks** and 2 apps whose images no longer resolve at all (plant-it, wanderer). | M | idea | Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in `scripts/check-image-pins.py` — a mechanical gate run on every change. Cheap first slice: a **resolvability gate** (`docker manifest inspect` every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. **Silent rot is the real risk** — an app can die upstream and nobody learns until a customer clicks Telepítés | **SLICE 1 SHIPPED 2026-07-21 — `app-catalog-felhom.eu/scripts/check-image-resolvable.py`** (+ 14 fixture tests, no network, resolver injected). Resolves every unique pin with `docker manifest inspect`, ONE image at a time; exit 0 / 1 (the registry says GONE) / 2 (inconclusive). **Two traps encoded, both hit live while building it:** (a) `docker manifest inspect` prints `toomanyrequests: …` and **still exits 0** — the same exits-0-on-failure shape as the ISO tooling's `validate-answer`, so stderr is inspected even on rc=0; (b) the inverse and more dangerous one — the first full sweep called **24 of 65 pins dead, including `postgres:16-alpine` and `redis:7-alpine`**, purely because Docker Hub throttled it partway through. Ambiguity therefore resolves to INCONCLUSIVE and never to an accusation: a gate that cries wolf gets ignored, and then it protects nothing. **The full sweep is still OWED** — DooPlex is not logged in to Docker Hub, so the 52-app table needs one re-run after `docker login`. Wired into `CLAUDE.md` + `REUSE.md` as a start-of-campaign / pre-publish-train step. It immediately paid for itself: it is what turned plant-it and wanderer from 'images do not resolve' into two DIFFERENT diagnoses (see the 2026-07-21 catalog entry). Full slice — the periodic deploy-all sweep on the demo box — remains open | -| R-45 | **[P2] Unified async-job feedback.** Every long operation invents its own progress surface, or none. Tonight produced three more one-off cards (v0.147.x: samba bring-up, offsite progress, restore result) on top of two existing patterns (deploy 3-step panel; storage-init/netstorage status poll). They agree on nothing: some use `{ok,data}` envelopes and some raw JSON, some poll 1 s / 1.5 s / 3 s, some are in-memory-only and lie after a restart, and each re-implements single-flight + snapshot + phase→Hungarian mapping. | M | idea | Origin: 2026-07-19 feedback slice 1 (controller v0.147.0). The cases to generalise from are all in-tree: `web/storage_init_job.go` (the best shape — acquire/release/set/snapshot), `web/netstorage_job.go`, `web/samba_ensure_job.go`, `backup/opstatus.go`, `backup/offbox_progress.go`. Shape: one job registry + one poll endpoint + one client-side renderer, phases declared per job. **Two lessons tonight that any framework must encode:** (1) a terminal state must be **probed, not inferred** — `compose up -d` exits 0 on a crash-loop; (2) a progress source that reports nothing is normal, not broken — restic reports 0 bytes for a whole incremental run, and a bar that sits at 0% is worse than no bar. Also fixes the restart hole: in-memory job state currently vanishes and the card silently disagrees with reality **2026-07-20 — the first bill for NOT having this arrived, and it was customer-facing.** The samba card's poll (`web/samba_ensure_job.go` + `sharing.html`) mixed a job EDGE and a service LEVEL on one JSON field, and `/sharing` reload-looped at ~1.2 s for every customer with sharing enabled until controller v0.151.0 (`audits/DIAG-sharing-2026-07-20.md`, S-1/S-4). v0.151.0 fixed THAT card's contract only — the framework is still this item. **Third lesson for it to encode, beside the two already listed:** a phase a client answers with a one-shot action must be an EDGE the registry SERVES ONCE, and must never be synthesised from a level; if it can be re-read, it will be re-acted on. | -| R-46 | **[P2] Verification copies need a customer-visible browse surface and an expiry.** v0.147.0 made them *visible* (listed with path/size/date, individually deletable) — but the customer still cannot LOOK INSIDE a verification restore to confirm the file they wanted is really there, which is the entire point of a verification restore, and nothing ever removes them. | S–M | idea | Origin: 2026-07-19 feedback slice 4a, registered as the explicit follow-up to it. Two gaps, deliberately designed together because they are the same object: (a) **the invisible-result gap** — a read-only browse of `backups/offsite-restore/` (the FileBrowser infra stack already exists and already serves scoped roots, so this may be a mount rather than new code); (b) **the disk-lifecycle gap** — auto-expiry after N days with the count/size surfaced before it fires, so a drive is never quietly filled by verification restores nobody remembers taking. Pairs with R-43: a browse surface is also how a customer would discover that a DB-indexed app's files came back but the app still cannot see them | -| R-48 | **[P2-HIGH] Restore controls are separable only by layout — and the difference between them is whether the data comes back.** The offsite restore row renders four buttons plus hint text into an overlapping, unreadable line, and the decisive second step („Teljes visszaállítás indítása") appears ONLY after „…előkészítése" was pressed, with no signposting that a second step exists or that the first one did nothing to live data. | M | idea | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` (finding 1) — this is not theoretical: it is the CAUSE of the round-2 incident.** An operator who had read the code pressed the missing-only button instead of the full restore; the controller log shows `/backup/offbox/reconstitute` was never hit at all. The rule this establishes, worth stating once and applying beyond this page: **two adjacent controls whose difference is "your data comes back" vs "your data cannot come back" must not be distinguishable only by layout.** Direction (ruled in principle, spec rides v0.149): collapse to a single „Visszaállítás…" guided dialog — one intent, visible phases, the escrow-wizard precedent. Pairs with R-45 (the phases are exactly the async-feedback surface) and R-46 **SHIPPED 2026-07-21 — controller v0.154.0** (`3a9d744`). Each app row on `/backups/restore` now carries ONE „Visszaállítás…" entry linking to a per-app wizard at `GET /backups/restore/app?name=`: three intent CARDS each with a consequence sentence (ellenőrzés külön mappába / hiányzó fájlok visszahozása / teljes visszaállítás), a visible phase strip so the sequence is legible *before* the first click, danger styling on the destructive card, and the R-43 double-confirm carried over verbatim with its pair-honesty facts. `deriveWizardStep` is a PURE function of (op running, size-gate flash, scratch ready) — the step is never taken from the request, and a running op outranks a stale `?full_prep=` so no commit button survives into a restore. While ANY op runs every mutation form is suppressed server-side rather than offered and then refused. **No new mutation endpoint** (one GET route; every card posts to the pre-existing `/backup/offbox/*` with unchanged field names and gates) and **no R-45 graft** — the wizard polls the two existing status surfaces as-is. Works with JavaScript disabled. Latent bug fixed on the way: `offboxRedirectTo` hardcoded `"?"` when appending its flash, which against the wizard's `?name=` target would have buried the flash inside the app name. 9 new tests + the Group-B red-proof (trivial always-INTENT impl → all 7 rows red). **Live click-through + one non-destructive Ellenőrzés still PENDING** (rides the operator's floor save). Evidence: `felhom-controller/REPORT.md` §3 (2026-07-21). | -| R-65 | **Buddy-box backup replication, cross-household — two Felhom boxes in different homes replicate backups to each other.** | L | idea (post-alpha, spike-first, 2026-07-22) | The natural big sibling of R-64: two households each hosting the other's encrypted backup tier. Explicitly **spike-first** — the transport is NOT SMB (R-64's live-share protocol is wrong for backup replication across the internet: no auth story between households, no resumability, cleartext LAN assumptions); candidates to spike: restic rest-server / rclone / syncthing over the existing WG/tailnet plumbing, encryption keyed so the buddy can never read the payload. Sits on top of the offsite tier's FILL/OVERSUB thresholds thinking (R-5 aggregate). Flips: would add a "cross-household buddy replication" capability row (currently unlisted). Pairs with R-64 (same topology, different transport + guarantees) | - -## Pre-invite checklist — what stands between here and the first remote tester - -Not roadmap items in their own right; the short list the 2026-07-18 rehearsal leaves behind. -Everything here is **remote-doable** — the N100 is packed, and none of it needs hands on the box. - -| Action | Owner | Note | -|---|---|---| -| ~~Rebuild the golden → 0.146.0~~ **BAKED + PUBLISHED 2026-07-18; awaiting the operator's two saves** | Viktor (saves) | Golden **0.146.0** baked on the drill VM and published to gitea — `felhom-golden/0.146.0/golden.tar.zst`, **sha256 `4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955`**, 612 696 588 B (584 MB archive). All pass markers green: `Result=success`/`ExecMainStatus=0`, **0** FATAL/exclusions, `docker OK (overlay2)`, **all three mounts included** (rootfs + mp0 `/var/lib/docker` + mp1 `/mnt/sys_drive`), pre-delete **404**, upload **HTTP 201**; controller **0.146.0** confirmed baked in. Integrity round-trip independent of the build host: anonymous `GET | sha256sum` **matches byte-for-byte**, ranged GET **206**, `content-length` matches. Teardown per GL-1: guest 9100 purged, token/script/log shredded in-VM, VM powered off, drill disk reverted to `virgin` exactly-as-found, **token-leak grep = 0**. Log retained `180:/mnt/5_hdd/felhom.eu/drill/bake-0.146.0.log`. Now selectable in the hub dropdown (versions: 0.136.0, 0.143.0, **0.146.0**). **REMAINING = operator, password-gated:** Day-0 manifest Golden → 0.146.0 (Agent stays 0.90.0, MinAgent stays 0.90.0 — the v0.146.0 CHANGELOG declares no new agent coupling) → save; **then** floor → v0.146.0 saved **LAST** | -| **Golden ≥ 0.147.x carries ALL FOUR infra images** | — (next bake) | `build-golden.sh` v2.1.0 (2026-07-19) now derives the pre-pull list from the controller binary it is about to bake (`--print-infra-images`) instead of a hand-maintained copy that had already drifted: `felhom-samba` was never added to it, so every golden so far baked **3 of 4** — which is why enabling Megosztás on a fresh box pulled from the registry with zero feedback. **No golden rebuild for this alone**; it takes effect at the next bake. Until then a fresh box still pulls felhom-samba at enable time, which controller v0.147.0's progress card now at least explains | -| **freemail.hu test-send** | Viktor | The open half of R-4; the gmail half closed on 2026-07-18 under `p=quarantine` | -| **C6 — customer performs a restore, unassisted** | Viktor as customer zero | The one open script step in R-3 and still MISSING as capability evidence. Remote-doable on the reborn box — the dashboard is remote | -| **R-11 rulings** | Viktor | Contact channel, tester agreement, alert thresholds (the R-5 gauge thresholds are still pending a ruling) | - -## Absorbed / superseded notes in this folder - -- `FOLLOWUP-nas-automount-guest-reboot-reassert.md` — **shipped** (agent v0.84/v0.85, CAMPAIGN-3); keep for history -- `FOLLOWUP-golden-default-controller-tag.md` — **FIXED** (agent `ceca355`); moved to `documentation/archive/` 2026-10-03 -- `FIX-M18-NOTES.md`, `FIX-M19-NOTES.md`, `DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md` — historical diagnoses; superseded by shipped fixes; moved to `documentation/archive/` 2026-10-03 \ No newline at end of file +Verdict per file: `backlog/README.md`. diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index b97447ca..f3654f0d 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,11 @@ +## one-register reads suffix ids and splits on pipes outside code (2026-10-03, backlog triage, Part D) + +- `one_register_gate.py`: the id pattern was `R-(\d+)`, so ROADMAP rows R-27b, R-27c and R-50b were never read; + and its naive split read R-50b's state from inside `create\|reconcile\|grant`. Both fixed (suffix ids; + `register_table.split_cells`). The fix convicted R-50b at once — a finding that lived only in `ROADMAP.md`; it + moved to `OPEN-ITEMS.md`. Decoy `one-register/suffix-id-row`: **passed the old gate — seen red**, convicts now. +- `instructions_gate.py`: `ROADMAP-HISTORY.md` counts as history, like the narratives archive. + ## Every open row has a category and a severity (2026-10-03, backlog triage, Part C) - `register_shape_gate.py` **RULES 5–8**: a register row splits (pipes outside backticks) into exactly the cells its diff --git a/scripts/instructions_gate.py b/scripts/instructions_gate.py index aa1f2054..e576a802 100644 --- a/scripts/instructions_gate.py +++ b/scripts/instructions_gate.py @@ -558,11 +558,14 @@ def register_state(workspace_root): # 2026-10-03: the register's narrative sections (campaign write-ups, rulings, ranking # paragraphs) moved word for word to this archive. An id cited only there — R-153, named in the # 2026-08-02 intake paragraph and in felhom.eu/CLAUDE.md — is history, not "a reference to nothing". - narr = os.path.join(workspace_root, "felhom.eu", "documentation", "archive", - "OPEN-ITEMS-narratives-2026-10-03.md") - if os.path.exists(narr): - with io_open(narr) as fh: - hist_text += "\n" + fh.read() + # ROADMAP-HISTORY.md is the same kind of record for the roadmap: an id the 2026-10-03 clean-up moved + # there (a shipped or killed intention) is history too. + for narr in (os.path.join(workspace_root, "felhom.eu", "documentation", "archive", + "OPEN-ITEMS-narratives-2026-10-03.md"), + os.path.join(base, "ROADMAP-HISTORY.md")): + if os.path.exists(narr): + with io_open(narr) as fh: + hist_text += "\n" + fh.read() mentioned = set(re.findall(r"\bR-\d+[a-z]?\b", reg_text)) | set( re.findall(r"\bR-\d+[a-z]?\b", hist_text) ) | set(re.findall(r"\bR-\d+[a-z]?\b", closed_text)) diff --git a/scripts/one_register_gate.py b/scripts/one_register_gate.py index 414a6f10..1bda5a55 100644 --- a/scripts/one_register_gate.py +++ b/scripts/one_register_gate.py @@ -37,11 +37,14 @@ import os import re import sys +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) +import register_table # noqa: E402 + ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) ROADMAP = os.path.join(ROOT, "documentation", "backlog", "ROADMAP.md") REGISTER = os.path.join(ROOT, "documentation", "backlog", "OPEN-ITEMS.md") -ROW = re.compile(r"^\|\s*\*{0,2}R-(\d+)\*{0,2}\s*\|") +ROW = re.compile(r"^\|\s*\*{0,2}R-(\d+[a-z]?)\*{0,2}\s*\|") # 2026-10-03: a suffix id (R-27b, R-27c, R-50b) was skipped unchecked # DONE and IDEA are matched against the row's STATE cell only, never the whole row: the body of a # finding routinely contains the word "shipped" while describing something else. # BANKED and PROVEN-LIVE are this project's own done-words and were found by running the gate: a row @@ -58,8 +61,11 @@ def rows(path): m = ROW.match(line) if not m: continue - cells = line.split("|") - state = cells[4].strip() if len(cells) > 4 else "" + # 2026-10-03: split on pipes OUTSIDE backticks (`register_table.split_cells`). R-50b's item + # carries `create\|reconcile\|grant` in code, and the naive split read its STATE from the + # middle of that — so the row was judged on words that were not its state at all. + cells = register_table.split_cells(line) or [] + state = cells[3].strip() if len(cells) > 3 else "" yield m.group(1), state, line diff --git a/scripts/test_gate_decoys.py b/scripts/test_gate_decoys.py index 70855d05..e3a6fdcf 100644 --- a/scripts/test_gate_decoys.py +++ b/scripts/test_gate_decoys.py @@ -240,6 +240,15 @@ decoy("closed-register/open-row-closed-word (BY DESIGN)", "closed_register_gate. append_to(_OPEN_REG, _open_row("R-907", "**READY — RE-RANKED UP 2026-08-03 (R-86 closed)**")), expect="accept") +# --- one-register (2026-10-03): an open ROADMAP row whose id carries a LETTER SUFFIX ------------ +# The gate's id pattern was `R-(\d+)`, so R-27b, R-27c and R-50b were never read. Fixing it convicted +# R-50b at once — a finding that had lived only in ROADMAP.md. This decoy is that shape. (one-register +# stays in decoy_coverage_gate's EXEMPT list for R-424's hole — a finding filed under `idea` — which +# this decoy does not close.) +decoy("one-register/suffix-id-row", "one_register_gate.py", + append_to(os.path.join(ROOT, "documentation", "backlog", "ROADMAP.md"), + u"\n| R-905b | **A finding filed only here.** | S | READY — 2026-10-03 | none |\n")) + # --- hub-copy: an ENGLISH retrieval promise, planted in the bundle (R-558) ----------------------- # # TWO SHAPES AT ONCE, and the second is why this decoy exists at all.