697c2a7b10
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
126 lines
38 KiB
Markdown
126 lines
38 KiB
Markdown
# ROADMAP — future features & open work
|
||
|
||
> **What this is:** the prioritized decision log of planned/open work. Items are *intentions*, not
|
||
> claims about live behavior — the capability map (`architecture/00-capability-map.md`) is the only
|
||
> place that states what the platform does today.
|
||
>
|
||
> **Lifecycle:** idea → spiked → spec'd → in-progress → **shipped** (item collapses to a one-liner
|
||
> with the version, and the corresponding capability-map row changes status with evidence). Items
|
||
> can also be **killed** (keep the one-liner + why — decisions are worth remembering).
|
||
>
|
||
> **Coupling rule:** every item names the capability-map row(s) it flips. Every map gap row points
|
||
> back here by ID. Neither file duplicates the other's content.
|
||
>
|
||
> ## ONE REGISTER — operator ruling, 2026-08-22 (R-369)
|
||
>
|
||
> **Open FINDINGS live in `OPEN-ITEMS.md`, not here.** This file keeps history and reasoning, which
|
||
> is what its first paragraph has claimed since 2026-07-27. On 2026-08-22, **16 rows were moved**
|
||
> to the register — every row that asserted something checkable about the shipped product, plus one
|
||
> owed operator decision. **Their copies remain below, marked `MOVED -> OPEN-ITEMS.md`, and are not
|
||
> deleted:** this file's job is history.
|
||
>
|
||
> **The sorting rule, so it need not be re-invented:** *does the item assert something about the
|
||
> shipped product that a reader could go and check, and find false?* If yes it is a FINDING and it
|
||
> belongs in the register. If it proposes something that does not exist yet — a feature, a spike, a
|
||
> curation task — there is nothing to be wrong about, and it stays here as an intention.
|
||
>
|
||
> **`scripts/one_register_gate.py` enforces it:** a row here that is neither an idea nor done, and
|
||
> has no counterpart in the register, fails the push.
|
||
>
|
||
> **Severity — ONE scale for this file and `OPEN-ITEMS.md` (2026-10-03; operator may reverse):** **P1** now (a
|
||
> household can lose or leak data, a box can stop or be taken over, or a promise is false — today) · **P2** before the
|
||
> first paying customer · **P3** during the first customers · **P4** later / nice to have. Each item's tag leads its
|
||
> Item cell; an older tag later in the text (`[P2-HIGH]`) is history, and a re-rank says why in one dated line.
|
||
>
|
||
> **Cleaned 2026-10-03.** Shipped, killed, ruled and moved items went to `ROADMAP-HISTORY.md` (compressed; full text
|
||
> `git show 9f77865:documentation/backlog/ROADMAP.md`). Rows that are FINDINGS live in `OPEN-ITEMS.md` only. The loose
|
||
> notes of this folder have their verdicts in `README.md`.
|
||
|
||
|
||
---
|
||
|
||
## Intentions — by severity (2026-10-03)
|
||
|
||
> One table, P2 first. **There is no P1 intention**: nothing on this page is a harm happening today — those are
|
||
> findings, and they live in `OPEN-ITEMS.md`. Inside a severity, the order is the operator's request first, then
|
||
> by how directly a household meets it.
|
||
|
||
### P2 — before the first paying customer
|
||
|
||
| ID | Item | Size | Status | Notes / map rows flipped |
|
||
|----|------|------|--------|--------------------------|
|
||
| R-808 | **[P2] Box system security updates.** Goal: every box receives operating-system security patches on a schedule, and a failed update is undone. Why: a box lives in a home for years and today keeps the packages it was installed with — nothing updates the Proxmox host, the guest's Debian or its Docker engine (the finding is **R-812**). Scope: (1) regular security patches, host and guest; (2) reboots and their timing — inside the night window, never across a backup; (3) host kernel updates (a kept fallback boot entry); (4) Docker engine updates in the guest (an engine restart stops every app — quiesce like a backup); (5) the Proxmox MAJOR upgrade path (PVE 9 → 10) as its own later step, drilled first; (6) how the household and the operator are told (an event, a dashboard line); (7) how a failed update is undone (host: the vzdump/PBS copy plus the boot entry; guest: a snapshot before the run). Spike first: what the agent may run under its sudoers fence, and what a half-applied `apt` run leaves behind. | L | idea — 2026-10-03 (operator request) | Flips `00` §G *"Box operating-system security updates"* (MISSING, added 2026-10-03). Finding half: **R-812** in `OPEN-ITEMS.md`. |
|
||
| R-809 | **[P2] Legal pages and business papers before the first paying customer.** Goal: Felhom may legally take money from a household. Pieces: the website's ÁSZF, adatkezelési tájékoztató and impresszum (none exist — finding **R-813**); the customer contract; a data-processing agreement (Felhom monitors boxes and holds encrypted off-site backups, so it processes household data); billing and invoicing; the lawyer's licence review already on STATUS's list (**R-802**). Connects to the old **R-11** rulings (contact channel — RULED 2026-07-21; the tester agreement — never written). Owner: **operator**; CC drafts a text on request. | M | idea — 2026-10-03 (operator request) | Flips `00` §H rows *"Website legal pages"*, *"Customer contract and data-processing agreement"*, *"Billing and invoicing"* (MISSING, added 2026-10-03). Finding half: **R-813**. |
|
||
|
||
### P3 — during the first customers
|
||
|
||
| ID | Item | Size | Status | Notes / map rows flipped |
|
||
|----|------|------|--------|--------------------------|
|
||
| R-810 | **[P3] Independence — "if the household leaves Felhom, or Felhom stops".** Goal: a written answer the data-sovereignty pitch can point at. Questions: does the box keep working without the hub (partly answered — `architecture/_recovery-inventory-2026-07-28.md` §D2.4 and `07` §8 row 11b cover a LOST hub: a day is invisible, a week loses alarms, resets and convergence); who owns the domain (the customer, `01` §7), the Cloudflare tunnel and the off-site storage account; how a household exports everything; what a hand-over to the household or another provider looks like. Status **spike**: nothing is known to be wrong; no document answers the leave/hand-over half. | M | idea — spike, 2026-10-03 (operator request) | Flips `00` §E *"The household can leave Felhom, or outlive it"* (MISSING, added 2026-10-03). |
|
||
| R-811 | **[P3] A second login step for the dashboard.** Goal: a stolen or guessed password alone does not open the dashboard, which controls the whole box and is on the internet. Today: one password, one bcrypt hash (`felhom-controller/controller/internal/web/auth.go:37-44`). Options: a TOTP code or a passkey; recovery when the phone is lost modelled on the existing reset-code and recovery-code designs. Relates to **R-15** (member accounts) and the family gate (`09` §3 decisions 63–65): the same door should carry both. | M | idea — 2026-10-03 (operator request) | Flips `00` §E *"A second login step for the dashboard"* (MISSING, added 2026-10-03). |
|
||
| R-19 | **[P3]** Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does | S | idea | Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE **Flips (2026-10-03):** `00` §E "LAN access when internet is down" IMPLEMENTED → PROVEN-LIVE. |
|
||
| R-34 | **[P3]** **Backup data lifecycle management.** An "inactive backups" section on „Távoli mentés": apps that have snapshots but no active backup — **disabled OR uninstalled** — listed with name / size / last snapshot / restorable, plus an explicit **double-confirmed per-app delete** via `restic forget --tag` + nightly prune. | M | idea | **RULING: the offsite toggle NEVER offers deletion — policy and destruction stay decoupled.** Turning backups off must never be a data-destroying act, and deletion must never hide behind a toggle. Origin: 2026-07-18 rehearsal. Pairs with R-32 (that one is the operator's view of dead bytes; this one is the customer's) **Flips (2026-10-03):** `00` §C "Offsite (restic → Hetzner Storage Box)" — adds the inactive-backups leg; a new §C row when built. |
|
||
| R-46 | **[P3]** **[P2] Verification copies need a customer-visible browse surface and an expiry.** v0.147.0 made them *visible* (listed with path/size/date, individually deletable) — but the customer still cannot LOOK INSIDE a verification restore to confirm the file they wanted is really there, which is the entire point of a verification restore, and nothing ever removes them. | S–M | idea | Origin: 2026-07-19 feedback slice 4a, registered as the explicit follow-up to it. Two gaps, deliberately designed together because they are the same object: (a) **the invisible-result gap** — a read-only browse of `backups/offsite-restore/<app>` (the FileBrowser infra stack already exists and already serves scoped roots, so this may be a mount rather than new code); (b) **the disk-lifecycle gap** — auto-expiry after N days with the count/size surfaced before it fires, so a drive is never quietly filled by verification restores nobody remembers taking. Pairs with R-43: a browse surface is also how a customer would discover that a DB-indexed app's files came back but the app still cannot see them **Flips (2026-10-03):** `00` §C "Offsite restore: local-preferred scratch …" — the browse-and-expiry leg. **Re-ranked 2026-10-03:** [P2] → P3: a verification restore works today; what is missing is a way to look inside it and its expiry — a household meets it, with a workaround (FileBrowser). **Finding-shaped** (it states a fact about the shipped product): check it against today's product before building — an R-424 instance. |
|
||
| R-58 | **[P3]** **[P2] Assisted disk-picker install mode — the installer should let the operator CHOOSE the target disk instead of requiring the serial up front.** Today an install is either unattended (the answer file pins one `ID_SERIAL_SHORT`, which you can only know by first booting the machine) or match-nothing safety (aborts by design). That forces a two-boot dance for every new box: boot the safety ISO to read the serial, rebuild the ISO armed, boot again. | S–M | **idea — operator ruling 2026-07-21** | **Operator's argument, verbatim:** *"the installer should list the available storage devices (excluding the installation media) and let us select one, and continue."* **Shape:** a THIRD ISO mode alongside the two that exist — unattended-serial and match-nothing-safety. It enumerates candidate disks with **size / model / serial**, excludes the installation media itself, takes a selection plus a confirm, and proceeds. **Unattended+serial REMAINS the appliance/factory mode** — it is the right shape when the machine is provisioned in bulk and nobody is standing there; the picker is for the case where somebody is. **Slice 1 (cheap, same code surface, do this first):** improve the abort screen. On filter-no-match the installer currently just fails safe and says nothing useful — it should print the candidate table (size/model/serial) plus the one-line hint naming which serial to put in the profile. That alone collapses the two-boot dance from "boot, guess, go read docs, rebuild" to "boot, copy the serial off the screen, rebuild", and it is the same enumeration code the full picker needs. **Why it matters beyond convenience:** it is the BYO / reinstall flow — a customer's existing hardware, or a rebuild of a box whose disk layout nobody recorded, is exactly where the serial is unknown and a wrong guess is destructive. The current fail-safe is correct but mute. Origin: TASK-G, arming the HP install ISO — the serial had to be read off the board by hand between two boots **Flips (2026-10-03):** `00` §A "Bare-metal Felhom ISO" — adds a third, assisted install mode. **Re-ranked 2026-10-03:** [P2] → P3: the operator installs every box today and the answer-file route works. |
|
||
| R-69 | **[P3]** **F14-full: an operator push channel that actually interrupts (ntfy / Telegram / similar), beyond mail-client priority flags.** F14-light (v0.71.0 headers + Gmail filter) nudges a mail client; a 15:29 node_down should reach the operator's pocket in seconds regardless of inbox hygiene. Needs: channel choice (self-hosted ntfy on k3s vs Telegram bot), dispatcher fan-out seam, per-severity routing, quiet hours. | M | idea | Origin: `AUDIT-power-outage-recovery-2026-07-22.md` F14. Deliberately NOT built in the v0.71.0 train (scope-forked per the task spec) **Flips (2026-10-03):** `00` §F "Operator alerting (Healthchecks → monitoring@felhom.eu)" — a channel that interrupts. |
|
||
|
||
### P4 — later / nice to have
|
||
|
||
| ID | Item | Size | Status | Notes / map rows flipped |
|
||
|----|------|------|--------|--------------------------|
|
||
| R-832 | **[P4]** **ep0's copy in a place outside both Hetzner and the operator's home.** Today DooPlex holds it (decision 71). Not a Hetzner Storage Box: same provider as ep0 and every household's file backups, and no PBS there to verify or restore from. | S–M | idea | register R-832 |
|
||
| R-3 | **[P4]** Friend-alpha onboarding runbook (generalized from `pilot/RUNBOOK-peti-return-2026-07-13`): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step | M | idea | Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of `PETI-tester-agreement.md`. **Next from-scratch rehearsal to include customer DELETE + re-create** — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). **S6b (manual stale-host delete before re-enroll) is OBSOLETE** — re-enrollment upserts the existing host row cleanly (`store.UpsertHost` ON CONFLICT DO UPDATE; `handleAdminCreateHost` no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. **NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE):** one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). **Remaining consistency gap:** the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b **Flips (2026-10-03):** `00` §C "A customer (not the operator) performs a restore via UI alone" (MISSING). **May be superseded** by `runbooks/day0-install.md` and the first-hour drills — check before building. |
|
||
| R-6 | **[P4]** **Spike: LAN service discovery from the guest** — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? | M | **spiked (2026-07-18)** | **VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker `--network host` (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP `MediaServer:1` advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: `vmbr0 multicast_snooping=1` worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps.** **S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS `0x80070035` — WSD gives no name resolution; the flat `\\FELHOM-SPIKE` resolved by no path. Adding `nmbd` (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone.** Doc: `audits/SPIKE-lan-discovery-2026-07-18.md`. **Flips (2026-10-03):** `00` §E "Media to TV via DLNA" (MISSING) — the spike half, done; R-8 is the build. |
|
||
| R-8 | **[P4]** DLNA (**gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net**): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails | S | idea (unblocked) | Don't add catalog weight before proving the cheap path. **R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)** **Flips (2026-10-03):** `00` §E "Media to TV via DLNA" (MISSING → IMPLEMENTED). |
|
||
| R-9 | **[P4]** Uninstaller trio (from 07-15 Peti session): cluster-aware `felhom_guests` guard (node-local `pct list` deletes cluster-wide pveum objects); saferemove detection + time estimate + opt-in `--quick-remove` (never mutate `storage.cfg`); smarter `restore_storage` default for BYO clusters (shared storage, not local-lvm) | M | idea | Second item's rejected alternative (temp-disable-and-restore) stays rejected — crash window silently downgrades cluster wipe policy **Flips (2026-10-03):** `00` §A "Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling" — the cluster/BYO half. |
|
||
| R-27c | **[P4]** **Customer self-bind, slice 2 — console-passphrase bind.** Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. | M | idea | **Security constraints from the session ruling, all load-bearing:** passphrase **issued at customer creation**; the global-lookup endpoint must be **spray-hardened** — per-appliance **and** per-IP caps, constant-time comparison, a **single generic failure** (no oracle), alerting on abuse; an **accent-free wordlist** (console keymaps are not Hungarian); the **web capability-link path is RETAINED**; **claim-by-email is RETAINED** as the delivery-channel proof. **Also under this item:** the self-bind email gains the **public universal-ISO download link + two-line instructions** (the DIY case). **Secret-bearing per-customer ISOs are ruled OUT.** Sibling of R-27b (second-box flow) — different axis, both build on the same `/bind/` page **Flips (2026-10-03):** `00` §A "Customer binds their own appliance (self-service)" — a second bind route. |
|
||
| R-56 | **[P4]** **[P3] Apps do not say how technical they are, so a beginner can be ambushed by a config-heavy one.** The catalog presents every app as equally approachable — one Telepítés button, the same Hungarian copy — but they are not. Glance needs a hand-written `glance.yml` before it does anything; some apps need a reverse-proxy or API concept to configure; others genuinely are install-and-use. A tester who picks the wrong first app concludes the PRODUCT is broken, not that they picked an advanced app. | S | **idea (filed 2026-07-21)** | Origin: TASK-E Part 3 — **filed, deliberately not implemented**. Shape: a `difficulty:` field in `.felhom.yml` (`kezdő` / `haladó` / `technikás`) surfaced as a catalog-card badge and repeated on the deploy screen. Cheap and incremental: one optional metadata field plus a badge, classifiable app-by-app with no migration — an app with no `difficulty:` simply shows no badge. **This is the constructive half of the glance ruling**: glance STAYS in the catalog (operator ruling 2026-07-21 — it is a legitimate app, not a broken one; its missing seeded `glance.yml` is a known pre-existing finding), and the honest fix is to LABEL it rather than hide it. Pairs with R-41: that gate proves an app CAN still deploy; this field tells a customer whether THEY should be the one deploying it. **Badge plumbing is ALREADY BUILT (controller v0.158.0)** — `web.MetaBadge` + the `meta_badge` template partial + the `lifecycleBadge` funcmap entry were written generic for exactly this: a `difficultyBadge` funcmap function returning the same `*MetaBadge`, plus a `difficulty:` field on `stacks.Metadata`, is the whole remaining job. No new markup, no new CSS. Re-sized accordingly **Flips (2026-10-03):** `00` §B "Deploy an app from the catalog" — a difficulty label on the card. |
|
||
| R-62 | **[P4]** **[P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape.** | XS | **idea (operator, 2026-07-22)** | Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; **wording only** — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). **Flips (2026-10-03):** `00` §G "Customer/host management" — the delete dialog. |
|
||
| R-64 | **[P4]** **„Felhom↔Felhom media pairing blessed" — the two-box SMB pairing (one box shares, the other mounts it as NAS storage) becomes a supported, documented flow.** | XS–S | idea (2026-07-22) | Origin: the operator ran the pairing drill on the live demo pair and it WORKS — the drill itself is the pending evidence leg (a written run-through with the R-66 surfaces in play). R-66 shipped the enabling visibility: the serving box's address is now on its own Beállítások → Rendszer „Hálózat" card, and the add form names the NetBIOS trap. Blessing = a short customer-facing recipe (`documentation/controller/network-storage-nas.md` naming-caveat paragraph is the seed) + one supported-path sentence in the capability map. Flips: would add a "Felhom↔Felhom media pairing" capability row (currently unlisted). Pairs with R-65 (same two-box topology, entirely different transport + guarantees) **Flips (2026-10-03):** `00` §E "Files from Windows Explorer / Mac Finder (SMB server)" — box-to-box pairing as a supported flow. |
|
||
| R-26 | **[P4]** **Guided old-history recovery via a retained superseded escrow + the recovery code.** Enabled by hub v0.60.0 (Part B) which now RETAINS superseded escrow blobs (`host_escrow_superseded`, `ListSupersededEscrow`). Build the flow that, given the customer's recovery code, unwraps a retained old blob → recovers the old repo passphrase → mounts/reads the moved-aside `.orphaned-<date>` repo for restore. | M | idea (enabled by v0.60.0) | Turns "history recoverable in principle" into a real customer-drivable path; pairs with the controller v0.142.0 orphaned-repo move-aside. Origin `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Flips (2026-10-03):** `00` §C "Offsite restore …" — old history via a retained escrow. **Parked by a decision:** R-312 (decided 2026-08-13) keeps any route to a set-aside store operator-only until a real customer asks (`07` §11). |
|
||
| R-27b | **[P4]** **Customer self-bind, second-box flow (controller side).** For a customer who ALREADY has a bound box and installs another, the controller shows a dismissable "bind another box" prompt (and a bind-later entry under settings) that walks to the hub `/bind/` page — so a returning customer isn't emailed a fresh operator-sent link for every box. Mechanism sketched in the hub v0.66.0 REPORT; NOT built (R-27 slice 1 deliberately did not touch the controller). | M | idea (minted by hub v0.66.0) | Origin: hub v0.66.0 slice-1 ship (first-box only). Reuses the same `/bind/` public page + tokenized-link machinery; adds a controller-side entry point + the operator "mint a link for an existing customer" affordance **Flips (2026-10-03):** `00` §A "Customer binds their own appliance (self-service)" — the second-box flow. |
|
||
| R-12 | **[P4]** Cluster mode: agent-follows-guest, bind-mount reconciliation on HA migration | XL | idea | Scoped 07-15; interim = HA-group pin to one node. Driven by Peti's two-node cluster **Flips (2026-10-03):** `00` §G — a new row "a guest follows its node in a cluster" when built. |
|
||
| R-14 | **[P4]** Headscale/WireGuard spike: Minecraft/gaming port connectivity (CGNAT-proof, sovereign DERP fallback) | M | idea | **Flips (2026-10-03):** `00` §E — a new row "game-server ports reachable behind CGNAT" when built. |
|
||
| R-15 | **[P4]** Multi-user dashboard accounts (household members, roles) | L | idea | Single password is a stated alpha limitation (R-11). **Launcher coupling — REVISED (controller v0.165.0):** the "share the launcher outside the household" need is now met WITHOUT member accounts — the **Indítópult megosztása** capability-URL guest link (`/s/<token>`, information-only, no account) shipped in v0.165.0. What remains for this arc is member-specific: **per-member tile visibility** (each member sees only their apps) and the launcher-as-member-landing-page — both live inside this SSO/members arc; the guest-link ruling explicitly SUPERSEDES the earlier "members are how you share the launcher" framing **Flips (2026-10-03):** `00` §E "Multiple household users / per-person accounts" (MISSING). Relates to **R-811** (a second login step): the same door should carry both. |
|
||
| R-72 | **[P4]** Curate `brand_color` for the top catalog apps | XS | idea | Parked follow-up to the v0.163.0 launcher. `.felhom.yml` `brand_color` (`#rgb`/`#rrggbb`) overrides the deterministic slug-hash tile color; no catalog app sets it yet. Pick brand-accurate colors for the most-installed apps so their launcher tiles match their real brand. Catalog-only change (`app-catalog-felhom.eu`), `brand_color` is already `omitempty` and consumed by the controller **Flips (2026-10-03):** `00` §E "Indítópult (app launcher)" — brand colours. |
|
||
| R-74 | **[P4]** **Island control plane on a CLUSTER (Peti's 2 nodes)** — bring R-50's island bridge to a multi-node PVE cluster. | M | idea (Phase C of R-50, parked) | R-50 shipped the island for the ONE-host fleet (demo-hp, demo-felhom). A cluster needs **bridge parity on every node**: either per-node identical `/etc/network/interfaces` `vmbr9` stanzas (simplest, drift-prone) or — preferred at ≥2 nodes — a Proxmox **SDN zone/vnet** defined cluster-wide (one definition, auto-applied per node). The guest island IP is per-guest + node-independent; the **agent-follows-guest** rule holds (each node's agent binds its own `vmbr9` `169.254.253.1`). Migration order per the spike: drill-proven → demo (done) → **Peti (this row)**. Its own supervised runbook, coordinated with Peti (a live customer). Completes the capability-map "site/network change" row for clustered installs. Source: `audits/SPIKE-island-bridge-2026-07-25.md` (cluster-parity finding) + `RUNBOOK-island-migration.md` (single-host procedure to generalise) **Flips (2026-10-03):** `00` §G — a new row "the island control plane on a multi-node cluster" when built. |
|
||
| R-41 | **[P4]** **[SLICE 1 SHIPPED 2026-07-21] The catalog has no standing "does every template still deploy?" check.** Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found **5 that had NEVER been deployable**: papra (missing required `AUTH_SECRET`), zipline (v4 renamed `CORE_DATABASE_URL` → `DATABASE_URL`), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the `v` tag prefix + new required env), glance (needs a seeded `glance.yml` the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus **7 broken healthchecks** and 2 apps whose images no longer resolve at all (plant-it, wanderer). | M | idea | Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in `scripts/check-image-pins.py` — a mechanical gate run on every change. Cheap first slice: a **resolvability gate** (`docker manifest inspect` every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. **Silent rot is the real risk** — an app can die upstream and nobody learns until a customer clicks Telepítés | **SLICE 1 SHIPPED 2026-07-21 — `app-catalog-felhom.eu/scripts/check-image-resolvable.py`** (+ 14 fixture tests, no network, resolver injected). Resolves every unique pin with `docker manifest inspect`, ONE image at a time; exit 0 / 1 (the registry says GONE) / 2 (inconclusive). **Two traps encoded, both hit live while building it:** (a) `docker manifest inspect` prints `toomanyrequests: …` and **still exits 0** — the same exits-0-on-failure shape as the ISO tooling's `validate-answer`, so stderr is inspected even on rc=0; (b) the inverse and more dangerous one — the first full sweep called **24 of 65 pins dead, including `postgres:16-alpine` and `redis:7-alpine`**, purely because Docker Hub throttled it partway through. Ambiguity therefore resolves to INCONCLUSIVE and never to an accusation: a gate that cries wolf gets ignored, and then it protects nothing. **The full sweep is still OWED** — DooPlex is not logged in to Docker Hub, so the 52-app table needs one re-run after `docker login`. Wired into `CLAUDE.md` + `REUSE.md` as a start-of-campaign / pre-publish-train step. It immediately paid for itself: it is what turned plant-it and wanderer from 'images do not resolve' into two DIFFERENT diagnoses (see the 2026-07-21 catalog entry). Full slice — the periodic deploy-all sweep on the demo box — remains open **Flips (2026-10-03):** `00` §B "Catalog sync … validation choke point" — a standing deploy-all check. |
|
||
| R-45 | **[P4]** **[P2] Unified async-job feedback.** Every long operation invents its own progress surface, or none. Tonight produced three more one-off cards (v0.147.x: samba bring-up, offsite progress, restore result) on top of two existing patterns (deploy 3-step panel; storage-init/netstorage status poll). They agree on nothing: some use `{ok,data}` envelopes and some raw JSON, some poll 1 s / 1.5 s / 3 s, some are in-memory-only and lie after a restart, and each re-implements single-flight + snapshot + phase→Hungarian mapping. | M | idea | Origin: 2026-07-19 feedback slice 1 (controller v0.147.0). The cases to generalise from are all in-tree: `web/storage_init_job.go` (the best shape — acquire/release/set/snapshot), `web/netstorage_job.go`, `web/samba_ensure_job.go`, `backup/opstatus.go`, `backup/offbox_progress.go`. Shape: one job registry + one poll endpoint + one client-side renderer, phases declared per job. **Two lessons tonight that any framework must encode:** (1) a terminal state must be **probed, not inferred** — `compose up -d` exits 0 on a crash-loop; (2) a progress source that reports nothing is normal, not broken — restic reports 0 bytes for a whole incremental run, and a bar that sits at 0% is worse than no bar. Also fixes the restart hole: in-memory job state currently vanishes and the card silently disagrees with reality **2026-07-20 — the first bill for NOT having this arrived, and it was customer-facing.** The samba card's poll (`web/samba_ensure_job.go` + `sharing.html`) mixed a job EDGE and a service LEVEL on one JSON field, and `/sharing` reload-looped at ~1.2 s for every customer with sharing enabled until controller v0.151.0 (`audits/DIAG-sharing-2026-07-20.md`, S-1/S-4). v0.151.0 fixed THAT card's contract only — the framework is still this item. **Third lesson for it to encode, beside the two already listed:** a phase a client answers with a one-shot action must be an EDGE the registry SERVES ONCE, and must never be synthesised from a level; if it can be re-read, it will be re-acted on. **Flips (2026-10-03):** `00` §B "Deploy an app from the catalog (… health-aware progress)" — one job surface for every long operation. **Re-ranked 2026-10-03:** [P2] → P4: polish; every long operation already shows a progress surface of its own. |
|
||
| R-65 | **[P4]** **Buddy-box backup replication, cross-household — two Felhom boxes in different homes replicate backups to each other.** | L | idea (post-alpha, spike-first, 2026-07-22) | The natural big sibling of R-64: two households each hosting the other's encrypted backup tier. Explicitly **spike-first** — the transport is NOT SMB (R-64's live-share protocol is wrong for backup replication across the internet: no auth story between households, no resumability, cleartext LAN assumptions); candidates to spike: restic rest-server / rclone / syncthing over the existing WG/tailnet plumbing, encryption keyed so the buddy can never read the payload. Sits on top of the offsite tier's FILL/OVERSUB thresholds thinking (R-5 aggregate). Flips: would add a "cross-household buddy replication" capability row (currently unlisted). Pairs with R-64 (same topology, different transport + guarantees) **Flips (2026-10-03):** `00` §C — a new row "backups replicated to a second household" when built. |
|
||
|
||
## UPDATE-ARC — what is still open (collapsed 2026-10-03)
|
||
|
||
The app-update arc (`audits/SPIKE-app-update-2026-09-01.md`; its living design is `architecture/09-update-architecture.md`).
|
||
**Parts 1–10 SHIPPED** (slices 1, 1b, 2, 3, 4 and 5 in controller v0.233.0 – v0.238.1, then the tested-steps,
|
||
ladder, hold, undo and night-update parts through controller v0.271.0 — `09` §6.4); **part 11 deferred by ruling**
|
||
(`09` §3 decision 18: the fleet sweep waits until the fleet grows). The full previous entry is in
|
||
`ROADMAP-HISTORY.md`. **The open work is in the register, by id:** R-450 (slice 6 — automatic within a major, never
|
||
across one), R-451 (slice 7 — the fleet sweep, DEFERRED), R-469 (lift the engine-major rule per app), R-440, R-446,
|
||
R-458, R-462, R-607, R-610, R-618, R-621, R-622, R-683, R-738, R-785 — all under **App updates** in `OPEN-ITEMS.md`.
|
||
Flips: `00` §B "Automatic updates at night — the update leg" and "Whether the box UPGRADES an app by itself".
|
||
|
||
## Gating candidates — Campaign 12, Part 4 (2026-08-08) — all **P4**
|
||
|
||
**Ranked by what a gate would be worth**, using Campaign 12's own instance counts as the evidence
|
||
(`audits/CAMPAIGN-12-class-sweep-2026-08-08.md` §4). G-1 was built (`scripts/wire_contract_gate.py`); G-6 and G-7
|
||
were recorded as a NO with evidence — all three are in `ROADMAP-HISTORY.md`. **Flips:** none — gates change no
|
||
capability; they keep the map's claims true.
|
||
|
||
| Rank | Item | Size | Status | Notes |
|
||
|----|------|------|--------|-------|
|
||
| **G-2** | **Gate C3 — a success verdict may not be set where an incompleteness signal is in scope.** Assert that every literal success-status assignment either has no gap/skip/missing signal available at that point, or consults it. | S | candidate | The whole population in the controller is **9 sites** — Campaign 12 read all of them, which is why this class is the one where "no others exist" is supportable. Small enough to gate by enumeration rather than by inference. **Known miss:** verdicts expressed as booleans, enum constants, or the absence of an error — and the controller does use those elsewhere. Instances: R-240 (open), R-258 (new). |
|
||
| **G-3** | **Gate C4 — every rendered count/size/percentage needs a `*Known` companion.** | M | candidate — **UNBLOCKED 2026-08-08: the convention decision it was waiting for has been made** | The decision owed was *"how does this codebase say 'we could not look'"*, and it is now ruled (`CONTEXT.md` **S-39**, shipped in controller v0.210.0 / R-259): **an explicit `…Known bool` companion beside the figures, checked in the template before anything is rendered** — the shape `Offbox.StatsKnown` already used, whose own comment carries the reasoning (*"a 0%-wide bar over an unread store is a picture of emptiness, and a picture is a claim"*). Pointers and separate error fields remain legitimate Go and both still exist here; the ruling is that **new** three-state figures use the companion, because a codebase with three dialects cannot be gated by a name-based check. **Existing call sites were deliberately NOT converted** — that conversion is the bulk of this item's M and is what remains before a gate can be turned on without a wall of false positives. **Next step is therefore a survey, not a gate:** count the rendered figures that lack a companion, decide which are genuinely three-state, convert those, then gate. Instances so far: R-225 (fixed, the pattern's origin), R-259 (fixed, the ruling). |
|
||
| **G-4** | **Complete C1's runtime body assertion — 4 of 27 pages today.** | L | candidate — the expensive one, and honestly so | `secret_in_markup_gate.py` covers all 36 templates on the NAME-based check and is **blind to a secret under a neutral page-data key** — verified 2026-08-08 by replaying the three pre-fix templates through it: it convicts 2 of 3 and not the third. The runtime assertion catches all three; extending it means constructing each remaining page's data in a test, which is a per-page cost and is the real reason it has not been done. **Do NOT adopt the Go-side mirror Campaign 12 wrote as a gate on its own** — 27 candidates, 0 findings is bad signal-to-noise in front of every push. → R-255 |
|
||
| **G-5** | **Gate C7, narrowly — uniqueness claims only.** A comment saying "X is the ONLY writer/place/caller of Y" is mechanically falsifiable; assert it. | S | candidate — narrow by construction | Covers ~60 of the 2652 production invariant comments. The other ~97% of the vocabulary (`never`, `always`, `must not`, `guarantees`) is not mechanical and a gate must not pretend otherwise. Instance: R-263. |
|
||
| **G-8** | **R-242's untouched half — catch a SKIPPED VOUCH.** | S then M | candidate — **the one with a live recurrence** | Measured during Campaign 12's own bake: `golden_currency_gate.py` **flipped red→green the moment the evidence DIRECTORY existed**, before the round-trip download finished and with no vouch near it. **(a) Cheapest, now:** have the bake session re-read `/configuration` after the operator's Save and write the observed `golden_version` into the evidence README as a machine-readable line; the gate then requires that line rather than the directory. Detects the FORGETTING, which is the actual failure mode. **(b) Loudest, when the hub is next touched:** a hub-side daily check comparing the vouched golden against the newest controller the fleet reports — it fires within a day and catches a silent ROLLBACK too, which nothing in git can ever see. **(c) A checklist item is not a fix** — R-242 already was a rule without a mechanism and it recurred the next day. → R-242 |
|
||
|
||
## Before the first paying customer
|
||
|
||
**One list, not two:** root `STATUS.md`, section "Before the first paying customer". The old "Pre-invite checklist"
|
||
(2026-07-18, the remote-tester era) is in `ROADMAP-HISTORY.md`; its open legal half — the **R-11** rulings
|
||
(the tester agreement was never written) — is now part of **R-809**.
|
||
|
||
## Loose notes in this folder
|
||
|
||
Verdict per file: `backlog/README.md`.
|