Files
felhom.eu/documentation/backlog/ROADMAP.md
T

38 KiB
Raw Blame History

ROADMAP — future features & open work

What this is: the prioritized decision log of planned/open work. Items are intentions, not claims about live behavior — the capability map (architecture/00-capability-map.md) is the only place that states what the platform does today.

Lifecycle: idea → spiked → spec'd → in-progress → shipped (item collapses to a one-liner with the version, and the corresponding capability-map row changes status with evidence). Items can also be killed (keep the one-liner + why — decisions are worth remembering).

Coupling rule: every item names the capability-map row(s) it flips. Every map gap row points back here by ID. Neither file duplicates the other's content.

ONE REGISTER — operator ruling, 2026-08-22 (R-369)

Open FINDINGS live in OPEN-ITEMS.md, not here. This file keeps history and reasoning, which is what its first paragraph has claimed since 2026-07-27. On 2026-08-22, 16 rows were moved to the register — every row that asserted something checkable about the shipped product, plus one owed operator decision. Their copies remain below, marked MOVED -> OPEN-ITEMS.md, and are not deleted: this file's job is history.

The sorting rule, so it need not be re-invented: does the item assert something about the shipped product that a reader could go and check, and find false? If yes it is a FINDING and it belongs in the register. If it proposes something that does not exist yet — a feature, a spike, a curation task — there is nothing to be wrong about, and it stays here as an intention.

scripts/one_register_gate.py enforces it: a row here that is neither an idea nor done, and has no counterpart in the register, fails the push.

Severity — ONE scale for this file and OPEN-ITEMS.md (2026-10-03; operator may reverse): P1 now (a household can lose or leak data, a box can stop or be taken over, or a promise is false — today) · P2 before the first paying customer · P3 during the first customers · P4 later / nice to have. Each item's tag leads its Item cell; an older tag later in the text ([P2-HIGH]) is history, and a re-rank says why in one dated line.

Cleaned 2026-10-03. Shipped, killed, ruled and moved items went to ROADMAP-HISTORY.md (compressed; full text git show 9f77865:documentation/backlog/ROADMAP.md). Rows that are FINDINGS live in OPEN-ITEMS.md only. The loose notes of this folder have their verdicts in README.md.


Intentions — by severity (2026-10-03)

One table, P2 first. There is no P1 intention: nothing on this page is a harm happening today — those are findings, and they live in OPEN-ITEMS.md. Inside a severity, the order is the operator's request first, then by how directly a household meets it.

P2 — before the first paying customer

ID Item Size Status Notes / map rows flipped
R-808 [P2] Box system security updates. Goal: every box receives operating-system security patches on a schedule, and a failed update is undone. Why: a box lives in a home for years and today keeps the packages it was installed with — nothing updates the Proxmox host, the guest's Debian or its Docker engine (the finding is R-812). Scope: (1) regular security patches, host and guest; (2) reboots and their timing — inside the night window, never across a backup; (3) host kernel updates (a kept fallback boot entry); (4) Docker engine updates in the guest (an engine restart stops every app — quiesce like a backup); (5) the Proxmox MAJOR upgrade path (PVE 9 → 10) as its own later step, drilled first; (6) how the household and the operator are told (an event, a dashboard line); (7) how a failed update is undone (host: the vzdump/PBS copy plus the boot entry; guest: a snapshot before the run). Spike first: what the agent may run under its sudoers fence, and what a half-applied apt run leaves behind. L idea — 2026-10-03 (operator request) Flips 00 §G "Box operating-system security updates" (MISSING, added 2026-10-03). Finding half: R-812 in OPEN-ITEMS.md.
R-809 [P2] Legal pages and business papers before the first paying customer. Goal: Felhom may legally take money from a household. Pieces: the website's ÁSZF, adatkezelési tájékoztató and impresszum (none exist — finding R-813); the customer contract; a data-processing agreement (Felhom monitors boxes and holds encrypted off-site backups, so it processes household data); billing and invoicing; the lawyer's licence review already on STATUS's list (R-802). Connects to the old R-11 rulings (contact channel — RULED 2026-07-21; the tester agreement — never written). Owner: operator; CC drafts a text on request. M idea — 2026-10-03 (operator request) Flips 00 §H rows "Website legal pages", "Customer contract and data-processing agreement", "Billing and invoicing" (MISSING, added 2026-10-03). Finding half: R-813.

P3 — during the first customers

ID Item Size Status Notes / map rows flipped
R-810 [P3] Independence — "if the household leaves Felhom, or Felhom stops". Goal: a written answer the data-sovereignty pitch can point at. Questions: does the box keep working without the hub (partly answered — architecture/_recovery-inventory-2026-07-28.md §D2.4 and 07 §8 row 11b cover a LOST hub: a day is invisible, a week loses alarms, resets and convergence); who owns the domain (the customer, 01 §7), the Cloudflare tunnel and the off-site storage account; how a household exports everything; what a hand-over to the household or another provider looks like. Status spike: nothing is known to be wrong; no document answers the leave/hand-over half. M idea — spike, 2026-10-03 (operator request) Flips 00 §E "The household can leave Felhom, or outlive it" (MISSING, added 2026-10-03).
R-811 [P3] A second login step for the dashboard. Goal: a stolen or guessed password alone does not open the dashboard, which controls the whole box and is on the internet. Today: one password, one bcrypt hash (felhom-controller/controller/internal/web/auth.go:37-44). Options: a TOTP code or a passkey; recovery when the phone is lost modelled on the existing reset-code and recovery-code designs. Relates to R-15 (member accounts) and the family gate (09 §3 decisions 63–65): the same door should carry both. M idea — 2026-10-03 (operator request) Flips 00 §E "A second login step for the dashboard" (MISSING, added 2026-10-03).
R-19 [P3] Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does S idea Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE Flips (2026-10-03): 00 §E "LAN access when internet is down" IMPLEMENTED → PROVEN-LIVE.
R-34 [P3] Backup data lifecycle management. An "inactive backups" section on „Távoli mentés": apps that have snapshots but no active backup — disabled OR uninstalled — listed with name / size / last snapshot / restorable, plus an explicit double-confirmed per-app delete via restic forget --tag + nightly prune. M idea RULING: the offsite toggle NEVER offers deletion — policy and destruction stay decoupled. Turning backups off must never be a data-destroying act, and deletion must never hide behind a toggle. Origin: 2026-07-18 rehearsal. Pairs with R-32 (that one is the operator's view of dead bytes; this one is the customer's) Flips (2026-10-03): 00 §C "Offsite (restic → Hetzner Storage Box)" — adds the inactive-backups leg; a new §C row when built.
R-46 [P3] [P2] Verification copies need a customer-visible browse surface and an expiry. v0.147.0 made them visible (listed with path/size/date, individually deletable) — but the customer still cannot LOOK INSIDE a verification restore to confirm the file they wanted is really there, which is the entire point of a verification restore, and nothing ever removes them. S–M idea Origin: 2026-07-19 feedback slice 4a, registered as the explicit follow-up to it. Two gaps, deliberately designed together because they are the same object: (a) the invisible-result gap — a read-only browse of backups/offsite-restore/<app> (the FileBrowser infra stack already exists and already serves scoped roots, so this may be a mount rather than new code); (b) the disk-lifecycle gap — auto-expiry after N days with the count/size surfaced before it fires, so a drive is never quietly filled by verification restores nobody remembers taking. Pairs with R-43: a browse surface is also how a customer would discover that a DB-indexed app's files came back but the app still cannot see them Flips (2026-10-03): 00 §C "Offsite restore: local-preferred scratch …" — the browse-and-expiry leg. Re-ranked 2026-10-03: [P2] → P3: a verification restore works today; what is missing is a way to look inside it and its expiry — a household meets it, with a workaround (FileBrowser). Finding-shaped (it states a fact about the shipped product): check it against today's product before building — an R-424 instance.
R-58 [P3] [P2] Assisted disk-picker install mode — the installer should let the operator CHOOSE the target disk instead of requiring the serial up front. Today an install is either unattended (the answer file pins one ID_SERIAL_SHORT, which you can only know by first booting the machine) or match-nothing safety (aborts by design). That forces a two-boot dance for every new box: boot the safety ISO to read the serial, rebuild the ISO armed, boot again. S–M idea — operator ruling 2026-07-21 Operator's argument, verbatim: "the installer should list the available storage devices (excluding the installation media) and let us select one, and continue." Shape: a THIRD ISO mode alongside the two that exist — unattended-serial and match-nothing-safety. It enumerates candidate disks with size / model / serial, excludes the installation media itself, takes a selection plus a confirm, and proceeds. Unattended+serial REMAINS the appliance/factory mode — it is the right shape when the machine is provisioned in bulk and nobody is standing there; the picker is for the case where somebody is. Slice 1 (cheap, same code surface, do this first): improve the abort screen. On filter-no-match the installer currently just fails safe and says nothing useful — it should print the candidate table (size/model/serial) plus the one-line hint naming which serial to put in the profile. That alone collapses the two-boot dance from "boot, guess, go read docs, rebuild" to "boot, copy the serial off the screen, rebuild", and it is the same enumeration code the full picker needs. Why it matters beyond convenience: it is the BYO / reinstall flow — a customer's existing hardware, or a rebuild of a box whose disk layout nobody recorded, is exactly where the serial is unknown and a wrong guess is destructive. The current fail-safe is correct but mute. Origin: TASK-G, arming the HP install ISO — the serial had to be read off the board by hand between two boots Flips (2026-10-03): 00 §A "Bare-metal Felhom ISO" — adds a third, assisted install mode. Re-ranked 2026-10-03: [P2] → P3: the operator installs every box today and the answer-file route works.
R-69 [P3] F14-full: an operator push channel that actually interrupts (ntfy / Telegram / similar), beyond mail-client priority flags. F14-light (v0.71.0 headers + Gmail filter) nudges a mail client; a 15:29 node_down should reach the operator's pocket in seconds regardless of inbox hygiene. Needs: channel choice (self-hosted ntfy on k3s vs Telegram bot), dispatcher fan-out seam, per-severity routing, quiet hours. M idea Origin: AUDIT-power-outage-recovery-2026-07-22.md F14. Deliberately NOT built in the v0.71.0 train (scope-forked per the task spec) Flips (2026-10-03): 00 §F "Operator alerting (Healthchecks → monitoring@felhom.eu)" — a channel that interrupts.

P4 — later / nice to have

ID Item Size Status Notes / map rows flipped
R-832 [P4] ep0's copy in a place outside both Hetzner and the operator's home. Today DooPlex holds it (decision 71). Not a Hetzner Storage Box: same provider as ep0 and every household's file backups, and no PBS there to verify or restore from. S–M idea register R-832
R-3 [P4] Friend-alpha onboarding runbook (generalized from pilot/RUNBOOK-peti-return-2026-07-13): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step M idea Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of PETI-tester-agreement.md. Next from-scratch rehearsal to include customer DELETE + re-create — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). S6b (manual stale-host delete before re-enroll) is OBSOLETE — re-enrollment upserts the existing host row cleanly (store.UpsertHost ON CONFLICT DO UPDATE; handleAdminCreateHost no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE): one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). Remaining consistency gap: the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b Flips (2026-10-03): 00 §C "A customer (not the operator) performs a restore via UI alone" (MISSING). May be superseded by runbooks/day0-install.md and the first-hour drills — check before building.
R-6 [P4] Spike: LAN service discovery from the guest — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? M spiked (2026-07-18) VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker --network host (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP MediaServer:1 advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: vmbr0 multicast_snooping=1 worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps. S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS 0x80070035 — WSD gives no name resolution; the flat \\FELHOM-SPIKE resolved by no path. Adding nmbd (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone. Doc: audits/SPIKE-lan-discovery-2026-07-18.md. Flips (2026-10-03): 00 §E "Media to TV via DLNA" (MISSING) — the spike half, done; R-8 is the build.
R-8 [P4] DLNA (gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails S idea (unblocked) Don't add catalog weight before proving the cheap path. R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7) Flips (2026-10-03): 00 §E "Media to TV via DLNA" (MISSING → IMPLEMENTED).
R-9 [P4] Uninstaller trio (from 07-15 Peti session): cluster-aware felhom_guests guard (node-local pct list deletes cluster-wide pveum objects); saferemove detection + time estimate + opt-in --quick-remove (never mutate storage.cfg); smarter restore_storage default for BYO clusters (shared storage, not local-lvm) M idea Second item's rejected alternative (temp-disable-and-restore) stays rejected — crash window silently downgrades cluster wipe policy Flips (2026-10-03): 00 §A "Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling" — the cluster/BYO half.
R-27c [P4] Customer self-bind, slice 2 — console-passphrase bind. Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. M idea Security constraints from the session ruling, all load-bearing: passphrase issued at customer creation; the global-lookup endpoint must be spray-hardened — per-appliance and per-IP caps, constant-time comparison, a single generic failure (no oracle), alerting on abuse; an accent-free wordlist (console keymaps are not Hungarian); the web capability-link path is RETAINED; claim-by-email is RETAINED as the delivery-channel proof. Also under this item: the self-bind email gains the public universal-ISO download link + two-line instructions (the DIY case). Secret-bearing per-customer ISOs are ruled OUT. Sibling of R-27b (second-box flow) — different axis, both build on the same /bind/ page Flips (2026-10-03): 00 §A "Customer binds their own appliance (self-service)" — a second bind route.
R-56 [P4] [P3] Apps do not say how technical they are, so a beginner can be ambushed by a config-heavy one. The catalog presents every app as equally approachable — one Telepítés button, the same Hungarian copy — but they are not. Glance needs a hand-written glance.yml before it does anything; some apps need a reverse-proxy or API concept to configure; others genuinely are install-and-use. A tester who picks the wrong first app concludes the PRODUCT is broken, not that they picked an advanced app. S idea (filed 2026-07-21) Origin: TASK-E Part 3 — filed, deliberately not implemented. Shape: a difficulty: field in .felhom.yml (kezdő / haladó / technikás) surfaced as a catalog-card badge and repeated on the deploy screen. Cheap and incremental: one optional metadata field plus a badge, classifiable app-by-app with no migration — an app with no difficulty: simply shows no badge. This is the constructive half of the glance ruling: glance STAYS in the catalog (operator ruling 2026-07-21 — it is a legitimate app, not a broken one; its missing seeded glance.yml is a known pre-existing finding), and the honest fix is to LABEL it rather than hide it. Pairs with R-41: that gate proves an app CAN still deploy; this field tells a customer whether THEY should be the one deploying it. Badge plumbing is ALREADY BUILT (controller v0.158.0) — web.MetaBadge + the meta_badge template partial + the lifecycleBadge funcmap entry were written generic for exactly this: a difficultyBadge funcmap function returning the same *MetaBadge, plus a difficulty: field on stacks.Metadata, is the whole remaining job. No new markup, no new CSS. Re-sized accordingly Flips (2026-10-03): 00 §B "Deploy an app from the catalog" — a difficulty label on the card.
R-62 [P4] [P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape. XS idea (operator, 2026-07-22) Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; wording only — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). Flips (2026-10-03): 00 §G "Customer/host management" — the delete dialog.
R-64 [P4] „Felhom↔Felhom media pairing blessed" — the two-box SMB pairing (one box shares, the other mounts it as NAS storage) becomes a supported, documented flow. XS–S idea (2026-07-22) Origin: the operator ran the pairing drill on the live demo pair and it WORKS — the drill itself is the pending evidence leg (a written run-through with the R-66 surfaces in play). R-66 shipped the enabling visibility: the serving box's address is now on its own Beállítások → Rendszer „Hálózat" card, and the add form names the NetBIOS trap. Blessing = a short customer-facing recipe (documentation/controller/network-storage-nas.md naming-caveat paragraph is the seed) + one supported-path sentence in the capability map. Flips: would add a "Felhom↔Felhom media pairing" capability row (currently unlisted). Pairs with R-65 (same two-box topology, entirely different transport + guarantees) Flips (2026-10-03): 00 §E "Files from Windows Explorer / Mac Finder (SMB server)" — box-to-box pairing as a supported flow.
R-26 [P4] Guided old-history recovery via a retained superseded escrow + the recovery code. Enabled by hub v0.60.0 (Part B) which now RETAINS superseded escrow blobs (host_escrow_superseded, ListSupersededEscrow). Build the flow that, given the customer's recovery code, unwraps a retained old blob → recovers the old repo passphrase → mounts/reads the moved-aside .orphaned-<date> repo for restore. M idea (enabled by v0.60.0) Turns "history recoverable in principle" into a real customer-drivable path; pairs with the controller v0.142.0 orphaned-repo move-aside. Origin DIAGNOSE-offbox-repo-orphaned-2026-07-17 Flips (2026-10-03): 00 §C "Offsite restore …" — old history via a retained escrow. Parked by a decision: R-312 (decided 2026-08-13) keeps any route to a set-aside store operator-only until a real customer asks (07 §11).
R-27b [P4] Customer self-bind, second-box flow (controller side). For a customer who ALREADY has a bound box and installs another, the controller shows a dismissable "bind another box" prompt (and a bind-later entry under settings) that walks to the hub /bind/ page — so a returning customer isn't emailed a fresh operator-sent link for every box. Mechanism sketched in the hub v0.66.0 REPORT; NOT built (R-27 slice 1 deliberately did not touch the controller). M idea (minted by hub v0.66.0) Origin: hub v0.66.0 slice-1 ship (first-box only). Reuses the same /bind/ public page + tokenized-link machinery; adds a controller-side entry point + the operator "mint a link for an existing customer" affordance Flips (2026-10-03): 00 §A "Customer binds their own appliance (self-service)" — the second-box flow.
R-12 [P4] Cluster mode: agent-follows-guest, bind-mount reconciliation on HA migration XL idea Scoped 07-15; interim = HA-group pin to one node. Driven by Peti's two-node cluster Flips (2026-10-03): 00 §G — a new row "a guest follows its node in a cluster" when built.
R-14 [P4] Headscale/WireGuard spike: Minecraft/gaming port connectivity (CGNAT-proof, sovereign DERP fallback) M idea Flips (2026-10-03): 00 §E — a new row "game-server ports reachable behind CGNAT" when built.
R-15 [P4] Multi-user dashboard accounts (household members, roles) L idea Single password is a stated alpha limitation (R-11). Launcher coupling — REVISED (controller v0.165.0): the "share the launcher outside the household" need is now met WITHOUT member accounts — the Indítópult megosztása capability-URL guest link (/s/<token>, information-only, no account) shipped in v0.165.0. What remains for this arc is member-specific: per-member tile visibility (each member sees only their apps) and the launcher-as-member-landing-page — both live inside this SSO/members arc; the guest-link ruling explicitly SUPERSEDES the earlier "members are how you share the launcher" framing Flips (2026-10-03): 00 §E "Multiple household users / per-person accounts" (MISSING). Relates to R-811 (a second login step): the same door should carry both.
R-72 [P4] Curate brand_color for the top catalog apps XS idea Parked follow-up to the v0.163.0 launcher. .felhom.yml brand_color (#rgb/#rrggbb) overrides the deterministic slug-hash tile color; no catalog app sets it yet. Pick brand-accurate colors for the most-installed apps so their launcher tiles match their real brand. Catalog-only change (app-catalog-felhom.eu), brand_color is already omitempty and consumed by the controller Flips (2026-10-03): 00 §E "Indítópult (app launcher)" — brand colours.
R-74 [P4] Island control plane on a CLUSTER (Peti's 2 nodes) — bring R-50's island bridge to a multi-node PVE cluster. M idea (Phase C of R-50, parked) R-50 shipped the island for the ONE-host fleet (demo-hp, demo-felhom). A cluster needs bridge parity on every node: either per-node identical /etc/network/interfaces vmbr9 stanzas (simplest, drift-prone) or — preferred at ≥2 nodes — a Proxmox SDN zone/vnet defined cluster-wide (one definition, auto-applied per node). The guest island IP is per-guest + node-independent; the agent-follows-guest rule holds (each node's agent binds its own vmbr9 169.254.253.1). Migration order per the spike: drill-proven → demo (done) → Peti (this row). Its own supervised runbook, coordinated with Peti (a live customer). Completes the capability-map "site/network change" row for clustered installs. Source: audits/SPIKE-island-bridge-2026-07-25.md (cluster-parity finding) + RUNBOOK-island-migration.md (single-host procedure to generalise) Flips (2026-10-03): 00 §G — a new row "the island control plane on a multi-node cluster" when built.
R-41 [P4] [SLICE 1 SHIPPED 2026-07-21] The catalog has no standing "does every template still deploy?" check. Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found 5 that had NEVER been deployable: papra (missing required AUTH_SECRET), zipline (v4 renamed CORE_DATABASE_URL → DATABASE_URL), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the v tag prefix + new required env), glance (needs a seeded glance.yml the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus 7 broken healthchecks and 2 apps whose images no longer resolve at all (plant-it, wanderer). M idea Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in scripts/check-image-pins.py — a mechanical gate run on every change. Cheap first slice: a resolvability gate (docker manifest inspect every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. Silent rot is the real risk — an app can die upstream and nobody learns until a customer clicks Telepítés
R-45 [P4] [P2] Unified async-job feedback. Every long operation invents its own progress surface, or none. Tonight produced three more one-off cards (v0.147.x: samba bring-up, offsite progress, restore result) on top of two existing patterns (deploy 3-step panel; storage-init/netstorage status poll). They agree on nothing: some use {ok,data} envelopes and some raw JSON, some poll 1 s / 1.5 s / 3 s, some are in-memory-only and lie after a restart, and each re-implements single-flight + snapshot + phase→Hungarian mapping. M idea Origin: 2026-07-19 feedback slice 1 (controller v0.147.0). The cases to generalise from are all in-tree: web/storage_init_job.go (the best shape — acquire/release/set/snapshot), web/netstorage_job.go, web/samba_ensure_job.go, backup/opstatus.go, backup/offbox_progress.go. Shape: one job registry + one poll endpoint + one client-side renderer, phases declared per job. Two lessons tonight that any framework must encode: (1) a terminal state must be probed, not inferred — compose up -d exits 0 on a crash-loop; (2) a progress source that reports nothing is normal, not broken — restic reports 0 bytes for a whole incremental run, and a bar that sits at 0% is worse than no bar. Also fixes the restart hole: in-memory job state currently vanishes and the card silently disagrees with reality 2026-07-20 — the first bill for NOT having this arrived, and it was customer-facing. The samba card's poll (web/samba_ensure_job.go + sharing.html) mixed a job EDGE and a service LEVEL on one JSON field, and /sharing reload-looped at ~1.2 s for every customer with sharing enabled until controller v0.151.0 (audits/DIAG-sharing-2026-07-20.md, S-1/S-4). v0.151.0 fixed THAT card's contract only — the framework is still this item. Third lesson for it to encode, beside the two already listed: a phase a client answers with a one-shot action must be an EDGE the registry SERVES ONCE, and must never be synthesised from a level; if it can be re-read, it will be re-acted on. Flips (2026-10-03): 00 §B "Deploy an app from the catalog (… health-aware progress)" — one job surface for every long operation. Re-ranked 2026-10-03: [P2] → P4: polish; every long operation already shows a progress surface of its own.
R-65 [P4] Buddy-box backup replication, cross-household — two Felhom boxes in different homes replicate backups to each other. L idea (post-alpha, spike-first, 2026-07-22) The natural big sibling of R-64: two households each hosting the other's encrypted backup tier. Explicitly spike-first — the transport is NOT SMB (R-64's live-share protocol is wrong for backup replication across the internet: no auth story between households, no resumability, cleartext LAN assumptions); candidates to spike: restic rest-server / rclone / syncthing over the existing WG/tailnet plumbing, encryption keyed so the buddy can never read the payload. Sits on top of the offsite tier's FILL/OVERSUB thresholds thinking (R-5 aggregate). Flips: would add a "cross-household buddy replication" capability row (currently unlisted). Pairs with R-64 (same topology, different transport + guarantees) Flips (2026-10-03): 00 §C — a new row "backups replicated to a second household" when built.

UPDATE-ARC — what is still open (collapsed 2026-10-03)

The app-update arc (audits/SPIKE-app-update-2026-09-01.md; its living design is architecture/09-update-architecture.md). Parts 1–10 SHIPPED (slices 1, 1b, 2, 3, 4 and 5 in controller v0.233.0 – v0.238.1, then the tested-steps, ladder, hold, undo and night-update parts through controller v0.271.0 — 09 §6.4); part 11 deferred by ruling (09 §3 decision 18: the fleet sweep waits until the fleet grows). The full previous entry is in ROADMAP-HISTORY.md. The open work is in the register, by id: R-450 (slice 6 — automatic within a major, never across one), R-451 (slice 7 — the fleet sweep, DEFERRED), R-469 (lift the engine-major rule per app), R-440, R-446, R-458, R-462, R-607, R-610, R-618, R-621, R-622, R-683, R-738, R-785 — all under App updates in OPEN-ITEMS.md. Flips: 00 §B "Automatic updates at night — the update leg" and "Whether the box UPGRADES an app by itself".

Gating candidates — Campaign 12, Part 4 (2026-08-08) — all P4

Ranked by what a gate would be worth, using Campaign 12's own instance counts as the evidence (audits/CAMPAIGN-12-class-sweep-2026-08-08.md §4). G-1 was built (scripts/wire_contract_gate.py); G-6 and G-7 were recorded as a NO with evidence — all three are in ROADMAP-HISTORY.md. Flips: none — gates change no capability; they keep the map's claims true.

Rank Item Size Status Notes
G-2 Gate C3 — a success verdict may not be set where an incompleteness signal is in scope. Assert that every literal success-status assignment either has no gap/skip/missing signal available at that point, or consults it. S candidate The whole population in the controller is 9 sites — Campaign 12 read all of them, which is why this class is the one where "no others exist" is supportable. Small enough to gate by enumeration rather than by inference. Known miss: verdicts expressed as booleans, enum constants, or the absence of an error — and the controller does use those elsewhere. Instances: R-240 (open), R-258 (new).
G-3 Gate C4 — every rendered count/size/percentage needs a *Known companion. M candidate — UNBLOCKED 2026-08-08: the convention decision it was waiting for has been made The decision owed was "how does this codebase say 'we could not look'", and it is now ruled (CONTEXT.md S-39, shipped in controller v0.210.0 / R-259): an explicit …Known bool companion beside the figures, checked in the template before anything is rendered — the shape Offbox.StatsKnown already used, whose own comment carries the reasoning ("a 0%-wide bar over an unread store is a picture of emptiness, and a picture is a claim"). Pointers and separate error fields remain legitimate Go and both still exist here; the ruling is that new three-state figures use the companion, because a codebase with three dialects cannot be gated by a name-based check. Existing call sites were deliberately NOT converted — that conversion is the bulk of this item's M and is what remains before a gate can be turned on without a wall of false positives. Next step is therefore a survey, not a gate: count the rendered figures that lack a companion, decide which are genuinely three-state, convert those, then gate. Instances so far: R-225 (fixed, the pattern's origin), R-259 (fixed, the ruling).
G-4 Complete C1's runtime body assertion — 4 of 27 pages today. L candidate — the expensive one, and honestly so secret_in_markup_gate.py covers all 36 templates on the NAME-based check and is blind to a secret under a neutral page-data key — verified 2026-08-08 by replaying the three pre-fix templates through it: it convicts 2 of 3 and not the third. The runtime assertion catches all three; extending it means constructing each remaining page's data in a test, which is a per-page cost and is the real reason it has not been done. Do NOT adopt the Go-side mirror Campaign 12 wrote as a gate on its own — 27 candidates, 0 findings is bad signal-to-noise in front of every push. → R-255
G-5 Gate C7, narrowly — uniqueness claims only. A comment saying "X is the ONLY writer/place/caller of Y" is mechanically falsifiable; assert it. S candidate — narrow by construction Covers ~60 of the 2652 production invariant comments. The other ~97% of the vocabulary (never, always, must not, guarantees) is not mechanical and a gate must not pretend otherwise. Instance: R-263.
G-8 R-242's untouched half — catch a SKIPPED VOUCH. S then M candidate — the one with a live recurrence Measured during Campaign 12's own bake: golden_currency_gate.py flipped red→green the moment the evidence DIRECTORY existed, before the round-trip download finished and with no vouch near it. (a) Cheapest, now: have the bake session re-read /configuration after the operator's Save and write the observed golden_version into the evidence README as a machine-readable line; the gate then requires that line rather than the directory. Detects the FORGETTING, which is the actual failure mode. (b) Loudest, when the hub is next touched: a hub-side daily check comparing the vouched golden against the newest controller the fleet reports — it fires within a day and catches a silent ROLLBACK too, which nothing in git can ever see. (c) A checklist item is not a fix — R-242 already was a rule without a mechanism and it recurred the next day. → R-242

Before the first paying customer

One list, not two: root STATUS.md, section "Before the first paying customer". The old "Pre-invite checklist" (2026-07-18, the remote-tester era) is in ROADMAP-HISTORY.md; its open legal half — the R-11 rulings (the tester agreement was never written) — is now part of R-809.

Loose notes in this folder

Verdict per file: backlog/README.md.