Files
felhom.eu/documentation/backlog/ROADMAP.md
T
admin 7ee25925f9
gates / gates (push) Failing after 17s
R-87 CLOSED: live evidence, capability row, architecture verdict, registers
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp.

LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level
through the exact route the debug button invokes):

- THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest
  declaring nothing - was pushed to the live store and the proof returned verdict "fail"
  with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE
  offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself
  the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes.
- THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden:
  stopping the app does NOT produce a failed dump leg, because the off-site run's own
  capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is
  therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no
  forget and no prune. State restored: the product's own run made a healthy snapshot the
  newest again and the proof then passed opengist.
- The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s
  each, matching the spike's measured band.
- The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear
  and vanish across a real restic check, and ZERO across the proof - including a direct 6x
  test of the snapshot-lookup argv, which settles that restic snapshots does not lock in
  0.14.0 either.
- Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the
  flag returned skipped:true duration_ms:0, no verdict, no alarm.
- The customer's own verification copies were untouched throughout, which is the safety
  property the separate proof root exists for.

ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at
19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the
proof; I did not establish what it was.

CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only -
the job is REGISTERED, which is not the same claim.

07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one
sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore
puts data back into a running app. Without that sentence the new green tick reads as
covering the drill.

REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 ->
152. No new rows minted. R-408 and R-409 stay open and are referenced by this work.

golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the
newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job.
A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses
--no-verify for that reason - bypass #8.
2026-08-31 21:32:06 +02:00

77 KiB
Raw Blame History

ROADMAP — future features & open work

What this is: the prioritized decision log of planned/open work. Items are intentions, not claims about live behavior — the capability map (architecture/00-capability-map.md) is the only place that states what the platform does today.

Lifecycle: idea → spiked → spec'd → in-progress → shipped (item collapses to a one-liner with the version, and the corresponding capability-map row changes status with evidence). Items can also be killed (keep the one-liner + why — decisions are worth remembering).

Coupling rule: every item names the capability-map row(s) it flips. Every map gap row points back here by ID. Neither file duplicates the other's content.

ONE REGISTER — operator ruling, 2026-08-22 (R-369)

Open FINDINGS live in OPEN-ITEMS.md, not here. This file keeps history and reasoning, which is what its first paragraph has claimed since 2026-07-27. On 2026-08-22, 16 rows were moved to the register — every row that asserted something checkable about the shipped product, plus one owed operator decision. Their copies remain below, marked MOVED -> OPEN-ITEMS.md, and are not deleted: this file's job is history.

The sorting rule, so it need not be re-invented: does the item assert something about the shipped product that a reader could go and check, and find false? If yes it is a FINDING and it belongs in the register. If it proposes something that does not exist yet — a feature, a spike, a curation task — there is nothing to be wrong about, and it stays here as an intention.

scripts/one_register_gate.py enforces it: a row here that is neither an idea nor done, and has no counterpart in the register, fails the push.

Priorities: P1 = closed-alpha blocker · P2 = close during alpha · P3 = post-alpha. Existing loose notes in this folder (FOLLOWUP-*, FIX-M*) are absorbed as references below.


P1 — closed-alpha blockers

ID Item Size Status Notes / map rows flipped
R-116 The drive-absent alarm and its recovery are a mismatched pair — generic on the way out, specific on the way back S idea — PROVEN LIVE 2026-07-29 Absent fires storage_disconnected; return fires backup_target_restored. backup_target_absent never fires at all (count 0 across a full Session-C run), so an operator gets an alarm they cannot match to its recovery — exactly what notifyDriveReturned's own comment forbids. Root cause: notifyDriveAbsent (intermediary.go:635-646) branches on isTarget[a.Path] with a.Path the GUEST path, and driveTargetByPath (:602-616) builds it as out[GuestPath] = d.BackupTarget — but the drive is TWO /disks rows and the flag and the guest path sit on different ones: the felhom-backup storage row has BackupTarget: true (felhom-agent/internal/localapi/disks.go:211) and gets a guest path only while classified user-data, while the registry union row has the guest path and never assigns BackupTarget (disks.go:265-267). Absent ⇒ the flagged row loses its guest path ⇒ the union row writes false ⇒ generic. On return the rows rejoin ⇒ specific. v0.184.1 fixed the KEYING, not this. Only reachable because R-113 made the gate fire at all. Fix likely agent-side; decide the repo first. Blocks E-2's C5. Evidence: audits/SESSION-C-2026-07-29.md §5
R-113 The drive-absent gate cannot fire on device loss — E-2b's alarm is wired to an unreachable condition M idea — PROVEN LIVE 2026-07-29 planDriveGates (felhom-controller/internal/web/intermediary.go:216-262) treats a path as present by OR-ing in d.BoundUnderParent, which the agent derives from GuestSeesMount() — "is this path a mount target in the guest's /proc/<pid>/mountinfo" (internal/localapi/disks.go:210). The raw drive mount is a device-bound systemd unit and dies with the device; the agent's own bind under the shared parent is not device-bound and its mountinfo entry outlives the device, so the gate reads it as present and notifyDriveAbsent is never called. Live on a fresh box: target drive hot-detached, agent said enrolled drive absent by UUID every 20 s for 4½ min, controller logged 0 [gate] lines, hub received zero events — neither backup_target_absent nor the generic storage_disconnected. Not a virtualisation artefact (device-bound-mount vs manual-bind is the same on metal); caveat: SCSI hot-detach, physical unplug not staged. Sixth instance of seam-built-but-never-wired — E-2b wired the seam to a condition that cannot occur. Evidence: audits/E2D-fresh-vm-2026-07-29.md §5.2
R-112 E-2's degraded banner and offer have no UI consumer — correct endpoint, invisible to the customer S idea — PROVEN LIVE 2026-07-29 GET /api/storage/backup-target returns byte-exact Hungarian copy (verified on a live box), and nothing in the product asks for it: grep 'backup-target' across every *.html/*.js/*.css → 0 hits; no template references OfferPath/Degraded/the copy; resolveBackupTargetState and degradedMessageFor are consumed only by the JSON handler, with no page handler injecting the state. Decisive contrast: the templates fetch 18 distinct /api/storage/* endpoints — backup-target and backup-target/assign are the only two with zero references. The handler's own comment calls itself "the dashboard's source for the degraded banner and the offer". v0.185.1 shipped as "the offer endpoints were mounted where nothing routed to them" and fixed the mount, stopping one layer short of the render; its test pins dispatch, not reachability. Fifth instance of the class. Fix R-114 first — wiring this alone starts showing customers a wrong message. Evidence: audits/E2D-fresh-vm-2026-07-29.md §5.1
R-114 On target-drive loss the customer is told the wrong story and offered the drive that vanished S idea — PROVEN LIVE 2026-07-29 With the assigned target absent the endpoint returned degraded:true, target:"felhom-backup" plus the "ugyanazon a lemezen van, mint a rendszer" message — false, the target is a missing drive, not the system disk — and an offer_path pointing at the drive that just disappeared. resolveBackupTargetState falls through to the generic degraded branch whenever no disk satisfies d.BackupTarget && d.MountPath != "", never distinguishing never configured from configured and now missing. Shares R-113's root cause (two disagreeing presence signals), different code path and fix. Invisible today only because of R-112. Evidence: audits/E2D-fresh-vm-2026-07-29.md §5.3
R-3 Friend-alpha onboarding runbook (generalized from pilot/RUNBOOK-peti-return-2026-07-13): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step M idea Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of PETI-tester-agreement.md. Next from-scratch rehearsal to include customer DELETE + re-create — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). S6b (manual stale-host delete before re-enroll) is OBSOLETE — re-enrollment upserts the existing host row cleanly (store.UpsertHost ON CONFLICT DO UPDATE; handleAdminCreateHost no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE): one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). Remaining consistency gap: the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b

P2 — during alpha

Sub-rank P2-HIGH = close before the first REMOTE tester. These are the 2026-07-18 N100 rehearsal's findings (tests/VALIDATION-n100-rehearsal-2026-07-18.md). They are not P1 — the rehearsal proved the product flow works — but each one either misleads the operator, misleads the customer, or hides a failure, and all of that gets materially worse the moment the box is somewhere you cannot walk over to.

ID Item Size Status Notes
R-6 Spike: LAN service discovery from the guest — SSDP multicast (UDP 1900, DLNA), WSD (Windows discovery), mDNS; host-network vs macvlan; is the customer LXC LAN-bridged in appliance deployments? M spiked (2026-07-18) VERDICT: appliance guest IS LAN-bridged (own DHCP lease on the household /24); multicast discovery works ONLY in the guest netns — guest-direct or Docker --network host (SSDP/mDNS/WSD all PASS both ways); the default docker bridge is categorically DEAF to LAN multicast (WSD/mDNS RX FAIL, unicast-publish PASS). Real samba+wsdd on host-net → Windows 11 ProbeMatch + FELHOM-SPIKE renders in Explorer + 445 + authenticated SMB round-trip all PASS; real SSDP MediaServer:1 advert reaches both LAN clients. → R-7 SMB stack MUST be host-network LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. Caveat: vmbr0 multicast_snooping=1 worked only because the household router is a live querier — customer LANs w/ snooping+no-querier, and Peti's BYO bridge, are UNTESTED gaps. S4b (human leg, the sharpest finding): wsdd makes the box VISIBLE but the Explorer double-click FAILS 0x80070035 — WSD gives no name resolution; the flat \\FELHOM-SPIKE resolved by no path. Adding nmbd (NetBIOS) fixed it live (flat name resolves + mounts). → R-7 needs smbd+wsdd+nmbd (+avahi/.local for modern clients), not wsdd alone. Doc: audits/SPIKE-lan-discovery-2026-07-18.md.
R-7b Share backup EXECUTION — put share data into the live tier-2 + offsite runs (the design fork reported by R-7 slice 1) M SHIPPED (controller v0.145.0, 2026-07-18) Viktor's ruling: Model B′ — a SIBLING shares source. New, additive job/leg code reusing the proven primitives (tier-2 mirror seam, restic wrappers, soft-quota/enlargement gate, status recorders) while leaving every per-app engine path byte-identical — NOT a synthetic recovery unit (breaks on multi-drive shares, wraps 1 KB of JSON in dump machinery) and NOT engine-loop surgery. The B′ invariant is enforced by test in both tiers, red-proofed. Tier 2 → RunSharesTier2 (legs grouped by SOURCE drive → backups/secondary/_shares/<driveKey>/<share>, payload at _payload/, layout marker LAST). Tier 3 → runOffboxSharesLeg: ONE extra restic backup --tag felhom-offbox --tag _shares placed after the app loop and BEFORE retention, so forget --group-by host,tags covers the new group with no flag change; a quota-blocked push degrades to the manifest only, never to nothing. Restore → „Megosztások" on /backups/restore: scratch, then a missing-only merge whose every destination is PREFIX-ASSERTED against live storage roots, definitions merged existing-wins, then ReconcileSamba, then the credential. The payload (_shares-manifest.json + a best-effort secret-bearing passdb.tar) is what makes DR return files + configuration + password rather than loose bytes. Fold-in: samba joins the liveness set — EffectiveProtected adds the CONTAINER felhom-samba exactly while sharing is on. FULLY PROVEN-LIVE on demo (2026-07-18), all four legs. (1) tier-2: real /api/backup/tier2 trigger → _shares tree + marker + payload on the cross-drive target, mirrored file md5-identical, payload 0600 preserved. (2) offsite: Viktor's manual run 12:18:16Z → snapshot e0b9d723 (tags felhom-offbox,_shares) with the payload dir + both share folders; a second run via the „Távoli mentés" button → 4e2b15ec, containing _shares-manifest.json (418 B) AND passdb.tar (855 040 B), both 0600, share files with uid 1000 preserved. (3) restore round-trip: probe file + the dokumentumok DEFINITION deleted via the real endpoints, then „Megosztások" restore + place → 1 file(s), 1 definition(s) re-added, 1 kept, 0 refused, credential=true; probe back md5-identical, the two pre-existing files NOT overwritten (missing-only proven on live data), definition back with its ORIGINAL flags and created_at, smb.conf re-rendered, filmek untouched. (4) liveness: samba stopped → health_critical pushed and hub-accepted (200) → self-healed. Remaining human leg: SMB positive auth with the real household password (never persisted by design). Correction: an earlier revision of this row and of the ship REPORT wrongly claimed the demo box had no offsite target — the verification read a guessed settings key (offbox_target) instead of the real one (offbox); root cause dissected in REPORT §7b. Findings: the reserved-name assumption was FALSE (nbNameRe accepted „_shares" as a share name — now refused); the alert/e-mail pipeline needed NO change and adds no new event type. Docs: controller/sharing.md; ship report felhom-controller/REPORT.md.
R-8 DLNA (gate input now exists — R-6 spiked 2026-07-18: SSDP reaches LAN clients from host-net): validate Jellyfin's built-in DLNA server first; only add minidlna to the catalog if Jellyfin-DLNA fails S idea (unblocked) Don't add catalog weight before proving the cheap path. R-6 confirmed the cheap path is physically viable — Jellyfin DLNA must run host-network (same multicast constraint as R-7)
R-9 Uninstaller trio (from 07-15 Peti session): cluster-aware felhom_guests guard (node-local pct list deletes cluster-wide pveum objects); saferemove detection + time estimate + opt-in --quick-remove (never mutate storage.cfg); smarter restore_storage default for BYO clusters (shared storage, not local-lvm) M idea Second item's rejected alternative (temp-disable-and-restore) stays rejected — crash window silently downgrades cluster wipe policy
R-202 The orphan card promises recoverability unconditionally, which after R-198 is true going forward and false for anything already orphaned S gate hit 2026-08-04 — card untouched, sentence still live Blocked on knowing which escrow generation an orphaned repo belongs to (R-199/R-201). A single ACK boolean can say a retained recoverable blob EXISTS but not that one COVERS this repo; a conditional promise that can still be false is worse on that surface than a hedged one
R-201 The wipe-and-recover drill L PREPARED, HALTED BEFORE THE WIPE (2026-08-04) Nothing irreversible done. Established live for the first time: a rebuilt box's off-site run REFUSES (the orphan card, not a silent fresh history — closes R-193 Q3), the orphan reset works move-aside-never-delete, and demo-hp's pre-rebuild key is permanently gone (superseded four hours before v0.93.0). Resume needs R-203
R-19 Internet-outage customer-experience drill: pull WAN on demo, verify lan_resolver path, document what the customer actually sees/does S idea Flips map row E "LAN access" IMPLEMENTED→PROVEN-LIVE
R-34 Backup data lifecycle management. An "inactive backups" section on „Távoli mentés": apps that have snapshots but no active backup — disabled OR uninstalled — listed with name / size / last snapshot / restorable, plus an explicit double-confirmed per-app delete via restic forget --tag + nightly prune. M idea RULING: the offsite toggle NEVER offers deletion — policy and destruction stay decoupled. Turning backups off must never be a data-destroying act, and deletion must never hide behind a toggle. Origin: 2026-07-18 rehearsal. Pairs with R-32 (that one is the operator's view of dead bytes; this one is the customer's)
R-27c Customer self-bind, slice 2 — console-passphrase bind. Viktor's direction: bind using a passphrase shown on the box console, alongside (not instead of) the emailed capability link. M idea Security constraints from the session ruling, all load-bearing: passphrase issued at customer creation; the global-lookup endpoint must be spray-hardened — per-appliance and per-IP caps, constant-time comparison, a single generic failure (no oracle), alerting on abuse; an accent-free wordlist (console keymaps are not Hungarian); the web capability-link path is RETAINED; claim-by-email is RETAINED as the delivery-channel proof. Also under this item: the self-bind email gains the public universal-ISO download link + two-line instructions (the DIY case). Secret-bearing per-customer ISOs are ruled OUT. Sibling of R-27b (second-box flow) — different axis, both build on the same /bind/ page
R-50b [P2] A root-owned privileged host artifact is delivered unversioned from main — "which wrapper is on this host?" is unanswerable. configs/felhom-pbs-apply installs to /usr/local/sbin/felhom-pbs-apply (0755 root:root) and is the pinned sudoers vector for create|reconcile|grant against /etc/pve/priv/storage. It is fetched by felhom-host-install.sh:1914 via fetch_raw, which hits raw/branch/main/<path> — no tag, no pin, no checksum, and no record in the Day-0 artifact manifest, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to main reaches the next install of every host with no review gate between commit and root-owned deployment. S–M (a) SHIPPED 2026-07-21; (b)/(c) open Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish (felhom-controller/REPORT.md §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least visible; (b) fetch_raw takes a pinned ref (tag or commit) supplied by the manifest rather than main; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. (a) is the cheap honest first step and would have caught this class already. Pairs with R-39 (whose remaining fleet half is specced separately) (a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2. ArtifactManifest.WrapperSHA256 + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. An unknown on EITHER side reads as quiet, never as drift — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed 2888f2ea…, matching no commit anyone could name; it now reports 104db0a4… against a vouchable manifest value. (b)/(c) REMAIN OPEN: the wrapper is still fetched unversioned from raw/branch/main — this makes drift visible, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible.

14:29:28 [gate] boot 1784525102-11906045: live bind confirmed — recreating drive-backed app calibre-web (state=stopped) onto /mnt/felhom-drives/hdd_1 14:29:29 [gate] boot 1784525102-11906045: 1 drive-backed app(s) left stopped — zero containers means the customer stopped them on purpose

immich is absent from the recreate list and came back STOPPED (0 containers) — on 2026-07-21 before this fix, the identical fixture brought it back RUNNING. calibre-web recreated, bookstack back, [bootrecon] no boot-orphaned apps (consistent — nothing was left orphaned for it to adopt), ZERO alerts, whole convergence ~15 s from reboot to steady state. The left stopped INFO line fired in production for the first time, so the honoured path is observable rather than silent | | R-56 | [P3] Apps do not say how technical they are, so a beginner can be ambushed by a config-heavy one. The catalog presents every app as equally approachable — one Telepítés button, the same Hungarian copy — but they are not. Glance needs a hand-written glance.yml before it does anything; some apps need a reverse-proxy or API concept to configure; others genuinely are install-and-use. A tester who picks the wrong first app concludes the PRODUCT is broken, not that they picked an advanced app. | S | idea (filed 2026-07-21) | Origin: TASK-E Part 3 — filed, deliberately not implemented. Shape: a difficulty: field in .felhom.yml (kezdő / haladó / technikás) surfaced as a catalog-card badge and repeated on the deploy screen. Cheap and incremental: one optional metadata field plus a badge, classifiable app-by-app with no migration — an app with no difficulty: simply shows no badge. This is the constructive half of the glance ruling: glance STAYS in the catalog (operator ruling 2026-07-21 — it is a legitimate app, not a broken one; its missing seeded glance.yml is a known pre-existing finding), and the honest fix is to LABEL it rather than hide it. Pairs with R-41: that gate proves an app CAN still deploy; this field tells a customer whether THEY should be the one deploying it. Badge plumbing is ALREADY BUILT (controller v0.158.0) — web.MetaBadge + the meta_badge template partial + the lifecycleBadge funcmap entry were written generic for exactly this: a difficultyBadge funcmap function returning the same *MetaBadge, plus a difficulty: field on stacks.Metadata, is the whole remaining job. No new markup, no new CSS. Re-sized accordingly | | R-58 | [P2] Assisted disk-picker install mode — the installer should let the operator CHOOSE the target disk instead of requiring the serial up front. Today an install is either unattended (the answer file pins one ID_SERIAL_SHORT, which you can only know by first booting the machine) or match-nothing safety (aborts by design). That forces a two-boot dance for every new box: boot the safety ISO to read the serial, rebuild the ISO armed, boot again. | S–M | idea — operator ruling 2026-07-21 | Operator's argument, verbatim: "the installer should list the available storage devices (excluding the installation media) and let us select one, and continue." Shape: a THIRD ISO mode alongside the two that exist — unattended-serial and match-nothing-safety. It enumerates candidate disks with size / model / serial, excludes the installation media itself, takes a selection plus a confirm, and proceeds. Unattended+serial REMAINS the appliance/factory mode — it is the right shape when the machine is provisioned in bulk and nobody is standing there; the picker is for the case where somebody is. Slice 1 (cheap, same code surface, do this first): improve the abort screen. On filter-no-match the installer currently just fails safe and says nothing useful — it should print the candidate table (size/model/serial) plus the one-line hint naming which serial to put in the profile. That alone collapses the two-boot dance from "boot, guess, go read docs, rebuild" to "boot, copy the serial off the screen, rebuild", and it is the same enumeration code the full picker needs. Why it matters beyond convenience: it is the BYO / reinstall flow — a customer's existing hardware, or a rebuild of a box whose disk layout nobody recorded, is exactly where the serial is unknown and a wrong guess is destructive. The current fail-safe is correct but mute. Origin: TASK-G, arming the HP install ISO — the serial had to be read off the board by hand between two boots | | R-62 | [P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape. | XS | idea (operator, 2026-07-22) | Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; wording only — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). | | R-64 | „Felhom↔Felhom media pairing blessed" — the two-box SMB pairing (one box shares, the other mounts it as NAS storage) becomes a supported, documented flow. | XS–S | idea (2026-07-22) | Origin: the operator ran the pairing drill on the live demo pair and it WORKS — the drill itself is the pending evidence leg (a written run-through with the R-66 surfaces in play). R-66 shipped the enabling visibility: the serving box's address is now on its own Beállítások → Rendszer „Hálózat" card, and the add form names the NetBIOS trap. Blessing = a short customer-facing recipe (documentation/controller/network-storage-nas.md naming-caveat paragraph is the seed) + one supported-path sentence in the capability map. Flips: would add a "Felhom↔Felhom media pairing" capability row (currently unlisted). Pairs with R-65 (same two-box topology, entirely different transport + guarantees) | | R-69 | F14-full: an operator push channel that actually interrupts (ntfy / Telegram / similar), beyond mail-client priority flags. F14-light (v0.71.0 headers + Gmail filter) nudges a mail client; a 15:29 node_down should reach the operator's pocket in seconds regardless of inbox hygiene. Needs: channel choice (self-hosted ntfy on k3s vs Telegram bot), dispatcher fan-out seam, per-severity routing, quiet hours. | M | idea | Origin: AUDIT-power-outage-recovery-2026-07-22.md F14. Deliberately NOT built in the v0.71.0 train (scope-forked per the task spec) |

Recovery-model gaps (2026-07-28, 07-backup-architecture.md §10.2)

Minted when 07-backup-architecture.md was rewritten as the recovery model. Every one of these is a divergence between that model and the system as it is, and each is cited there. They are filed at P2 as the neutral default, not ranked — ranking them needs the per-scenario RTO/RPO targets that 07 §11-C records as never having been stated. Flips: 00-capability-map.md §C rows, which now cite the matrix rather than restating the route.

ID Item Size Status Notes / map rows flipped
R-127 data_key: true is unreliable (4+ encryption keys unflagged, contradicting the catalog's own labels), and O4 can regenerate a DB password that no longer matches the restored data directory S/M READY — NEW 2026-07-30 Found by D5's Part 0, and the reason D5's boundary became type: secret rather than data_key. Leg (a): flag the missing keys (catalog-only) + pin flag-vs-label agreement; the residual risk after D5 is that the fail-closed gate keys on data_key, so an unflagged key missing from both sources lets the restore proceed onto undecryptable data. Leg (b): a regenerated DB password is silently wrong — POSTGRES_PASSWORD is ignored once PGDATA is non-empty, so the app cannot authenticate while the replay still succeeds over the local trust socket (proven live on postgres:16-alpine). v0.188.0 corrected the false "stored data is unaffected" WARN but added no guard. Flips: 07 §7.4
R-133 The vaulted break-glass console credential is PLAINTEXT AT REST, so every hub DB backup is a fleet-wide console-credential dump. host_recovery.secret holds each managed box's root@pam password verbatim (hub/internal/store/host_recovery.go — "a hub-held secret, operator-retrievable (NOT zero-knowledge like escrow)"), so anything that copies the SQLite DB — a Longhorn snapshot, a PBS backup of the hub PVC, a hand-taken copy during a diagnosis — carries root console access to every Felhom host in one file, with no second factor and no key to withhold M READY — NEW 2026-07-31 The DEFERRED LEG of hub v0.84.0, filed as its own ID because v0.84.0 changed only WHO can retrieve the secret, never how it is stored — the at-rest shape predates it and is untouched by it. v0.84.0 made it more worth doing, not more broken: putting retrieval behind the hub session means the hub login password alone now unlocks console root fleet-wide, so the DB and the login are the whole of the protection. Shape: envelope-encrypt the host_recovery.secret column under a KEK held OUTSIDE the DB (k8s Secret / out-of-band file, the way manifests/ already keeps the bearer out of git), so a DB copy is opaque the way escrow blobs already are — the contrast is the argument, since the hub already proves it can hold a secret it cannot itself read. Constraints the design must respect: the credential must stay retrievable when the box is unreachable (that is the entire point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Flips: architecture/00-capability-map.md "Break-glass management-plane recovery" — the row that today reads IMPLEMENTED with a plaintext-at-rest caveat
D5 Move app secrets into the LOCAL recovery unit so Tier-1/Tier-2 restore stop needing the guest and stop needing R M SHIPPED + PROVEN-LIVE — controller v0.188.0, 2026-07-30 The arc's architectural centrepiece. Tier-1/2 no longer depend on the whole-guest tier — a customer needs the DRIVE AND NOTHING ELSE. Part 0 tested this row's own premise and rejected it: data-keys-only is both insufficient and unsafe, because data_key is unreliable (→ R-127) and a DB password is not resettable in practice (POSTGRES_PASSWORD is ignored once PGDATA is non-empty, so a regenerated value leaves the app unable to reach its own restored rows while the dump replay still reports success — proven on postgres:16-alpine). Operator ruling: type: secret travels (45 fields), type: password never (7) plus a code register (vaultwarden/ADMIN_TOKEN); plaintext, because withholding the internet-reachable class is what licenses it — the two are coupled. stacks.PortableSecretEnvVars is the single boundary; the register is code, not a catalog flag (R-97a). Precedence: the UNIT WINS (its secrets match the data being restored, not merely the newest), pinned both directions. Fail-closed data-key gate UNCHANGED. Manifest schema 2; schema-1 units still restore. Proven live on a scratch drill guest: AdventureLog restored with the guest app.yaml moved aside (secrets recovered=2/2, 27.6 s) and the app read the seeded row over TCP with its own credential; Grafana's admin password withheld with 0 hits across the backup namespace. 4 red-proofs each verified to land. audits/D5-drive-alone-restore-2026-07-30.md Flips 07 §3/§7.1/§7.3/§7.4/§8/§10.1 + a new capability-map row
R-126 A .fab bundle — plaintext secrets, optional password — can be exported ONTO a NAS. storageDriveList() (internal/web/handler_export.go) does not filter network paths S READY — 2026-07-30 Split out of R-108 on its closure. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (07 §7.3 records the reasoning). Fix = filter network paths from the export destination list, or force the bundle password when the destination is a share. Flips: 07 §5
E-2 Drive-role machinery around the moved vzdump target. The 2026-07-28 runbook proved the architecture change by hand on both demo boxes; this is the machinery: a backup-target role on StoragePath beside Schedulable/IsDefault/Kind; assignment in the storage wizard (suggest by attribute, refuse the absurd, never decide by transport or removable — on the reference hardware demo-felhom's target IS a USB HDD and BOTH drives report removable=0); unassigned drives do nothing automatically; stickiness (never silently retarget); felhom-host-install.sh creating the target with --is_mountpoint 1 and issuing the FelhomAgentStore ACL; absent-target policy; retention/space accounting on a drive the customer shares; the honest single-drive label; remaining fleet migration M READY — 2026-07-28 Full scope + rationale in runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md §7. Two traps already paid for live: the storage path must BE the mountpoint or the agent reports the target disconnected forever (internal/storage/observe.go:321), and the per-storage FelhomAgentStore grant is mandatory or every backup 403s. Absent-drive behaviour today is fail-loudly, no silent retarget (is_mountpoint 1 proven live) — which is NOT the intended fall-back-and-alarm design. Flips: matrix row 4

Gating candidates — Campaign 12, Part 4 (2026-08-08)

Ranked by what a gate would be worth, using Campaign 12's own instance counts as the evidence. Nothing here was built; §4 of audits/CAMPAIGN-12-class-sweep-2026-08-08.md carries the reasoning and the measurements. The recurring lesson these rank against: a pattern found three times is not closed by looking a fourth time.

Rank Item Size Status Notes
G-1 Gate C5 — the cross-repo tag-reachability check. S BUILT AND CLOSED 2026-08-08 — scripts/wire_contract_gate.py, --fast, registered in repo_gates.py Shipped as ranked. Built BEFORE the fixes and seen failing on 40 fields (documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md) — the order was the method, because deadcode had been rejected for C6 the night before precisely for failing that test. 210 tags checked across 3 declared wires, 51 skipped (generic / opaque / allowlisted, each with a reason). Carries a --selftest that plants an unreachable tag on a real root and asserts conviction, and publishes its blind spots in both its docstring and its output. Two instrument defects the control caught before it was trusted: a substring false negative (grep -F healed_at matched privsep_healed_at), and treating dr_recipe as wholly opaque when its top-level section keys ARE decoded through an allow-list that already cost offsite_restic (R-122) — it is now opaque only BELOW depth 1. Estimate held: the size guess was right and the --fast judgement was right. Not covered, and stated in the gate itself: the hub's desired-state (served as raw stored JSON, no typed emitter) and the agent's local API (no single root). → R-260 CLOSED, R-247 CLOSED, leftover appetite R-264
G-2 Gate C3 — a success verdict may not be set where an incompleteness signal is in scope. Assert that every literal success-status assignment either has no gap/skip/missing signal available at that point, or consults it. S candidate The whole population in the controller is 9 sites — Campaign 12 read all of them, which is why this class is the one where "no others exist" is supportable. Small enough to gate by enumeration rather than by inference. Known miss: verdicts expressed as booleans, enum constants, or the absence of an error — and the controller does use those elsewhere. Instances: R-240 (open), R-258 (new).
G-3 Gate C4 — every rendered count/size/percentage needs a *Known companion. M candidate — UNBLOCKED 2026-08-08: the convention decision it was waiting for has been made The decision owed was "how does this codebase say 'we could not look'", and it is now ruled (CONTEXT.md S-39, shipped in controller v0.210.0 / R-259): an explicit …Known bool companion beside the figures, checked in the template before anything is rendered — the shape Offbox.StatsKnown already used, whose own comment carries the reasoning ("a 0%-wide bar over an unread store is a picture of emptiness, and a picture is a claim"). Pointers and separate error fields remain legitimate Go and both still exist here; the ruling is that new three-state figures use the companion, because a codebase with three dialects cannot be gated by a name-based check. Existing call sites were deliberately NOT converted — that conversion is the bulk of this item's M and is what remains before a gate can be turned on without a wall of false positives. Next step is therefore a survey, not a gate: count the rendered figures that lack a companion, decide which are genuinely three-state, convert those, then gate. Instances so far: R-225 (fixed, the pattern's origin), R-259 (fixed, the ruling).
G-4 Complete C1's runtime body assertion — 4 of 27 pages today. L candidate — the expensive one, and honestly so secret_in_markup_gate.py covers all 36 templates on the NAME-based check and is blind to a secret under a neutral page-data key — verified 2026-08-08 by replaying the three pre-fix templates through it: it convicts 2 of 3 and not the third. The runtime assertion catches all three; extending it means constructing each remaining page's data in a test, which is a per-page cost and is the real reason it has not been done. Do NOT adopt the Go-side mirror Campaign 12 wrote as a gate on its own — 27 candidates, 0 findings is bad signal-to-noise in front of every push. → R-255
G-5 Gate C7, narrowly — uniqueness claims only. A comment saying "X is the ONLY writer/place/caller of Y" is mechanically falsifiable; assert it. S candidate — narrow by construction Covers ~60 of the 2652 production invariant comments. The other ~97% of the vocabulary (never, always, must not, guarantees) is not mechanical and a gate must not pretend otherwise. Instance: R-263.
G-6 C2 — NOT mechanically gateable. — recorded as a no "Names a route" is a judgement, not a predicate. The most a check could do is enforce a convention (e.g. every customer-visible refusal string ends in an imperative clause), which would be gamed rather than followed. Better served by the UI-copy review the felhom-ui-design skill already governs. Instances: R-256, R-257.
G-7 C6 — NOT gateable, and the measurement is the finding. — recorded as a no, with evidence The class looks the most mechanical of the seven and is the one where the off-the-shelf tool measurably fails: golang.org/x/tools/cmd/deadcode re-found neither known instance, and a planted probe showed why — it reports an unreachable exported FUNCTION but not an unreachable exported METHOD on a widely-used type, and both known instances are methods. A bespoke gate would be conservative by construction and would spend most of its output on inert dead accessors (4 of the 5 residual candidates were inert). Recommendation: gate C5 instead — it catches a strict subset of the same "the answer was available and discarded" family with none of the ambiguity. → R-261
G-8 R-242's untouched half — catch a SKIPPED VOUCH. S then M candidate — the one with a live recurrence Measured during Campaign 12's own bake: golden_currency_gate.py flipped red→green the moment the evidence DIRECTORY existed, before the round-trip download finished and with no vouch near it. (a) Cheapest, now: have the bake session re-read /configuration after the operator's Save and write the observed golden_version into the evidence README as a machine-readable line; the gate then requires that line rather than the directory. Detects the FORGETTING, which is the actual failure mode. (b) Loudest, when the hub is next touched: a hub-side daily check comparing the vouched golden against the newest controller the fleet reports — it fires within a day and catches a silent ROLLBACK too, which nothing in git can ever see. (c) A checklist item is not a fix — R-242 already was a rule without a mechanism and it recurred the next day. → R-242

P3 — post-alpha

ID Item Size Status Notes
R-26 Guided old-history recovery via a retained superseded escrow + the recovery code. Enabled by hub v0.60.0 (Part B) which now RETAINS superseded escrow blobs (host_escrow_superseded, ListSupersededEscrow). Build the flow that, given the customer's recovery code, unwraps a retained old blob → recovers the old repo passphrase → mounts/reads the moved-aside .orphaned-<date> repo for restore. M idea (enabled by v0.60.0) Turns "history recoverable in principle" into a real customer-drivable path; pairs with the controller v0.142.0 orphaned-repo move-aside. Origin DIAGNOSE-offbox-repo-orphaned-2026-07-17
R-27b Customer self-bind, second-box flow (controller side). For a customer who ALREADY has a bound box and installs another, the controller shows a dismissable "bind another box" prompt (and a bind-later entry under settings) that walks to the hub /bind/ page — so a returning customer isn't emailed a fresh operator-sent link for every box. Mechanism sketched in the hub v0.66.0 REPORT; NOT built (R-27 slice 1 deliberately did not touch the controller). M idea (minted by hub v0.66.0) Origin: hub v0.66.0 slice-1 ship (first-box only). Reuses the same /bind/ public page + tokenized-link machinery; adds a controller-side entry point + the operator "mint a link for an existing customer" affordance
R-25b RULED: customer DELETE becomes a guided full-teardown cascade. The middle-tier Customer RESET (hub v0.61.0) runs the full external teardown (Hetzner sub-account/box + PBS namespace/groups/token) and refuses while any host row exists. The Danger-zone DELETE still (a) leaves host rows and (b) does NOT run that teardown. M (was S) SHIPPED hub v0.69.0 (2026-07-21) operator ruling 2026-07-21: DELETE subsumes the whole cascade, behind explicit consent. Three separate acknowledgements, each its own checkbox — (1) the host(s) will be deleted, (2) the customer will be RESET including external teardown and offsite data destruction, (3) the customer record and escrow will be purged — plus a typed customer-name confirmation before the button arms. Internal order is host-delete → RESET → delete, which preserves every existing invariant rather than relaxing any: RESET keeps its no-hosts precondition (hosts are already gone by then), and escrow keeps its demote-then-purge custody rule (host delete DEMOTES to retained custody, the final delete PURGES — the one true purge point). Re-sized S → M: this is a multi-step destructive wizard with three acks and a typed confirmation, not a checkbox. Implementation is explicitly NOT part of TASK-E; the row carries the ruling and awaits its own spec. It no longer blocks R-3 — the model is decided, so the friend-alpha runbook can be written against it. IMPLEMENTED per the ruling (TASK-I, hub v0.69.0): POST /configs/{id}/delete now runs hosts → RESET → purge; three acks + typed customer-id + a stale-preview check + the ONLINE-host refusal, all gates before any write (zero side effects on refusal); custody purged exactly ONCE in leg 3 (leg 2 runs with purgeEscrow=false); ruling-3 preserved BY CONSTRUCTION and asserted from inside leg 2; failed legs retain the journal and the dialog offers Resume. Standalone RESET byte-identical. 5 red-proofs. Offboarding guidance: runbooks/RUNBOOK-onboarding-draft-v4.md §G. v0.70.0 follow-up (same day, found validating against the live hub): a completed delete still left the customer on the Customers list and still ALERTING, because GetCustomers() is report-derived and no tier ever deleted a report — new residue leg (reports/telemetry/log-tails/notif-prefs + the credential-bearing appliance_registrations/selfbind_tokens), and ghost customers are now deletable (404 = nothing here, not no-config-row). v0.70.1 (2026-07-22): the ghost delete was implemented but UNREACHABLE — the Danger-zone card (and the customerDeleteOpen script) sat inside {{if .HasConfig}}, so a ghost rendered no Delete button at all (the fourth inert-seam defect; handler tests POST directly and proved nothing about reachability). Render gate split: RESET stays HasConfig-gated, Danger zone gates on Deletable (the exact negation of the preview's 404 predicate), Block/Unblock stay config-only; render tests per branch + 2 red-proofs. Operator live leg: the demo-vm-felhom ghost delete click — PENDING (doubles as the v0.70.0+v0.70.1 live validation; expect residue=ok customer_delete=ok with skipped_no_config Hetzner/descriptor legs, staleness emails stop)
R-12 Cluster mode: agent-follows-guest, bind-mount reconciliation on HA migration XL idea Scoped 07-15; interim = HA-group pin to one node. Driven by Peti's two-node cluster
R-14 Headscale/WireGuard spike: Minecraft/gaming port connectivity (CGNAT-proof, sovereign DERP fallback) M idea
R-15 Multi-user dashboard accounts (household members, roles) L idea Single password is a stated alpha limitation (R-11). Launcher coupling — REVISED (controller v0.165.0): the "share the launcher outside the household" need is now met WITHOUT member accounts — the Indítópult megosztása capability-URL guest link (/s/<token>, information-only, no account) shipped in v0.165.0. What remains for this arc is member-specific: per-member tile visibility (each member sees only their apps) and the launcher-as-member-landing-page — both live inside this SSO/members arc; the guest-link ruling explicitly SUPERSEDES the earlier "members are how you share the launcher" framing
R-72 Curate brand_color for the top catalog apps XS idea Parked follow-up to the v0.163.0 launcher. .felhom.yml brand_color (#rgb/#rrggbb) overrides the deterministic slug-hash tile color; no catalog app sets it yet. Pick brand-accurate colors for the most-installed apps so their launcher tiles match their real brand. Catalog-only change (app-catalog-felhom.eu), brand_color is already omitempty and consumed by the controller
R-74 Island control plane on a CLUSTER (Peti's 2 nodes) — bring R-50's island bridge to a multi-node PVE cluster. M idea (Phase C of R-50, parked) R-50 shipped the island for the ONE-host fleet (demo-hp, demo-felhom). A cluster needs bridge parity on every node: either per-node identical /etc/network/interfaces vmbr9 stanzas (simplest, drift-prone) or — preferred at ≥2 nodes — a Proxmox SDN zone/vnet defined cluster-wide (one definition, auto-applied per node). The guest island IP is per-guest + node-independent; the agent-follows-guest rule holds (each node's agent binds its own vmbr9 169.254.253.1). Migration order per the spike: drill-proven → demo (done) → Peti (this row). Its own supervised runbook, coordinated with Peti (a live customer). Completes the capability-map "site/network change" row for clustered installs. Source: audits/SPIKE-island-bridge-2026-07-25.md (cluster-parity finding) + RUNBOOK-island-migration.md (single-host procedure to generalise)
R-87 The restic (app-data offsite) tier is NEVER restore-tested M SHIPPED 2026-08-31 — controller v0.231.0, RE-SCOPED by its own spike. Not built as written: the spike measured that an unattended scratch restore would have caught ONE of five drill-found restore defects, so it proves the snapshot CONTAINS a recoverable unit rather than testing the restore code. See audits/SPIKE-restic-restore-test-2026-08-31.md. R-85 covers whole-guest vzdump tiers only (local, felhom-pbs); the agent has no restic surface at all. restic is the CONTROLLER's app-data offsite backup to the Hetzner Storage Box, a separate mechanism — so the tier that is arguably most important to a customer is the one nothing verifies. It is the only tier that survives losing the box and carries their actual app data: the whole-guest snapshot deliberately excludes the bind-mounted data drives (/mnt/felhom-drives). Restore code exists and has been exercised BY HAND (the immich destroy-and-recover drill, PROVEN-LIVE), but nothing tests it unattended — exactly the state PBS was in before R-85: it works when someone tries it, and nobody would know if it stopped. Needs its own design: a restic restore-test is controller-side, has no scratch-guest analogue, and would verify into a scratch dir rather than a booted guest, so R-85's machinery does not transfer.

Each attempt runs the full quiesce cycle, so every customer app stack is STOPPED and RESTARTED for a backup that cannot succeed. Measured on demo-felhom: 07:07:58 quiescing 4 stack(s): [bookstack calibre-web docmost immich] → 07:08:17 unquiescing (backup failed) → 07:08:45 failed — ~19 s of app downtime per cycle (~50 s per full cycle), every 5 minutes.

The amplifier, and the part worth designing against: the agent answers Due: true, Reason: "no successful backup recorded yet", AgeSecs: nil, and that nil age does double duty. scheduledRunAllowed (quiesce.go:466-480) returns true whenever lastAgeSecs == nil — "no recorded backup yet — never withhold the first one" — so the same nil that makes every poll due also bypasses the time-of-day gate [W+2h, W+6h). On the live box the gate was [04:30, 08:30) and the cycles ran at 09:02–09:12 Budapest, i.e. outside the backup window entirely. So the fault stops customer apps every 5 minutes at any hour, including business hours — the one protection specifically built to prevent that is switched off by the same missing value. A safety valve written for a genuine first-ever backup is being triggered by an unreachable storage read, which is not the same thing at all.

Self-resolves the moment the target answers (the storage read succeeds, sees the archive, tier stops being due) — which is why it can hide indefinitely: it needs an offsite outage to appear at all. PHASE-0 ROOT CAUSE, established at source 2026-07-27 — it is AGENT-side, case (a). The storage read errored (could not read the backup storage for the due-check … err=… at 09:02:57/09:07:58/09:12:57 CEST), so this was never an empty-success. The failure is a type boundary: newestArchiveOn (localapi/server.go:1095-1111) documents "Errors and unsupported services degrade to unknown, never to 'no backup'" — but its (time.Time, bool) signature cannot represent unknown, so an error and a genuinely-empty storage both collapse to (zero, false), and handleBackupDue (server.go:934-941) then emits a POSITIVE claim: Due: true, Reason: "no successful backup recorded yet", AgeSecs: nil. The fail-safe that does exist — targetStoragePresent's "a storage-view error must never be read as 'not there'" (server.go:1131-1151) — answers a different question (does the storage exist) and behaved correctly. Decisive for scoping: the errored path and the genuine-never path are BYTE-IDENTICAL on the wire — same Due, same Reason string, same nil AgeSecs — so the controller has nothing to discriminate on and Part 2 CANNOT be done controller-side. Two further P0 findings: the agent restarted 4× on 2026-07-27 (07:36:39, 07:54:06, 08:50:16, 11:31:52 CEST) — all deliberate (NRestarts=0, Restart=on-failure, Result=success), zero self-update — so the trigger is armed by ordinary operator/config work far more often than "only when ep0 is down"; and the loop alerted NOBODY — zero backup_failed events despite the hub allowlist carrying that type, because internal/quiesce does not import internal/notify at all. Its only trace was 07:13:27 info app_start_failed "Telepített alkalmazás nem fut: BookStack" — a customer-tier, Hungarian, info-severity SYMPTOM of the third cycle catching BookStack mid-restart. The whole-guest backup tier R-82 built has no failure signal to the hub → its own item. Shape: distinguish storage unreachable from storage readable and empty. Unreachable is UNKNOWN — defer the due-verdict rather than resolving it either way, exactly as R-81 made the hub do with a missing report. Only a target that is reachable AND has no archive is genuinely due. Fix the window bypass in the same slice: AgeSecs == nil must stop meaning "run now regardless of the hour". Either the agent distinguishes never backed up from cannot tell in what it reports, or scheduledRunAllowed gates on the former only — otherwise any future nil-age path re-opens the same hole. Note this does NOT weaken R-84's fail-safe intent: a tier whose storage is merely slow or briefly unreadable should still err toward backing up — it is specifically the cold-store + unreachable pair that must defer, because there the fallback has no information at all, only an empty default that looks like a fact. | | R-95 | The restic offsite tier's credential CAN DELETE — R-89's "parallel question", now ANSWERED | M | idea — established read-only 2026-07-27 | The exposure closed on the weekly PBS tier is fully open on the daily restic tier, which holds the customer's actual documents and photos and is the only tier that survives losing the box. Established without mutating anything: (1) Identity — a per-customer subaccount on storage-box-pool-1 (box 611714, bx11, u629488): u629488-sub1 home felhom-demo-felhom, sub2 peti-felhom, sub3 demo-hp, each labelled felhom-customer. Auth is an SSH key stored ON THE BOX (…/felhom-controller-data/_data/data/offbox/ssh_key, 0600, beside repo_password + a pinned known_hosts) — customer-side, not hub-side, so a compromised guest holds it. (2) Read-write: YES — the API reports readonly=False on all three subaccounts, and it is not merely latent: the controller runs restic forget --group-by host,tags --keep-daily 7 --keep-weekly … --prune from the box (backup/offbox.go:984, also :1070). Delete rights are exercised on every run. (3) Append-only: NO, and not expressible — the repo is built as sftp: (offbox.go:482); restic's append-only mode requires the REST server backend, which plain SFTP cannot provide. (4) A zero-code mitigation exists and is unused: the box type carries snapshot_limit=10 and the API reports snapshot_plan=null with 0 snapshots and size_snapshots=0. Hetzner Storage Box snapshots are taken server-side, outside the SFTP namespace — an SFTP subaccount cannot delete them — so they are a genuine immutability layer at no extra cost and with no code change. Rule once for both tiers, per R-89. Options, cheapest first: enable a snapshot plan (operator click, immediate); split backup-write from prune so pruning runs somewhere the box cannot reach; or move the repo to restic's REST server with --append-only. Flips the capability-map row for offsite immutability | | R-162 | docker diff is the gate's only witness and its failure mode is quiet | XS | WATCHING — 2026-08-02 | A limitation, not a defect. The gate's power is docker diff excluding mounted paths; on a driver where it is unsupported or lies, the gate degrades to mount-occupancy + writability and would not say so. It fails closed (the canary self-test stops reporting BROKEN and the gate then refuses to report), but the message blames the prober rather than the driver. Revisit only if a non-overlay driver ships | | R-164 | C2's chain — the DB volume tar cannot be dropped until a sound dump predicate exists | S | BLOCKED — on the predicate (2026-08-02) | The unit holds a volume tar and a SQL dump and the restore uses both: the dump is authoritative and replayed after the tar so it WINS (F17), with only the DB service up (R-47) — restore_unit.go:262-266. Dropping the DB tar would halve DB-app units and close R-127(b)'s initdb-skip trap. The obvious gate is dead, measured: ValidateDump's empty-accounts warning was correct (the DB truly had 0 rows; seeding one stopped the warning and put the row in the dump) — but a fresh appliance legitimately has zero accounts, so gating on it blocks every new customer's first backup. Order: sound predicate (dump vs live per-table counts) → warn→gate → tar-drop. Pairs with R-127 | | R-91 | The old 13 GB datastore copy is still on ep0's root disk | XS | WATCHING — gated on demo-felhom's first post-migration PBS backup | The datastore moved to a Hetzner Cloud Volume on 2026-07-27 (/dev/sdb, 100 GiB, attached 06:29:40 UTC, now /mnt/pbs-datastore, 13 G used of 98 G). The pre-migration copy survives at /srv/pbs-felhom, 13 G, on / (38 G total, 16 G used, 21 G free). Do not delete yet: demo-hp has landed two post-migration snapshots (07-27 08:25:47Z, 09:37:29Z) but demo-felhom's newest is 2026-07-26T12:21:48Z — before the migration, so the new volume has not yet proven a write for that namespace. Delete once it has. Doc drift to fix in the same commit: CONTEXT.md:1018 still records the datastore at /srv/pbs-felhom | | R-92 | The hub's PBS-DR gauge is 0.1 GB-granular, so small deltas are unverifiable | XS | idea — 2026-07-27 | The PBS-DR box card rounds to 0.1 GB, which is coarser than the changes an operator wants to confirm after a prune or a GC — a successful prune of a small namespace moves the number by less than one displayed digit, so the UI cannot distinguish "it worked" from "nothing happened". Cosmetic today; it becomes load-bearing the moment retention (R-89) is customer-visible and someone needs to see that a policy change took effect | | R-93 | drill-r50 is both a blocked customer and the only drift fixture | XS | idea — 2026-07-27 | The drill customer is blocked in the hub (so it stops alarming) yet it is also the only record exercising the endpoint-drift path R-77 added. Blocking hides it from GetActiveCustomerIDs, so the fixture it provides is silently inert — a monitor with no live subject reads exactly like a monitor that passes. Decide: retire it and build a synthetic fixture, or unblock it and silence per-customer instead (the operator has a per-alert silencing feature planned). Related to the R-50 drill VM, now shut down | | R-29 | The design-v2 green gates are not enforced anywhere — one has been RED for 16 releases. controller/scripts/docker_run_volume_path_gate.py has failed continuously since 2026-07-14 (v0.129.0) and nobody noticed until R-7b's close-out ran it by hand at v0.145.0. Two separable parts. (a) The finding itself is benign and the fix is 3 lines. The flagged call is internal/appexport/estimate.go:179 docker run --rm -v <volumeName>:/vol:ro alpine du — a NAMED-VOLUME mount, i.e. daemon-side with no host path, which is the safe shape and byte-for-byte the same pattern as three entries already on the gate's ALLOWLIST (export.go volName+":/vol", backup.go volName+":/vol:ro", restore.go volName+":/vol"). It is NOT the v0.124.0 path-strand class the gate exists to catch — the author of the v0.129.0 F-A fix explicitly avoided that class (see the function's own comment) and simply never added the allowlist entry. So the fix is an ALLOWLIST addition WITH ITS WHY, not a docker-cp rewrite; anyone who 'fixes' this by rewriting the call has misread the gate. (b) The systemic half is the real item: the gates run only when a human remembers to run them, so a gate can sit red across 16 releases while every REPORT says 'green'. This is the SECOND instance of the class — cf. the v0.123.0 note 'Windows green gate silently red (read-only fsync)'. Decide where they run (pre-push hook, build.sh step, or a CI job) and make a red gate block the train the way the Go green gate does. | S (a) / M (b) | idea | Origin: R-7b close-out, felhom-controller REPORT §4(f) — CC correctly left it alone as out-of-scope and pre-existing, and verified by stashing that it fails identically on the unmodified tree. Flips no capability-map row (engineering hygiene, no customer-visible behaviour). Affected gates to audit for the same rot: controller template_id_gate / emoji_gate / native_confirm_gate / offbox_rename_gate / mojibake_gate / app_row_dedup_gate / docker_run_volume_path_gate, hub hub_confirm_gate, manifests manifest_bearer_gate, website site_gates. Do not bundle (a) into an unrelated feature commit — it is a one-line behavioural claim about a mount's safety and deserves its own reviewed diff. 2026-07-18 rehearsal note: the run's finding list independently re-raised "assign the pre-existing docker_run_volume_path_gate failure its ID so red stops normalizing" — that is this item; no second ID was minted. 2026-07-29 — audit list extended, and a THIRD independent re-raise absorbed under the same rule (again no new ID): add scripts/hostinstall_gates.py, which postdates this item (it comes from drill F-1, 2026-07-12) and is therefore not a design-v2 gate — but it is the identical failure shape and is tracked as R-94 leg (b). It is RED as of 2026-07-29: hub Setup-tab hostInstallVersion=1.19.0 != SCRIPT_VERSION=1.22.0, exit 1, with its nine other assertions green. scripts/hub_confirm_gate.py, already on the list above, was verified orphan on the same date. Both confirmed by repo-wide grep across all file types plus sibling repos, ~/.claude settings/skills/hooks, .git/hooks (no non-sample hooks exist), a Makefile/justfile/Taskfile find (only hub/Makefile, zero gate occurrences) and a CI-directory find (felhom.eu has no CI configuration at all) — all 19 hits are docstrings, code comments or prose; zero are invocations. Only site_gates.py is mandated (CLAUDE.md:153); manifest_bearer_gate.py is named in runbooks/secrets.md:76. Now also filed in OPEN-ITEMS.md — this item predates the 2026-07-27 register rebuild and was never carried across, so an open item about work not getting done was itself missing from the page that decides what gets done. 2026-07-30 — THE FIRST ENTRY ON THE OTHER SIDE OF THE LEDGER, recorded so the contrast is not lost: the R-120 golden-staleness gate (hub v0.82.0, hub/internal/web/configs.go handleSetArtifacts) IS enforced. It is not a script in scripts/ that someone must remember; it sits inside the only UI path that writes SetArtifactManifest, so it runs on every vouch whether or not anyone chose to run it, and it refuses (operator ruling, 2026-07-30) rather than warning — because this row's whole finding is that a non-blocking check reads as coverage it is not providing. It compares the submitted golden against the newest controller any box has reported (store.NewestReportedControllerVersion) and is pinned by four tests driven through the production handler over httptest, not an injected seam, plus a red-proof: deleting the block makes the stale golden vouchable again. Note the near-miss worth keeping: the first draft read guests.controller_version, a column that exists in the schema and that nothing writes — it would have been an inert gate, i.e. this row's exact failure shape, caught by grepping for a writer before trusting the column. The three orphans above are unchanged and still orphaned — this entry proves the pattern is available, not that the backlog moved UPDATE 2026-08-02 — leg (a) CLOSED (felhom-controller c432f70, its own reviewed diff); leg (b) HALF-SHIPPED: every repo now has ONE entry point wired to .githooks/pre-push --fast, each mandated in its CLAUDE.md. The census that drove it: 13 gates, and every gate a CLAUDE.md names was green while two of the four unnamed ones were red. Stays open for the automatic half → R-168 CLOSED 2026-08-02 on the demonstrated alarm, not on a green run: both halves are live (local hook refuses; CI notices a bypass and emails). Remaining is a working-style choice → R-169 | | R-169 | CI can only report, because there is no gate in the road | S | idea — minted 2026-08-02, WAITING-ON-OPERATOR | Making CI blocking needs branch protection on main plus a pull-request workflow instead of direct-to-main pushes — both change how the operator works, so neither was done. Current arrangement is two nets: the pre-push hook refuses locally, R-168's runner notices a --no-verify bypass and emails. Decide only if that window ever costs something. Detail: OPEN-ITEMS.md R-169 |

| R-41 | [SLICE 1 SHIPPED 2026-07-21] The catalog has no standing "does every template still deploy?" check. Campaign 7 was the first thing that ever tried to deploy all 53 apps, and found 5 that had NEVER been deployable: papra (missing required AUTH_SECRET), zipline (v4 renamed CORE_DATABASE_URL → DATABASE_URL), wishlist (Docker Hub image gone; upstream moved to ghcr.io), homebox (upstream dropped the v tag prefix + new required env), glance (needs a seeded glance.yml the template never provides — PROVEN pre-existing: the pre-campaign v0.7.4 pin fails identically). Plus 7 broken healthchecks and 2 apps whose images no longer resolve at all (plant-it, wanderer). | M | idea | Origin: CAMPAIGN 7 (§7 F5/F6). The repo already has the right pattern in scripts/check-image-pins.py — a mechanical gate run on every change. Cheap first slice: a resolvability gate (docker manifest inspect every pin) would alone have caught plant-it, wanderer, wishlist and homebox, and needs no box. Full slice: a periodic deploy-all sweep on the demo box reusing the campaign's engine. Silent rot is the real risk — an app can die upstream and nobody learns until a customer clicks Telepítés | SLICE 1 SHIPPED 2026-07-21 — app-catalog-felhom.eu/scripts/check-image-resolvable.py (+ 14 fixture tests, no network, resolver injected). Resolves every unique pin with docker manifest inspect, ONE image at a time; exit 0 / 1 (the registry says GONE) / 2 (inconclusive). Two traps encoded, both hit live while building it: (a) docker manifest inspect prints toomanyrequests: … and still exits 0 — the same exits-0-on-failure shape as the ISO tooling's validate-answer, so stderr is inspected even on rc=0; (b) the inverse and more dangerous one — the first full sweep called 24 of 65 pins dead, including postgres:16-alpine and redis:7-alpine, purely because Docker Hub throttled it partway through. Ambiguity therefore resolves to INCONCLUSIVE and never to an accusation: a gate that cries wolf gets ignored, and then it protects nothing. The full sweep is still OWED — DooPlex is not logged in to Docker Hub, so the 52-app table needs one re-run after docker login. Wired into CLAUDE.md + REUSE.md as a start-of-campaign / pre-publish-train step. It immediately paid for itself: it is what turned plant-it and wanderer from 'images do not resolve' into two DIFFERENT diagnoses (see the 2026-07-21 catalog entry). Full slice — the periodic deploy-all sweep on the demo box — remains open | | R-45 | [P2] Unified async-job feedback. Every long operation invents its own progress surface, or none. Tonight produced three more one-off cards (v0.147.x: samba bring-up, offsite progress, restore result) on top of two existing patterns (deploy 3-step panel; storage-init/netstorage status poll). They agree on nothing: some use {ok,data} envelopes and some raw JSON, some poll 1 s / 1.5 s / 3 s, some are in-memory-only and lie after a restart, and each re-implements single-flight + snapshot + phase→Hungarian mapping. | M | idea | Origin: 2026-07-19 feedback slice 1 (controller v0.147.0). The cases to generalise from are all in-tree: web/storage_init_job.go (the best shape — acquire/release/set/snapshot), web/netstorage_job.go, web/samba_ensure_job.go, backup/opstatus.go, backup/offbox_progress.go. Shape: one job registry + one poll endpoint + one client-side renderer, phases declared per job. Two lessons tonight that any framework must encode: (1) a terminal state must be probed, not inferred — compose up -d exits 0 on a crash-loop; (2) a progress source that reports nothing is normal, not broken — restic reports 0 bytes for a whole incremental run, and a bar that sits at 0% is worse than no bar. Also fixes the restart hole: in-memory job state currently vanishes and the card silently disagrees with reality 2026-07-20 — the first bill for NOT having this arrived, and it was customer-facing. The samba card's poll (web/samba_ensure_job.go + sharing.html) mixed a job EDGE and a service LEVEL on one JSON field, and /sharing reload-looped at ~1.2 s for every customer with sharing enabled until controller v0.151.0 (audits/DIAG-sharing-2026-07-20.md, S-1/S-4). v0.151.0 fixed THAT card's contract only — the framework is still this item. Third lesson for it to encode, beside the two already listed: a phase a client answers with a one-shot action must be an EDGE the registry SERVES ONCE, and must never be synthesised from a level; if it can be re-read, it will be re-acted on. | | R-46 | [P2] Verification copies need a customer-visible browse surface and an expiry. v0.147.0 made them visible (listed with path/size/date, individually deletable) — but the customer still cannot LOOK INSIDE a verification restore to confirm the file they wanted is really there, which is the entire point of a verification restore, and nothing ever removes them. | S–M | idea | Origin: 2026-07-19 feedback slice 4a, registered as the explicit follow-up to it. Two gaps, deliberately designed together because they are the same object: (a) the invisible-result gap — a read-only browse of backups/offsite-restore/<app> (the FileBrowser infra stack already exists and already serves scoped roots, so this may be a mount rather than new code); (b) the disk-lifecycle gap — auto-expiry after N days with the count/size surfaced before it fires, so a drive is never quietly filled by verification restores nobody remembers taking. Pairs with R-43: a browse surface is also how a customer would discover that a DB-indexed app's files came back but the app still cannot see them | | R-48 | [P2-HIGH] Restore controls are separable only by layout — and the difference between them is whether the data comes back. The offsite restore row renders four buttons plus hint text into an overlapping, unreadable line, and the decisive second step („Teljes visszaállítás indítása") appears ONLY after „…előkészítése" was pressed, with no signposting that a second step exists or that the first one did nothing to live data. | M | idea | Evidence: audits/DIAG-immich-restore-round2-2026-07-19.md (finding 1) — this is not theoretical: it is the CAUSE of the round-2 incident. An operator who had read the code pressed the missing-only button instead of the full restore; the controller log shows /backup/offbox/reconstitute was never hit at all. The rule this establishes, worth stating once and applying beyond this page: two adjacent controls whose difference is "your data comes back" vs "your data cannot come back" must not be distinguishable only by layout. Direction (ruled in principle, spec rides v0.149): collapse to a single „Visszaállítás…" guided dialog — one intent, visible phases, the escrow-wizard precedent. Pairs with R-45 (the phases are exactly the async-feedback surface) and R-46 SHIPPED 2026-07-21 — controller v0.154.0 (3a9d744). Each app row on /backups/restore now carries ONE „Visszaállítás…" entry linking to a per-app wizard at GET /backups/restore/app?name=<app>: three intent CARDS each with a consequence sentence (ellenőrzés külön mappába / hiányzó fájlok visszahozása / teljes visszaállítás), a visible phase strip so the sequence is legible before the first click, danger styling on the destructive card, and the R-43 double-confirm carried over verbatim with its pair-honesty facts. deriveWizardStep is a PURE function of (op running, size-gate flash, scratch ready) — the step is never taken from the request, and a running op outranks a stale ?full_prep= so no commit button survives into a restore. While ANY op runs every mutation form is suppressed server-side rather than offered and then refused. No new mutation endpoint (one GET route; every card posts to the pre-existing /backup/offbox/* with unchanged field names and gates) and no R-45 graft — the wizard polls the two existing status surfaces as-is. Works with JavaScript disabled. Latent bug fixed on the way: offboxRedirectTo hardcoded "?" when appending its flash, which against the wizard's ?name=<app> target would have buried the flash inside the app name. 9 new tests + the Group-B red-proof (trivial always-INTENT impl → all 7 rows red). Live click-through + one non-destructive Ellenőrzés still PENDING (rides the operator's floor save). Evidence: felhom-controller/REPORT.md §3 (2026-07-21). | | R-65 | Buddy-box backup replication, cross-household — two Felhom boxes in different homes replicate backups to each other. | L | idea (post-alpha, spike-first, 2026-07-22) | The natural big sibling of R-64: two households each hosting the other's encrypted backup tier. Explicitly spike-first — the transport is NOT SMB (R-64's live-share protocol is wrong for backup replication across the internet: no auth story between households, no resumability, cleartext LAN assumptions); candidates to spike: restic rest-server / rclone / syncthing over the existing WG/tailnet plumbing, encryption keyed so the buddy can never read the payload. Sits on top of the offsite tier's FILL/OVERSUB thresholds thinking (R-5 aggregate). Flips: would add a "cross-household buddy replication" capability row (currently unlisted). Pairs with R-64 (same topology, different transport + guarantees) |

Pre-invite checklist — what stands between here and the first remote tester

Not roadmap items in their own right; the short list the 2026-07-18 rehearsal leaves behind. Everything here is remote-doable — the N100 is packed, and none of it needs hands on the box.

Action Owner Note
Rebuild the golden → 0.146.0 BAKED + PUBLISHED 2026-07-18; awaiting the operator's two saves Viktor (saves) Golden 0.146.0 baked on the drill VM and published to gitea — felhom-golden/0.146.0/golden.tar.zst, sha256 4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955, 612 696 588 B (584 MB archive). All pass markers green: Result=success/ExecMainStatus=0, 0 FATAL/exclusions, docker OK (overlay2), all three mounts included (rootfs + mp0 /var/lib/docker + mp1 /mnt/sys_drive), pre-delete 404, upload HTTP 201; controller 0.146.0 confirmed baked in. Integrity round-trip independent of the build host: anonymous `GET
Golden ≥ 0.147.x carries ALL FOUR infra images — (next bake) build-golden.sh v2.1.0 (2026-07-19) now derives the pre-pull list from the controller binary it is about to bake (--print-infra-images) instead of a hand-maintained copy that had already drifted: felhom-samba was never added to it, so every golden so far baked 3 of 4 — which is why enabling Megosztás on a fresh box pulled from the registry with zero feedback. No golden rebuild for this alone; it takes effect at the next bake. Until then a fresh box still pulls felhom-samba at enable time, which controller v0.147.0's progress card now at least explains
freemail.hu test-send Viktor The open half of R-4; the gmail half closed on 2026-07-18 under p=quarantine
C6 — customer performs a restore, unassisted Viktor as customer zero The one open script step in R-3 and still MISSING as capability evidence. Remote-doable on the reborn box — the dashboard is remote
R-11 rulings Viktor Contact channel, tester agreement, alert thresholds (the R-5 gauge thresholds are still pending a ruling)

Absorbed / superseded notes in this folder

  • FOLLOWUP-nas-automount-guest-reboot-reassert.md — shipped (agent v0.84/v0.85, CAMPAIGN-3); keep for history
  • FOLLOWUP-golden-default-controller-tag.md — verify against current golden flow; close or promote to an item
  • FIX-M18-NOTES.md, FIX-M19-NOTES.md, DIAGNOSIS-f9-storage-registration-gap-2026-06-14.md — historical diagnoses; superseded by shipped fixes