Files
felhom.eu/documentation/backlog/OPEN-ITEMS.md
T

459 KiB
Raw Blame History

OPEN-ITEMS — the single source of truth for open work

Rebuilt 2026-07-27 by read-only triage. ROADMAP.md keeps the full history and reasoning; this page keeps only what is open, and it is the file to read first. Root REPORT.md is per-session and overwritten — nothing durable may live only there; a session that must not clobber it writes a non-overwritten REPORT-<topic>.md sibling instead (CLAUDE.md:82-87), of which 14 now exist.

How a row is filed (2026-10-03 — the triage; operator may reverse the scale and the list)

One table shape: | ID | Category | Sev | What | State | Blocked on | Next action | Owner |. scripts/register_shape_gate.py refuses a row that does not split into these eight cells (a literal | belongs inside backticks), whose Category is not one of the eleven below, or whose Sev is not P1–P4. A new row goes into its category's section, in severity order.

Fix small, do not file (the size rule — operator brief 2026-10-05, the burn-down). The register grew because every session closed a few rows and filed a few small new ones. So: a finding that is cosmetic or small — fixable in the session in about 30 minutes, in a repo the session may change — is FIXED in that session, with a test where it changes behaviour, and NOT filed. It is recorded in that repo's CHANGELOG.md and in the session report under „fixed without a row". Only a finding that needs a decision, a design, a larger build, or a change the session may not make (a protected machine, a repo out of scope, a release budget already spent) becomes a row. This narrows — it does not repeal — „an enumerated gap becomes a row": a small gap leaves the session fixed, which is a record too.

Every session report states four numbers: rows before, rows after, rows opened, rows closed — counted the way register_shape_gate.py counts (| **R-n** | lines in this file).

Sev — one scale for this file and ROADMAP.md:

  • P1 — now. A household can lose or leak data, a box can stop or be taken over, or a promise we make is false — today. A P1 row says in one sentence what the harm is.
  • P2 — before the first paying customer. Must be done before anyone pays (business and legal work too).
  • P3 — during the first customers. A real defect or gap a customer may meet, with a workaround or a low chance.
  • P4 — later / nice to have. Ideas, polish, tooling comfort, post-alpha features.

Older tags stay readable in the row text: P2-HIGH/P2-MEDIUM → P2, P3-LOW → P3. The Sev cell is the rank. A row whose rank moved says so in one dated sentence at the end of its State.

Category — eleven, following architecture/00-capability-map.md's sections where they fit: Install & onboarding · Apps & catalog · App updates · Backup & restore · Storage & devices · Security & access (logins, gates, the tunnel, secrets) · Box system & updates (host, guest, OS, controller/agent self-update, golden) · Monitoring & notifications · Hub & operator · Business & legal · Process & tooling (gates, docs, the workflow).

State leads with one of: READY · OPEN · BLOCKED · WATCHING · WAITING-ON-OPERATOR · NARROWED · DEFERRED · VERIFY (= may already be finished; check against live source before closing). Every row has an Owner. A row whose state LEADS with a finished word (CLOSED, SHIPPED, FIXED, DECIDED, …) does not belong here — it moves to CLOSED-ITEMS.md in the same commit that finishes it; scripts/closed_register_gate.py RULE 3 refuses the push otherwise.

2026-10-03 — the triage. 125 finished rows moved to CLOSED-ITEMS.md (and 20 old rows that never had an id); the campaign write-ups, the 2026-08-04 rulings and the old ranking paragraphs that sat between the tables moved, word for word, to documentation/archive/OPEN-ITEMS-narratives-2026-10-03.md. The full previous text of this file: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md. What to work on next: documentation/audits/backlog-triage-2026-10-03/RECOMMENDATION.md.

DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01)

Written in the voice this file uses for a settled decision: what was decided, and the condition that reopens it. It exists because nobody ever said the arc was finished, and an arc nobody closed gets picked up again in a month by someone who reads the blanks in 07 §8 as unfinished work.

CLOSED FOR BETA at controller v0.232.0 / hub v0.111.1 (2026-09-01).

What is finished, and proven live: everything a customer does for themselves — losing files, losing an app's data, losing a whole app, losing a drive. The restore states what it returned, refuses without room, cannot be fed a part-copy, and puts the customer's own data back if it fails. The second drive's copy is a route. A poorer copy cannot delete a richer one. The off-site store is verified weekly at full depth and proved nightly to still contain something. Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of 07 §8 — every one PROVEN.

What is deliberately deferred until after beta — [BETA-DEFERRED], and these are the row numbers, not a description:

07 §8 row the failure status today
4 primary drive dies — the drive-loss journey (the route is proven; no disk has ever died under it) PARTIAL
8 host dies, drives intact — a host rebuilt as itself IMPLEMENTED, never executed
9 whole box lost (fire/theft) IMPLEMENTED / UNPROVEN
10 ransomware / malicious deletion — a ransomware-shaped recovery PARTIAL
11 (and 11b) hub lost — a hub restore UNPROVEN; it has never been performed
12 off-site provider lost (Hetzner) [FACT] only

These are real, they are recorded, and none of them is a beta blocker. Every one is invocable by the operator, not the customer; every one needs hardware, a provider, or a destructive rehearsal that beta does not.

A NUMBER IN THE BRIEF FOR THIS SECTION WAS WRONG AND IS CORRECTED HERE. It said "six rows of §8 still have no measured time". Six rows are DEFERRED; ELEVEN rows carry a blank RTO — counted, not estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not deferred work, and collapsing them into one number is how a blank stops meaning anything:

  • row 5 — PROVEN by construction; the secondary is a derived copy, so there is no recovery to time.
  • row 13 — NONE for host-loss by design; R exists in zero system copies, so there is no route.
  • row 14 — PROVEN; the break-glass route works and has simply never been stopwatched.
  • row 15 — an open DEFECT (R-104, the stale-lock path), not a deferred recovery. It is NOT inside this stopping line and must not be read as parked by it.
  • row 11b — a consequences note attached to row 11, not a recovery row of its own.

What is NOT deferred and stays open: R-95 and R-433, both BLOCKED-ON-PROVIDER behind documentation/runbooks/provider-questions-2026-09-01.md; R-435, documentation only; and R-104, row 15's defect. The alarm text (R-434) was fixed the same day and is closed.

THIS REOPENS IF: a customer-facing recovery path is found broken; or Hetzner's answers change what the snapshots are worth (either answer in Part 3's file can do it — "full restore only" makes row 10 urgent, "append-only enforced" makes R-95's root cause cheap to remove); or a real customer's data is at stake in one of the deferred rows.

THE MARKER. Every deferred row carries the literal string [BETA-DEFERRED] in 07 §8, so grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md returns the set as a group. It returns EIGHT lines, not seven — the seven tagged rows plus the one line in §8's own header that defines the marker. That is stated rather than hidden, because a count that does not match what the reader sees is how an instrument stops being believed (R-421). It is a marker, not a status — the status cells are unchanged, because nothing was proven on the day this line was drawn and a stopping line that moves a status is a stopping line that lies.

Install & onboarding — 15 rows (P3 9, P4 6)

ID Category Sev What State Blocked on Next action Owner
R-130 Install & onboarding P3 A "hard min" that only warns. A fresh box's local-lvm was ~75 GiB against HARD_MIN_LVM_GIB=120 (scripts/felhom-host-install.sh); the installer logged [WARN] local-lvm free ~75 GiB < hard min 120 GiB and went on to a fully successful install READY (S) — Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 CC
R-179 Install & onboarding P3 --uninstall leaves the NAS network-storage systemd units behind, with the automount in failed state and the parent bind still mounted. The teardown's residue-diff provenance (day0-install.md Part E: "a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers") is from v1.9.1, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps /etc/systemd/system/mnt-felhom\x2ddrives-<share>.mount and .automount after a full uninstall READY (S) — NEW 2026-08-03 — Observed on demo-hp 2026-08-03 after --uninstall --vmid 9201: mnt-felhom\x2ddrives-Felhom\x2dShare.automount loaded failed failed, its .mount loaded inactive dead, and mnt-felhom\x2ddrives.mount still active mounted — the uninstall's own output had warned /mnt/felhom-drives/Felhom-Share is busy — NOT forcing and /mnt/felhom-drives root bind left mounted, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, daemon-reload, unmount the autofs then the parent. NEGATIVE CONTROL, same day: demo-felhom's uninstall left nothing (`ls /etc/systemd/system grep -i felhom→ only the unrelatedfelhom-bootstrap.service; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** felhom-bootstrap.serviceis NOT residue — it is the ISO first-boot unit,disabled+inactive`, exactly-once and already fired
R-180 Install & onboarding P3 --archive-storage is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated. felhom-host-install.sh validates the archive storage EXISTS (pvesm status --storage, :1583) and that the golden volid RESOLVES on it (:1661), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default local local-lvm felhom-pbs (--acl-storages, which runbooks/day0-install.md tells the operator not to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step READY (S) — NEW 2026-08-03 — Hit live on demo-hp 2026-08-03 during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on felhom-backup (the enrolled NVMe, where the box's vzdumps live) and --archive-storage felhom-backup passed. Pre-flight passed; steps 1–7 ran; step 8 returned reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace). The cost is the ORDER, not the error — by the time it fires, step 2 has minted the PVE token, step 4b has rotated root@pam and vaulted it (so the old console password is already dead), and step 5 has installed the agent. Recovery was --resume after moving the golden to local, which worked cleanly. This is statically checkable in pre-flight: ARCHIVE_STORAGE ∈ PVE_STORAGES is a one-line assertion over two variables both known at :1583. Same class as R-29 — the checkable thing that nothing checks CC
R-250 Install & onboarding P3 A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time. Found 2026-08-07 creating the fifth walk's venue. POST /configs/new with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — fail-closed by design (offsite.go:111-121, "don't serve a descriptor the controller can't verify"), with defaultScanBackoff = 2+4+8+16+30 ≈ 60 s, sized by its own comment to "the observed DNS propagation lag". Measured, both halves: the first create exhausted the ladder — five no such host, then dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable (the name had just begun resolving, AAAA-first, into a pod with no IPv6 route) — and the hub logged [ERROR] offsite provision for walk5. An identical second POST succeeded ~70 s later on its own final rung (shared already provisioned … (subaccount 285351)). Total settle ≈ 100 s against a 60 s budget. The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, SSH-2.0-OpenSSH_9.6p1. Two distinct things are wrong and should not be merged: (1) the budget is sized against DNS existence, but what actually bit is the AAAA-before-A window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) the operator is told nothing actionable — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — one_time_secrets 1, one sub-account, no double-provision) but that safety is invisible to the person deciding whether pressing again will double-charge them. Severity LOW-MEDIUM: self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. READY — owner Viktor — — operator
R-283 Install & onboarding P3 After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page. customer_claims for demo-hp still read claimed_at 2026-07-21 16:29:25, generation 2, issued_at 2026-08-03 while the freshly provisioned guest — whose settings.json is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ R-282), and any previously issued code fails with „Hibás vagy lejárt kód" — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of R-214/R-235 (an already-paired box still told to pair itself) READY (S) — NEW 2026-08-09 — Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page CC
R-306 Install & onboarding P3 --preflight-only says "no state written" and writes state — with an answer that can be wrong. _state_put short-circuits on DRY_RUN only (felhom-host-install.sh:418), so a preflight-only run creates /var/lib/felhom-install/state.json. Observed live: after a run whose banner read PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed, the file existed containing {"completed": [], "dnsmasq_preexisting": "yes"}. Both the banner and the flag's own comment at line 226 assert the opposite. The harm is not the file, it is the value: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins READY (S) — NEW 2026-08-12, RANK 3 R-300, R-305 Either make _state_put a no-op under PREFLIGHT_ONLY (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it CC
R-317 Install & onboarding P3 The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns. EnsureDnsmasq (felhom-agent/internal/lanresolver/lanresolver.go:105) does os.Stat("/usr/sbin/dnsmasq") and skips the apt install when it exists — but that path is shipped by dnsmasq-base, while the systemd unit comes from dnsmasq (confirmed on the box: dpkg -S /usr/sbin/dnsmasq → dnsmasq-base; dpkg -S /usr/lib/systemd/system/dnsmasq.service → dnsmasq). So on any host carrying dnsmasq-base without dnsmasq, the agent skips the install and then runs systemctl enable --now dnsmasq against a unit that is not there: the resolver never comes up and the failure is a retried WARN in the journal rather than anything a customer or the install sees. Pre-existing, NOT introduced by R-316 — but R-316 makes the shape reachable, because a host whose dnsmasq-base pre-dated Felhom now keeps it while dnsmasq is removed. R-316's uninstall says so explicitly instead of leaving it to be found from a silent resolver. Ranked 2 (costs time), not 1: the box installs fine, only LAN name resolution is missing READY (S) — NEW 2026-08-13 R-316 Probe what is actually needed — the unit or the dnsmasq package — rather than a path a sibling package owns. One-line change in the agent; deliberately NOT made here to keep this session to one repo CC
R-516 Install & onboarding P3 [P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night. MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item „Debug"; (2) the dashboard CPU tile „Load: 0.29 / 0.39 / 0.37"; (3) the launcher tile „Filebrowser" opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) Uptime Kuma 2.4 opens on „Which database would you like to use?" (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), wiki.DOMAIN (R-498). Fix shape: rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed db-config.json for SQLite in the template (catalog). Added by F4 (20:04:54Z): (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. Added by F7 (20:50–21:00Z, system disk at 95 %): (10) a banner on every page in English, „SSD disk usage high: 90%"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (90%)" while df reports 95 % (reserved blocks ignored), and „(/)" labels the data volume /mnt/sys_drive; the deploy page says nothing about free disk. Added by the i18n spike (2026-09-17, controller v0.247.0): (12) six formal („ön") forms in the converted dashboard copy — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by controller/scripts/i18n_missing_gate.py (HU_FORMAL_CEILING = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (audits/I18N-INVENTORY-2026-09-17.md) is the list this row closes against in localisation slice 6 (R-561). Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0): the converted copy now counts 16 formal forms (HU_FORMAL_CEILING 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least 22 keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. NARROWED 2026-09-20 by localisation slice 6's walk, item by item (audits/i18n-slice6-2026-09-20/R-516-item-by-item.md): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. What this row is now waiting for is a HUNGARIAN walk on a box with a second drive, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk — — CC
R-554 Install & onboarding P3 [P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable. OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). 02-controller-module-map.md L56 calls internal/setup/ obsolete; cmd/controller/main.go L322 still enters it when setup.NeedsSetup(cfg) — customer.id empty after bootstrap ingestion, or a .needs-setup marker (internal/setup/setup.go L17-25). Ingestion leaves customer.id empty on a missing/invalid bootstrap.json, a failed hub pull, or a failed merge/write/reload (internal/bootstrap/bootstrap.go L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). Fix shape: decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete internal/setup/ and runSetupMode; red-proof that a failed ingestion renders the waiting page, not a 404. Check first whether any drill/golden path still relies on .needs-setup. READY - rank P3-LOW; owner: CC — — CC
R-310 Install & onboarding P4 Two small edges on the installer, neither costing more than a moment. (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences ("…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."). (2) --uninstall reads its typed vmid confirmation from /dev/tty and --force deliberately does not bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as line 891: /dev/tty: No such device or address and an rc=1 that looks like a failure rather than a refusal to proceed unattended READY (S) — NEW 2026-08-12, RANK 4 R-297 Drop the duplicated sentence; add one runbook line naming the pty requirement CC
R-494 Install & onboarding P4 NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (architecture/01-topology-and-trust.md). Original finding, kept: [P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead. MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention I1): the claim mail points at https://felhom.drill0242.felhom.eu; that name has no A and no AAAA record (dig @1.1.1.1, control felhom.enkisfelhom.hu resolves); the hub has no tunnel- or DNS-creation code (hub/internal/cloudflare/ holds only geo-rule removal; cf_tunnel_token is a pasted, optional form field, configs.go:1478) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered google.com but not the dashboard name at 13:27:39Z; the agent applied the record at 13:27:44Z (lanresolver: applied split-horizon record … ip=192.168.0.158, 3 m 46 s after the controller started), so the box CAN answer the name — but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (curl --resolve …:443:192.168.0.158). A volunteer could not have done that. What it needs: an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook. — — CC
R-503 Install & onboarding P4 [P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian". Offered to the operator 2026-09-14 and not chosen: the 2026-07-31 ruling (a person chooses the disk) stands. Recorded so a reversal starts from measurements, not from the offer. What must be measured first: (1) whether the Proxmox auto-installer's HTTP answer mode can serve a per-machine answer from posted system info without network being a precondition a volunteer can miss; (2) whether USB transport is reliably visible in sysfs (/sys/block/*/device path, removable) where udev properties were measured blind (SPIKE-universal-iso-1 §3.2); (3) whether any refusal can be shown in Hungarian without modifying the Proxmox installer squashfs. Reverses two rulings if built — needs an operator word. WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (ruling), CC (spike) Re-ranked 2026-10-03: P3→P4: an unchosen idea that would reverse two rulings. — — operator
R-504 Install & onboarding P4 [P3-LOW] iso.felhom.eu cannot show an index page on its own — its root returns 404, and the download page lives on the website instead. MEASURED 2026-09-14: https://iso.felhom.eu/ and /index.html → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded index.html at / was not measured (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at felhom.eu/letoltes (published with the ISO, after the operator's yes). Remaining: a redirect from iso.felhom.eu/ to that page needs a Cloudflare rule the session has no credential for. WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule) Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic. — — operator
R-725 Install & onboarding P4 [P3-LOW] Small copy slips on the first-hour path. MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). FIXED 2026-09-30: the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). NARROWED — remaining: the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). NARROWED — three small copy items; owner: CC Re-ranked 2026-10-03: P3→P4: three small copy slips left; nothing blocks the household. — — CC
R-881 Install & onboarding P4 The installer's uninstall does not remove /usr/local/sbin/felhom-priv-apply (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (scripts/felhom-host-install.sh ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. READY — Add the file to the uninstall list and fix the comment at the next installer tag CC

Apps & catalog — 33 rows (P3 13, P4 20)

ID Category Sev What State Blocked on Next action Owner
R-498 Apps & catalog P3 [P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal wiki.DOMAIN — the placeholder is never filled in. MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a wiki.DOMAIN címet a böngészőben", PrivateBin's „Nyisd meg a paste.DOMAIN címet". grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml → 52 of 53 templates carry it in first_steps; the controller renders the string as text (internal/stacks/metadata.go). A stranger reading their first instruction meets a word that is not an address. Fix shape: substitute the stack's real SUBDOMAIN.DOMAIN at render time (one place in the controller), with a render test per template that fails on a literal DOMAIN. READY — rank P3-LOW; owner: CC — — CC
R-562 Apps & catalog P3 [P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves. FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (2006. 01. 02. 15:04 Hungarian vs 2006-01-02 15:04 ISO); 10 layout literals in internal/web Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (%.1f GB, 4 helpers) where Hungarian uses a comma; timeAgo/nextRunLabel/pruneLabel produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). Fix shape: one date and one size formatter per language in internal/i18n, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). READY - rank P3-LOW; owner: CC — — CC
R-575 Apps & catalog P3 [P3-LOW] The soft memory-overcommit warning is returned as a Hungarian STRING, so it renders Hungarian on an English page. FOUND 2026-09-18 by localisation slice 2 release B (R-557, controller v0.253.0): memoryVerdict now returns its REFUSAL as an error carrying a key (util.MsgErrorf(ErrNotEnoughMemory, …)), but its WARNING is a plain string with no error to carry one, and no language is known where it is built — so it goes through msgHU and is always Hungarian. The copy is in the bundle (a translator sees it; i18n_go_parity.py pins it), only the render is fixed to hu. The named helper and this consequence are stated in internal/stacks/deploy_errors.go. Fix shape: the caller carries the key the way Alert and UpdateRefusal now do — memoryVerdict returns (refusal error, warningKey string, warningArgs []any) and the deploy answer renders it — not a helper guessing a language it cannot know. READY - rank P3-LOW; owner: CC — — CC
R-593 Apps & catalog P3 [P3-LOW] papra's session-signing key is described as „the app's subdomain". FOUND 2026-09-20 translating the catalog (R-560 slice 5, batch 1). templates/papra/.felhom.yml deploy_fields[AUTH_SECRET].description reads „Az alkalmazás aldomainje" — the sentence that belongs on SUBDOMAIN, on a field that signs sessions. SUBDOMAIN itself has NO description at all in that file, so this is a copy-paste that landed one field too low and took the original with it. A customer opening papra's install page reads a wrong explanation under a key they must not regenerate. A localisation release may not change Hungarian bytes (10 §1), so it was not fixed there; and translating a wrong sentence faithfully would have shipped the error in a second language, so that one field was left untranslated — it falls back to the Hungarian exactly as today, and it is the ONE string keeping papra at 13/14 and the catalog ceiling off zero. Fix shape: move the sentence to SUBDOMAIN and give AUTH_SECRET its own („A munkamenetek aláírásához használt kulcs" or similar), re-capture that app's two entries in scripts/copy_freeze/hu.json in the same commit with the reason, then add the English. Owned by R-516 as a Hungarian-words change. READY - rank P3-LOW; owner: CC (catalog) — — CC
R-612 Apps & catalog P3 [P1-HIGH] wishlist cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE. MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot pnpm prisma db seed is Killed — OOM at the catalog's mem_limit: 128M. Without it the Role and Group rows are absent, so every signup fails. The message the user is shown is User with username or email already exists while the container log says the real cause: FOREIGN KEY constraint violated. A household would conclude the account already exists and try to recover a password that was never created. The controller reports the app running and HEALTHY throughout, and the deploy reported successful — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. The catalog was NOT changed — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. Needs: the actual peak RSS of that seed, then a mem_limit that clears it, plus a check that the seed's failure is not silent. — FIXED 2026-09-23 night (catalog a5a729a): 512M. Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel oom_kill 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (The seed command has been executed), and wishlist then moved v0.66.0 → v0.67.1 with its test record. Still open: a failed first-boot seed is invisible to the box — nothing reads the seed's exit. READY — P3, narrowed to making a failed seed visible; owner: CC (catalog) Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains. — — CC
R-613 Apps & catalog P3 [P2-MEDIUM] uptime-kuma parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY. MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at SETUP-DATABASE (Waiting for user action...) and its main socket.io server never starts. The controller's http :3001 probe sees the wizard's 302 and records the app as running and healthy. So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with POST /setup-database {"dbConfig":{"type":"sqlite"}}. This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. Needs: a healthcheck for this template that fails while the wizard is up (the catalog REUSE.md maps the families), and a sweep for other templates whose probe would pass on a setup wizard. — UPDATE NIGHT 2026-09-21: the update night could not seed uptime-kuma for the same reason and left it out rather than faking it. — FIXED 2026-09-23 night for uptime-kuma (catalog a5a729a): UPTIME_KUMA_DB_TYPE=sqlite — the database choice is made, the wizard never appears, the real server starts (/api/entry-page → entryPage, /metrics 401), the database lands in /app/data (the backed-up volume). Red-proofed on 9202 through the product: before, the box read running over setup-database; after, the real server. (A first "after" run used a stale template — R-607's lag — and is kept.) Still open: the sweep for other templates whose probe passes on a setup wizard. READY — P3, narrowed to the sweep; owner: CC (catalog) Re-ranked 2026-10-03: P2->P3: uptime-kuma fixed; only a sweep of other templates remains. — — CC
R-676 Apps & catalog P3 [P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it. From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. audits/night-2026-09-24/A3/40-first-start-restarts.txt 2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression — Deploying clears when compose up -d returns (deploy.go "Clear deploying flag"), and ObserveUnhealthy then samples the app; an automatic update's step, verify and undo ARE covered (Updating, pinned by TestD28_NoCrashLoopStopDuringAnAutomaticStep). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. -- 2026-09-30: the first-start restarts are explained. immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (56c4888, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. OPEN — P3; owner: CC (watch) — — CC
R-682 Apps & catalog P3 [P3-LOW] A Remove interrupted by a controller kill leaves the app half-removed: containers gone, the app still listed as installed (and held). MEASURED 2026-09-24 on 9202 (chaos round 9): the kill 2 s after the Remove press answered the household 502 Bad Gateway; after the restart chaoscrash read deployed, stopped, unhealthy_stop, with NO container left. Pressing Remove again completed it cleanly (200, only the catalog template left). Recoverable by the household's own second press; nothing tells them to press it. Fix direction: the remove journals its intent and finishes (or says it was interrupted) at boot, as the update does. audits/night-2026-09-24/E/round-09*.json, E/round-09b-remove-again.txt READY — P3; owner: CC (controller) — — CC
R-704 Apps & catalog P3 [P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held. Measured 2026-09-28 on 9202: calcom crash-looped at 08:22 and 08:28 (unhealthy_stop, crash_loop, trip 2, recorded 08:28:59Z); it was then REMOVED through the product twice and installed fresh twice (09:14:51Z the last). At 09:45 the new, healthy install's Update was refused 409 held with the crash-loop sentence („…újra és újra összeomlott…"), and GET /api/stacks/calcom carried the old hold_reason while state=running. Start lifted it (the unhealthy stop is LIFTED by Start). So a household that removes a crash-looping app and installs it again (the obvious fix) finds its updates refused for a crash of a previous install. Not measured: whether the nightly update leg also skips it; whether other holds (restore hold) behave the same. Fix direction: the remove clears the app's box-set holds, as DeleteAppBackupPrefs clears its backup preferences (R-474). Evidence: audits/pg-calcom-claper-2026-09-28/box/calcom/hold.txt, …/box/calcom/move.txt. -- 2026-09-28 later: SECOND and worse instance, then FIXED in controller v0.278.0. demo-hp's fresh nextcloud (installed 10:13) carried an UPDATE hold from a nextcloud of 2026-09-13 (set before v0.242.0 made removals clear update holds; nothing ever swept it). At the manual off-site run (15:17) the backup leg logged Skipping volume dump for nextcloud — the app is HELD stopped, captured no unit, and pushed a snapshot that carried NO database dump and NO volume tar — a freshly installed app silently NOT backed up. Fix: a removal also clears the crash-loop stop (settings.ClearUpdateHold), and a new install (plain or "use my kept data") drops a leftover update/crash-loop hold of an app that is not installed (Router.dropLeftoverHold); restore holds (R-379) untouched. Red-proofed RP4–RP6 (audits/kept-offsite-2026-09-28/redproofs/). Floor 0.278.0. STILL OPEN: live proof of the install-time drop (a box with a leftover hold on an uninstalled app). WATCHING — P2; owner: CC (install-time drop, live) — — CC
R-757 Apps & catalog P3 [P3-LOW] A template that gains a generated secret field makes the box INVENT that value for apps already installed — for calibre-web a login name the app never got. MEASURED 2026-10-01 on demo-hp: 9 minutes after catalog e9f50b5 (decision 61) synced, the controller logged InjectMissingFields … injected missing fields: ADMIN_USER (deploy.go:1337) and the app page's reveal returned a 10-character name that was NOT calibre-web's login (the app had kept its own name; after_install runs only after a fresh install). The page did not list the field, but the reveal answers it, and the password field's text now says "the user name above". demo-hp was fixed by renaming the app's user to the box's recorded name (credentials file updated). Any other installed calibre-web gets the same made-up name at its next sync while its login stays admin (no other box has one today: N100 none, Tester-2 not registered). Needs: InjectMissingFields must not invent a value an app has to have been GIVEN (a field consumed only by after_install), or such a field needs an "installed apps: ask" path. audits/calibre-name-and-prune-2026-10-01/A/A2…, A3… OPEN — rank P3-LOW; owner: CC (controller) — — CC
R-758 Apps & catalog P3 [P3-LOW] Eight templates declare a mem_limit that is not the sum of their compose limits, and no gate checks it. FOUND 2026-10-01 by scripts/onboarding_gaps.py (the new-app checklist's gap page, row 5.3): adventurelog 384M vs 896M, bookstack 512M vs 768M, calcom 768M vs 1792M, claper 384M vs 640M, kimai 384M vs 640M, nextcloud 1024M vs 1664M, outline 768M vs 1152M, zipline 512M vs 768M — every one UNDER the sum, so the deploy screen and the box's capacity figure (09 decision 22) read less memory than the app may take. REUSE.md §2 says mem_limit = the sum. The compose limits are what Docker enforces, so no app is starved by this; the figure the household and the capacity check read is wrong. Needs: the eight figures corrected (a template change: needs its own session, no image move), and a static gate (--fast) with a decoy, so a new app cannot repeat it. app-catalog-felhom.eu/onboarding/EXISTING-APPS-GAPS.md READY — rank P3-LOW; owner: CC (catalog) — — CC
R-762 Apps & catalog P3 [P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every /media/ file answers 404. MEASURED 2026-10-01 on 9202 (drill catalog, the live template 82fff32, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links /static/css/workout-manager.css, /static/bootstrap-compiled.css — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to /api/v2/gallery/ answered 201 and the file is on the media volume, but GET /media/gallery/…png answers 404 signed in, without a session, and straight at the app inside the container. Cause, read in the image: the entrypoint runs collectstatic only when DJANGO_DEBUG == "False" and the template sets no DJANGO_DEBUG; and wger serves /media/ only in development (urls.py:393 „served like this during development only”) — upstream's production setup puts nginx in front for /static and /media. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). Needs: DJANGO_DEBUG=False (collectstatic) and something that serves /static + /media (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt, C/C4-seed-photo-size.txt -- 2026-10-01 (operator): wger is lifecycle: hidden until this and its twin are fixed (catalog 55b8c8a; read back on 9202: not on the app list, mealie control present). Merged 2026-10-05 from R-755 (duplicate): templates/wger/docker-compose.yml still sets no WGER_USE_GUNICORN — the gunicorn switch is the same server question. READY — rank P2-MEDIUM; owner: CC (catalog) Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again. — — CC
R-774 Apps & catalog P3 [P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers. READ/MEASURED 2026-10-01: Karakeep's smtp_mapping (plaintext :2526, as Cal.com) was proven only with mail OFF (a fresh install boots) — 9202 has no hub, so the relay cannot be exercised there; the official mobile app ships Sentry crash reporting with a hard-coded DSN (apps/mobile/app/_layout.tsx, FIT.md). Needs: one password-reset mail from Karakeep on a hub-enabled box (demo-hp 9201, a throwaway install); and a sentence on the page about the phone app (copy, freeze). audits/new-apps-2026-10-01/bench/karakeep-mail-off-boot.txt READY — rank P3-LOW; owner: CC (catalog) — — CC
R-76 Apps & catalog P4 FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state idea (surfaced by the R-75 spike, 2026-07-26). Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — Two related findings from audits/SPIKE-catalog-data-paths-2026-07-26.md P3/P5, both pre-existing and deliberately left alone by that spike. (a) FileBrowser Quantum 1.3.3 creates files 0644 and folders 0755 and does not propagate the setgid bit — even though the entrypoint wrapper's umask 002 really is in effect (/proc/1/status Umask: 0002). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's group half holds and only its mode half is lost. The consequence is proven with a control: inside a UI-created 0755 folder a gid-1000 process's file landed group 1000, while the identical write into the 2775 parent landed group 100. So any folder a customer creates through FileBrowser breaks the shared-group chain one level down. Latent today — every userdata-touching catalog app that declares an identity declares uid/gid 1000, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at infra/infra.go:156 is right that the image ignores -e UMASK but does not say t CC
R-284 Apps & catalog P4 „A kiválasztott tárhely majdnem megtelt." on a store that is 93 % FREE — an apparent inverted threshold. Calibre-Web's deploy page rendered <option value="/mnt/sys_drive" data-free-percent="93"> alongside „Tárhely (sys_drive) — 64.2 GB szabad" and the warning „A kiválasztott tárhely majdnem megtelt." 93 % free read as 93 % used is the obvious candidate, and checkStorageSpace(this) is the function to look at. Not confirmed by reading the code — reported as measured output only. A capacity warning that cries wolf on an empty disk is one a customer learns to click past VERIFY (2026-10-03 triage: Probably not a defect: the warning is hidden by default and shown only when free space is below 20%; the measurement likely read raw HTML without running scripts — felhom-controller/controller/internal/web/templates/deploy.html:622,808) — READY (XS) — NEW 2026-08-09 — Check checkStorageSpace's comparison against data-free-percent; add a render test per branch CC
R-527 Apps & catalog P4 [P3-LOW] The catalog flag locked_after_deploy is read by no controller code — every setting is read-only after install whatever the catalog says. FOUND 2026-09-15: stacks/metadata.go parses it; grep -rn LockedAfterDeploy finds no reader; deploy.html renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in 02-controller-module-map.md; the flag is a seam never wired. Fix shape: remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). VERIFY (2026-10-03 triage: the row says nothing reads the flag, but deploy.go reads LockedAfterDeploy into LockedFields; the row may be stale or describe only the page — felhom-controller/controller/internal/stacks/deploy.go:400,1348-1371) — READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: every field is already read-only; no harm. — — CC
R-532 Apps & catalog P4 [P3-LOW] Vaultwarden's /api/config still says disableUserRegistration:false with signups off, so the web vault shows a register form that the server then refuses. MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. READY — rank P3-LOW; owner: CC (catalog/upstream note) Re-ranked 2026-10-03: P3->P4: cosmetic; the server refuses correctly. — — CC
R-567 Apps & catalog P4 [P3-LOW] The two drive wizard pages (/storage/init, /storage/attach) do not highlight the Tárhely menu group — the sidebar reads as if the household left the storage section. FOUND 2026-09-17 by slice 1 release C: the release C parity cases first used page name storage for the wizards; re-captured with the handler's real page name (storage_handlers.go storageWizardPageHandler passes the TEMPLATE name, storage_init/storage_attach, as Page) the fixtures lost nav-group is-open and the active link. Present since the wizards shipped; not caused by localisation. Fix shape: pass Page storage (or teach the layout's storage group both names), with a render assertion that /storage/init carries the open storage group. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: cosmetic navigation polish. — — CC
R-577 Apps & catalog P4 [P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them. FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (launcher_shared, launcher_share_password) deliberately do NOT, and TestGuestSharePagesHaveNoGlobe pins that so it stays a decision rather than an oversight. Why it is the operator's and not CC's: a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The felhom_lang cookie already built would fit them exactly (display-only, their own browser, never the household's setting). Fix shape, if the operator says yes: add {{template "lang_globe" .}} to both shells with the anonymous form, and one render case per page per language. READY - rank P3-LOW; owner: operator (the decision), CC (the change) Re-ranked 2026-10-03: P3->P4: feature decision for the operator; Hungarian default works today. — — operator
R-619 Apps & catalog P4 [P3-LOW] A type: password deploy field is MANDATORY however required reads, and the deploy-fields contract says the opposite — so any caller that trusts it is refused. MEASURED 2026-09-21 on guest 9202 while widening the update drill. GET /api/stacks/grafana/deploy-fields serves {"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}; a deploy carrying only the two required:true fields is refused 400 „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". The BEHAVIOUR is right and is a decision, not a bug: deploy.go:305-312 refuses a password field with no caller value on purpose — "We never silently auto-generate — the user needs to know their password" — which is the opposite of the secret case one branch above, where a generated value the customer never sees is exactly correct. The defect is the CONTRACT. .felhom.yml declares required: false, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (secret) from "optional in the template and mandatory in the code" (password). A person using the deploy page never meets this because the page renders a Generálás button; anything that is not that page does, which now includes this drill harness and would include 09 §6.2's unattended caller the day it deploys anything. Fix shape (smallest that keeps the decision): serve required: true for type: password in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a password field always reaches the wire as required. Alternatively state it in the field's description, which is weaker because it is prose. Evidence: audits/update-night-2026-09-21/apps/grafana/log.txt (the refusal) and batchA.log. READY — rank P3-LOW; owner: CC (controller) Re-ranked 2026-10-03: P3→P4: only non-page callers (test harness) meet it; the deploy page works. — — CC
R-644 Apps & catalog P4 [P3-LOW] gokapi on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, password does not appear to be a SHA-1 hash — and the controller still lists it deployed. OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a gokapi container left by R-633 by name; a gokapi is running again, recorded deployed: true. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. OPEN — P3; owner: CC; investigate before the next drill on 9202 Re-ranked 2026-10-03: P3→P4: seen only on a scratch test guest; the row itself calls it not a customer fact. — — CC
R-654 Apps & catalog P4 [P3-LOW] opengist 1.15 moved every page under /-/ — a household's /login bookmark answers 404 after the update. MEASURED 2026-09-23 night: 1.13 serves /login, /register, /all; 1.15.2 answers 404 on all three and serves /-/login, /-/register, /-/all; / redirects to /-/all. The app, its data and its probe (/healthcheck) are fine, and a household arriving at the root lands correctly — only a deep link breaks. 1.15 also marks its session cookie Secure. Needs: a line in opengist's app_info if the operator wants households told; nothing in the product. Evidence: apps/opengist-oldfixture/, apps/opengist/. READY — P3; owner: operator (copy decision) / CC (writes it) Re-ranked 2026-10-03: P3→P4: only an old bookmark breaks; the app's front page still works. — — CC + operator
R-707 Apps & catalog P4 [P2] 37 apps still start with a login a stranger can take (09 §3 decision 45). Audit of all 53 apps: app-catalog-felhom.eu/FIRST-ADMIN.md (class, fix route, status, measured or read). Open: 3 hard-coded defaults — calibre-web (admin / admin123, measured working on demo-hp and 9202; its own cps.py -s route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (changeme@example.com / MyPassword), wger (admin / adminadmin); 34 open first-run screens (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). Stale notes: romm's default_creds admin / admin answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. Measured on demo-hp 2026-09-28 (read-only): bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via after_install:, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). 2026-09-29 (controller v0.280.0, catalog d0e7e2e): every class-3 app fixed — mealie, wger, calibre-web by after_install (calibre-web with the new password:24:special), proven on 9202 fresh installs (audits/login-gate-2026-09-29/D/); the setup gate (decision 46, spike PASSED) built and live on immich, n8n, audiobookshelf (probes measured) and uptime-kuma (button) (…/C/); romm's and zipline's stale notes removed. Left: 30 class-4 apps — gate each (probe measured on 9202 where one exists — 11 upstream candidates listed in …/B/B-VERDICT.md §3; the button otherwise). 2026-09-29 afternoon (controller v0.281.0, catalog 6faf432): 28 more class-4 apps gated — 32 of 34 — each proven on 9202 (audits/gate-rollout-2026-09-29/B): stranger → gate page / 401, household reached the first-setup screen, the gate opened (9 by a measured probe, the rest by the press), the app answered after. seerr, outline, rallly: gated, their opening needs a media server / e-mail (not proven). Left: wanderer (R-714); plant-it is not installable. NARROWED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — seerr, outline and rallly are gated, but the gate OPENING is not proven (needs a media server / e-mail)) — CLOSED — 2026-09-29 (the rest → R-714) — — CC
R-718 Apps & catalog P4 [P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so. MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. Fix direction: the close card and the gate-open moment say "the app restarts once" where after_setup.env exists. ALSO MEASURED 2026-09-29 (new-household drill): the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. OPEN — P3; owner: CC Re-ranked 2026-10-03: P3→P4: a short unannounced restart; copy polish only. — — CC
R-760 Apps & catalog P4 [P3-LOW] vikunja's compose has no healthcheck, and nothing says why. FOUND 2026-10-01 by scripts/onboarding_gaps.py (row 4.1): two services in the catalog have no compose healthcheck: — adventurelog-frontend (deliberate, a comment cites R-655: the image brings its own) and vikunja (no comment). Not measured: whether the vikunja image declares its own HEALTHCHECK. The controller's probe still runs (healthcheck.checks in .felhom.yml), so the badge is not blind; Docker's own health state is. Needs: read the image's config; either a compose healthcheck of the family the image has (REUSE.md §2), or a comment saying why none. app-catalog-felhom.eu/onboarding/EXISTING-APPS-GAPS.md READY — rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3→P4: the box's own health probe still runs; only Docker's health state is blind. — — CC
R-764 Apps & catalog P4 [P3-LOW] wger sends no mail: its mail backend is the console, so a password-reset mail never leaves the box, and the template maps no SMTP. READ 2026-10-01 inside the app on 9202 (checklist row 7.1): EMAIL_BACKEND django.core.mail.backends.console.EmailBackend; the settings read ENABLE_EMAIL, EMAIL_HOST, EMAIL_PORT, FROM_EMAIL …; the template carries no smtp_mapping. The household's admin can reset another member's password in the app; a member who forgets theirs and asks wger by e-mail gets nothing, and the page does not say so. Not measured: what wger shows after a reset request. Needs: an smtp_mapping (vaultwarden's shape, a fresh install with mail OFF booting — REUSE.md §2), or the page saying mail is not available. audits/new-app-checklist-2026-10-01/C/C2-static-reads.txt READY — rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3→P4: wger is hidden and no box runs it. — — CC
R-768 Apps & catalog P4 [P3-LOW] Grimoire is not built: upstream rules out public exposure and publishes no image for its current line. READ 2026-10-01: SECURITY.md „Public-network exposure is not a supported Grimoire mode”; docs/06-remote-access.md „General REST routes remain loopback-only and tokenless”; v1.x has no published image (the compose builds from source; Docker Hub is the old 0.x). Karakeep covers the bookmark need. Reopens if upstream publishes an image and supports authenticated remote use. audits/new-apps-2026-10-01/FIT.md WATCHING — rank P3-LOW; owner: CC (re-read at the next catalog campaign) Re-ranked 2026-10-03: P3→P4: a declined app idea, watched only. — — CC
R-769 Apps & catalog P4 [P3-LOW] Pinchflat is not built: upstream is paused (no release in 2026, last push 2025-12-16) and its last release has no image tag. READ 2026-10-01: ghcr's newest version tag v2025.6.6, latest amd64 only; runs as root by default; an unanswered 30 GB yt-dlp memory report (#866); third parties call it unmaintained (community-scripts #15968). Forks with images exist (Pinchflat-NGX, MorganKryze). Needs: the operator's word on a fork (a new upstream), after the YouTube sentence (R-767). audits/new-apps-2026-10-01/FIT.md WAITING-ON-OPERATOR — rank P3-LOW; owner: operator Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision. — — operator
R-770 Apps & catalog P4 [P3-LOW] Invidious — fit check only; the recommendation is not to build it. READ 2026-10-01: playback needs invidious-companion (rolling latest, no version tags); PostgreSQL 14 (EOL 2026-11); registration_enabled: true by default; upstream: a bot check means „your IP is blocked from YouTube”, a 429 can last 24 h, triggered by „someone on your network” — on our boxes that IP is the household's. One bad period in 2026 (March, ~2 weeks). No report found of a family's other devices being bot-checked (inference). Needs: the operator's go / no-go. audits/new-apps-2026-10-01/FIT.md WAITING-ON-OPERATOR — rank P3-LOW; owner: operator Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision. — — operator
R-771 Apps & catalog P4 [P3-LOW] moonlight-web — fit check only; not buildable through an HTTP-only tunnel at usable latency. READ 2026-10-01: two unrelated projects (MrCreativ3001/moonlight-web-stream, the original; linckosz/moonlight-web); both need Sunshine/Apollo/Wolf on a gaming PC on the LAN and WebRTC over UDP (40000-40100/udp; linckosz recommends host networking and sends telemetry by default); both have a WebSocket fallback (high latency, all video through the tunnel); a logged-in user controls the PC's desktop. Needs: the operator's go / no-go (LAN-only use would need a different publishing model). audits/new-apps-2026-10-01/FIT.md WAITING-ON-OPERATOR — rank P3-LOW; owner: operator Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision. — — operator
R-796 Apps & catalog P4 [P3-LOW] MeTube's "send to MeTube" helpers (browser extensions, bookmarklets, phone apps) cannot work behind the family gate. READ 2026-10-02 (audits/family-gate-2026-10-02/B/metube-reads/reads-2026.09.29.txt): they call /add from another site and carry no family login; the template sets no CORS origin and MeTube has NO exception by design (an exception on an app with no login would be an open door). The household pastes links in the page (first_steps). Needs: nothing unless households ask; a way would be a per-member token the gate accepts on /add only — a design question, not a fix. WATCHING — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: a design idea only if households ask. — — CC
R-798 Apps & catalog P4 [P3-LOW] Grimmory's template sets SWAGGER_ENABLED=false, which v3.5.0 does not read (it reads API_DOCS_ENABLED, default false). READ 2026-10-02 (audits/family-gate-2026-10-02/B/grimmory-reads/upstream-v3.5.0.txt). Harmless today — the API docs are off by default. Needs: remove the dead line or set the right name, on the next Grimmory step. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: harmless dead setting. — — CC
R-804 Apps & catalog P4 [P3-LOW] plant-it's image is gone from Docker Hub: msdeluise/plant-it:0.10.0 answers "pull access denied … repository does not exist". MEASURED 2026-10-02 on the bench, logged in (audits/persistence-sweep-2026-10-02/A/sweep/batch-*.txt, plant-it). The app is already lifecycle: abandoned (not offered for new installs); a box that runs it could no longer reinstall or restore it from images. No box runs it (read 2026-10-02). Needs (operator): keep the abandoned template as it is, or hide it entirely. WAITING-ON-OPERATOR — rank P3-LOW; owner: operator Re-ranked 2026-10-03: P3→P4: the app is already abandoned and no box runs it. — — operator

App updates — 15 rows (P3 9, P4 6)

ID Category Sev What State Blocked on Next action Owner
R-440 App updates P3 [P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible. compose pull on a moving tag fetches whatever upstream published that day. MEASURED 2026-09-01 over app-catalog-felhom.eu @ 29edad9c5bf4: 79 image: lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version. postgres:16-alpine (8 apps), redis:7-alpine (6), mariadb:11.6 (2), plus one each of postgres:15-alpine, postgis/postgis:16-3.5-alpine, mariadb:11.4, mariadb:12.3, ghcr.io/claperco/claper:2.5, ghcr.io/thomiceli/opengist:1.13, wger/server:2.6. A 24th is arguable and is recorded rather than rounded away: ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0 pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control. Running digests on demo-hp compared against what the registry serves for the same tag today: mariadb:11.4 MOVED (sha256:4f1d8d20... -> sha256:611a2fcc...) and mariadb:12.3 MOVED (sha256:a02fe89c... -> sha256:dd9b303a...), while postgres:16-alpine, redis:7-alpine, mariadb:11.6 and opengist:1.13 were SAME — and both fully-pinned CONTROLS (rommapp/romm:5.0.0, privatebin/pdo:2.0.5) were SAME. So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under romm and bookstack, with no catalog change and no record. Compounding fact found while reading: the recovery unit records ImagePins but the manifest comment says "image NOT stored - re-pulled on restore", so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): app.yaml.installed_images now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade. What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). The row therefore stays OPEN and its rank is unchanged — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. audits/SPIKE-app-update-2026-09-01.md — NIGHT 2026-09-23: every image moved tonight (and the 21 backfilled moves) carries its resolved digest in the catalog's test record; the floating pins themselves still float until the box pulls by digest (09 §6.4 part 6). -- MEASURED 2026-09-30 on demo-hp 9201: a fresh install or a guarded Update renders name:tag@sha256 from the ladder (digest.go RenderWithLadderDigests) — 9 of 20 services there run such a definition (adventurelog, docmost, paperless-ngx). The other 11 run TAG-ONLY definitions: bookstack, kimai, opengist, privatebin, romm were installed before v0.269.0 and not updated since, and CarryDigests (digest.go:141) keeps only a digest the running definition already has; bentopdf and calibre-web have no ladder at all. So a re-pull of those (a restore) takes whatever the tag serves that day. audits/pg-last-six-2026-09-30/E/E2-digests-demo-hp.txt. -- 2026-09-30 (evening): calibre-web now has a ladder entry (53a4a1d), so a new install or a guarded Update renders its tested digest; bentopdf still has none. audits/more-night-apps-2026-09-30/ Merged 2026-10-05 from R-446 (duplicate): the household-visible consequence — for these apps the „Naprakész" badge can never turn „behind" (felhom-controller controller/internal/stacks/updateorder.go:96 returns false with no catalog digests / no test date); R-740 (same-tag re-test) is closed, so floating-tag drift for laddered apps is handled. NARROWED 2026-09-30 — floating only for (1) apps with no ladder entry and (2) apps installed before v0.269.0 until their next guarded Update; owner: CC. Re-ranked 2026-10-03: P2→P3: narrowed to apps with no ladder entry and pre-v0.269 installs; each guarded update fixes the latter. — — CC
R-458 App updates P3 [P3-LOW] .felhom.yml keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running. The v0.235.0 render freezes docker-compose.yml for a pinned app once the catalog moves past its version, but copies .felhom.yml verbatim in every case (Syncer.copyTemplates). The asymmetry is deliberate and both directions were considered: .felhom.yml carries no image, and it carries catalog_since — the single input the update badge uses to say „Frissítés elérhető — N napja" — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. What it costs: the file also carries the controller-side healthcheck: block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. Not fixed now, and the reason is that the cheap fix is wrong: freezing the whole file breaks the badge, and freezing only the healthcheck: key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. What would settle it: whether any catalog healthcheck: has ever been changed in the same commit as an image: line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: CC. architecture/09-update-architecture.md §5.4, §8.5 — UPDATE NIGHT 2026-09-21: MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states. A .felhom.yml-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN bentopdf — installed v2.8.6, catalog ahead. §5.4's asymmetry is confirmed live: the new .felhom.yml reached the box while the compose image: line stayed v2.8.6. But no false alarm was produced: ten samples over two minutes all read state=running with the front door at 200. The reason is the probe's own semantics, not luck — healthprobe.go:258-261 treats any response as healthy for type: http, and the bogus path answers 404, which is a response. So this row's false-alarm risk exists only for type: api probes carrying an expect block, where the status is compared; for every type: http template and every type: api without expect, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md. READY — rank P3-LOW; owner: CC — — CC
R-462 App updates P3 [P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time. The R-449 harness works and is proven by a red negative control (audits/SPIKE-upgrade-test-2026-09-06.md §1). Costed with this run's REAL numbers rather than an estimate: a successful edge takes 6.4 s – 305.1 s, median 71.8 s; a FAILING edge takes 556 s, roughly 8×, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost 5.07 GB, so 53 apps naively extrapolate to ~90 GB and, at the median, about an hour of harness time for one edge each. THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row. Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). Fixture time scales with apps and does not amortise. The decision this row is really asking for is scope, not schedule: all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. Recommended shape, NOT a design — the operator picks: start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: VIKTOR rules on scope, CC implements. audits/SPIKE-upgrade-test-2026-09-06.md §5 ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was. It read "VIKTOR rules on scope"; the operator ruled on 2026-09-13 (09 §3 decision 6) that the upgrade test goes to ALL apps through the nightly rotation, explicitly not "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in 09 §6.4: legs A–E ≈ 21–34 CC-hours plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL pg_upgrade rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. -- UPDATE NIGHT 2026-09-21: The count moved from 3 apps to 21 EDGES ACROSS 19 APPS. The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: 14 proven, 3 failed, 4 inconclusive. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. Ten of the fourteen printed a verbatim migration line, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at audits/update-night-2026-09-21/fixtures.py, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into app-catalog-felhom.eu/scripts/upgrade_fixtures.py with seven new EDGES (U1-U7) so the same edges can be run on the harness venue with their ABORT step, which the box deliberately does not offer. OWED, stated so it is not mistaken for done: the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). -- 2026-09-23, the RomM lesson is IN THE HARNESS: upgrade-test.py v2 runs every edge that read back under light load for --soak seconds (default 600) and reads the kernel's oom_kill counter host-side; a kill or restart turns proven into failed, a peak over 80 % of the limit adds memory_tight. Red-proof: romm 5.0.0 → 5.3.0 on the template AS PROMOTED (512M, four workers) — seeded, migrated, read back, and then OOM-killed at +76 s under four light callers, verdict failed (Docker's own OOMKilled read true here; restarts stayed 0, which is why the walk never saw it). Ten minutes is ample for this failure; demo-hp's first kill came at two hours only because nothing was loading it. Evidence audits/update-rulings-2026-09-23/harness/. — NIGHT 2026-09-23: the bench and the box now share ONE fixture set — the box walk's fixtures run on the bench through upgrade_boxport.py (ported verbatim into the catalog), plus a new wishlist fixture and fixes for opengist 1.15, komga (/api/v2/users/me) and nextcloud (wait for occ status). 14 apps / 15 edges tried; 12 proven on both venues and published with their test records (audits/DRILL-night-2026-09-23.md Part C). -- 2026-09-30 (day brief): three new front-door fixtures (sparkyfitness, rallly, outline; catalog e6f3ec2) and five apps moved on both venues — the three PostgreSQL conversions (R-463) plus bookstack 26.05.5 → 26.09.1 and kimai apache-2.57.0 → 2.67.0 (upstream dropped the apache- prefix; the plain tag's digest equals apache's, measured). The catalog currency and the next-apps ranking: audits/catalog-currency-2026-09-30.md — 25 of 53 apps behind upstream inside a major, 19 across one (morning); after the session 31 apps can update themselves at night (+ nextcloud with a whole copy), 32 carry a proven step, 21 have no ladder. -- 2026-09-30 (evening, more-night-apps brief): ten more steps published on both venues — first ladders for calibre-web, gitea, wger, crafty-controller, uptime-kuma, zipline (new front-door fixtures; gitea's installer form), within-major steps for emby, ghost, home-assistant, outline, rallly. Night-updatable by the audit's method (a): 30 + 2 conditional at the session's start → 35 + 3 conditional (calibre-web conditional: files_may_change). zipline got a two-step ladder 4.6.1 → 4.7.0 → 4.8.0 (the direct jump fails, R-742). Stopped: wanderer (the bench cannot run it, R-739). Fixture/harness defects fixed on the way: R-735, gitea's installer, the HTTPS backend on the bench, per-file names for files_may_change. audits/more-night-apps-2026-09-30/ READY — rank P2-MEDIUM; owner: CC (scope already ruled, 09 §3 decision 6) Re-ranked 2026-10-03: P2→P3: 35+ apps already update by proven steps; an app without a ladder simply does not move, which is safe. — — CC
R-469 App updates P3 [P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse. Since 2026-09-13 app-catalog-felhom.eu CLAUDE.md rules that until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version (four MariaDB, eleven PostgreSQL services), and scripts/check-engine-major.py (fourth row of catalog_gates.py, run by .githooks/pre-push with the push range) refuses one, naming the rule and this expiry. Why the rule: every mariadb: sidecar now carries MARIADB_AUTO_UPGRADE=1 (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. Honest limit, not re-filed: the gate needs a parent commit and CI fetches at --depth 1 — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. When R-448 ships: delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. 2026-09-13 — UNBLOCKED, NOT LIFTED. R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. The rule stays in force until someone deliberately removes it, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). HALF-LIFTED 2026-09-21, catalog 5ff36d098cbc. Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — for MariaDB: the four mariadb: sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and MARIADB_AUTO_UPGRADE=1 whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. PostgreSQL and MySQL stay refused — postgres performs no pg_upgrade and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. R-450's second half is enforced in its place: a MariaDB major must be the ONLY image move in its template in that commit (the bookstack 0b73e5e shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. What remains of this row: the PostgreSQL half, which is R-463's to clear — see 09 §3b Q5. -- NARROWED 2026-09-25 (evening), catalog 6a4a5f0, 09 §3 decision 35: the PostgreSQL half now passes ONE app at a time — only a template whose ladder entry for the step is proven on BOTH venues and carries engine_conversion (the box converts it, controller v0.273.0), as the only image move in its commit. Every other PostgreSQL app stays refused; the postgis family is judged now (it was not). CLAUDE.md rule text updated the same commit. Decoys + red-proof: audits/night-2026-09-26/C/. NARROWED — PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463 — — CC
R-607 App updates P3 [P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and catalog_images stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds. MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, POST /api/sync was invoked, and it answered „nincs változás" — yet the box's own cache file <data>/catalog-cache/templates/uptime-kuma/docker-compose.yml had moved to the new tag. Stack.CatalogImages as served by the API stayed at the OLD value until a separate POST /api/stacks/rescan. WHY IT MATTERS AND WHY IT IS NOT COSMETIC: CatalogImages is the single input stacks.CatalogOrder compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of 09 was written to prevent, arriving by a different door. It also cost a measurement: the session that found it nearly recorded a tag-ok badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. An instrument that can report an old value as a current one is not a measurement. TWO SEPARATE QUESTIONS, and the row does not conflate them: (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether CatalogImages is refreshed by the sync at all or only by ScanStacks on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. Neither was isolated — this row records the observation, not a diagnosis. Also observed in the same run, NOT diagnosed and folded in here rather than filed twice: a removed app leaves applied-compose.yml behind in its stack directory. Stated as observed; it was not established whether that is intended. Fix shape: first reproduce with a loop that pushes a tag, syncs, and reads catalog_images on a timer, so the window is a NUMBER before anything is changed. Evidence: audits/update-arc-2026-09-21/06-sync-after-bump.txt and 16-r524-sync-box-ahead.txt. — UPDATE NIGHT 2026-09-21: Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one. On mealie the bump was pushed, POST /api/sync AND POST /api/stacks/rescan were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached catalog_images yet. The guarded Update was then pressed and reported „Frissitve" after 2.1 seconds having moved nothing at all: pinned, installed, the live compose line and docker inspect all still read v3.20.1. That is honest given a stale cache — the pin is written from the catalog's current definition, which was still the old one — but what the household sees is a button that says it updated them and did not. A NUMBER, at last, which is what this row asks for: the night's harness was changed to poll catalog_images until the pushed reference appears and to report how long that took; those figures are each edge's badge_catchup_seconds, and here they are: 4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s. The four fast ones are one sync+rescan round; the 29-second one (nextcloud, an engine-sidecar bump) needed additional sync+rescan rounds before catalog_images carried the pushed reference. So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by POST /api/sync answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with catalog_images staying stale until a separate POST /api/stacks/rescan. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. The window was still never measured as a NUMBER — that is what the row asks for and what remains owed. — NIGHT 2026-09-23, measured twice more: for an INSTALLED app the stack directory's compose is the applied one, so waiting on it for a drill commit took ~7 min (chaos rounds 2 and 4); the badge's own catalog_images caught up in seconds. And twice a fresh box walk deployed a STALE template (the uptime-kuma "after" run, the opengist re-walk) because the deploy ran before the sync reached the file — the walks now wait for the box's template to show FROM. READY — rank P3-LOW; owner: CC (controller) — — CC
R-615 App updates P3 [P3-LOW] Pointing a box at a different app catalog by git.repo_url alone is INERT — the box keeps fetching from the repository it first cloned. FOUND 2026-09-21 by reading sync.go before running it, which is the only reason the update night's drill catalog worked at all. Syncer.gitCloneOrPull (controller/internal/sync/sync.go:274-306) clones only when <data>/catalog-cache/.git is absent; on every later cycle it runs git fetch --depth 1 origin <branch> + git reset --hard origin/<branch> against the remote stored in the clone, which buildRepoURL wrote at clone time. Changing git.repo_url in controller.yaml and restarting therefore changes nothing: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, git -C <data>/catalog-cache remote -v still read app-catalog-felhom.eu; the box only followed the drill repo once the cache directory was removed. Why it matters beyond a drill: this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. Nothing is wrong with the CACHING, which is right; what is missing is that a changed repo_url must invalidate the clone. Fix shape: on start, compare git.repo_url with the clone's origin and re-clone when they differ (or git remote set-url + a full fetch); log which happened. A test that changes repo_url under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: audits/update-night-2026-09-21/04-9202-config-pre.txt, 05-9202-follows-drill.txt. -- 2026-10-01 (night, new apps): the other direction measured on 9202: after repoint restore a DRILL-only template (grimmory) stayed OFFERED on the app list, because its /opt/docker/stacks/grimmory folder (written by the drill sync) remained; removing the folder by hand ended it (audits/new-apps-2026-10-01/B/B9-restore-live.txt). Older drill folders (chaoscrash, chaosoom, chaosoomb) sit there too. The drill teardown should list and remove stack folders the live catalog does not have. READY — rank P3-LOW; owner: CC (controller) — — CC
R-683 App updates P3 [P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made. 2026-09-24 chaos round 3 (nextcloud, backup_max_age: 1m): no backing-up phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. audits/night-2026-09-24/E/round-03*.json, E/round-11-controller-pre.log OPEN — P3; owner: CC (watch) — — CC
R-738 App updates P3 [P2-MEDIUM] Every wger update that brings database migrations leaves wger broken, and the guarded Update reports done. MEASURED 2026-09-30 on scratch guest 9202 (drill catalog): wger 2.6 → 2.7 through the product's guarded Update ended done in 72.8 s (health: the front page answers); the web login then answered 500 (no such column: core_userprofile.time_zone), still 500 a minute later, 12 migrations unapplied (manage.py showmigrations). Cause, read from the image: entrypoint.sh runs manage.py migrate only when DJANGO_PERFORM_MIGRATIONS=True; the template does not set it (a fresh install works because wger bootstrap builds a new database). The harness caught it: the box verdict is failed, and no ladder entry was written. No box reporting to the hub runs wger (hub /apps, same day). Also a product gap: the guarded Update's health check cannot see an app whose front page serves while its data is unusable — the fixture's read-back is what saw it. Fix: DJANGO_PERFORM_MIGRATIONS=True in the template (the image's own switch for it), then the step again on both venues. audits/more-night-apps-2026-09-30/box/wger/ -- FIXED 2026-09-30: catalog 7a4ff48 sets DJANGO_PERFORM_MIGRATIONS=True; the same 2.6 → 2.7 step then ran the migrations (Applying … in the log) and read back on 9202 (46.2 s) and on the bench (proven, 0 kills); published 4ad32aa. The product-side gap (a health check that sees only the front page) remains, as a note: the fixture's read-back caught it, the box could not. NARROWED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the guarded Update's health check reads only the front page, so it cannot see an app whose front page serves while its data is unusable) — CLOSED 2026-09-30 — catalog 7a4ff48 + 4ad32aa — — CC
R-785 App updates P3 [P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24). READ 2026-10-01 (audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt). Needs: an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. OPEN — rank P3-LOW; owner: CC (after R-784) — — CC
R-460 App updates P4 [P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE. MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: APP_URL comes from the template as https://${SUBDOMAIN}.${DOMAIN}, so the app marks its session and XSRF cookies secure; curl over plain http stores neither and every login POST returns 419 Page Expired, which looks exactly like a wrong password. The container serves no TLS. The database half IS provable — the harness seeds with php artisan bookstack:create-admin and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. What is unprovable is an uploaded image or attachment, i.e. exactly the half a customer would notice. THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. Deliberately NOT worked around: planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. What would remove it: a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: CC. audits/SPIKE-upgrade-test-2026-09-06.md §6 -- UPDATE NIGHT 2026-09-21: bookstack's edge was walked again on 2026-09-21 and is again half-proven: the database half read back through php artisan with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: a limit of test coverage for one app, not a defect a household meets. — — CC
R-610 App updates P4 [P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only. R-520 (CLOSED 2026-09-21) cut in pulling, where nothing had run: the pin goes back and that is the easy case. Its own last lines said the starting cut was "NOT measured and does not re-open this row", and no row carried it. The dangerous case is the cut AFTER the new version has started and may already have migrated the customer's data — where RecoverUpdates (stacks/update.go:909) marks the app Updating and RESUMES rather than rolling back. That behaviour was READ from the source and never observed. CLOSED 2026-09-21 — measured THREE times on guest 9202, controller v0.260.0, with three different apps and two different cut mechanisms: vikunja 2.3.0→2.6.0 and uptime-kuma 2.4.0→2.5.0 by pct stop (a real power cut), and wishlist v0.66.0→v0.67.0 by restarting ONLY the controller container (exactly what a self-update does). All three ended HONEST: the recovery line appeared, the update resumed, each app came up on the NEW version, and in every case all FOUR version observables agreed — pinned_images, installed_images, the live compose image: line, and docker inspect of the running container (digests matched too). No hold, no stuck Updating, no surviving journal, no retry loop. The seeded data read back through each app's own front door for the two where a post-cut read-back was taken. THE DANGEROUS CASE WAS GENUINELY EXERCISED, and the proof is a log line, not an assumption: vikunja's own log shows Ran all migrations successfully and Vikunja version v2.6.0 at 12:28:26.881 UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering. The 2.6.0 schema migration had ALREADY been applied to the customer's SQLite database when the power went. Recovery resumed FORWARD, so old-binary-on-migrated-database never happened — but this branch is one step from it: had the cut landed a second earlier, in pinning or pulling, the 2.3.0 pin would have been put back onto a 2.6.0 database. That is not a defect today; it is the reason §4's "no automatic rollback" ruling is right, and it is now evidence rather than argument. INSTRUMENT LIMIT, stated because it bounds the claim: starting lasts well under a second on this box. Three attempts across two cut mechanisms (pct stop returning in 3.0–3.8 s; a controller restart in 1.7 s) ALL landed in verifying. No phase was faked. RecoverUpdates handles starting and verifying in ONE branch, so all three runs exercise the same recovery arm — the arm under test. A cut that lands inside starting itself remains unmeasured and would need an in-process fault injector. Evidence: audits/update-arc-gaps-2026-09-21/ 04, 05, 07. NARROWED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — a power cut landing INSIDE the starting phase is unmeasured (all three cuts landed in verifying); needs an in-process fault injector) — CLOSED 2026-09-21 — measured three times — — CC
R-618 App updates P4 [P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need. RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point: tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered verifying at 21:17:46. At 21:18:28 the NEW version was Up 25 seconds and answering HTTP 200 on /accounts/login/ through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. verifying therefore cannot pass, the full update.health_timeout is spent, Manager.failAndHold runs compose down, and the app is STOPPED. Nothing is lost — the data is in the volumes and the restore works — but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it. Evidence: audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog f5f6a152b513). Two shapes, one class: (a) tandoor — the WRONG PORT. .felhom.yml probes port: 8080; the container listens on 80 and nothing else (ss -ltn inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads healthy, and /accounts/login/ answers 200 through the household's real front door. GET /api/stacks/tandoor nevertheless reads state: "unhealthy". (b) zipline — the WRONG PATH. .felhom.yml probes /api/health, which zipline 4.6.1 answers 404 Route GET:/api/health not found; the compose healthcheck in the very same file uses /api/healthcheck and is correct and green. /dashboard answers 200. The controller reads unhealthy. This is the MIRROR of R-613 — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). IT DOES NOT ALARM, AND THAT SETS THE RANK: 08-alarm-ladder.md §4 puts unhealthy deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on state. IT ALREADY COST A MEASUREMENT TONIGHT: this drill's harness waited for state == "running" and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE. Both halves of the answer live in the same template: compare the .felhom.yml probe's port and path against the compose's own healthcheck: test: URL. A sweep of all 53 templates on that rule was run tonight and returns five disagreements: tandoor (PORT — CONFIRMED live), zipline (PATH — CONFIRMED live), wger (PORT, probe 80 vs compose 8000 — CONFIRMED live the same night), home-assistant (PATH, /api/ vs /manifest.json — NOT MEASURED), and adventurelog (a FALSE POSITIVE of the sweep's own regex — it reads running live). So the rule finds both real defects, with two candidates and one false positive out of 53 — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the traefik port) is strictly worse: it clears zipline and convicts adventurelog. AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE: in both confirmed cases the container's OWN docker healthcheck was green the whole time. A verifying phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports healthy, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. Needs: fix tandoor's port (80), zipline's path (/api/healthcheck) and wger's port (8000) — all three are now CONFIRMED live, none is a guess; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. THE GATE'S RULE WAS THEN SHARPENED BY READING healthprobe.go RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction. type: http treats any response as healthy (healthprobe.go:258-261), and type: api with no expect block does the same (:265-268); only type: api WITH expect.status cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns four candidates — tandoor, zipline and wger (all three CONFIRMED live — probe type: http, port: 80; inside the container port 80 is refused and port 8000 ANSWERED; docker's own healthcheck green; front door 302; the box reads unhealthy), and adventurelog (a false positive: its compose lists two containers' ports and the probe targets the backend; measured running). home-assistant is correctly CLEARED by the sharpened rule — type: api, no expect, so its /api/ answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check. Evidence: audits/update-night-2026-09-21/10-probe-port-sweep.txt, 12-probe-vs-compose-healthcheck.txt and 13-probe-sweep-sharpened.txt. CLOSED 2026-09-22. All three fixed in one commit (app-catalog-felhom.eu@793c4fb): tandoor 8080 -> 80, wger 80 -> 8000, zipline /api/health -> /api/healthcheck. No image: line moved, so no catalog_since moved. RED-PROOFED LIVE ON 9202 THROUGH THE PRODUCT, BOTH DIRECTIONS. Before the fix, at the LIVE pin, all three read Nem egészséges / Not healthy on their own app page while docker reported every container healthy and each front door served a real page through the household's own route — tandoor 200 Login / Sign In, zipline 200 Zipline, wger 200 wger Workout Manager. The fix was applied through the REAL sync (POST /api/sync answered frissítve: tandoor, wger, zipline) and all three read Fut / Running at the next poll, with no redeploy and no restart. AND THE EDGE THAT FAILED WAS RE-WALKED AND PASSED: tandoor 2.6.13 -> 2.6.15 via the drill catalog ended done at +41.1 s with the seed read back through tandoor's own front door and both containers running with zero restarts — where the identical edge on 2026-09-21 entered verifying at +58.4 s and ended failed at +361.9 s with the app stopped. Same app, same versions, same button; the only change is one port number. tandoor's verdict moved failed -> proven and it is now on the live catalog. THE GATE SHIPPED WITH IT: scripts/check-probe-matches-compose.py, a --fast row in catalog_gates.py, comparing the probe against the SAME service's own compose healthcheck. The rule was read out of healthprobe.go rather than guessed: a wrong PORT refuses for every check type; a wrong PATH refuses only for type: api WITH an expect block and WARNS otherwise, which is why home-assistant is warned about and not convicted. Four red-proofs and five decoys, plus five more with PyYAML shadowed out (R-630's sibling problem — see below). Residual, filed separately: R-630 (paperless-ngx's probe can never run at all) and R-631 (five templates the static rule cannot judge). NARROWED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the controller-side idea — let verifying accept docker's own healthy before it stops a working app — was raised and never decided) — CLOSED 2026-09-22 — three probes fixed, red-proofed live both ways, gate shipped with decoys; tandoor re-walked failed -> proven — — CC + operator
R-621 App updates P4 [P2-MEDIUM] A held update DESTROYS the evidence of why it failed: failAndHold runs compose down, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it. MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: adventurelog v0.12.1 → v0.13.0. The new backend applied nine Django migrations successfully and then never listened; the update held after the full 5-minute health wait. Manager.failAndHold (stacks/update.go:723) calls updateCompose(dir, env, "down"), which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded Logs result for adventurelog: 0 bytes returned (empty) and docker ps -a held nothing at all. What survives is the WHAT and not the WHY: the controller line update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy) and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. Why it matters beyond a drill: the hold sentence sends the household to a restore, and after the restore the only remaining question is should I press Update again? — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. Fix shape: capture compose logs --no-color --tail N into the stack directory (beside applied-compose.yml, which already travels with the stack) IMMEDIATELY before the down, and surface it on the app page's hold panel or at least through the existing /api/stacks/<n>/logs fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt, state-after-hold.txt. FIXED in v0.262.0. failAndHold now writes each service's log into <stackdir>/hold-logs/<ts>/compose-logs.txt (compose logs --no-color --tail 400) before the down that destroys them, best-effort by design: a hold must never fail because its evidence could not be written. This is what adventurelog cost — nine migrations ran, the app never bound its port, and the only log that could have said why was gone before anyone looked. VERIFY (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fix shape also said to show the captured hold log on the app page's hold panel; the row does not say that was done) — CLOSED 2026-09-22 — v0.262.0: the hold keeps the app's own log before stopping it — — CC
R-687 App updates P4 [P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap. (1) W+5h reached with steps left is proven by unit test only (TestLeg_NoStepAtOrAfterW5h) — the leg starts at W+105m and would need a 3-hour leg live; (2) the off-site leg FAILING before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are TestChainUpdateLeg_EveryPath; (3) a files_may_change step WITHOUT a whole copy: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) the full-system gate waiting cannot run on 9202 (no agent), and did not occur on the demo boxes' real night either (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (TestD20_GateWaitsForTheLeg). Also cosmetic: a leg with no steps reports "steps": null to the hub, not []. Observability: when the leg TAKES a files_may_change step it does not log which whole copy allowed it (only the skip says why). audits/night-2026-09-25/C/ -- NARROWED 2026-09-25 (controller v0.273.0): the cosmetic "steps": null → [] and the taken files_may_change step's missing log line are FIXED (red-proofed, audits/night-2026-09-26/F/). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). -- 2026-09-28 (night 27/28): (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt). -- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE. The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, quiesce.poll_interval 1m: [quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's backup: completed 9.98 GB. Config and window put back and read back (audits/pg-last-six-2026-09-30/C/). Found, cosmetic, manual chain only: the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text. — — CC
R-734 App updates P4 [P3-LOW] The harness marks immich files_may_change because immich rewrites six 13-byte .immich folder markers at every start. MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of appdata/immich changed; the only changed files were {encoded-video,library,backups,profile,thumbs,upload}/.immich, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (files_changed []). Needs: a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. -- 2026-09-30 (evening): the harness now NAMES the files behind the mark (files_changed_detail, catalog 5b1972b); on immich's step 0b82… re-proof it named exactly the six .immich markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: media/books/metadata.db, metadata.db-shm, metadata.db-wal changed; the book file did not (bench names re-run, A/calibre-names/). That is household data (the template's backup class mandatory holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. audits/more-night-apps-2026-09-30/ READY — rank P3-LOW; owner: CC (harness); the rule change needs a word Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk. — — CC + operator

Backup & restore — 51 rows (P2 8, P3 21, P4 22)

ID Category Sev What State Blocked on Next action Owner
R-32 Backup & restore P2 [P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible. The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's "hetzner":"ok" leg destroys the sub-account, but a Hetzner sub-account is an access-control object, not a data object — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — Ruling from the run (three parts, deliberately separate): (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable BY DESIGN → RESET gains a main-account purge of the customer base dir (the existing operator ack already covers it); (2) the move-aside guard STAYS for reinstall-without-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator Restic tab shows per-customer directory bytes vs attributed snapshot bytes, so dead data cannot hide. Measured on the pool box that night: 49 M attributed (2 snapshots, 48.717 MiB) against 1.4 G + 3.0 M unattributed across TWO .orphaned-* dirs. Evidence restic-and-pool.txt CC
R-105 Backup & restore P2 Three hub-held DR records are empty on the entire live fleet. hosts.dr_record_json = {} on all 3 hosts; host_escrow.directive_json = {} on both escrowed hosts; dr_recipe.host_half.drives = [] on every customer including two with enrolled data drives (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state READY — 2026-07-28. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. PARTLY FIXED BY ITS OWN UPDATE. The drives third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so isUserDataDrive never saw them). The other two thirds — hosts.dr_record_json and host_escrow.directive_json — were NOT re-verified this session and are carried as written. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — These are exactly the fields a host-loss recovery reads: 05-hub-architecture.md:175-176,186 names the slim DR record as one of four durable sources; 06-offsite-connectivity.md:148-150 says the escrow upload carried the DR directive; felhom-agent/internal/dr/plan.go:34-35 makes PlannedDrive the re-attach-by-durable_id wrong-disk guard. The three may have different causes — isUserDataDrive (internal/hub/dr_recipe.go:129-136) requires type usb/local-dir and a non-empty DurableID and MountPath, and which of the three fails was not traced. Evidence: architecture/_recovery-inventory-2026-07-28.md Part D2.3. UPDATE 2026-07-28 (vzdump-target move): the drives third is TRACED and now POPULATED on both demo boxes. Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered report.StorageTargets and isUserDataDrive never saw them. Giving each drive a dir storage at its own mountpoint supplied all three required fields at once (type local-dir, fs-UUID durable id, mount path), and the recipe now emits uuid:91d2dc2d-…//mnt/nvme-1tb on demo-hp and uuid:47a3361a-…//mnt/hdd_1 o CC
R-232 Backup & restore P2 DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails. Surveyed read-only 2026-08-06 (audits/RECON-dooplex-backup-2026-08-06.md). What works: five sets, 14/14 successful runs in 14 days; a file was restored from the data repo and matched the live original byte for byte; every set except two is cross-disk; k3s is integrity-checked on every run. What the matrix exposes, ranked: (a) notify_failure is a no-op — NOTIFY_ON_FAILURE=true but NOTIFY_WEBHOOK_URL is commented out, so a failed backup notifies nobody; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) Nothing leaves the box — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is nfs://192.168.0.180: pointing at DooPlex itself, and the only outbound-looking cron pulls inbound from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) The backup tree is a single writable path and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) Two same-disk sets: .claude-memory and the PostgreSQL dumps, whose source directory sits inside the backup tree. (e) Longhorn retain=1 — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) /opt/backup/docs/BACKUP-RESTORE.md does not exist though the systemd unit advertises it. (g) secrets/restic-repo has never held a snapshot — backup-secrets.sh contains no restic call; the secrets are GPG files on sda1 only. (h) No restore has ever been run beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. Not a finding: the restic passphrase. The on-box copy is on sdb1, a different disk from the backups, and the operator holds an offline copy out of band — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. Nothing was changed by the recon. NARROWED 2026-10-05 — owner Viktor. (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (HubDBBackupStale); notify_failure is still a no-op for the rest. (c)–(h) unchanged. READY for the rest — — operator
R-304 Backup & restore P2 The retained escrow key works, and the customer is told their correct code is wrong. DRILL 2026-08-12 answered the three questions separately, on demo-felhom, with planted data. (a) retention: WORKS — the first retained row in fleet history to carry material (host_escrow_superseded id 11, identity_blob 572 B), byte-identical (sha256 a10032341c8584ed…) to the pre-supersession host_escrow row. (b) the material opens the old store: YES — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (sha c60c8bc737a6b7c6…), and restored three planted files byte-identical from a store the box itself could no longer open (negative control first: Fatal: wrong password or no key found), including a Hungarian accented filename verified as raw bytes. (c) the customer's route: DOES NOT EXIST, and misinforms. ListSupersededEscrow (store.go:2841) is the only reader of a retained identity_blob and has zero production callers — five call sites, all _test.go; the product path (POST /escrow/recover-offsite-password → FetchIdentityEscrow → GetHostDRBundle, store.go:3152) selects FROM host_escrow — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered "the recovery code did not open the sealed bundle — nothing was written". This is the R-224 class again: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. Consequence: the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; any capability-map claim that the customer can recover the old history with their recovery code is false today and must move READY (L) — NEW 2026-08-12, RANK 1 R-198, R-199, R-224, R-241 Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. Until one of those, the honest position is that retention is an operator-only capability. At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked operator + CC
R-366 Backup & restore P2 The 21 August reinstall orphaned demo-hp's PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure. Hub event 3016, 2026-08-21 21:59:28Z, unprompted: Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b. The archive predates the reinstall by three days. This is the PBS-tier analogue of R-193 (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. Credit: the restore-test caught it and said so precisely — the mechanism works. The gap is what it is called: it is reported as a restore test that failed, which reads as a flaky verification, not as every whole-guest backup you took before the reinstall is unreadable on this machine. Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it. OPEN — HIGH related: R-193 Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. CC
R-518 Backup & restore P2 [P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre". MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (phase4/guest-backup-quiesce-log.txt) — ≈ 7 m 45 s with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. Fix shape: state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0): a tier whose storage the agent reports absent is skipped before anything stops (backup_tier_skipped, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. Still open: quiesce per tier, so a slow second tier does not keep every app down. — NIGHT 2026-09-23 (controller v0.267.0): the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (audits/night-2026-09-23/A5-*). The brief's „csak néhány másodpercre" had already gone in v0.243.0. Still open: quiesce per tier, so a slow second tier does not keep every app down. READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, audits/hub-safety-2026-10-05/partE/. — — CC
R-638 Backup & restore P2 [P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind. MEASURED 2026-09-23 on 9202. ImportDump (appbackup/dbdump.go:719, psql -v ON_ERROR_STOP=1 --single-transaction) replays a pg_dump --clean --if-exists file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey … — the new version created six tables whose foreign keys point at old ones, and --clean only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (mariadb-dump, FOREIGN_KEY_CHECKS=0) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left 12 base tables of the new version behind; RomM 5.0.0 happened to ignore them. What worked: DROP SCHEMA public CASCADE; CREATE SCHEMA public; + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. Why this is a row of its own and not only part of R-637: the SAME loader backs shipped paths — rollbackSafetyDump (off-site restore's undo) and the dump replay of the restores — so any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED: whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: audits/update-rulings-2026-09-23/README.md Part 1, docmost-45, romm-44. -- NARROWED 2026-09-23: the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). What stays open is the part about SHIPPED paths: rollbackSafetyDump and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first — — CC
R-822 Backup & restore P2 An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one. MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone --append-only): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6 keep only the fakes and select all 3 real snapshots for removal (--dry-run). Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this. Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by MaxRemove and the hub's count check, not prevented. — Before any forget: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) CC
R-49 Backup & restore P3 [P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup". Measured 2026-07-19: immich_ml_cache.tar 823 660 032 B (~60%) — re-downloadable ML model weights; immich_postgres_data.tar 308 251 136 B (~23%) — a raw tar of the postgres data dir that DUPLICATES the logical .sql dump captured beside it; upload/backups/ 18 MB — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 dccc13fe… tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: 72 MB of originals. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker. — Evidence: audits/DIAG-immich-restore-round2-2026-07-19.md §4 (full byte breakdown). This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. Recorded, deliberately not changed — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) immich_ml_cache — pure cache, strongest case; (b) the postgres_data volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) upload/backups/. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app CC
R-127 Backup & restore P3 The catalog's data_key: true flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory READY (S/M) — Found by D5's Part 0, and it is why D5's boundary is type: secret rather than data_key. Two separable legs. (a) The misclassification. Only 5 fields across 4 apps set data_key: true (adventurelog/SECRET_KEY, homebox/HBOX_AUTH_API_KEY_PEPPER, papra/AUTH_SECRET, sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}), yet n8n/N8N_ENCRYPTION_KEY („Titkosítási kulcs"), wanderer/POCKETBASE_ENCRYPTION_KEY („Adatbázis titkosítási kulcs"), calcom/CALENDSO_ENCRYPTION_KEY and bookstack/APP_KEY are unflagged — the catalog's own Hungarian labels contradict the flag. D5 makes this non-urgent but not harmless: everything type: secret now travels, so the keys DO reach the drive; what stays wrong is the fail-closed gate, which only refuses for data_key names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (app-catalog-felhom.eu, a catalog-only change) + a gate/test that the flag set and the label set agree. (b) The regenerated-DB-password trap. internal/backup/restore_unit.go O4 generates a replacement for any missing non-data-key secret. Proven on postgres:16-alpine: with PGDATA restored from the volume tar, POSTGRES_PASSWORD is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local trust socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that "stored data is unaffected" and scoped it, but did not add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or ALTER USER to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half CC
R-231 Backup & restore P3 /opt/backup/scripts/ on DooPlex is unversioned host state — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (CLAUDE_MEMORY_DIR in backup-config.sh, multi-path restic call in backup-data.sh) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in felhom.eu/workspace/README.md so it is at least recorded. Two related facts, both understating current safety: the backup destination (/mnt/5_hdd/backup) is on the same physical disk as the workspace it protects, and the DooPlex backup set has no off-site leg (sync-hetzner-backups.sh is jarrs.eu and pulls from Hetzner to DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. READY — owner Viktor — — operator
R-240 Backup & restore P3 A backup that covered nothing calls itself „Sikeres". On a configured box with no app selected for off-site backup, a run reports status ok with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — successful immediately beside nothing is selected. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. It is the same rhetorical shape the project has spent a fortnight removing — R-203's a warning beside a success is read as a success, R-234's „✓ Rendben" over an app that was skipped, R-225's unknown rendered as zero — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay ok: an unconfigured box reporting incomplete forever is its own defect, pinned by a test. The defect is the word „Sikeres", not the verdict. Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. READY — owner Viktor — — operator
R-251 Backup & restore P3 The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice. Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags felhom-offbox,calibre-web; the listing renders two rows — calibre-web · 2026-08-07 14:57 · 12.8 MB and felhom-offbox · 2026-08-07 14:57 · 12.8 MB. felhom-offbox is the tier's own marker tag, not an application. The screen's whole job is to let the customer check that what is in the store is what they expect ("Nézd át, hogy tényleg azt találod-e itt, amire számítasz"), and it shows them a stranger's name beside their own data and a total that is double the truth. Cosmetic, not a data defect — the restore page correctly offers only calibre-web. Fix: filter the marker tag out of the listing, or key the rows on the app tag. READY — owner Viktor — — operator
R-257 Backup & restore P3 C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route. web/offbox_handlers.go:270 (the Go error it mirrors is backup/offbox.go:343). „Offsite" is untranslated; „elárvult állapot" is the codebase's own OffboxOrphaned() predicate surfacing verbatim. A customer who pressed a button and got this cannot tell whether something failed, whether they did something wrong, or what to do instead. This is a refusal that is CORRECT and fail-closed and still a dead end — the same shape R-241 recorded for --recover-offsite-install. Fix shape, not a decision: say what the customer tried to do, why it does not apply right now, and where to look — or, since this is a state they cannot reach deliberately, do not offer the action at all READY — owner Viktor — — operator
R-314 Backup & restore P3 StopAbandon has no web route — a customer who telephones is served by a command line. --abandon-stop exists on the controller binary (cmd/controller/main.go:86) and refuses rather than silently no-opping, which is right. But there is no handler: the only in-product way to cancel a countdown is the customer finding their recovery code (Scenario G cancels it at the moment the code proves they still have it). An operator who is telephoned instead has to reach a shell on the customer's machine. Used this session on the operator's ruling, container stopped first so the running controller could not overwrite settings.json from memory — a sequencing subtlety that is itself an argument for a route READY (S) — NEW 2026-08-12, RANK 3 R-241, R-307 An operator-authenticated POST that calls the same StopAbandon, so the telephone path and the code path converge CC
R-362 Backup & restore P3 A data drive detached mid-restore is reported as „permission denied". Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (IsDisconnected, used by both backup legs) and the restore path never consults it. A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist. Creditable in the same test: the agent re-bound the drive 5 s later, unaided. OPEN — MEDIUM — Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. CC
R-401 Backup & restore P3 Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store. Controller v0.228.0 (R-399) made --read-data-subset=100% the default for every box. The whole justification is a single data point: demo-hp, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% 39.2 s — four seconds. Re-proven live 2026-08-31 at 38.7 s. It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA. A 50 GB store is ~370x the data and this curve says nothing about it. Nothing was invented from that one point — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. readDataSubsetRe already accepts n/m, so a rotating schedule (1/7 on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. THE TRIGGER IS AN EVENT, NOT A DATE: the slow-check WARN from v0.228.0 firing on any box (integritySlowNoticeThreshold, 5 min) — that line names the duration, the depth and this row. WHAT HAPPENS IF NOBODY ACTS: every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. Whoever acts must also revisit integrityCheckTimeout (30 min), which is now the number a large store meets first. OPEN — WATCHING — When the WARN fires: measure the curve on that store, then choose between a rotation (n/m), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. CC
R-409 Backup & restore P3 Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it. MEASURED on demo-hp 2026-08-31 against kimai's restored unit: manifest.json's checksums object carries sha256 for .felhom.yml (2 235 B), app.yaml (488 B) and docker-compose.yml (2 195 B) — 4 918 bytes of a 213 231 242-byte unit. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. And nothing else supplies one: restic 0.14.0's restore --verify is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and restic ls --json file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and no content hash. So "the restore produced correct files" is currently unanswerable by any automated means. What is NOT claimed here: restic check --read-data-subset=100% already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. OPEN — MEDIUM R-87, R-361 Cheapest fix, and it is already half-built: extend the capture's checksums to cover db_dumps and volume_dumps — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: audits/SPIKE-restic-restore-test-2026-08-31.md §Q2, §Q3. CC
R-412 Backup & restore P3 A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success. CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it. As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". That is wrong, and measuring it overnight is what showed it: the off-site run has its OWN pre-push dump leg — "Stopping calibre-web for safe volume dump", "Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB" — so a unit that is hollow when a run starts is REPAIRED before it is pushed. Proven twice: opengist (2026-08-31 21:0x) and calibre-web (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (6fee3b5a) holds the volume tar and all 17 userdata files. WHAT REMAINS REAL, and it is narrower: the one hollow snapshot that DID reach the store (35ba9fe7, opengist) was created when the unit was destroyed inside a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged "backed up opengist (… 0 mandatory path(s))", a success line over a backup holding none of the app's data. That race is real, it was observed, and the success wording is wrong either way. The R-403 mirror guard holds throughout — proven live: "unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one", secondary byte-identical. NARROWED — LEG 1 CLOSED 2026-09-01 (controller v0.232.0) — LEG 2 STILL OPEN, LOW R-403, R-87, R-413 LEG 1 IS DONE: a per-app push whose unit carried no database dump and no volume tar now logs at WARN and says what it did not carry, using the existing unitIsHollow predicate. Wording only — no guard, and the capture is untouched (08 §8.2). Pinned by TestR412a_EmptyPushDoesNotReadAsAPlainSuccess, which asserts the hollow line carries the words, the SOUND line does not, and neither is at INFO; red-proofed by restoring the single unconditional line. LEG 2 IS STILL OPEN and is the remaining work on this row: whether the push should RE-READ the unit it is about to send, or whether the race window is small enough to accept. Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. Do NOT guard the capture (08 §8.2). Evidence: audits/DRILL-soak-2026-08-31/phase2-guard-interactions/ and phase5-mutated-cycle/09-what-reached-the-store.txt. CC
R-433 Backup & restore P3 A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED. MEASURED 2026-09-01 on demo-hp over the credential the box already holds, read-only, no delete verb issued. The sweep: a batched stat -c %n over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried 777,600 names of the vendor form YYYY-MM-DDTHH-MM-SS across nine full days at second granularity, and 126 alternative shapes — zero resolved. The control is what makes the zero mean anything: the identical 600-name batch with one real path appended returned it in 6 of 6 batches. The structural cause: /home (the customer data, u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev 0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a different dataset than the one holding felhom-repo. What still stands: clause (a) — the box can delete its live repository but cannot WRITE into /.zfs/snapshot — is unchanged and re-confirmed. What is now open again: the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and hub/internal/hetznerapi/hetznerapi.go has no snapshot method at all, so it needs new code regardless). NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone: whether the MAIN account can see the snapshots. No main-account credential exists in this project. The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed. audits/evidence-drill-r95-recovery-2026-09-01/ BLOCKED-ON-PROVIDER 2026-09-01. The one question that can move this is drafted and ready to send: Question 1 of documentation/runbooks/provider-questions-2026-09-01.md — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore? Neither answer leaves this row where it is: "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — and R-95 becomes urgent. Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. ⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS. Storage SHARE (a managed Nextcloud, NOT us) documents "Currently, we only support restores for the full backup ZFS snapshot to a specific point in time" (docs.hetzner.com/storage/storage-share/faq/backup-snapshot/). Storage BOX (ours) documents the opposite — "You can download individual files or entire directories as usual" (docs.hetzner.com/storage/storage-box/snapshots/). A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO. Tell them apart by the giveaways: the Share page talks about Nextcloud's data cache, a database dump and the konsoleH interface, and never mentions Storage Box. Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason. Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it. Nothing in this repository ever leaned on the Share claim — verified by grep at the time; the only vendor line we cite is the Storage Box one. RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate): the operator mailbox read through the Gmail connector ((from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01) holds no Hetzner reply — one match, our own offsite_snapshots_dropped alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. -- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628). Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box? Hetzner: "With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation", citing docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots; and separately "A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken." So the Storage BOX documentation governs, not the Storage Share FAQ - which is exactly the trap provider-questions-2026-09-01.md warned the reader about, and the answer came back on the right side of it. File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on u629488. That is now the measurement this row needs, and it is cheap. Q2 - is --append-only enforced by Hetzner, or taken from what the client sends? Hetzner: "You can use --append-only and it would look like this: command="rclone serve restic --stdio --append-only path/to/repo" <public key>", with fluix.one/blog/hetzner-restic-append-only/. So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's authorized_keys - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. This is the answer R-95 has been blocked on and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof. What is owed is a real forced-command key on u629488, a restic forget --prune through it that is REFUSED, and a backup through it that still succeeds. BLOCKED — BLOCKED-ON-PROVIDER — Question 1 of runbooks/provider-questions-2026-09-01.md — — CC + operator
R-540 Backup & restore P3 [P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills. Read from source 2026-09-16 while making off-site the default: HETZNER_POOL_BOX_ID is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, monitor/offsite.go) tells the operator it is filling but nothing says which box a new customer should land on. Needs a selection rule (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. READY — rank P3-LOW; owner: CC (hub) — — CC
R-545 Backup & restore P3 [P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository. FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: POST /backup/offbox/config configures a target and can disable it (enabled unchecked), but nothing removes it. POST /backup/offbox/reset refuses unless OffboxOrphaned() is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in data/offbox/. Why P3 and not higher: a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. Fix shape: a „Távoli cél törlése" action beside the config form that clears the target and shreds data/offbox/, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. READY — rank P3-LOW; owner: CC — — CC
R-548 Backup & restore P3 [P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever. MEASURED 2026-09-17 (chaos night) on tester-1-022354: a whole-guest backup wrote a ~29 GB source (mp0 = local-lvm:vm-9201-disk-1, 70 G provisioned, 40.58 % used, backup=1) into pve-root, which on a 32 GB system disk is 14 GB total with ~4.9 GB free. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — ~16 MB/s, i.e. under four minutes to a full / on the nested PVE. The product’s behaviour is correct and legible throughout: it failed the tier and said which one — whole_guest_backup_failed (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (target_id:"local", success:false, size_bytes:0), and the off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can never succeed, and it keeps retrying on a backoff for ever, burning I/O and risking / each time. Honest caveat: the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. Fix shape: compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: audits/evidence-chaos-night-2026-09-17/round-6.txt. -- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27 („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's local holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). audits/pg-last-six-2026-09-30/C/. READY — rank P3-LOW; owner: CC — — CC
R-552 Backup & restore P3 [P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever. FOUND 2026-09-17 by CC reviewing its own controller v0.246.0 (R-550) during live validation. The per-app notice (Manager.opInterrupted, persisted in restore-status.json) is cleared in exactly one place — BeginRestoreOp for that app (internal/backup/opstatus.go) — and removeStack (internal/api/router.go) never touches the restore record. So a household that answers „A visszaállítás megszakadt … indítsd el újra" by REMOVING the app instead of restoring it keeps a „Megszakadt visszaállítás" card about an app that no longer exists. Measured shape, not hypothetical: on 9201 the notice cleared only when homebox was restored again (08:54:50Z, card count 0) — the teardown deliberately took that path before removing it. Fix shape: removeStack clears the app's notice (a ClearInterruptedRestore(stack) beside the existing update-hold clear, R-491's precedent), with a wiring test. Not fixed in v0.246.0: found after the release was built; one release per repo per session. READY - rank P3-LOW; owner: CC — — CC
R-645 Backup & restore P3 [P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten. MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; --clear-restore-hold docmost + the controller restart it requires ran at ~07:59:23Z, and at 07:59:26Z the controller logged Recovery unit captured for docmost — the unit's compose/docker-compose.yml now named docmost/docmost:0.96.0, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (isHeld, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. Who it hits: an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt (the invalid run) and README §"Three things". -- 2026-09-23 (controller v0.263.0): the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. OPEN — P3; owner: CC — — CC
R-675 Backup & restore P3 [P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore. missingFileLegsRefusal predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. READY — P3; owner: CC (controller) — — CC
R-698 Backup & restore P3 [P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start. RecoveryManifest.image_pins ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running ref@digest, and a restore brings the data back AT ITS OWN VERSION (07 §6.6) — so a restore asks for exactly the old image. Measured 2026-09-26 (audits/version-travel-2026-09-26/A7/, registry HEADs, no pulls): the catalog's 42 ladder ref@digest pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. Options (decide nothing yet): (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) docker save into the unit — hundreds of MB per app per copy on every tier. -- 2026-09-30 late (decision 53): a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. OPEN — P3; owner: operator (a decision), CC measures — — CC + operator
R-706 Backup & restore P3 [P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive. Measured 2026-09-28 on demo-hp: after a full off-site restore of nextcloud (which leaves the downloaded copy in backups/offsite-restore/nextcloud, ~1 GB, by design, for the household to inspect), POST /api/stacks/nextcloud/remove with remove_backups: true removed the unit and listed backup_paths_removed WITHOUT the verification copy; it stayed until the restore page's own delete (POST /backup/offbox/verify-copy/delete, 302 scratch_deleted). A household that removes an app to free space keeps 1 GB it cannot see on the app list. Fix direction: the removal with backups also deletes the app's verification copy (the same DeleteOffsiteRestoreCopy). Evidence: audits/kept-offsite-2026-09-28/E/E9-teardown.txt. -- 2026-09-28 evening: FIXED in controller v0.279.0 — a removal with its backups also deletes the verification copy and lists it among the removed paths (TestR706_…, red-proofed RP7). Not yet seen live (needs a full off-site restore then a removal). WATCHING — P3; owner: CC (live) — — CC
R-729 Backup & restore P3 [P3-LOW] An off-site target, once saved on the page, cannot be removed through the product. MEASURED 2026-09-30 on 9202: /backup/offbox/config refuses an empty address and no route clears the target; the session removed its throwaway target from settings.json by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. Fix direction: a „Távoli mentési cél törlése" press that clears the target (never the repository). READY — rank P3-LOW; owner: CC (controller) — — CC
R-10 Backup & restore P4 T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. T-6E-1, confirmed in CAMPAIGN-6E. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — One-line hardening; batch with the next controller task CC
R-91 Backup & restore P4 Old 13 GB datastore copy at /srv/pbs-felhom on ep0's root disk WATCHING demo-felhom's first post-migration PBS backup Delete once it lands; fix CONTEXT.md:1018 same commit CC
R-99 Backup & restore P4 Server-side prune never removes a phantom snapshot. Confirmed it does NOT count them toward keep-last (dry-run kept 2 real + the phantom) so there is no retention/data-loss bug — but one accumulates per aborted upload, forever READY (S) — Decide a cleanup path. Deletion on a customer datastore is a separate ruling — detection shipped, removal deliberately not automated CC
R-104 Backup & restore P4 An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach. resticStep has unlock --remove-all (internal/backup/offbox.go:634-648) but ensureOffboxRepo's probe fails first, classifyResticProbe (:77-93) has no lock case → "other" → fail-fast; ClassifyOffsiteFailure likewise, so the operator is told „A távoli mentés ismeretlen okból nem sikerült" for a precisely-known, self-healable condition MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state READY — 2026-07-28. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away. The self-heal this row calls unreachable was built: resticStep escalates to unlock --remove-all and retries once (internal/backup/offbox.go:763-768), and unlockStale runs before every off-site run and restore (:1274, offbox_restore.go:261). Its premise that the probe fails first is also doubtful: the probe is restic cat config, a read that takes no lock. What REMAINS true: ClassifyOffsiteFailure (offbox.go:179-193) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — Was C9-F3. Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs restic unlock --remove-all. Flips: the offsite row in map §C; 07 §8 row 15 CC
R-124 Backup & restore P4 The recipe spells PBS's root namespace "root", but the PBS API spells it "" and no namespace is literally named root — an operator pasting the field into pct restore --ns root gets a failure READY (XS) — Pre-existing wire convention (ToHub has normalised empty→"root" since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at hub.PBSRootNamespace. Affects only a box with no namespace line — no real customer today, all three are per-customer. Fix = emit "" + rely on namespace_state, or emit a --ns-ready form CC
R-164 Backup & restore P4 C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists. The unit carries both a volume tar and a SQL dump; the restore uses both — the dump is authoritative and replayed after the tar so it WINS (F17), with only the DB service up (R-47) — internal/backup/restore_unit.go:262-266. Dropping the DB container's tar would halve DB-app units and close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ POSTGRES_PASSWORD ignored). BLOCKED — on the predicate a dump-validity predicate that is not accounts has rows The obvious gate is DEAD, measured: ValidateDump warns when the accounts table is empty, and that warning was correct — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But a fresh appliance legitimately has zero accounts, so promoting that predicate to a gate would block every new customer's first backup. Order: (1) a sound predicate — dump vs live per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. Until (1), the tar is load-bearing — not because dumps are bad, but because nothing can yet prove one is good. Pairs with R-127 CC
R-213 Backup & restore P4 Putting files back in place — the half the recovery screen deliberately does not do. The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: a screen that unlocks and then offers to overwrite is two decisions wearing one button. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. The operator named its requirement: a live-versus-backup comparison — the customer must be able to see what would change before anything is overwritten OPEN — not started, deliberately the comparison design (nothing exists for it yet) Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen Operator + CC
R-246 Backup & restore P4 A leftover staleness flag has silently disabled the new recovery discriminator on demo-hp since 2026-08-04, and the flag is WRONG. Found by a read-only spike, 2026-08-08. Q1 — traced to an act, to the second: at 2026-08-04 20:15:49 the hub emitted offsite_reissued and escrow_stale in the same second — an operator Re-issue pressed during the R-201 drill, three minutes after escrow_blob_served at 20:12:40/20:12:54. That was offsite.ReissueCredentials's precautionary MarkEscrowStale call, which hub v0.95.0 REMOVED the very next day (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. Q2 — the flag is wrong, measured on both sides: the hub's blob seals restic_pw_sha256 = 8a9e33aa4da6769c…d080a, and the key the box is actually using hashes to the identical value. The blob covers the key. Q3 — nothing clears it by itself: the ONLY writer of stale_at = NULL is SaveHostEscrow's ON CONFLICT — i.e. a fresh escrow ceremony, which is the one act that would supersede the good blob. So the only exit from a false alarm is the destructive act the false alarm recommends. Q5 — the blast radius, enumerated rather than assumed: (1) GetEscrowStatusForCustomer withholds restic_pw_sha256 from the ACK; (2) a pending box can never auto-confirm, so (3) every off-site run is refused indefinitely — neither bites demo-hp, which was already escrowed and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) NEW — v0.206.0's shape (c) is inert, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. Q6 — a fresh box CANNOT reach this state: MarkEscrowStale has no production caller anywhere in the tree (census: only its own definition, two comments and two test references). The next walk cannot meet it. NOT CLEARED, deliberately — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. ✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08. One row, identity-matched on host_id and guarded on stale_at IS NOT NULL; changes() returned 1. Verified end to end, not just in the database: the hub now serves the hash again, the box recorded hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a at 11:10:19Z, and that is byte-identical to the key it is using — so shape (c) compares, matches, and correctly stays silent. The false stale warning is gone, proven with a positive control rather than an absent line: 0 escrow-confirm lines since the restart while 5 scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. (Method note: the hub pod is Alpine with no sqlite3; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.) STILL OPEN under this ID: the ruling on whether stale_at keeps a live setter at all. It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; do not leave it as a trap that only a database read can spring. READY — flag cleared; the column ruling is still owed — owner Viktor Folded R-248 2026-10-03 (the same ruling: give stale_at a visible, evidence-bearing setter, or retire it). — — operator
R-256 Backup & restore P4 C2 — „A mentéskezelő nem elérhető." names no route at all. web/offbox_handlers.go:47 and :181 flash this to the customer on the off-site backup surface. It states an internal component's unavailability in the operator's vocabulary („mentéskezelő" = the backup Manager object), gives no reason the customer can act on, and names no next step — not "try again in a few minutes", not "contact support", not a page to go to. Contrast, in the same subsystem and shipped the same week: R-252's fix reads „Meghajtók", „Meglévő meghajtó csatolása". Utána gyere vissza ide. — a route. Severity is low and stated so it is not over-ranked: the condition is a nil backup manager, which on a healthy box does not occur; this is about the copy, not a broken path. Found by the C2 sample (19 refusals on the recovery/restore/offbox surface; ~202 of the repo's 221 refusal strings were NOT examined) READY — owner Viktor — — operator
R-279 Backup & restore P4 There is no operator-triggerable off-site backup. The only route to POST /backup/offbox/run is the customer's own dashboard session; signed_jobs carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of R-177 (no operator-triggerable fill check) READY (XS) — NEW 2026-08-09 — Same shape as R-177; solve both together CC
R-336 Backup & restore P4 The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage. ep0's PBS proxy served ~85,000 requests/day — a flat 3,538/hour, every hour, from two boxes: 74,445 GET /api2/json/admin/datastore (libwww-perl, i.e. PVE's pvestatd) and 73,171 GET /admin/datastore/felhom-offsite/status (proxmox-backup-client). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in 14 days, wedging the offsite tier for 9½ hours (audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md). The LimitNOFILE=65536 drop-in applied that morning raises the ceiling but does not fix the leak — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second READY (M) — NEW 2026-08-18 — CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist. This cell used to read "PVE storage status is the prime suspect, and its interval is tunable". The first half is right and the second half is false. pvestatd stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (pvesm set <id> --disable 1) around the backup window, and that is substantially more than a tuning knob: it collides with felhom-agent/internal/pbsdr/manager.go's health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. Doc-only correction — no agent code was changed. The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. Baseline measured 2026-08-18, and the FIRST measurement published was WRONG. The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was one descriptor — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h) — ~2.6x the published figure, putting the runway to the 65536 ceiling at ~357 days, not the ~2 years first claimed. And the named mechanism is the minority one: across that window CLOSE-WAIT held flat at 1 while ESTAB grew 45→49 — all the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. The fix must target connections the proxy never reaps, not just CLOSE-WAIT sockets. The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it. audits/SPIKE-ep0-established-connections-2026-08-20.md. The leak is OURS, and the poll rate is not what feeds it. Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by felhom-agent on the boxes — 194 on each, ss -tnp naming a single PID per box, and zero held by pvestatd or proxmox-backup-client. Confirmed independently from ep0's access log over the same 46.18 h window: libwww-perl (pvestatd) 81,192 requests -> 0 descriptors, proxmox-backup-client 80,061 requests -> 0 descriptors, Go-http-client/1.1 (the agent) 387 /snapshots calls -> 388 sockets — one per call, within one. So 162,404 requests, 99.5% of the traffic, produce 0% of the leak. Mechanism, named from source: felhom-agent/internal/pbs/client.go:56-60 builds &http.Transport{TLSClientConfig: tlsCfg} — a composite literal, so IdleConnTimeout is the zero value = no limit (http.DefaultTransport sets 90 s; a literal does not inherit it) — and cmd/felhom-agent/main.go:1486 (pbsTargetsFromPVE) builds a fresh client every cycle, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; CloseIdleConnections/IdleConnTimeout/MaxIdleConns appear nowhere in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h DefaultVerifyCadence (7.7 cycles) = 192.4 predicted vs 194 observed per box. CONSEQUENCE — RE-RANK. The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") would have produced a null result and read as a failed fix. Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and zero descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a scaling/cost item, not the leak fix, and the leak fix is R-344. Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED — Phase C is held at STOP 1 with its prediction pre-registered in evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt. Do not record a proportionality verdict here until that window has run. RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX. That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing. Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — ours, in felhom-agent, R-344 — would have been further from being found, not closer. Measured 2026-08-20: pvestatd (libwww-perl) and proxmox-backup-client made 162,404 requests in a 46 h window and leaked zero descriptors; the agent made 811 and leaked 388. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its baseline of 17, with the poll rate completely unchanged — 85,000/day before and after. WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits: ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — pvestatd stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with felhom-agent/internal/pbsdr/manager.go's health model. NEW ACCEPTANCE CRITERION, since the old one is void: the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. CC
R-365 Backup & restore P4 An overdue abandonment countdown renders its past due-date in the future tense. With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet 2026-08-20 napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). OPEN — LOW — Say "due, will run at the next daily sweep" once the date has passed. CC
R-367 Backup & restore P4 The database dumps already written under the wrong name are stranded, and nothing will ever collect them. R-355's fix sends paperless-ngx's dump to the right place from now on; it does not move the ones already written. On demo-hp that is /mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). Nothing deletes them and that is by design, not by luck: the F5 stale-primary prune (backup.go:1248) skips any directory whose name is not a deployed app, under the guard "an undeployed app's last backup is still its restore point" — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. They CAN be adopted, by hand: move the file to …/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql and it becomes a readable restore point for that app. It is deliberately not automatic. The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. Filed rather than done, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. OPEN — LOW follows R-355 Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. Viktor rules, CC executes
R-375 Backup & restore P4 A PBS datastore signal was noted and explicitly not filed. audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:171: "Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here." Recorded so the note has a number and stops depending on someone re-reading that report. Age when filed: 4 days. OPEN — LOW — Confirm the cause on ep0 the next time it is touched; it is a read-only check. CC
R-526 Backup & restore P4 [P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too. MEASURED 2026-09-15 from source: tenantsync.Deprovision „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build) Re-ranked 2026-10-03: P3->P4: operator teardown op on a protected box; adopt path covers the need. — — operator
R-541 Backup & restore P4 [P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated. Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (shared already provisioned for tester-1 (subaccount 311327)), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. Needs: a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. READY — rank P3-LOW; owner: CC (hub) — design first Re-ranked 2026-10-03: P3->P4: needs a second pool box or an outgrown customer first; operator-only and far off (0.3% full). — — CC
R-570 Backup & restore P4 [P3-LOW] The off-site stale-note display still has a Hungarian-text fallback, for boxes that have not run off-site since v0.251.0. OPENED 2026-09-17 by R-553's fix: offboxWarningDisplay (controller/internal/web/handlers.go) decides on LastWarningKind, but a box upgraded to 0.251.0 carries the PERSISTED old sentence with no kind until its next off-site run rewrites it, so the substring test survives under kind == "". Close when every fleet box has completed one off-site run on ≥ 0.251.0 (the hub's reports carry the controller version; the off-site anchor is offbox.last_success), then delete the fallback, its constant and its legacy test rows. Hard dependency: localisation slice 2 (R-557) must NOT translate the producer "Sikeres — nincs mentésre jelölt alkalmazás" (controller/internal/backup/offbox.go) until this row closes — translating it while the fallback is load-bearing strands exactly those boxes. WATCHING - rank P3-LOW; owner: operator (the fleet condition), CC (the deletion) Re-ranked 2026-10-03: P3->P4: cleanup waiting on a fleet condition; no defect today. — — operator
R-691 Backup & restore P4 [P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build. (1) The read-only file-browser view cannot open a folder another user owns with mode 0770 — nextcloud's appdata/nextcloud is www-data drwxrwx--- (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) „Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2); an app whose only database copy is off-site gets "no backup". Controller 43e99d1. audits/night-2026-09-26/E/ -- 2026-09-25 live proof: (1) confirmed on 9202 — the view mounts nextcloud's kept folders :ro but its files are www-data 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). -- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay :ro, nothing on disk changes (CC-unattended decision, 07 §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, audits/version-travel-2026-09-26/D3/. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy. -- 2026-09-27 (second session): (2) NOT built on purpose — it composes the unit-only off-site download (RestoreOffboxScratch(full=false)) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** -- 2026-09-28 (controller v0.277.0): (2) BUILT — KeptBestCopy offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; LoadKeptOffsite downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (07 §6.5/§6.6), then restores. Red-proofed (audits/kept-offsite-2026-09-28/redproofs/). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. STILL OPEN: the live off-site proof — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). -- 2026-09-28 afternoon: LIVE-PROVEN on demo-hp 9201 (controller 0.278.0), endpoint level. A throwaway nextcloud, seeded through its own front door, joined the off-site copy; the off-site run-now pushed snapshot 6cb379a8. (a) The full off-site restore (prepare → download → reconstitute): 3 volumes + the database replayed, the seed read back, a marker user written after the snapshot read ABSENT, same versions. (b) Remove keeping the data + deleting the local copies → reinstall: the choice and the kept list both named "távoli mentés, 2026-09-28 15:40" / "the off-site copy, 2026-09-28 15:40"; "use my kept data" downloaded the unit alone and loaded 3/3 volumes + 1/1 database in 55 s; the seed and the kept files read back. App removed with its data; demo-hp's app list equals the list before. audits/kept-offsite-2026-09-28/E/. NARROWED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the second-drive (Tier 2) path of Use/Load is not proven live — 9202 has one drive) — CLOSED — controller v0.277.0, live 2026-09-28 — — CC
R-815 Backup & restore P4 First-ever GC on felhom-offsite (armed today 13:11 UTC, never run) VERIFY (2026-10-03 triage: a July watch row with no id; given R-815. WATCHING — no completion record found; schedule sun 04:30 still present 2026-09-30.) — WATCHING schedule Sun 2026-08-02 04:30 UTC — confirm it completes CC
R-816 Backup & restore P4 No off-site failure class has ever been seen live. F-DIAG (controller v0.182.0, 2026-07-28) split off-site failures into six causes — quota, orphaned, no_repo, no_units, transport, unknown — each with its own Hungarian message. None of the six has been exercised by a real failure on a box; the recovery inventory records it only as a known limit (documentation/architecture/_recovery-inventory-2026-07-28.md:955). Filed 2026-10-03 from the F-DIAG row's residue when that row moved to CLOSED-ITEMS.md. READY — filed 2026-10-03 (triage); owner: CC. Exercise each class once on a scratch guest (a full quota, a missing repository, a blocked transport) and read the message the household sees. — — CC
R-832 Backup & restore P4 ep0's copy in a place outside both Hetzner and the operator's home (roadmap). Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. DEFERRED — later, if the product grows — — operator
R-878 Backup & restore P4 A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer. MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of health: starting; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. audits/catchup-2026-10-05/partA/live-demo-felhom.txt READY — owner: CC — measure a large volume first CC

Storage & devices — 11 rows (P3 7, P4 4)

ID Category Sev What State Blocked on Next action Owner
R-118 Storage & devices P3 An absent drive's union row advertises the ROOT filesystem's capacity as its own. In the absent-state payload the registry-union row reports total_bytes: 49675956224 / used_bytes: 4584579072 — byte-identical to the local row (durable_id: path:/var/lib/vz, i.e. pve-root) in the same response. The real drive is 4 GB READY (XS) — NEW 2026-07-30 — Cause: statfsCapacity(d.MountPath) (disks.go:335-338) statfs's /mnt/cel, which with the device gone is a bare directory on the root filesystem. observe.go:176-183's comment warns about exactly this trap and guards the Observe path ("an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id"); the union path has no equivalent guard. Not a DR mis-id — durable_id on that row is still the correct uuid:…, so re-attach identity is safe. It is a false capacity reaching every consumer of total_bytes/used_fraction (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as role.go:180-181 — an absent drive's fields decaying to the root filesystem's. Evidence: audits/DIAG-r116-disks-payload-2026-07-30.md §12 CC
R-298 Storage & devices P3 The /storage page's unregistered list is filtered by role==='user-data', so a drive that is also the backup target can never be registered from it. storage.html:363 routes anything not user-data into the read-only protected group with NO actions. On the rebuilt demo-hp the NVMe is deliberately BOTH the user-data drive and the felhom-backup target (/etc/pve/storage.cfg: dir: felhom-backup → /mnt/nvme-1tb), so it renders locked. This is the SECOND reason that page was empty during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action READY (S) — NEW 2026-08-10 R-280 Split the role gate: user-data keeps destructive actions; any mounted role may be REGISTERED CC
R-330 Storage & devices P3 Disk health Phase 2 — the three SMART attributes the wire does not carry. The failing drive's most telling counter was 187 Reported_Uncorrect, sitting at normalized 1 against threshold 0 with a raw count of 1001 — one point from failing and structurally unable to get there. Also wanted: 199 UDMA_CRC_Error_Count (cabling) and 188 Command_Timeout. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART samples (the right home is metrics.MetricsStore, NOT the Phase-1 state file, which is one record per disk and must stay that way). This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change READY (M) — NEW 2026-08-14 R-328 (closed) Add 187/199/188 to the agent's SmartSummary + hub model in one session; then persist samples CC
R-332 Storage & devices P3 The new Hiba-from-counters path has never fired on real hardware. v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the negative — three healthy disks still read Rendben across the deploy, no false alert — and the severity wire end to end, but no live disk has actually reached Hiba. This is the honest gap and it must not be closed by pointing at the fixture tests: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy WATCHING — NEW 2026-08-14, NARROWED same day. One item originally in this gap is now PROVEN LIVE: the persisted state surviving a controller restart. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a changed_at written by the PREVIOUS version (2026-08-14T07:23:14.640216851Z, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting a real degrading disk, or an injection harness Closing condition: a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent /disks → controller check → hub event), not a hand-set verdict CC
R-542 Storage & devices P3 [P3-LOW] /api/disks/candidates offers a REGISTERED, in-use drive under „initialize". MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after /dev/sdb was formatted, mounted at /mnt/felhom-drives/adatlemez and registered as the default data drive, the endpoint still listed it under initialize (and again under attach with already_mounted: null). Customer-invisible today: the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. It also misleads a session: I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt). Fix shape: exclude paths the controller has registered from initialize, and set already_mounted from the real mount state rather than null. READY — rank P3-LOW; owner: CC (controller/agent) — — CC
R-700 Storage & devices P3 [P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder. Found 2026-09-27 reading the code for R-697 (not seen on a box): doFlipRedeploy (the per-app and whole-drive move) persisted through the restore's fresh app.yaml write, which drops pinned_images, desired_state, installed_images, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (sync.renderSource's table) and the next up runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. -- 2026-09-27 (controller v0.276.0): FIXED — persistDriveFlip changes HDD_PATH and nothing else; red-proofed (audits/records-carried-2026-09-27/redproofs/RP3, RP4). STILL OPEN: the live proof — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading pinned_images before and after and the running image after the next sync. WATCHING — P2; owner: CC (live proof) Re-ranked 2026-10-03: P2→P3: the fix shipped and is red-proofed; only a live proof on a two-drive box is left. — — CC
R-756 Storage & devices P3 [P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder. MEASURED twice on 2026-10-01 (audits/lockouts-2026-10-01/B/B1…, audits/calibre-name-and-prune-2026-10-01/A/A1…): POST /api/stacks/calibre-web/remove with remove_hdd_data → 409; docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. -- 2026-10-01 (night, new apps): the same 409 for Grimmory (remove_hdd_data), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged drive … not mounted — skipping ensure (held by drive gate); the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (audits/new-apps-2026-10-01/box/grimmory/restore-why.txt). OPEN — rank P3-LOW; owner: CC — — CC
R-25 Storage & devices P4 Device-node TOCTOU hardening (drive init). Graduate the controller v0.141.0 Observation: the format → resolveEnrollUUID(path) → AssignDisk(uuid) sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — From the v0.141.0 F6 commit's security-review finding (felhom-controller REPORT). Low real risk (single-operator, agent-guarded), but cheap to close CC
R-331 Storage & devices P4 Disk health Phase 3 — growth-rate detection, and retiring the static 64. The v0.215.0 count backstop (64 unreadable sectors → Hiba) is a judgement from ONE drive: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — is this count climbing, and how fast — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists READY (M) — NEW 2026-08-14 R-330 Growth-rate rule over persisted samples; re-derive or delete the static 64 CC
R-352 Storage & devices P4 Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data. Measured on demo-hp 2026-08-21. (1) 40 of 53 catalogue templates declare no data path (grep -rl 'env_var: HDD_PATH' --include='.felhom.yml' → 13; total 53); those apps get no storage field and no default — their data lands in a named Docker volume on the system drive. (2) GetDefaultStoragePath() has exactly three non-test callers — the metrics collector (cmd/controller/main.go:410), the dashboard SystemInfo panel (web/server.go:733) and .fab import landing (handler_export_upload.go:154). The deploy route never reads it. Its field comment // new apps use this by default (internal/settings/settings.go:453) has never been true — an invariant with no test pinning it. (3) The first-tier backup follows the data onto the same disk (backup/backup.go:324-334 → systemDataPath), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at tier2.go:329. (4) „1 alkalmazás használja" on the Drives page counts only Env["HDD_PATH"] == path (web/handlers.go:2118), so it can never include the 40-class; it truthfully means „1 of the apps that CAN use a drive does". NARROWED — PARTLY CLOSED 2026-08-21 — visibility shipped; placement OPEN — Shipped tonight (visibility only, no placement change, nothing migrated): the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. ⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT. (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places hot data (DB/config/cache) on fast storage inside the guest and states that placement is ENFORCED (documentation/architecture/01-topology-and-trust.md:150-152). The 40 are all-hot apps; the 13 are the ones with bulk content, which belongs on an attached drive. There is no choice being denied. (3) overstated one risk and understated a distinction. Since R-165 the guest carries a small OS rootfs plus ONE data volume at /var/lib/felhom; /var/lib/docker and /mnt/sys_drive are two binds of that same volume (felhom-agent/configs/build-golden.sh:29-40, 99) — the mp0/mp1 split assumed here was retired 2026-08-03. Real risk: a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. Overstated risk: a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (00-capability-map.md:94), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. The comparison to Tier 2's same-disk refusal (tier2.go:329) is withdrawn: Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. (2) and (4) are untouched and remain correct — (2) is now filed on its own as R-368 with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. The specification for the rest is filed at documentation/backlog/SPEC-app-data-placement-2026-08-21.md (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, IsDefault must become true or go away with a test, and the Drives-page count). An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN — it assumed the customer had failed to choose; they had no choice to make. Viktor rules, CC executes
R-568 Storage & devices P4 [P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart. MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: /dashboard fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt). diskHealthRows (disk_health.go L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. Fix shape: sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: cosmetic. — — CC

Security & access — 31 rows (P2 2, P3 26, P4 3)

ID Category Sev What State Blocked on Next action Owner
R-777 Security & access P2 [P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped. READ in source (audits/visitors-2026-10-01/A/sweep/sweep-1.md, sweep-2.md), not measured live: Jellyfin with KnownProxies empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. Needs: measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: KnownProxies 172.16.0.0/12 in network.xml (no env — an after_install or a seed file); Emby: no setting fixes it (its LocalNetworkSubnets still counts private ranges) — a page sentence or a decision. READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route) — — CC + operator
R-861 Security & access P2 The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (03 §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary. READ 2026-10-04 from felhom-agent/configs/felhom-agent.sudoers (not exploited): FELHOM_GUESTHOOK installs /tmp/felhom-guest-hook-*.sh as a hookscript Proxmox runs as root at guest start, and pct reboot is granted; FELHOM_INTERMEDIARY installs a script + a systemd unit that run as root at boot; FELHOM_ESCROW runs /usr/local/bin/felhom-agent as root, and FELHOM_SELFUPDATE apply accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like felhom-os-apply; delivered by the config bundle. 11 §5.4.2, 03 §11. NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: sudo -l 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design 03 §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline). — the operator decides whether (a)–(c) are accepted or need work before the first paying customer CC
R-126 Security & access P3 A .fab bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS. storageDriveList() (internal/web/handler_export.go) does not filter network paths READY (S) — Split out of R-108, which closed without it: this is an explicit customer-chosen export destination, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (07 §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share CC
R-132 Security & access P3 curl -w '%{redirect_url}' reconstructs the request URL WITH its basic-auth credential — so a -u ":$HUB_PW" call that never put the password in a URL still printed it WAITING-ON-OPERATOR — ACTION: rotate HUB_PW Folded R-580 2026-10-03 (the same curl -w %{redirect_url} credential echo, seen again 2026-09-18). — Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. -u is safe; the reporting was not. Rule: read the redirect from -D - and grep ^Location:, never %{redirect_url}, on any authenticated call. Rotate the hub password (/configuration → Login password; ConfigMap auth.password_hash is the reset path) and update ~/.config/credentials Viktor
R-136 Security & access P3 Rename hub_session → __Host-hub_session — makes cookie tossing structurally impossible READY (XS, one line) — Verified on the live production response that all three prefix preconditions already hold: Path=/, Secure, no Domain. Caveat for the ticket: browsers reject a __Host- cookie without Secure, and isSecure is conditional on r.TLS/X-Forwarded-Proto, so plain-HTTP browser access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: r.Cookie returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 CC
R-137 Security & access P3 Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults. globalRuleDesc = "[felhom-geo] Global" (waf.go:18) is one literal description per ZONE; appRuleDescPrefix keys by app name with no customer (waf.go:21); BuildGlobalExpression has no positive hostname scoping (waf.go:241); applyDiff deletes every [felhom-geo] rule not in THIS box's desired set (geosync.go:320) READY (M) — blocks shared-zone onboarding — With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's RemoveGeoRules) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by customer_id + add http.host ends_with "<domain>" to both expressions — a TWO-REPO change (controller + hub RemoveGeoRules). Same audit §5.1 CC
R-138 Security & access P3 A shared-zone cf_api_token is a zone-wide DNS-write capability on a customer's box — written 0600 to /opt/docker/stacks/traefik/.env (controller/internal/infra/infra.go:123) READY (S) — Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (traefik.yml.tmpl) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 CC
R-255 Security & access P3 The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped. Filed 2026-08-08 while closing R-254, because a partial guard reported as complete is worse than no guard — it stops the next person looking. Two nets, both measured. (1) scripts/secret_in_markup_gate.py reads all 36 templates and convicts any {{ … }} naming a secret unless allowlisted with a reason. It catches {{.RetrievalPassword}} and {{.InitialCreds.Password}}, and it catches a launder through a local variable because the assignment itself names the secret ({{$v := .InitialCreds.Password}} is convicted — verified). It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY — data["Tagline"] = creds.Password then {{.AppInfo.Tagline}} passes it cleanly, also verified. That is exactly the shape of R-254 site two (value="{{$val}}" inside an {{if eq .Type "secret"}} branch), so the gate would not have caught one of the three instances it was written for. (2) The runtime body assertion — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and only 4 of 27 page templates have that today: settings_security, app_info, deploy, backups_restore — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. The other 23 pages have no runtime coverage at all. What closing this needs, so the cost is not re-estimated: a per-page data fixture for the remaining 23 (most need a wired Server — stackMgr, backupMgr, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. That is real scaffolding, which is why it was NOT built inside R-254's session rather than half-built and declared done. READY — owner Viktor — — operator
R-269 Security & access P3 A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER. localapi.TokenStore.Mint documents "last-write wins — any previous token for this guest is revoked". Across processes that is FALSE until something unrelated forces a reload: the long-lived agent serves Lookup from an in-memory index and re-reads the store only on a miss (the B3 reload-on-miss optimisation, tokenstore.go), so a superseded token is a direct map hit and returns (vmid, true). Red-proved twice. (a) A unit probe — TestTokenStore_ReloadOnMiss_RemintCoherence with the two lookups swapped, i.e. present the rotated-out token FIRST — fails; the shipped test passes only because it looks up the NEW token first, and that miss is what evicts the old hash. (b) On hardware, 2026-08-09: after the on-disk rotation the old token returned HTTP 200, then 401 only once a new-token lookup had forced the reload, and reliably 401 after systemctl restart felhom-agent. This is the CLAUDE.md invariant-comment case exactly — the comment reads as settled and the test that looks like its pin is order-dependent. Fix options: pin the reversed order with a test, or make eviction not depend on an unrelated miss. Until then, an operator rotating a leaked token MUST restart the agent — the runbook step is not optional READY (S) — NEW 2026-08-09 — Found by doing R-268's rotation rather than reading about it CC
R-270 Security & access P3 R-268's stated rotation recipe is incomplete: the controller never re-reads bootstrap.json's local_api, so a rotation leaves the agent channel dead across restarts. bootstrap.ensureLocalAPI returns early when cfg.LocalAPI.Endpoint != "" — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own controller.yaml, not in the mount. Proved live 2026-08-09: two controller restarts after a correct bootstrap.json rotation, still HTTP 401; the channel came up only once local_api.token was written into controller.yaml. The neighbouring DetectEndpointDrift compares the ENDPOINT and deliberately does not compare the token ("a token mismatch is a different failure"), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is R-78 READY (S) — NEW 2026-08-09 — Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way CC
R-275 Security & access P3 --uninstall leaves five 0600 agent.json.* credential backups, and the reinstall hands them to the new service account. /etc/felhom-agent/ survives with agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}, each carrying a 64-char hub.api_key and a 59-char proxmox.token. The teardown claims to remove "config (+ its .bak backups)" and scripts/CHANGELOG F1 records "uninstall now purges the agent config's .bak* siblings (one held a live hub api_key)" — that fix does not match the filenames in use, and it misses agent.json.pre-prunegate.bak, a file that literally ends in .bak. Exposure assessed, not assumed: these are SUPERSEDED — the orphaned key hashes to a5d2222a…, the hub's current demo-hp key to 8c59d1b6…, and the Proxmox token was deleted by the same uninstall. But the reinstall recreates felhom-agent at uid 999, the same uid the deleted account had, so three of the backups become the new account's files — verified readable as felhom-agent. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). Also left, undeclared: /etc/felhom/{.bootstrap-done,appliance-pairing-code}, felhom-bootstrap.service + /usr/local/sbin/felhom-bootstrap.sh, the vmbr9 stanza in /etc/network/interfaces, and /etc/sudoers.d/felhom-agent.bak-pre-e2a (21 KB — INERT: sudo skips dotted filenames, verified with sudo -l -U felhom-agent; visudo -c -f parsing it OK is NOT evidence sudo loads it) READY (S) — NEW 2026-08-09 — Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets CC
R-276 Security & access P3 RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing. After --uninstall on demo-hp, wg-quick@wg-felhom is enabled and active, /etc/wireguard/wg-felhom.conf present, handshake to 167.233.158.164:443 52 s old, counters 5.86 GiB in / 2.48 GiB sent. It appears in neither the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in ("an OUTBOUND WireGuard tunnel to the Felhom hub"). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told READY (S) — NEW 2026-08-09 — Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command CC
R-338 Security & access P3 demo-hp is not on the R-50 island at all, and operations/nodes.md states that it is. The page records both fleet boxes as island-migrated 2026-07-25. True of felhom-pve; false of demo-hp, whose agent.json has listen_addr: 192.168.0.87:8443 — the customer LAN address — and no island_bridge/island_guest_addr keys at all, whose guest 9201 has net0 only (no eth1), and whose vmbr9 exists with zero members. The controller's controller.yaml points at the LAN address, so the box works; this is inventory drift, not breakage. Two costs. A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is bound to the customer LAN on this box rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it READY (S) — NEW 2026-08-18 — Decide which is true: migrate demo-hp to the island, or correct nodes.md. Leaving both is the one option that keeps the doc lying Viktor decides; CC executes
R-350 Security & access P3 SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended. What happened: vouching the artifact manifest used curl -w '%{redirect_url}' for confirmation. The hub answers POST /configuration/artifacts with a 303, and curl renders the redirect target with the basic-auth credentials re-attached — so the URL it printed contained http://:<HUB_PW>@10.43.52.34:8080/configuration?flash=artifacts_set. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only ${#HUB_PW}; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. Blast radius, stated precisely rather than minimised: the value is not in git, not in CHANGELOG.md/REPORT*.md/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under ~/.claude/projects/ on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file. READY (S) — NEW 2026-08-20 — Operator decides whether to rotate. The hub's own /configuration password form does it (current_password/new_password/confirm_password), and per hub-password-ui-2026-07-13 the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation file-to-file without printing the new value (the operator-present-one-time-secrets convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. The reusable half, which matters more than this one password: never use curl's %{redirect_url} (or -v, or --libcurl) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with %{http_code} and read the flash from a follow-up GET. Viktor decides, CC executes
R-584 Security & access P3 [P2-MED] Credential-bearing probe scripts were left in a live guest's /tmp for hours, across three releases. FOUND 2026-09-18 while cleaning up after slice 3's live proof: probeB.sh, probeB2.sh, probeB3.sh, probeC.sh and probeD.sh were still in demo-hp guest 9201's /tmp from slice 2's releases B, C and D earlier the same session, and all five carried the controller password INLINE (-d "password=$PW" with the value substituted). The standing rule in .claude/rules/ui-hungarian.md says exactly this: delete any credential-bearing helper from /tmp (host AND guest) when done. They were not deleted at the end of each phase; they were found by grepping my own litter for secret-shaped strings before removing it, which is a check that happened only because a later phase went looking. All five are shredded. Why it is a row and not a note: the rule existed, was loaded, and was still not followed three times running — so the rule is not the mechanism. The password is the shared demo one (Tier-0 boxes, ~/.config/credentials), unchanged; rotating it is cheap and is the operator's call. Fix shape: never inline a credential in a helper — write it to a mode-0600 file and read it with curl's @file form, which this run did and which is why this run's own scripts were clean; and make the last act of a phase that pushed a script to a box the shred -u of it, the same way R-320 made evidence-copying the last act of a phase. READY - rank P2-MED; owner: CC (the mechanism), operator (whether to rotate) Re-ranked 2026-10-03: P2->P3: demo boxes and a shared demo password only; files shredded; the fix is process. — — CC + operator
R-587 Security & access P3 [P3-LOW] Two root-password files sit in the directory the public ISO is published FROM. FOUND 2026-09-18 running the ISO release gate's credential criteria for slice 4: /mnt/5_hdd/felhom.eu/felhom-iso/out/ holds felhom-pve-9.2-1-v1.24.0-nested-probe-generic.iso.rootpw.txt and felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso.rootpw.txt from 2026-07-22/23, mode 0600, left by two appliance-mode builds. They are NOT on the bucket — both https://iso.felhom.eu/<name> return 404, against a control (felhom-installer-1.28.0-pve9.2-1.iso.sha256) that returns 200, so the check is real and not a dead probe. The only thing keeping them off a PUBLIC bucket is the --include "felhom-installer-<VER>*" pattern in the publish command, and they are named felhom-pve-*, so the pattern misses them by an accident of naming rather than by design. A publish typed without the include, or an include widened to felhom-*, uploads root passwords to a world-readable bucket. Fix shape: shred the two files (they are three months old and their VMs are long gone), and make the release build refuse to run — or the publish step refuse to start — while any *.rootpw.txt exists in the out directory. A pattern that protects by coincidence is not a control. The two files were SHREDDED 2026-09-18 from the publish source directory, immediately before the 1.29.0 upload ran from it. The directory now holds no *.rootpw.txt. The mechanism half is still open: nothing stops the next appliance build leaving one there, and nothing refuses a publish while one exists — the --include pattern still protects by coincidence. OPEN — PARTLY DONE 2026-09-18 (files gone; the guard is not built) - rank P3-LOW; owner: CC — — CC
R-616 Security & access P3 [P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary git remote -v. FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. Syncer.buildRepoURL injects username:token into the HTTPS URL, and git clone persists that URL as the clone's origin, so <data>/catalog-cache/.git/config holds the token in the clear and any diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (curl -w '%{redirect_url}'). maskRepoURL exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. INERT ON THE FLEET TODAY — the live catalog is public and git.token is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs git remote -v prints it. Fix shape: store the remote WITHOUT credentials and supply them per-fetch (a credential helper, http.extraHeader, or GIT_ASKPASS), and a test asserting the clone's stored origin contains no @. Operator action from tonight, unrelated to the fix: the Gitea admin token used for the drill repo was printed by that command and must be rotated. Evidence: audits/update-night-2026-09-21/05-9202-follows-drill.txt (redacted). READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token) — — CC + operator
R-717 Security & access P3 [P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone. MEASURED 2026-09-29: opengist disable-signup is an admin-panel setting (no env, no CLI); wishlist system_config.enableSignup (Prisma). Their blocks are case-insensitive and refused every trick shape (audits/signup-lock-2026-09-29/B/). Fix direction: an after_setup command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). OPEN — P3; owner: CC — — CC
R-747 Security & access P3 [P3-LOW] A stranger can lock the household out of mealie with five wrong logins. MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. Needs: measure the lock's length; decide whether the page tells the household what to do. audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt -- 2026-10-01: Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for SECURITY_USER_LOCKOUT_TIME hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via POST /api/admin/users/unlock, but the household's only admin is the locked account. Both login names are public (admin, changeme@example.com). Fixed (09 §3 decision 57, decided by CC unattended — operator may reverse): SECURITY_USER_LOCKOUT_TIME=1, catalog a4597cd; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (audits/rulings-2026-10-01/C/). Left: a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC — — CC
R-763 Security & access P3 [P2-MEDIUM] On wger a stranger can make an account after the household's setup, and every anonymous visit to the dashboard creates a guest account. MEASURED 2026-10-01 on 9202 (live template 82fff32), found by checklist row 3.4: after the admin existed, a stranger with no dashboard session POST /en/user/registration → 302, signed in with it → 302, read the API → 200; two anonymous GET /en/dashboard raised the user count from 2 to 4 (wger's middleware create_temporary_user, utils/middleware.py:69). Settings read inside the app: ALLOW_REGISTRATION True, ALLOW_GUEST_USERS True (the image defaults; the template sets neither). GET /en/user/demo-entries as a stranger answered 500. wger is FIRST-ADMIN class 3 (a known default login, fixed by after_install), so it never got decision 47's sign-up lock — it was not in R-711's list. Every crawler visit adds a user row to the household's database. Needs: per decision 47, close it after the first admin: ALLOW_REGISTRATION=False and ALLOW_GUEST_USERS=False (env switches the settings read — measure that the admin can still add family members, row 3.7), proven on 9202 as a stranger. audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt -- 2026-10-01 (operator): wger is lifecycle: hidden until this and its twin are fixed (catalog 55b8c8a; read back on 9202: not on the app list, mealie control present). READY — rank P2-MEDIUM; owner: CC (catalog) Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again. — — CC
R-775 Security & access P3 [P2-MEDIUM] Grimmory: a stranger's 5 wrong sign-ins lock EVERY visitor out of the web login for 15 minutes — so Grimmory was not published. MEASURED 2026-10-01 on 9202 (drill catalog, v3.4.1, through traefik): after 5 wrong tries for admin every further sign-in answered 429 — the household's right password AND a different name — and stayed 429 for 10+ minutes of retries. Read in the jar: AuthRateLimitService — Caffeine expireAfterWrite(ofMinutes(15)), MAX_ATTEMPTS 5, keys login:ip: and login:user:; Spring forward-headers-strategy: native takes the address from X-Forwarded-For, and behind the tunnel every visitor is the tunnel container's address (R-753) — the wger shape (R-752), with no setting to change it. Everything else in the checklist passed (bench + box step v3.4.1 → v3.5.0, gate by its own probe, OPDS through traefik); two smaller findings for the publishing session: on a reinstall over the first install's kept books, a new upload was saved to the drive but not added to the library (box/grimmory/reinstall-c1.txt, not investigated); and the remove + restore round trip (2.5) cannot be shown on 9202 for a drive app — its backup lives on the scratch drive, which is not a registered drive (R-756). The template waits in audits/new-apps-2026-10-01/wip/grimmory/. Needs (operator): (A) publish with a sentence on the page that wrong guesses by others can lock the login for 15 minutes (MEASURED: during the lock an e-reader's OPDS feed still answered 200 with its own login, wrong 401 — box/grimmory/opds-under-lock.txt), or (B) wait until the box passes each visitor's real address (R-753). audits/new-apps-2026-10-01/box/grimmory/throttle.txt -- 2026-10-01 (evening): option B's precondition SHIPPED (controller v0.286.1, R-753): Grimmory's Tomcat RemoteIpValve walks from the right and counts 172.16.0.0/12 as a proxy (READ in source, Spring Boot 4.1.1 — not yet measured with Grimmory's own lock), so login:ip: becomes per visitor; login:user: still lets a stranger lock the public name admin 15 min. A third route was spiked and passed: Grimmory behind the permanent family gate with its e-reader paths excepted (R-780). Recommendation: publish behind the family gate if R-780 is built; otherwise B with a measured 3.6. UPDATE 2026-10-02 — NARROWED, Grimmory PUBLISHED behind the family gate: a stranger cannot reach Grimmory's web sign-in at all (6 tries through the simulated tunnel: the gate's 401, then the household signs in 200 — audits/family-gate-2026-10-02/A/items.txt), and the e-reader exceptions keep Grimmory's own login. The 2.5 round trip now WORKS on 9202 (audits/family-gate-2026-10-02/B/box/life.txt). What is left: a family member past the gate can still lock a NAME (the admin's) for 15 minutes with 5 wrong tries — hard-coded in Grimmory; and the reinstall-over-kept-books finding (new-apps-2026-10-01/box/grimmory/reinstall-c1.txt) is still not investigated. WATCHING — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P2→P3: Grimmory is now behind the family gate; only a family member can still lock a name. — — CC
R-776 Security & access P3 [P3-LOW] Right-walking catalog apps need ONE setting to see each visitor since v0.286 (R-753); without it they keep the tunnel's one address (a stranger can still trip their per-address limits for everyone). READ in source (audits/visitors-2026-10-01/A/sweep/): kimai TRUSTED_PROXIES=127.0.0.1,172.16.0.0/12; zipline CORE_TRUST_PROXY=true + CORE_TRUSTED_PROXIES=172.16.0.0/12; vikunja VIKUNJA_SERVICE_IPEXTRACTIONMETHOD=xff; nextcloud TRUSTED_PROXIES=172.16.0.0/12; n8n N8N_PROXY_HOPS=2 (optional). BookStack APP_PROXIES measured this session (see its catalog commit). Needs: per app, the setting + checklist 3.6 re-measured on 9202 through the simulated tunnel (the BookStack shape, audits/visitors-2026-10-01/tools/bookstack_36.py). Count readers (calibre-web, tandoor, wger) stay as they are — a fixed count is wrong for one of the two paths. READY — rank P3-LOW; owner: CC (catalog) — — CC
R-778 Security & access P3 [P3-LOW] A box that rolls back to a controller ≤ 0.285 after v0.286 keeps the new traefik (an old controller never rewrites a running traefik) — and the old clientIP believes the LEFTMOST X-Forwarded-For, which a stranger then writes: the dashboard's login counter becomes dodgeable until the box moves forward again. Reasoned from the code (old claim.go clientIP + v0.286 traefik trust), not measured. The floor never moves back; the window is the self-update's crash roll-back. Needs: decide whether that window matters (it closes at the next floor); if it does, a 0.285.x patch that reads the rightmost hop, or the self-update refusing to roll back across v0.286. OPEN — rank P3-LOW; owner: CC — — CC
R-782 Security & access P3 [P3-LOW] Two side observations of the R-753 sweep, inferred, not measured: glance's seeded glance.yml has no auth: block (the dashboard is public to anyone with the address), and homepage's /api/* refuses a Host not in HOMEPAGE_ALLOWED_HOSTS, which the template does not set (widgets may 400). Needs: measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). READY — rank P3-LOW; owner: CC (catalog) — — CC
R-783 Security & access P3 [P3-LOW] SparkyFitness: three wrong sign-ins by anyone shut EVERY visitor out of sign-in for ~10 s — a stranger retrying every 10 s keeps the household out. MEASURED 2026-10-01 on 9202 through the simulated tunnel (audits/visitors-2026-10-01/C/box/sparky-box.txt): better-auth's sign-in limit (3 per 10 s) is keyed on one address — it reads the LEFTMOST X-Forwarded-For, whose chain the R-753 router reset removes, so its frontend nginx hands it traefik's address; the household from another address got 429 at 3 s and 10 s, in at 16 s. Not forgeable (a rotating forged address did not escape). better-auth's header setting is not exposed as an env by SparkyFitness. Needs: accept (10 s), or an upstream setting for better-auth's ipAddressHeaders + a right-walking reader. OPEN — rank P3-LOW; owner: CC — — CC
R-831 Security & access P3 The Hetzner storage API token (HETZNER_TOKEN, the storage project's token in Secret/storagebox) was printed into the 2026-10-03 session transcript — CC read the gitignored manifests/storagebox.secret.yaml and its redaction pattern missed the quoted value. It can create, reset and delete Storage Box sub-accounts. Not rotated by the operator's choice (decision 73). Rotation, whenever chosen (3 steps): create a new token in the storage project in the Hetzner console → patch Secret/storagebox key HETZNER_TOKEN in felhom-system and kubectl rollout restart deployment/hub → delete the old token in the console. Rule for sessions: never print a file that holds secrets — read the one field needed. WAITING-ON-OPERATOR — rotation is his call — rotate when chosen operator
R-870 Security & access P3 Tester 1's two Cloudflare credentials — the zone API token (infrastructure.cf_api_token) and the tunnel token (infrastructure.cf_tunnel_token) of the hub's customer_configs row tester-1 — were printed into the 2026-10-04 night session's transcript (not into any file): a read-only query selected substr(config_json,1,400), and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (enkicsifelhom.hu). Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B). Rotation, whenever chosen (3 steps): in the Cloudflare dashboard create a new API token for the enkicsifelhom.hu zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → tester-1 → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (docker ps health healthy) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never config_json without json_extract of a non-secret field. WAITING-ON-OPERATOR — rotation is his call (ruled: not now) — rotate when chosen operator
R-134 Security & access P4 Two zone-resolvers disagree on depth. The controller strips labels progressively (controller/internal/cloudflare/zone.go:18); the hub's resolveZone tries the exact name then parentDomain, which strips exactly ONE label (hub/internal/cloudflare/unblock.go:115,136) READY (XS) — For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 CC
R-525 Security & access P4 [P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured. Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. What it needs: a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. READY — rank P3-LOW; owner: CC (spike) Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed. — — CC
R-779 Security & access P4 [P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box. Measured 2026-10-01 19:51 UTC (audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. Needs: one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up. — — CC + operator
R-879 Security & access P3 A copy of hub.db still holds secrets readable without any key: each box's hub API key (hosts.api_key), each household's owner passphrase and API key (customer_configs.retrieval_password, api_key), and the PBS-DR token values (host_pbs_secrets.value, kept after use). Found 2026-10-05 while answering "what does a hub database backup now contain" (R-133 sealed the console passwords; the off-site passwords were sealed by R-821). A stolen database copy lets an attacker report as any box and read every owner passphrase. 05 §16.2 READY — Seal or hash each with the same seal (the box keys and the passphrases are compared, so a hash may fit; the PBS token is served once, so it could be deleted after use) CC

Box system & updates — 15 rows (P2 1, P3 13, P4 1)

ID Category Sev What State Blocked on Next action Owner
R-812 Box system & updates P2 [P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine. SEARCHED 2026-10-03 (read-only): felhom-controller, felhom-agent, app-catalog-felhom.eu and felhom.eu hold no apt-get upgrade, apt full-upgrade, unattended-upgrades, pveupgrade or needrestart that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription "so the box can pull security updates" and then says plainly "No upgrades are run" (scripts/felhom-host-install.sh:2133-2136). The guest's Docker engine is installed when the golden is BAKED (felhom-agent/configs/build-golden.sh:124-125), so a fresh install gets that week's engine and an installed box keeps it forever. The only apt full-upgrade in the project is a by-hand step for the off-site endpoint ep0 (documentation/runbooks/offsite-endpoint.md:41), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is R-808 in ROADMAP.md. NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; 11 §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840). Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; 11 §8.2, audits/os-host-lane-2026-10-04/): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; 11 §8.1, audits/os-guest-lane-2026-10-04/). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (11 §7.1, corrections C1–C12, audits/os-updates-spike-2026-10-04/); no product code yet. LEFT: the build steps of 11 §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** — — CC + operator
R-862 Box system & updates P3 Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its felhom-os-apply (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it. FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; felhom-op cannot become root). The steps: runbooks/config-bundle.md "Tester 2". Then CC signs the bundle and reads it back. WAITING-ON-OPERATOR — the operator said (2026-10-04 ~19:05) he will try through his tunnel the operator's WireGuard tunnel the operator runs the three bootstrap commands; CC sends the bundle operator
R-35 Box system & updates P3 Config-apply should not end the customer's session. The offsite config push bumped config_version 10→11 at 16:54:58 and the controller self-restarted (container StartedAt 16:54:59Z, back up 16:55:02); in-memory sessions died with it and customer zero was force-logged-out mid-flow. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — Direction: hot-apply the offbox target (no restart for a config the running process can adopt), or persist sessions across restart. The restart itself is by design — the collateral is not. Evidence controller-log-full.txt CC
R-78 Box system & updates P3 local_api authority ruling — auto-reconcile vs detect-only MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state idea (deferred OUT of R-77 on purpose). Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — R-77 ships detection because the fix is genuinely undecided, and both directions can lose customer-visible function. Direction 1 (today): controller.yaml wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. Direction 2 (bootstrap.json wins, auto-reconcile on boot): a guest whose controller.yaml is CORRECT and whose bootstrap.json is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a working channel clobbered on the next restart, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — mergeLocalAPI replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. CC
R-190 Box system & updates P3 A storage ACL that demonstrably WORKED in the morning was gone by mid-morning, and nothing recorded its removal. On demo-felhom, a vzdump by felhom-agent@pve!agent with --storage felhom-backup completed OK at 04:44:50 CEST 2026-08-03 (task log read in full). From 09:24:56 the same path returned HTTP 403 … missing privilege Datastore.Allocate at /storage/felhom-backup, six times through the day, until the grant was re-applied by hand at 18:54. By ~14:50 pveum acl list showed no row at all for that path NARROWED — MITIGATION SHIPPED 2026-08-04 (agent v0.124.0 → v0.124.1) — MECHANISM STILL OPEN — Why this is not just R-185 restated: R-185's mechanism (the installer's Scenario-F arm resolves a pre-existing target without granting) explains a box that NEVER had the grant. This box HAD it and lost it, inside five hours, with the machine up throughout. Ruled out, each by measurement: a host reinstall (uptime = 12 days); any pveum/ACL/user.cfg activity in syslog between 04:00 and 10:00 (none); any ACL entry in /cluster/log (none). Correlated, not established: host_leaf_changed at 09:15 and controller_started at 09:19 — guest 9201 was reprovisioned nine minutes before the first 403. PVE removes ACLs at /vms/<vmid> when a guest is destroyed (AccessControl::remove_vm_access, the F-LEAK mechanism); whether any path can take a /storage/<id> row with it has NOT been established and is the first thing to check. Why it matters more than the grant did: a permission that can vanish silently makes every ACL-based guarantee on these hosts provisional, and the agent's new store-grant probe (v0.123.0) now detects the STATE but says nothing about the TRANSITION. Worth pairing with: whether the probe should report a grant it once had and no longer has as a distinct, louder signal than one it never had THE ROW NOW REFLECTS THE MITIGATION, NOT THE CAUSE — stated plainly because the two are different things. The box is resilient; the loss is still unexplained. Mitigation: when the store-grant probe finds the grant absent, the agent runs the EXISTING root wrapper (felhom-backup-target-apply grant <id>) and re-reads once to confirm — the pbsdr R-22 self-grant shape, including its restraint. No new privileged surface: the sudoers vector grant * already covers any storage id (confirmed in configs/felhom-agent.sudoers, not assumed), and the verb already grants BOTH user and token. The verb existed, was permitted, and had only ever been called at storage CREATION — the built but never wired shape in a verb rather than a seam, this project's seventh instance. Bounded at one attempt per tier per hour (a storage can be unreadable for reasons an ACL cannot fix; re-granting every cycle is a repair loop wearing a fix's clothes). THE RECORD IS THE HALF THIS ROW IS ABOUT, and v0.124.0 got it wrong in production while every unit test passed. It reported degraded for one cycle — meaning the probe call that repaired. But probeAll is invoked INDEPENDENTLY by the self-check log and by the collector building a host-report: on the box the repairing call was the log's (09:39:34, journal shows the repair and degraded=1) and the report three seconds later found the grant present and sent ok. The agent's journal had the record, the hub had nothing, and the operator would have learned nothing — the exact silence this row exists for, re-created inside its own mitigation. v0.124.1 replaces it with a latch on TIME (20 min > the 900 s report interval), so at least one report must carry it. PROVEN LIVE, twice, on demo-felhom (grant deleted by hand, both rows): agent logs store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190) target=felhom-backup privilege=Datastore.AllocateSpace action="felhom-backup-target-apply grant felhom-backup" confirmed_by=re-read; the ACL rows return; and on v0.124.1 the host-report at 08:00:30Z carried status=degraded with the explanation and the hub raised agent_capability_degraded and e-mailed the operator at 08:00:40. Nothing new was built to carry it — the hub's existing ok→degraded→ok edge is the channel, and the text rides Feature because that is the field the hub interpolates into the e-mail (Reason does not travel). PART 3 — the single bounded pass at the mechanism, with the negatives named. The lead §8.6 nominated is real as a CLASS and is documented in our own installer: "pveum user token remove purges the token's ACL, so re-applying post-rotate is mandatory". It does NOT fit this box. A rotation purges ALL of the token's ACLs and mints a NEW secret; demo-felhom's token still authenticates with the same secret (--selftest OK), it retained its other three storage grants throughout, and only felhom-backup was refused. No installer run is evidenced (no 2026-08-03 install log; host uptime 12 days at the time). Previously ruled out and unchanged: a host reinstall, any pveum/ACL/user.cfg activity in syslog 04:00–10:00, any cluster-log ACL entry. Ruled out on THIS box; NOT ruled out fleet-wide — any installer run still purges and re-grants only the hardcoded PVE_STORAGES set, though installer 1.24.0's reuse-arm fix now re-grants the backup target on that path. A NEW OBSERVATION FROM THE LIVE RUNS, relevant to the timeline: PVE caches permissions — after deleting both ACL rows the probe still read the privilege as present for ~40 s in one run and ~16 min in another. Detection is only as prompt as that cache, and a cache expiry could equally explain why a box kept working for hours after a grant was removed. → R-194. The alert pair CLOSED on its own at 10:30:40 (degraded → ok, agent_capability_recovered) once the 20-minute latch expired — one lost grant, one e-mail, one recovery, nothing further. CC
R-242 Box system & updates P3 A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that. R-239 is the symptom; this is the mechanism, recorded 2026-08-07 and deliberately NOT built (the task that found it scoped it as a record-only item). Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. Nothing in the release path knows a golden exists. The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's golden_version is edited by a separate operator act, in a different repo, with no link back. Proposed shapes, cheapest first — the choice is the operator's and is not taken here. (a) A release-path checklist step — one line in the controller's end-of-session checklist: a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not. Costs nothing, catches nothing mechanically. (b) A gate in repo_gates.py comparing the manifest's golden_version against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) A hub-side checker — the hub already knows every box's running controller version from /hosts and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. Earliest catch: (b). It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. (c) catches strictly more but only after boxes exist. (a) is worth doing regardless because it is free. Not built. No gate was written this session. ⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT. Controller v0.206.0 shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried 0.205.0 — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). ✅ SHAPE (b) BUILT 2026-08-08 — scripts/golden_currency_gate.py, registered in repo_gates.py as gate 7. It was shown FAILING against that exact state before anything was baked, which is its red-proof and the reason its own introducing push needed --no-verify (stated in the session report rather than worked around): newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED. IT IS --fast, AND THAT FORCED ITS DESIGN: both .githooks/pre-push AND CI run repo_gates.py --fast, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH, because the vouched version lives only in the hub's hub_settings with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. A bake without a vouch still passes: that half is NOT closed and stays on this row. It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. VOUCHED 2026-08-08 with the operator's approval — golden 0.206.0 / sha c85230b4…108e; agent_version and min_agent both stayed 0.127.0, and wrapper_sha256 was carried through explicitly because the handler clears it when omitted. The gate was CONVICTED before the bake and OK after it — red→green on the same command, which is its proof that it measures something real. ⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right. Controller v0.207.0 (R-249/R-252/R-253) is released, tested and pushed, and no golden carries it — the newest bake is 0.206.0 — so golden_currency_gate.py FAILED, saying exactly the true thing: a machine installed right now would receive v0.206.0. The felhom.eu push therefore used git push --no-verify, declared here, in the commit message and in the session report. A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. Owed: bake golden 0.207.0 and vouch it (RUNBOOK-manual-build.md §4.1; the vouch is a three-field change). This row's own remaining half is unchanged — nothing gates the VOUCH itself. ✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08). The gate went from red to green, and the --no-verify bypass declared above is now historical rather than standing. Round trip is the evidence, not the build log: the published bytes were downloaded back — 656 879 192 B, sha256 20ec9602…22995, both identical to what the bake reported — and ./etc/felhom-controller-image read OUT of the downloaded archive says felhom-controller:0.207.0, which is the delivered artifact naming the controller it will start. The vouch was a three-field change with all three checked deliberately (MinAgent 0.127.0 read from the golden's controller CHANGELOG header, not assumed; agent_version already ≥ it; min_agent not above agent_version, so not the R-216 shape) and verified by re-reading the manifest rather than trusting the flash. This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: tests/golden-0.207.0-2026-08-08/. ⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0. golden_currency_gate.py FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. The felhom.eu push used git push --no-verify, declared in the commit, the CHANGELOG and the session report — a bypass, not a waiver, on the same reasoning as yesterday: the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. Owed: bake golden 0.208.0 and vouch it (RUNBOOK-manual-build.md §4.1; three-field change, MinAgent 0.127.0 unchanged). Note the cadence this is establishing: two releases, two bakes owed within 24 h. That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. ⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0. golden_currency_gate.py FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. The felhom.eu push used git push --no-verify, declared in the commit message, in hub/CHANGELOG.md and in REPORT.md — a BYPASS, not a waiver, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that deliberately needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later on 2026-08-30, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md §4.1; three-field change — golden_version + agent_version + min_agent; MinAgent is 0.129.0 per both CHANGELOG headers). Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one. ⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0. The felhom.eu push carrying that release's documentation used git push --no-verify on the operator's standing ruling from earlier the same day, declared in the commit and in REPORT.md. The day-0 ground still holds for all three and was re-checked rather than assumed: R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore. The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence. Owed: ONE bake carrying 0.226.0 covers all three (RUNBOOK-manual-build.md §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. ✅ PAID THE SAME DAY — golden 0.226.1 baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30). Evidence: documentation/tests/golden-0.226.1-2026-08-30/. golden_currency_gate.py went red → green on the same command, which is its proof that it measures something real. The three declared bypasses above are now HISTORICAL rather than standing. The round trip is the evidence, not the build log: the published bytes were downloaded back — 657 197 592 B, sha256 70ed8e93…baefe69, both identical to what the bake reported — and ./etc/felhom-controller-image read out of the downloaded archive says felhom-controller:0.226.1, i.e. the delivered artifact naming the controller it will start. The three-field vouch was checked deliberately, not assumed: MinAgent 0.129.0 read from the golden's controller CHANGELOG header, agent_version 0.130.0 ≥ min_agent 0.129.0 (so NOT the R-216 shape), and the result verified by re-reading the manifest rather than trusting the flash — golden option 0.226.1 SELECTED, all four shas matching. The floor is proven ACTING, not merely set: demo-felhom self-updated within 30 s, logging [selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1). ⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN. v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that felhom.eu push used --no-verify and declared it, and golden 0.227.1 was baked, published, round-trip verified, VOUCHED and the floor RAISED to 0.227.1 within the hour. Evidence: documentation/tests/golden-0.227.1-2026-08-30/. THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day. Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. The floor was proven ACTING both times: demo-felhom self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new offsite-integrity job by itself, on a box nobody deployed to, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. This row's OTHER half is still open and untouched: nothing gates the VOUCH itself — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. 2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day. v0.230.0 shipped in the morning with the newest golden at 0.229.0, which is the build R-403 says deletes a good copy, so the gate was red across four commits (dddcc80, 6e550ae, 130f7a6, 32a4c35). Golden 0.230.0 baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; demo-felhom moved itself off the defective build unattended (controller-swap: new controller healthy, 16:21:40 CEST). Evidence: documentation/tests/golden-0.230.0-2026-08-31/. The gate did its job and its own weakness surfaced doing it — R-410. READY — the vouch half only — owner Viktor ⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0. golden_currency_gate.py FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The felhom.eu push carrying this release's documentation used git push --no-verify, declared in the commit message and in felhom-controller/REPORT.md - a BYPASS, not a waiver, on the same reasoning as the five entries above: the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. The day-0 ground was RE-CHECKED rather than reused: R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and MinAgent is unchanged at 0.129.0. The ground still expires the moment a release changes first-boot behaviour. OWED: bake a golden carrying 0.229.0 and vouch it (RUNBOOK-manual-build.md §4.1; three-field change - golden_version + agent_version + min_agent, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. Golden and fleet delivery are the operator's (this row). ✅ PAID THE SAME DAY — golden 0.229.0 baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31). Evidence: documentation/tests/golden-0.229.0-2026-08-31/. golden_currency_gate.py went red to green on the same command, which is its proof that it measures something real. The --no-verify bypass declared above is now HISTORICAL rather than standing. The round trip is the evidence, not the build log: the published bytes were downloaded back - 656 864 331 B, sha256 39aa886d…d7bdae87, both identical to what the bake reported - and ./etc/felhom-controller-image read out of the downloaded archive says felhom-controller:0.229.0. A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. The three-field vouch was checked deliberately, not assumed (MinAgent 0.129.0 read from the golden's controller CHANGELOG header; agent_version 0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a field it is not sent), and verified by re-reading the manifest rather than trusting the flash. The R-120 gate passed rather than being bypassed - fleet newest 0.229.0, golden 0.229.0. The floor is proven ACTING: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0) - nobody deployed to that box. Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour. This row's OTHER half is still open and untouched: nothing gates the VOUCH itself - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. ⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused. Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here. The felhom.eu push carrying this release's documentation used git push --no-verify, declared in the commit message and in felhom-controller/REPORT.md - a BYPASS, not a waiver. OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor. demo-hp was updated by hand; demo-felhom is still on 0.229.0 and still carries the defect. See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported. 2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN. R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. Nothing gates the VOUCH. A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's hub_settings table, there is no copy in git, and a hub-reading gate could not be --fast so it would run in neither the hook nor CI. Do not read R-404's closure as closing this. 2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468). The docstring's "honest fix is a recorded waiver in the register, never a habit of bypassing" is now a mechanism: golden_currency_gate.py reads documentation/tests/golden-waiver.yml (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. What stays open on THIS row is exactly one thing: nothing gates the VOUCH. The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). — — operator
R-274 Box system & updates P3 A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind. felhom-host-install.sh step 7: if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi — the hub manifest's golden.sha256, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path. A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst, whose baked marker reads felhom-controller:0.192.0, against a vouched golden of 0.210.0 — 18 releases stale. The sharp consequence: 0.192.0 is below 0.200.0, where R-193's off-site recovery SCREEN shipped, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. It compounds with the teardown, which deliberately keeps the old golden ("golden vzdump left in place"). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above. On the RESUME path step 7 fetched the vouched 0.210.0 correctly (fetching golden v0.210.0 from Gitea), because --resume skips preflight and preflight is where local auto-discovery sets GOLDEN_VOLID (the GL6-F4 comment says so). So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one. Discovery is `… sort tail -1, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on local` (07-21, 08-03, 08-09), because the teardown keeps them. Still not observed: a FRESH (non-resume) install taking a stale local golden. VERIFY (2026-10-03 triage: Local golden is now checked against the hub manifest's version and sha before use (GOLDEN_CHECK_WHY, R-297) — felhom.eu/scripts/felhom-host-install.sh:2855-2905) — READY (S) — NEW 2026-08-09, NARROWED same day —
R-349 Box system & updates P3 "Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice. Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. The mechanism: a proof deploy is a hand build (go build -ldflags "-X main.version=0.130.0"), while scripts/release-agent.sh deliberately builds with -trimpath -buildvcs=false so the published artifact is reproducible (R-186). Same source, same version string, different bytes: 256e0829... on the boxes vs a56a92a7... published and vouched. Nothing corrects it automatically, and that is the sharp edge: the boxes already report 0.130.0, so the self-update path sees the vouched version as already installed and does nothing, forever. The divergence is invisible to every version check in the system — the hub, --version, and the artifact manifest all agree, because they all compare the version STRING. Consequence if unnoticed: the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard publish-agent.sh already carries a comment about for CGO_ENABLED; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. Corrected here by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report sha256 a56a92a7..., matching the vouch. READY (S) — NEW 2026-08-20 — Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched agent_sha256 and flag drift — exactly the mechanism wrapper_sha256 already implements for the PBS wrapper (R-50b), whose manifest help text says it "makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page". The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. CC
R-444 Box system & updates P3 [P3-LOW] Nothing runs pct fstrim on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed. MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that local-lvm did not reclaim on delete (68.97% -> 70.91%); fstrim INSIDE the unprivileged container is refused (FITRIM ioctl failed: Operation not permitted, all three mounts); pct fstrim 9201 from the PVE host then trimmed 30.2 GiB + 57 GiB and took local-lvm to 26.78% — 23.8 GB BELOW this run's own starting point, i.e. the surplus was long-standing, not ours. Why it is not merely housekeeping: a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. Not urgent, and the row says so — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: CC. audits/SPIKE-app-update-2026-09-01.md OPEN — rank P3-LOW; owner: CC — — CC
R-468 Box system & updates P3 [P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13). 25 goldens in 26 days in August, almost one per release, because golden_currency_gate.py trips on every release by design and the only honest ways past it were a bake or a declared --no-verify (thirteen by 2026-09-01, R-404/R-417). The ruling: bake WEEKLY, and always before any drill or fresh install. Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. The mechanism (built 2026-09-13): documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml, four lines (issued, expires, reason, register_row: R-468), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. The 14-day cap is enforced by the gate, not the runbook — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. It never covers a golden that is UNRECORDED (R-385) — that is not a cadence choice. A dated waiver cannot be forgotten; it just expires — the difference from R-242's original rule, which recurred the day after it was written. Tests: scripts/test_golden_currency_gate.py cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only expires — 2). This is a PRE-CUSTOMER arrangement: the first external install retires it (delete the file in that commit). Cadence written into RUNBOOK-manual-build.md §4.2 and the felhom.eu end-of-session checklist. Does NOT touch R-242's open half (nothing gates the VOUCH). WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install — — CC
R-531 Box system & updates P3 [P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no controller_restarted_by_agent; deliberate operator kills spend the crash-loop budget. MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, A4-kill-middeploy-9201.txt). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). What it needs: a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget MEASURED 2026-09-16 on the drill box, both halves. (1) Timing during a deploy (F9'): the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again 37 s after the kill; the interrupted app ended not_deployed, not stuck. (2) The budget's shape (F9''): three further kills at idle, 20 minutes apart, recovered in 61 s / 41 s / 61 s - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. So the brake catches a FAST loop and is blind to a SLOW one: a controller dying every 20 minutes is restarted forever and the only trace is an info event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt and phase2-f9dprime.txt. READY — rank P3-LOW; owner: CC (measure) · operator (budget rule) — — CC + operator
R-194 Box system & updates P4 PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement. Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for /storage/felhom-backup were deleted, and GET /access/permissions continued to report Datastore.AllocateSpace present — for ~40 s in one run and ~16 minutes in another. During that window the capability probe reads healthy and the self-repair does not fire OPEN — Why it matters beyond the delay: it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — it is a candidate contributor to R-190's own timeline: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of worked at 04:44, refused at 09:24. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. Not a defect in our code — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). What is worth deciding: whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately ({"data":[]} while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read CC
R-836 Box system & updates P3 A new host kernel that hangs before userspace stays the GRUB default: --next-boot is not a one-shot on these hosts. MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; kernel pin <new> --next-boot writes an ordinary GRUB_DEFAULT, and proxmox-boot-cleanup.service clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only softdog runs (useless before userspace); demo-hp's sp5100_tco ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (GRUB_DEFAULT=saved + grub-reboot) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). audits/os-updates-spike-2026-10-04/partH/ NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either. GRUB_DEFAULT=saved (old kernel saved) + grub-reboot <new>: boot 1 → new kernel, Secure Boot ON and fine; but GRUB could not clear next_entry (grub-reboot itself warns: environment block on lvm device … will remain the default until manually cleared; /boot is ext4 on LVM pve-root), so boot 2 (no command) → the new kernel again. kernel.panic = 0: a panic leaves the host stopped (R-851). sp5100_tco LOADS and answers (SP5100 TCO timer, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. LEFT (fix direction): a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming sp5100_tco; to be measured before the kernel slow lane. audits/os-host-lane-2026-10-04/partE/ READY — owner: CC + operator (reboots). — — CC
R-853 Box system & updates P3 After a boot the box's versions and crash facts reach the hub up to ~15 minutes late. MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no system.facts — the facts read needs a RUNNING customer guest (firstGuest), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields unknown), and a failed read is not cached. audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt READY — owner: CC — — CC
R-839 Box system & updates P3 After a Docker restart that stopped every container, the boot sweep HELD an app whose app.yaml HDD_PATH names its user folder instead of the drive. MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): bootrecon logged drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it, although the drive /mnt/felhom-drives/scratch_hdd IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. Not diagnosed: which writer put a per-app path in HDD_PATH on 9202, and whether a household box can get it. audits/os-updates-spike-2026-10-04/partG/SUMMARY.md READY — diagnose; owner: CC — — CC

Monitoring & notifications — 25 rows (P2 3, P3 15, P4 7)

ID Category Sev What State Blocked on Next action Owner
R-243 Monitoring & notifications P2 A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES. Found by the R-241 spike (2026-08-07) as a by-product; not part of the walk's finding and not previously filed. The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: escrow_state is stuck pending forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and runOffboxBackup returns at the escrow gate (offbox.go:743) before touching anything. So off-site backups never run again — and the hub never notices. All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: offsite_stale — isStale (monitor/offsite.go:135) returns false unless EscrowState == "escrowed", and its own comment reads "Pending/disabled = normal onboarding, never stale", so the box is classified as still being set up, forever; offsite_delivery_stuck — monitor/offsite_delivery.go:91 skips the applied shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); backup_failed — never fires, because nothing fails: the run returns nil before it starts. Three correct exclusions leaving one state unobserved. This is the same class as the workspace CLAUDE.md "presence is not success" rule, one level up: the absence of a failure is being read as the presence of a working tier. Partly subsumed by R-241's fix — a box that recovers leaves this state — but not for a box that does not, and the alarm gap is what makes "does not" survivable indefinitely. Not fixed; no code written. ⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open. R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. What replaces it is a state that is VISIBLE rather than silent: the box declares offsite.state=awaiting_recovery_key and the customer is offered the recovery screen. But the hub still raises nothing for it, and for the same three reasons: isStale needs escrowed, the delivery checker skips the applied shape, and backup_failed needs a run that never happens. So a box whose customer never acts still stops backing up off-site with no operator signal — the difference is that the customer can now see it and act, where before nobody could. The remaining work is an operator-side signal for a box held in awaiting_recovery_key past some age, and it is deliberately not bundled into R-241's fix. ⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces. 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted offsite_delivery_stuck (warning) and wrote an operator-channel notification_log row recording offsite_credential_restaged / status REFUSED with an accurate reason — "the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193". So on the regressed-apply shape the operator IS told, promptly and correctly, and this row's "skips the applied shape" does not apply. The gap stands for a box that reaches the held state without a prior working tier in its report history. Recorded so the row is not read wider than it measures. READY — owner Viktor — — operator
R-528 Monitoring & notifications P2 [P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: OOMKilled stays false and no oom event fires, so the v0.243.0 OOM line is not proven live. MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with OOMKilled=false and zero docker events --filter event=oom; a memory hog inside the running container was killed (rc 137) with the same silence (E2-oom-signal-measure-9202.txt). BIGNIGHT VM 333 did read oomkilled=true, so the shape differs by case. Fix shape: the agent reads the guest container cgroups' memory.events oom_kill counters (host-side, reliable), or the controller alarms on a restart-count trend RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine: the Paperless webserver was capped at 128 M with docker update --memory; it restarted 9-10 times, and all three signals stayed silent - OOMKilled=false on every inspect, docker events --filter event=oom EMPTY for the whole window, the container's cgroup not visible from inside the guest, and dmesg unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt. READY — rank P2-MEDIUM; owner: CC 2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely. On tester-1-022354 (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed app_oom (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was write CONNECTION_CLOSED immich-postgres:5432 and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: audits/evidence-chaos-night-2026-09-17/round-2.txt. — — CC
R-79 Monitoring & notifications P3 report.Issues / report.Warnings are English on customer-facing surfaces MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged — Whole-surface, not a one-off (DIAG §6): every producer is English — "SSD/HDD disk usage critical", "Docker: %v", "Protected container not running: %s", and all six Warnings strings. They render on the customer's Hungarian dashboard, and the health_critical path has reached the customer email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. CC
R-211 Monitoring & notifications P3 Prometheus has no config-reloader — a rules change reaches the pod and is never read READY (S) — NEW 2026-08-05 — Found while verifying R-205 rather than by looking for it. The mon-system/prometheus Deployment runs one container (prom/prometheus:v3.12.0) with no configmap-reload/prometheus-config-reloader sidecar. After the ArgoCD sync the updated node-housekeeping-alerts.yml was present inside the pod (grep -c "and on(instance)" → 3 on the mounted symlink) while the Prometheus rules API still served the old expression — for 4+ minutes, with no error anywhere. It only took effect after an explicit POST /-/reload. The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod — so "committed and synced" has never meant "in force", and ArgoCD reporting Synced/Healthy is true and beside the point. --web.enable-lifecycle IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a checksum/config pod annotation so a rules change rolls the pod. Same class as the four built-but-never-wired seams — the control exists, nothing walks it CC
R-271 Monitoring & notifications P3 The agent_channel_unauthorized alarm can never be closed, because its own prescribed remedy is what silences the recovery. channelhealth.Checker.Check's UP branch notifies only when prev != "" && prev != "up"; a controller restart resets state to "", so an unseeded→up transition is silent by construction. The alert text says "token stale/rotated (re-bootstrap)" — i.e. restart the controller — so following the instruction guarantees no recovery event. Observed live 2026-08-09: two agent_channel_unauthorized errors on the hub (one sent, one suppressed by the 1 h operator cooldown) and nothing afterwards, though the channel came up 3 minutes later and stayed up. The down side is deliberately asymmetric (F2: a born-down channel alerts on cycle 1); the up side never got the matching treatment. Customer dashboard is fine — SetDashboard reflects current state every cycle. It is the OPERATOR's trail that ends on "down" READY (S) — NEW 2026-08-09 — Notify on unseeded→up when the previous persisted state was down, or seed from the hub's last event CC
R-333 Monitoring & notifications P3 Two disk-health questions the deploy raised and did NOT act on. (a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe. They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as Hiba — the single worst outcome this feature can produce. (b) The agent runs bare smartctl -a -j with no -n standby (felhom-agent/internal/storage/hostops.go:368), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only 3375 load cycles in 60505 hours (~one per 18h), i.e. that duty cycle barely spins down at all READY (S each) — NEW 2026-08-14 — (a) split the temperature bands by device class, or drop them for NVMe and rely on critical_warning; (b) add -n standby to the agent's smartctl invocation (an agent change, so fold it into R-330's session) Viktor decides (a); CC does (b)
R-340 Monitoring & notifications P3 The new reachability check does not touch the surface that actually failed. R-339 reports when the hub cannot READ ep0 — but the read it performs is the usage op, which is proxmox-backup-manager plus df over SSH, and therefore rides the local API daemon. The 2026-08-18 incident explicitly CLEARED that daemon: proxmox-backup.service was healthy throughout, and it was the HTTPS proxy on 8007 that was wedged with a full accept queue. So R-339's check would have returned green for all 9 h 37 m of that outage. It closes the case where ep0 is unreachable as a host; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier READY (M) — NEW 2026-08-18 a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout ErrUsageUnsupported already models) Add a health op to scripts/felhom-tenantsync.sh that probes https://127.0.0.1:8007/ on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time. Available in audits/evidence-ep0-established-connections-2026-08-20/: the proxy fd count and its type breakdown (lsof + /proc/<pid>/fd), the listen-queue depth (ss -lnt — Recv-Q 0, Send-Q 1024), the ESTAB/CLOSE-WAIT split, the per-peer connection histogram, a 31-minute persistence diff of full 4-tuples, and a 46.18 h slope with Poisson bounds. What the health op would still add beyond these: a loopback GET https://127.0.0.1:8007/ probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. And this spike sharpens what the op should report: a rising ESTAB count is the live signal (CLOSE-WAIT was 0, not merely flat), and per R-344 the fd ceiling that matters may be the agent's, not only ep0's. CC
R-363 Monitoring & notifications P3 The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps. sched.Daily("fill-watch", "03:30", …) (cmd/controller/main.go:1092) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused kimai per app and the hub received recovery_unit_capture_failed (error) naming the filesystem, and the fill watcher said nothing at all. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. OPEN — MEDIUM — The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. CC
R-388 Monitoring & notifications P3 PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector. The operator's framing, recorded verbatim 2026-08-23: "A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. A failed backup is our incident, not theirs. The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call." Today's page is the opposite shape — one switch per detector, and it grew from 12 to 15 in a single session (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. OPEN — DIRECTION, operator's call a decision on scope; nothing here is a bug Recorded as a dated [DESIGN — DIRECTION] entry at documentation/architecture/08-alarm-ladder.md §8, marked plainly as not current behaviour. Deliberately NOT implemented in the session that recorded it. app_start_failed defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. Viktor
R-435 Monitoring & notifications P3 The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce. snapshotDropFraction = 0.5 and snapshotDropFloor = 5 (hub/internal/monitor/offsite.go) require a fall of MORE than half the previous count. demo-hp's baseline is 69 across 9 apps, so ~35 snapshots must go before it speaks; one app's tag is ~9 and is invisible. offbox.go:1388 runs forget --prune grouped by host,tags — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. This is deliberate, not accidental: the constant's own comment argues the insensitivity, and "a detector that cries wolf is switched off within a fortnight" is a lesson this project paid for. So this row is NOT a demand to lower the threshold. It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and STATUS.md) is true only of falls above half. Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked: Phase 2 deletes one app and Phase 3 expects the alarm to fire. LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1 — the comment above snapshotDropFraction in hub/internal/monitor/offsite.go now states what the detector does NOT see, with the demo-hp arithmetic and the forget --prune grouping that makes the blind spot sit on the most likely single-app failure. It also says explicitly that the numbers must NOT be lowered to "fix" this and that per-app detection needs a SECOND signal keyed on the per-tag count. The row stays OPEN because documenting a blind spot is not covering it — and because STATUS.md and R-431 both still say "noticed within a day", which is true only of falls above half. OPEN — documented in code v0.111.1; the coverage gap itself is unclosed — — CC
R-521 Monitoring & notifications P3 [P3-LOW] One unplugged drive sends the operator five e-mails and the household none. MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): storage_disconnected (error) at 21:58:02 CEST plus app_start_failed (warning) for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five Operator email sent). The customer's mailbox (tester1@felhom.eu, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. Fix shape: suppress app_start_failed for apps stopped by a storage_disconnected (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. F6, 40 min later, the opposite failure: a second, separate drive loss (20:38:33Z) produced storage_disconnected (error) and four app_start_failed, and the hub logged Operator email suppressed … cooldown for all five — no mail at all for the second unplug; only health_degraded (warning) mailed. A per-key cooldown that outlives the recovery (storage_reconnected came between them) silences a new incident. F7: the system disk at 95 % produced only health_degraded (warning), whose operator mail was suppressed by the cooldown left by F6's health_degraded 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy) — — CC + operator
R-522 Monitoring & notifications P3 [P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline. MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · Fut · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged [report] Push failed … context deadline exceeded and Job hub-report failed: hub push failed after 3 attempts. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". Fix shape: the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. READY — rank P3-LOW; owner: CC (controller) — — CC
R-547 Monitoring & notifications P3 [P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: disk_critical is defined at ≥95 % used, but the fill-watch runs once a day. MEASURED 2026-09-17 (chaos night) on a fresh box (tester-1-022354, controller 0.245.0): the customer guest’s root filesystem was held at 96 % for ten minutes (29 G used, 1.5 G free) and no alarm of any kind fired — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: fillwatch runs daily at 03:30 plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about twenty seconds before the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. This is the ladder working as designed, not a missed alarm — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is no, unless the controller happens to restart while it is full, and that answer is not written down anywhere. Fix shape (one of): sample the fill more often than daily (a cheap statfs on the 5-minute health pass would do it); or say plainly in 08-alarm-ladder.md that a transient full disk is out of scope. Evidence: audits/evidence-chaos-night-2026-09-17/round-3.txt. READY — rank P3-LOW; owner: CC — — CC
R-585 Monitoring & notifications P3 [P3-LOW] Six event producers still send Hungarian only, so an English household can see one Hungarian line in some mails. FOUND 2026-09-18 by localisation slice 3 Part B (R-558, controller v0.256.0), which converted 19 of the customer-facing producers and left these: backup_failed, db_dump_failed, backup_integrity_ok, backup_integrity_failed, offbox_enlarge_blocked and local_api_endpoint_drift. Why they were left: each receives its sentence already FINISHED from another package, so the key and its arguments no longer exist by the time the notifier sees it — converting them means changing their callers, not the notifier. The 15 operator-tier types are deliberately excluded and are NOT part of this row: the operator reads Hungarian. Why it matters more than it looks: offbox_enlarge_blocked has no customerMessages entry on the hub, so its raw sentence IS the household's mail rather than an extra line under a translated headline — for that one type an English household gets a wholly Hungarian mail, not a mostly-English one. Fix shape: push the key and its arguments down from each caller (the shape slice 2 release B already used for errors, util.MsgError), then add each to the convertedProducers table in internal/notify/message_customer_test.go, which is the list both language tests walk. READY - rank P3-LOW; owner: CC — — CC
R-723 Monitoring & notifications P3 [P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong. MEASURED 2026-09-29: Operator email sent for tester-1/node_recovered 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed) → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. Fix direction: node_recovered not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is info, not a mailed warning. FIXED 2026-09-30 (hub v0.126.0): no node_recovered when the customer's host was enrolled after the outage began; backup_tier_skipped in a box's first hour recorded, not mailed. Real recoveries and old boxes' skips still mail (controls). Unit + red-proof (RP40, RP41); the live proof is Tester-2's first hour. WATCHING (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — FIXED in hub v0.126.0; the live proof is Tester-2's first hour) — FIXED — hub v0.126.0; live proof at the first real install; owner: CC — — CC
R-724 Monitoring & notifications P3 [P3-LOW] The status pages disagree with each other in small ways a household notices. MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw 2026-09-29T19:22:30Z; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). FIXED 2026-09-30 (controller v0.283.0): Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". NARROWED — remaining: „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). NARROWED — the LAN/gateway read only; owner: CC — — CC
R-177 Monitoring & notifications P4 There is no operator-triggerable "run the fill check now" path. fill-watch is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller READY (S) — NEW 2026-08-02 — Noticed while live-validating R-167 on 9201, not by a failure. It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. Partially mitigated already — v0.191.2 makes every run log a positive observable (checked N filesystem(s), M unreadable/skipped, K notification(s)), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has GetJobs but no run-now, so this is a general affordance, not a fill-watch one — scope it as "run a named scheduler job now", operator-gated. ID established free: grep -ro "R-177\b" documentation/ *.md → 0 hits CC
R-266 Monitoring & notifications P4 A failed root statfs still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one. Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. report/builder.go:93-95 copies sysInfo.DiskTotalGB / DiskUsedGB / DiskPercent into r.Storage[0] (Mount: "/"), and those are exactly the zeros a failed statfs leaves behind — the controller now KNOWS the measurement failed (SystemInfo.DiskKnown, controller v0.210.0) and the report still does not carry it. Deliberately not fixed here, for a reason that is now structural rather than a preference: adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (scripts/wire_contract_gate.py refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. RANKED LOW, and the reason is that the consequence is bounded: the hub bands host storage on disk_percent, so a failed read presents as 0% used — the quiet direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. Fix shape when it is taken: carry disk_known on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 READY — owner Viktor — — operator
R-285 Monitoring & notifications P4 A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere. During the 2026-08-09 rehearsal the hub sent, all status: sent to the operator channel: host_stale 08:58 UTC, node_stale 09:00, host_down 09:28 (error), node_down 09:30 (error), host_leaf_changed 09:31, host_recovered 09:31, node_recovered 09:34, offsite_delivery_stuck 09:34 — eight operator mails for work that was deliberate, attended and announced. This is the OPPOSITE gap from the one R-281 filed: the alarms are not missing, they are indiscriminate. host_stale at 30 min and host_down at 60 min (monitor/host_staleness.go:22-23, downAfter = 2 * threshold) cannot distinguish a wiped-on-purpose box from a dead one, and host_leaf_changed firing on a reinstall is correct-but-expected. Note the interaction with the mute used on 2026-08-09 evening: blocking a customer silences everything, so today the only two settings are page me for planned work and tell me nothing at all. What is owed is a middle: a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible READY (M) — NEW 2026-08-09 — The evidence is the operator's mailbox plus events/notification_log for 2026-08-09 CC
R-337 Monitoring & notifications P4 /backup/status lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect. During the R-336 recovery on 2026-08-18, demo-hp's snapshot landed on ep0 at 03:58:43Z (complete manifest; the host's own task index says OK) — yet GET /backup/status was still serving the superseded 03:27:00Z failure at ~04:03Z, four-plus minutes later. demo-felhom showed its new result within ~40 s of completion. The lag cleared on its own: demo-hp's 04:07:35Z host report carries felhom-pbs success=true, 4.29 GB, and the hub is green for both boxes. The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value. It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained WATCHING — NEW 2026-08-18 another observation, ideally during an incident rather than constructed Do not open a fix on this as written. First establish the intended refresh path for /backup/status after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so CC
R-371 Monitoring & notifications P4 The off-site tier is the only backup tier that announces nothing on success. Written down 2026-08-05 in audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513 and explicitly "recorded, not filed": the off-site run emits no hub event at all, while both lesser tiers do (db_dump_completed, crossdrive_completed). Failures are covered by backup_run_failures and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. Still true 2026-08-22 — the 2026-08-21 drill's own event dump shows db_dump_completed and six crossdrive_completed rows and no off-site success event. Age when filed: 17 days. OPEN — LOW — Either emit one, or record deliberately that the highest-value tier is silent on success and say why. CC
R-856 Monitoring & notifications P4 A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails. 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent app_start_failed (operator) and app_stopped_unhealthy (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (08 §5). A design question for the operator, not a defect yet. audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt READY — operator decision — — operator
R-872 Monitoring & notifications P2 A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is down, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is node_down. MEASURED 2026-10-05 05:00 Budapest, hub log: Deadline check: … 0 backup missed … 1 skipped (down) — the skipped one is Tester 2, off since 18:06 UTC (hub/internal/monitor/deadline.go ~360: if st == "down" || st == StateDisabled { skipped++; continue }). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. audits/night-fixes-2026-10-05/partF/FINDINGS.md NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, audits/catchup-2026-10-05/partC/): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (08 §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1 (if Tester 2 is still off at 05:00), and the two events in events. Holds → close; does not → a new row. R-871 — CC
R-886 Monitoring & notifications P3 DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z — every 15 min Running maintenance failed … open /alertmanager/nflog.…: permission denied (and the same for silences), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (alertmanager_notifications_total{integration="email"} 5 → 6, failed_total 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. OPEN — Compare the volume's file owner with the pod's securityContext (fsGroup/runAsUser); fix in homelab-manifests; prove with a silence that survives a pod restart operator
R-884 Monitoring & notifications P4 ArgoCD app monitoring shows Deployment/prometheus OutOfSync (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. OPEN — argocd app diff monitoring (or the CR's resource diff) to see what differs, then decide git or live operator

Hub & operator — 24 rows (P2 1, P3 9, P4 14)

ID Category Sev What State Blocked on Next action Owner
R-173 Hub & operator P2 The hub's SQLite PVC is excluded from every Longhorn backup job. pvc/hub-data carries recurring-job-group.longhorn.io/default: disabled, and backup-daily + backup-weekly (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the default group — so the 128 MB /data/hub.db has no volume-level backup. That database holds host_recovery (every managed box's break-glass root password), host_escrow + host_escrow_superseded (escrow custody), host_pbs_secrets, customer_configs, dr_recipe and the wg endpoints/peers — i.e. the material several documented recovery routes depend on NARROWED 2026-10-05 (evening) — option A IN FORCE (09 decision 125). The hub writes a nightly VACUUM INTO snapshot at 02:00 (hub v0.136.0, 05 §16.3, keep 2; the volume grew to 2 Gi); DooPlex checks it (integrity_check, size, ≥1 host, ≤26 h old), encrypts it with a key ep0 never sees and pushes it at 02:30 to ep0's operator namespace with a write-only token; a read-only token restore-tests it every Sunday 04:30 (and refuses a readable console password); HubDBBackupStale/HubDBRestoreTestStale alarm on success-only timestamps, absent() included. Both keys are off DooPlex (operator, 2026-10-05). PVC label fixed (enabled). Proven live: first push 7 s, restore test, token limits, a key rebuilt from the paper data field decrypts, runbook §3 steps 1–3 (4/4 console passwords open with the saved seal key, 0/4 with a random one). audits/hub-db-offsite-2026-10-05/ — LEFT: runbook §3 steps 4–5 (the copy into a live PVC) are not exercised — they need the hub down; do them at the next planned hub maintenance or a DR drill on a scratch k3s. Close then. CC
R-30 Hub & operator P3 [P2-HIGH] Liveness presence should come from the wait channel, not the report clock. The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (host_stale 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk. — Direction: derive presence from Dir-2 long-poll connectedness (~90 s grace), decoupled from notification hysteresis (the hysteresis is right for alerting, wrong for presence); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. (Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.) CC
R-31 Hub & operator P3 [P2-HIGH] Offsite provisioning is synchronous with no status affordance. Save runs the Hetzner sync in-request, so the request can hit the nginx 504 while succeeding server-side: the operator cannot tell failed from slow, and a retry races the first attempt. MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state idea. Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. OPEN — migrated from ROADMAP 2026-08-22, rank unchanged Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists. — Direction: make it async + a status card, reusing the proven awaiting-card/poll idiom (v0.138.0 escrow card). Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click. CC
R-244 Hub & operator P3 The customer DELETE cascade leaves app_log_issues behind, and it is systematic across every venue ever torn down. Found 2026-08-07 while verifying the finalwalk teardown with a full census (every table, every column) rather than a per-table query. After a cascade that logged COMPLETE … full teardown, 61 rows still matched finalwalk. Four of the five sources are deliberate and correct — the cascade's own header states "Provenance/events are NEVER wiped — audit outlives every tier": events 16, notification_log 14, host_deletions 1, customer_resets 1. The fifth is a gap: app_log_issues 29 rows, which the residue purge does not touch (its logged leg covers reports/app_telemetry/app_log_tails/log_tail_requests/notif_prefs/selfbind_tokens/appliance_registrations — not this table). It is not a finalwalk quirk: rows still reference c11 40, rewalk 20, part4 24 — all three torn down 2026-08-06, whose ledger recorded "0 occurrences". That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone. Why it was probably never written, established rather than assumed: the table is a fleet-wide aggregate keyed on app_name+fingerprint with an affected_customers JSON list — of the 29 finalwalk rows, 12 reference only finalwalk (orphans, safely deletable) and 17 are shared with LIVE customers (demo-felhom, peti-felhom, …) and must not be deleted, only de-referenced. A naive DELETE … WHERE customer LIKE would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. Severity is LOW and stated plainly: no secret material is involved — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's identifier inside an aggregate row. Proposed shape: a residue leg that (a) removes the customer id from affected_customers/context_customer, and (b) deletes rows whose affected_customers becomes empty; plus a one-off sweep for the four already-torn-down venues. The general lesson is the reusable part: a per-table absence query is not a census. The teardown verification is now a full-schema sweep, and that is what found this. Not fixed — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: tests/teardown-finalwalk-2026-08-07.md. ⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation). app_log_issues holds 1309 rows; 71 reference a torn-down venue (finalwalk, c11, rewalk, part4); of those 44 are ORPHANS — they name only torn-down customers and are safely deletable — and 27 are SHARED with a live customer (demo-felhom, peti-felhom, …) and must be de-referenced, never deleted. 1238 rows are untouched. The 27 are exactly why the leg was never written, and why a DELETE … WHERE customer LIKE would destroy a live customer's issue history. What it needs, precisely: a cascade leg that (a) removes the customer id from affected_customers / context_customer, and (b) deletes only rows whose affected_customers becomes empty; plus a one-off sweep for the four venues already gone. Why it was NOT done on 2026-08-08: the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. It accumulates one venue at a time, so the next walk adds to it; the numbers above mean the next session starts from data rather than a guess. ⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08. The fifth walk's venue was torn down with a full-schema census taken before and after: 168 rows → 67. Of the 67, 37 are by design (events 21, notification_log 14, host_deletions 1, customer_resets 1) and 30 are app_log_issues — this row's gap, and the count was predicted in the pre-run enumeration rather than discovered afterwards, which is the difference from the ledger that once recorded "0 occurrences" from a narrower query. The running total across torn-down venues therefore rises from 71 to ~101 rows (finalwalk, c11, rewalk, part4, now walk5) — the shared-with-a-live-customer subset must still be de-referenced, never deleted. It accumulates one venue at a time and it did so again. Evidence: tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md. 2026-09-25: peti-felhom is no longer a live customer (deleted through the cascade, journal #20); 8 app_log_issues rows still name it — the same gap. READY — owner Viktor — — operator
R-277 Hub & operator P3 Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run. For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, last_status: ok), yet: (a) the customer page's Backup panel read Snapshots 0 / Repo Size 0 MB / Integrity Unknown — it renders the local disk tier, while the healthy offsite object sits in the same report unrendered on that panel; (b) the Offsite page read 0.0 GB — true, but a 162 KB repo rounds to nothing; (c) a stale offsite_delivery_stuck event from 2026-08-07 10:19 (not recurring) reads as current state. Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed". It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. Note the true half: demo-felhom IS genuinely stuck (offsite.state=needs_credential, no run has ever succeeded) → R-278 READY (S) — NEW 2026-08-09 — Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history CC
R-581 Hub & operator P3 [P2-MED] ORDER BY received_at cannot answer "the newest report" — the column has SECOND granularity. FOUND 2026-09-18 building hub v0.118.0 (R-558): Store.CustomerLanguage read the household's language from the newest report ordered by received_at, and TestNewestReportedLanguageWins failed — four reports written in the same test tick all carry the same datetime('now') string, so the "newest" was whichever row SQLite felt like returning. On a real box the same shape appears whenever two reports land in one second (a settle burst, a restart race), and the symptom would have been a household switching language on their dashboard and getting the old language back at random. FIXED in the same release by ordering on the autoincrement id, which is the real insertion order. What is still open, and it is the reason this is a row rather than a note: GetCustomers() has the same shape — INNER JOIN (SELECT customer_id, MAX(received_at) …) with no tie-break — and it is what the whole operator dashboard and countBoxesBelowFloor read. A same-second tie there picks an arbitrary report's health, version and vitals. Not observed in the wild; not looked for either. Fix shape: tie-break every newest-report query on id DESC, or give reports a monotonic ordering column and use it everywhere; then a test that writes two reports in one tick and asserts which one wins. READY - rank P2-MED; owner: CC Re-ranked 2026-10-03: P2->P3: operator dashboard only; the household-facing language case was fixed; never observed. — — CC
R-600 Hub & operator P3 [P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0. FOUND 2026-09-20 by the slice-6 drill's teardown, measured on ep0 rather than inferred from the hub. The customer delete cascade finished at 19:41:54 with customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). Three minutes later wg show wg0 allowed-ips on ep0 still listed 10.77.0.5/32 — the drill box's peer — because wgsync pushes on its own cycle. Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the full teardown line by about 6 minutes. (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. A period read off two log lines is not a measurement.) The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push. The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. Why it is P2 rather than P3: a session that tears down, reads full teardown, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads full teardown and leaves has no way to know it is six and not six hours. Fix shape (smallest first): the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. 2026-09-25 (Peti's retirement): nothing to remove for peti-felhom-86d37d — its host was deleted 2026-07-15 and ep0's live wg show carries no peer beyond the demo boxes', drill-r50 and the operator OOB (audits/retire-peti-2026-09-25/A1-ep0-before.txt); whether that July delete removed a peer, or none existed, is not recorded. -- 2026-09-28: the Day-0 test install's peer (drill-g0276, key Ly0yjK…, 10.77.0.5) was checked on ep0 the day after its customer DELETE: gone from wg show, absent from /etc/wireguard/*.conf (control: a live key found), no namespace or file names the customer — the hub's sync removed it; nothing was removed by hand. audits/logins-nvme-2026-09-28/E/. MEASURED AGAIN 2026-09-30 (a HOST delete, not a customer delete; new-household drill): the host was deleted 07:23:03Z and wgsync: pushed 4 peers at 07:23:51Z — 10.77.0.5/32 gone from ep0 48 s later, the other four peers unchanged (audits/evidence-drill-new-household-2026-09-30/teardown/layer4-ep0.txt). READY - rank P2-MEDIUM; owner: CC (hub) Re-ranked 2026-10-03: P2->P3: measured as seconds to minutes and asynchronous by design; operator-only log wording. — — CC
R-728 Hub & operator P3 [P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link. MEASURED 2026-09-30 on Tester-2: the hub logged Customer config created: Tester-2 twice in the same second and two self-bind mints (hashes c40df008…, 6a1cbef4…); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). Fix direction: make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. READY — rank P3-LOW; owner: CC (hub) — — CC
R-882 Hub & operator P3 Longhorn on DooPlex could not grow a volume online: its instance-manager (116 days up) called a host process that no longer existed — nsenter: cannot open /host/proc/196610/ns/mnt on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing hub-data to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back attached/healthy in 110 s. Why the cached PID went stale was not established (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. audits/hub-db-offsite-2026-10-05/partA/step1-*.txt OPEN — Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart operator
R-883 Hub & operator P3 8 DooPlex workloads run an image by a moving tag (:latest or none), so any pod restart is a silent upgrade. Measured 2026-10-05: the Longhorn restart restarted zipline on ghcr.io/diced/zipline:latest (pull Always), which pulled 4.8.0; 4.8.0 refused its database (cannot safely migrate from prisma to drizzle: expected migration 20260508022000 … was not applied) and crash-looped. Fixed for zipline by pinning 4.7.0 (homelab-manifests 90f60e4, 4c8ec7a; 4.7.0 applied the four missing migrations; a dump from before is kept out-of-band). The other 7 were counted, not named or changed. OPEN — List the 7 (kubectl get deploy,sts -A images without a fixed tag), pin each to the running version, and let Renovate move them operator
R-92 Hub & operator P4 Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable READY (XS) — Widen precision when retention becomes customer-visible CC
R-264 Hub & operator P4 Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed". Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (operator_key_configured) could not be mistaken for having decided what the hub should do with the rest. The list, grouped by what a consumer would be for. (a) Guest-network health — guest_net and its seven children (checked_at, has_route, dhclient_alive, heal_succeeded, heals_last_hour, last_heal_at, damped). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed dhclient took a tunnel down for 1 h 15 m (audits/INCIDENT-guest-dhclient-killed-2026-07-20.md); a recurring-heal signal is exactly what would have surfaced it. This is the strongest candidate of the twenty-one. (b) selfupdate_pending + selfupdate_pending_version — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. (c) mgmt_plane.healed_recently — bounded: the hub DOES alarm on the privsep_healed_at timestamp beside it, so the recurring-clobber signal is not lost, only this flag. (d) restore_tests.mount_parity + mount_inventory — R-262's subject; the verdict is not lost (a mismatch fails before Pass is set) but the hub cannot tell a full-fidelity pass from a boot-only one. (e) pbs_dr.applied_at. (f) Controller-side: config_hash, reporting_disabled, stacks, storage.migrated_to, backup.last_db_dump, backup.last_integrity_check — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. Deliberately not decided in the G-1 session, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. ⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today. The rulings were made on 2026-08-12; this row still read READY — owner Viktor and the gate's twenty entries still all said "arguably owed", so a session told to "re-read the dispositions from the register" would have found none. They are written down now, which is the point of writing them down. THE COUNT WAS ALSO WRONG: this row says twenty-one; the gate's allowlist held twenty, measured. Twenty is the number the dispositions below account for, exactly. (1) BUILD A READER — four groups, fourteen facts. (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. (2) NO READER WANTED — five facts, now recorded as not consumed, DELIBERATELY with the ruling and its date in scripts/wire_contract_gate.py, each with its own reason rather than a bare refusal: mgmt_plane.healed_recently (the hub already alarms on the timestamp beside it), pbs_dr.applied_at (pbs_dr.state is the verdict; the timestamp alone is the attempt-read-as-result trap), config_hash (the hub authors the config and knows its own generation), stacks (the app view is built from the purpose-built app_telemetry wire), storage.migrated_to (box-local bookkeeping with no hub-side intent to reconcile against). The emitters are deliberately left alone — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. (3) reporting_disabled, decided on its own merits: RECLASSIFIED redundant — health.status = "disabled" travels in the same minimal report, is decoded into reports.health_status, and IS rendered. The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321. PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319. Its eight allowlist entries are removed (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose 182 → 190 and skipped fell 88 → 80, which is the positive control that the wiring is real. WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader (selfupdate_pending, selfupdate_pending_version, restore_tests.mount_parity, restore_tests.mount_inventory, backup.last_db_dump, backup.last_integrity_check) — counts measured from the allowlist, not estimated. Only ONE reader was built on purpose: four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319) — owner Viktor — — operator
R-292 Hub & operator P4 The artifact-save flash conflates three different facts, and a failing test found it rather than a reading. artifact_sha_invalid reads "the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid" — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting artifact_sha_invalid where it expected artifact_unverifiable: resolveArtifactSHA ran first and swallowed the distinction. Worked around in v0.102.0 by ORDERING — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — but the underlying message is untouched and still conflates on its own paths READY (XS) — NEW 2026-08-09 — Split it into "version not found", "registry unreachable" and "invalid sha" CC
R-372 Hub & operator P4 A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator. Written down 2026-07-15, in audits/CAMPAIGN-6E-2026-07-15.md:128 (F-6E-1), whose disposition ends "Optional product idea: surface 'tier-2 has never produced a copy (source missing)' more prominently in the operator UI — not filed." The finding it sits on was correctly judged demo-data churn rather than a product defect (the code warns loudly and does not silently succeed), but the surfacing idea was never carried anywhere. Age when filed: 38 days — the oldest gap this sweep recovered. OPEN — LOW — Decide whether "never produced a copy" deserves its own operator surface, distinct from "last copy failed". CC
R-402 Hub & operator P4 The off-site integrity verdict and its depth are published to the hub and NO hub surface reads either. offsite.last_integrity_ok has been on the wire since controller v0.227.0 and offsite.last_integrity_depth since v0.228.0; both are allowlisted in scripts/wire_contract_gate.py with their reason, which is why the gate is green rather than silent. The order is deliberate and is the opposite of the one that produced R-331: publish the value first, build the display when someone decides what the screen should say. R-331 removed a hub Backup card that rendered Integrity Unknown for every customer forever from fields nothing wrote. The depth is not decoration: "checked, OK" means two different things at structure depth and at 100%, so a card showing the verdict without the depth shows the same words for a check that re-read every byte and one that only read the index. WHAT HAPPENS IF NOBODY ACTS: the operator can only answer "was this customer's off-site store verified, and how deeply?" by reading that box's own log. OPEN — SMALL, needs a HUB decision first — Decide what the hub screen should say, then model both fields hub-side and delete the two allowlist entries together. offsite.last_integrity_check is already decodable and is not allowlisted. Viktor decides, CC builds
R-445 Hub & operator P4 [P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation. MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's /apps/nextcloud page still reports Deployments, Avg Memory 208 MB, P95 Memory 280 MB and Suggested Limit (P95x1.2) = 352 MB, plus three MariaDB io_uring rows under Known Issues attributed to demo-hp. The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere — and Nextcloud is a real catalog app whose limit someone may act on. RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row: the hub offers POST /apps/nextcloud/reset-telemetry whose own confirm reads "Delete all telemetry data for nextcloud? This cannot be undone." — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. The one-line command is recorded in the audit doc so it is a decision, not a task. The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: VIKTOR rules, CC implements. audits/SPIKE-app-update-2026-09-01.md OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements Re-ranked 2026-10-03: P3→P4: operator-only recommendation surface; no household meets it. — — CC + operator
R-451 Hub & operator P4 [P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is. Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. architecture/09-update-architecture.md §6, §8.4 BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies. The controller's payload carries no image (internal/report/types.go L98–103) and the hub's Store.SaveReport (hub/internal/store/store.go:965) denormalises only container counts — but the hub stores the raw report JSON whole, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in 09 §6.2–6.3; the payload question is 09 §3b Q7. -- RULED 2026-09-23 (09 §3 decision 18): the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. Built later, when the fleet grows — Q7's recommendation, confirmed. DEFERRED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fleet sweep (report field, hub denormalisation, fleet page) is ruled and not built — deferred until the fleet grows (09 §6.3)) — RULED — build deferred until the fleet grows; owner: CC — — CC
R-544 Hub & operator P4 [P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody. MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged host deleted: tester-1-33b6a9 (escrow deleted: true). The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". Nothing is broken; the log is. An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. Fix shape: log what happened — escrow custody demoted to retained (host delete) — and keep the boolean's name out of operator-facing text. READY — rank P3-LOW; owner: CC (hub) Re-ranked 2026-10-03: P3->P4: operator log wording only; no household meets it. — — CC
R-599 Hub & operator P4 [P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so. FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused both POST /configs/<id>/delete (409, host … is ONLINE) and the host delete (deletable:false) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a report-staleness window, and the window is 45 minutes — manifests/hub.yaml sets alerting.stale_threshold: "45m", which hostStatus() reads (ok under it, stale over, down at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. Measured the boring way, and worth recording: this row first said 30 minutes, because monitor/host_staleness.go's literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. The consequence is not theoretical: a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. Fix shape (smallest first): the 409 body says how long it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and runbooks/target-selection.md's drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. READY — rank P3-LOW; owner: CC (hub) Re-ranked 2026-10-03: P3->P4: operator teardown comfort; refusal is correct. — — CC
R-688 Hub & operator P4 [P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare. The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (hub/internal/web/customer_delete.go deleteCascadeAcks), while commitCustomerReset has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring peti-felhom, whose config carried a Cloudflare tunnel token and API token (sajatfelhom.hu): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. Fix direction: either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. audits/RETIRE-peti-2026-09-25.md -- HALF DONE 2026-09-25 (hub v0.125.0): the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (audits/night-2026-09-26/F/). The Cloudflare leg itself is NOT built. NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build) Re-ranked 2026-10-03: P3→P4: the dialog no longer promises it; what is left is operator comfort. — — CC + operator
R-719 Hub & operator P4 [P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it. MEASURED 2026-09-29 (new-household drill, tester-1): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (selfbind_mint.go callers hosts.go:908, configs.go:850, customer_reset.go:162) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. Fix direction: send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: audits/evidence-drill-new-household-2026-09-30/ phase0/operator-steps.txt. CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable: a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: audits/evidence-fixes-first-tester-2026-09-30/``partD/. WAITING-ON-OPERATOR (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another) — — operator
R-814 Hub & operator P4 PBS-storage-1 (u629193, box 611421) still status=active, 19.9 MB VERIFY (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR operator console Delete the box operator
R-844 Hub & operator P4 The household's OS-update line exists only on the hub's customer timeline. 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so os_update_applied is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (mail.event.os_update_applied) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt READY — owner: CC — — CC
R-855 Hub & operator P4 The hub's start log prints "after -1 healthy ring-0 night(s)" for the TEST override OS_DOCKER_APPROVE_NIGHTS=0 (the internal "none" value; the WARN line before it is right). Cosmetic, TEST configuration only. Fix: print 0. audits/os-docker-crash-2026-10-04/partB/ (hub log 16:27:42) READY — owner: CC — — CC
ID Category Sev What State Blocked on Next action Owner
R-784 Business & legal P2 [P2-MEDIUM] SparkyFitness's licence forbids commercial use: "may not be used, directly or indirectly, in any product, service … intended for … commercial advantage … without prior written permission from the author" — and Felhom is a paid service that offers it in its catalog. READ 2026-10-01 (audits/visitors-2026-10-01/C/C0-license.txt): a custom licence (GitHub: NOASSERTION), the same at the pinned tag v0.17.3 and at main; termination clause 7 ("cease all use"). SparkyFitness was named as wger's replacement for fitness. Needs (operator): (A) hide it from new installs (lifecycle: hidden) until the author gives written permission, and ask; (B) ask first and keep it offered meanwhile; (C) keep it. Recommended A. If nothing is decided it stays offered. UPDATE 2026-10-02 — DECIDED (09 §3 decision 65, option B): SparkyFitness stays offered; the operator asks the author. The request is drafted (NOT sent): audits/licences-2026-10-02/EMAIL-DRAFT-sparkyfitness.md — the author publishes no e-mail; the routes are the project's Discord (private, recommended) or a GitHub Discussion. Trigger: no written permission before the first paying customer → lifecycle: hidden (a STATUS standing item). WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (send the request; the answer) — — operator
R-789 Business & legal P2 [P2-MEDIUM] Tandoor's licence is AGPL-3.0 WITH the Commons Clause: it forbids selling "a product or service whose value derives, entirely or substantially, from the functionality of the Software" — fees for hosting or support included. READ 2026-10-02 at 2.6.15 (audits/licences-2026-10-02/TABLE.md). Felhom charges for installing and caring for the household's apps; whether that value comes "substantially" from Tandoor is the question. Needs (operator): keep (the fee is for the box, not Tandoor), hide for new installs, or ask the authors. Nothing changed meanwhile. RULED 2026-10-02 afternoon (09 §3 decision 66): treated like SparkyFitness — stays offered; the operator asks the authors for written permission; without it before the first paying customer Tandoor is hidden (lifecycle: hidden). STATUS "Before the first paying customer". WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (the request; trigger: first paying customer) — — operator
R-802 Business & legal P2 [P2-MEDIUM] A lawyer reviews the non-OSI licence list before the first paying customer. Operator ruling 2026-10-02 (09 §3 decision 66): Tandoor (R-789), SparkyFitness (R-784), Emby, n8n, Plex (kept — R-790..R-792), the EE/BUSL parts (R-793), redis 7.4 (R-794); the table is audits/licences-2026-10-02/TABLE.md. STATUS "Before the first paying customer". WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (trigger: first paying customer) — — operator
R-813 Business & legal P2 [P2] The website collects personal data but publishes no privacy notice, no terms and no imprint. CHECKED 2026-10-03 (read-only): website/ holds nine Hungarian pages and one English page; none is an ÁSZF, an adatkezelési tájékoztató or an impresszum, and no page links to one (ASCII-fragment search aszf, adatkezel, impresszum, impressum, privacy over website/; positive control: the same search finds adatkezel in the contact form). The contact form makes the visitor tick a data-processing consent (website/kapcsolat.html:123-128) whose text names no controller, no retention and no rights, and links nowhere. The papers around it — contract, data-processing agreement, billing — are the intention R-809 in ROADMAP.md. WAITING-ON-OPERATOR — the operator writes or commissions the texts; CC drafts on request; owner: operator — — operator
R-89 Business & legal P4 Retention as a per-customer commercial policy on the hub READY (increment 2) — Policy object + reconciler → ep0 prune job; keep box tokens write-only CC
R-793 Business & legal P4 [P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules). READ 2026-10-02 (audits/licences-2026-10-02/TABLE.md). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. Watch: never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the -enterprise image, and re-read on each major. WATCHING — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them. — — CC
R-794 Business & legal P4 [P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm. READ 2026-10-02 (audits/licences-2026-10-02/TABLE.md). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. Needs: a ladder step per app to valkey or redis 8, through the harness — no hurry. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted. — — CC

Process & tooling — 65 rows (P3 4, P4 61)

ID Category Sev What State Blocked on Next action Owner
R-565 Process & tooling P3 [P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page. FOUND 2026-09-17 by slice 1 release C (R-556, controller v0.250.0): after the extractor and the tests were green, a by-eye review of the English renders found six Hungarian fragments still in JavaScript strings — „, majd a(z)” and „FIGYELEM:” in the storage decommission dialog, „jelenlegi:” on the drive-init list, the uptime units „mp” and „p” and the count word „ db” on the debug page. All six were converted by hand; no test failed on any of them, because TestI18nEnglishPages looks for Hungarian letters and the extractor's ASCII word list (i18n_extract.py ASCII_HU) is used by neither test nor gate. Release B's review had found more of the same kind (Konfig, Megtartva, helyi, pl., Befejezve, automatikus, jelenleg:, kedd/szerda/szombat, szint). Fix shape: run the ASCII word list over the English renders in TestI18nEnglishPages (after the data mask), with a negative control on an English sentence and a decoy planting „mp” in an English value; extend the list with the words releases B and C found. READY - rank P3-LOW; owner: CC — — CC
R-578 Process & tooling P3 [P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package. FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: UpdateOffboxStatus holds the settings WRITE lock while it runs its callback, boxLang() reads the language through the READ lock, and sync.RWMutex is not reentrant — so the off-site run's final status write DEADLOCKED, holding the settings lock, which would wedge everything else on that box that touches settings.json. The only symptom was go test ./internal/backup/ going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; TestNoteHelpersAreNotCalledUnderTheSettingsLock (internal/backup) now names the file and line in a second. What is still open: that test covers internal/backup only, and it knows only the note/noteErr/boxLang helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. Fix shape: promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a settings.Update* function"; or give Settings a re-entrant read path and remove the class. READY - rank P3-LOW; owner: CC — — CC
R-586 Process & tooling P3 [P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw. FOUND 2026-09-18 starting slice 4 (R-559): running scripts/iso/test/bootstrap-modes.sh unchanged at 183727db9c44 reported FAIL: R-496: banner painted to the console seam and FAIL: R-496: banner names the Tulajdonosi jelmondat. Cause: the script paints each banner with > "$CONSOLE_DEV". On a real console that is a character device and truncation is a no-op, so every paint appears; pointed at a plain file, as the harness did, each paint TRUNCATES. Commit c033b3b (ISO 1.28.0, R-535, 2026-09-16) added print_bound_banner, which paints immediately after the pairing banner in the same invocation — from that commit the pairing banner was wiped before the check read it. c033b3b did not touch the harness. Why it survived: the harness is in NO gate and NO CI run — not in repo_gates.py, not in .gitea/; it runs only when a person runs it, and between 09-16 and 09-18 nobody did. FIXED in the same session: the harness points FELHOM_CONSOLE_DEV at a FIFO with a background reader, which restores device semantics (opening a FIFO with > truncates nothing) and lets a test see EVERY paint — which slice 4's golden checks then needed anyway. Production code untouched. What is still open: the harness remains outside every gate. It needs a container, so it cannot join repo_gates.py --fast, which is what both the pre-push hook and CI run — meaning a non-fast entry would still never execute. Fix shape: either give CI a container-capable job that runs it, or make the ISO release gate's G16 the place it is required (done for G16 this session — so it now runs at least once per ISO release, which is better than never but later than a push). READY - rank P2-MED; owner: CC Re-ranked 2026-10-03: P2->P3: test harness gap; now required at each ISO release gate. — — CC
R-733 Process & tooling P3 [P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap. MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read swap: 512; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges anon against the limit and never reads memory.swap.current. Needs: the golden's swap read and recorded; the box walk and the harness report memory.swap.peak beside anon; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). READY — rank P3-LOW; owner: CC (harness + golden evidence) — — CC
R-93 Process & tooling P4 drill-r50 is both a blocked customer and the only drift fixture FACT 2026-09-13 (R-461): the fixture is GONE — qm list is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture". READY (XS) — Retire it for a synthetic fixture, or unblock + silence per-customer CC
R-129 Process & tooling P4 Every doc says demo-hp has "no baked SSH key" and needs the G1 break-glass password — but ssh -o BatchMode=yes demo-hp authenticated by key, first try, 2026-07-31 READY (XS) — Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct CLAUDE.md, runbooks/target-selection.md:41-42, runbooks/workspace-CLAUDE.md and felhom-agent/CLAUDE.md together — or remove the key if it was not deliberate CC
R-161 Process & tooling P4 The volume-persistence gate is enforced by CONVENTION, not automatically. The catalog repo has no CI of any kind (.gitea/workflows, .github, drone/woodpecker — searched, none exists). OPEN — REDUCED SCOPE — open (operator ruling 2026-08-02) a second person touching templates RULED. Both obvious enforcement points were rejected for measured reasons. Controller-side at template load: rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean including papra — it would pass on the exact defect it exists to catch; the property is decidable only at runtime. CI: rejected for now — neither repo has any, and there are no users yet. SHIPPED instead (app-catalog-felhom.eu fd7747d): scripts/catalog_gates.py, ONE entry point running all three gates, non-zero exit on any failure, mandated in the catalog's CLAUDE.md the way site_gates.py is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — site_gates.py is run, R-29's three orphans are named nowhere and have stopped nothing. What remains open is only the automatic half: this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does UPDATE 2026-08-02: catalog_gates.py gained --fast (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and .githooks/pre-push now runs it. The automatic half now has a designated successor row: R-168 (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run UPDATE 2026-08-02 (second): the automatic half now EXISTS — R-168's runner executes catalog_gates.py --fast on every push to this repo (measured: run #1, image-pin gate OK — 53 templates, with the two runtime gates announced as skipped and their own output absent from the log). This row's original scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run operator
R-162 Process & tooling P4 docker diff is the gate's only witness, and its failure mode is quiet. The gate's power comes from docker diff excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where docker diff is unsupported or lies, the gate degrades to the mount-occupancy and writability legs and would not say so. WATCHING — a limitation, not a defect — It fails closed: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships CC
R-169 Process & tooling P4 CI can only report, because there is no gate in the road. Every felhom repo pushes straight to main with no pull request, so there is no merge for a status check to stand at. R-168's runner therefore notices a broken push after it has landed WAITING-ON-OPERATOR (a working-style decision, not a defect) an operator ruling Making CI blocking requires two things this task deliberately did NOT do, because both change how the operator works and that is not a task's call: (a) branch protection on main, and (b) a pull-request workflow instead of direct-to-main pushes. The cost is real — every change would need a PR, which for a single-operator project may be worse than the disease. The current arrangement is two nets, and it is not nothing: .githooks/pre-push REFUSES locally, and R-168's runner NOTICES when that hook was skipped or was never armed in a clone, and emails. The honest gap is the window between a --no-verify push landing and the operator reading the alarm. Decide only if that window ever actually costs something operator
R-206 Process & tooling P4 The build-cache cap and the weekly prune exist only as a hand-edited /etc/docker/daemon.json on DooPlex — not in Ansible, so a rebuild loses them. The node_housekeeping role must also carry the prune, which today it is forbidden to run READY (M) — NEW 2026-08-05 — The spike validated the recipe; this row builds it. Three parts. (a) Template /etc/docker/daemon.json with the policy array form — the flat form ({"gc":{"reservedSpace":…}}) is SILENTLY IGNORED, measured: the daemon starts, logs nothing, and docker buildx inspect still reports the built-in defaults. The oracle is docker buildx inspect, never dockerd --validate — the validator returned configuration OK for a bogus key AND for the config that then crashed the daemon (filter takes one value per policy entry, not an array; error initializing buildkit: filters expect only one value). (b) Narrow the role's Docker ban (node-housekeeping.sh.j2:10-14) to permit exactly docker builder prune -af and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. The measured prune is SYNCHRONOUS (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — unlike containerd's image GC, so it needs no settle_imagefs equivalent, but it MUST measure the filesystem rather than trust the command: prune claimed 156.9 GB and the filesystem returned 150.35 GB, the 6.5 GB gap being layers still shared with images. (c) A restart-safety note in the role: a bad daemon.json takes the daemon down AND leaves the unless-stopped dev containers stopped — they needed a manual docker start — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: audits/SPIKE-dooplex-buildcache-2026-08-05.md CC
R-208 Process & tooling P4 Every Felhom Go build re-downloads its modules because ARG VERSION sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN — Measured, not inferred. All 208 retained go mod download records carried Usage count: 1 — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the same --build-arg VERSION → RUN go mod download CACHED; the same tree with a new VERSION → executed. COPY go.mod ./ stays CACHED either way, which is the tell: a COPY's key is content-based, while a RUN's key includes the stage environment, and ARG VERSION/ARG GIT_COMMIT are declared before the download in felhom-controller/controller/Dockerfile. Since every real build passes a fresh version, the layer is invalidated every single time. felhom.eu/hub/Dockerfile has the identical defect (ARG VERSION/ARG BUILD_TIME above COPY go.mod go.sum* → RUN go mod download) — and because both Dockerfiles produce byte-identical buildx du description strings, the 208 records are a COMBINED count and must not be attributed to one project. Fix shape (one line each, not applied here): move the ARG VERSION/ARG GIT_COMMIT/ARG BUILD_TIME declarations down to just above the final go build. Worth more than the cap and the move combined — the cap bounds the symptom, this removes the source. build.sh's rm -rf + cp -a and its host-side go mod tidy were ruled out by fingerprinting: the tree is byte-identical across runs and tidy is a no-op CC
R-209a Process & tooling P4 The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated WATCHING — NEW 2026-08-05 the next DooPlex reboot Operator ruled explicitly: do NOT reboot DooPlex. Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). The distinction is stated rather than glossed: the MECHANISM is proven — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — but the CONSEQUENCE is not: that a real boot mounts /mnt/ssd_2 before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (RequiresMountsFor RE-MOUNTS rather than refusing — the ep0 lesson), and CLAUDE.md prefers a consequence assertion over a mechanism one. Two deliberate consequences: (1) the rollback copy /var/lib/containerd.pre-move-2026-08-05 (34.3 GB on /) STAYS until a reboot validates — which is why / sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. (2) validation is automatic and needs no one to remember it: felhom-store-postboot-check.service (oneshot, enabled, dry-run PASS at install) runs at every boot and writes RESULT: PASS/FAIL to /var/log/felhom-store-postboot-check.log, asserting positively that /mnt/ssd_2 is mounted, that containerd's root is on it, that /var/lib/containerd does NOT exist (the empty-store trap), that ≥100 images are visible and that both dev containers run. Next action: after the next reboot — planned or not — read that file; on PASS, rm -rf /var/lib/containerd.pre-move-2026-08-05 returns ~34 GB to / P3's prune already removed the urgency: / went 86% → 53% used and SSD1's Longhorn disk went Schedulable=False (DiskPressure) → Schedulable=True (18.85% → 50.32% available). The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. The numbers, measured (Crucial-SSD-240G, /mnt/ssd_2/data/longhorn, storageMaximum 235,148,750,848): available today 214,958,080,000 (91.41%); 25% floor 58,787,187,712. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves 63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured. But storageScheduled on SSD2 is 139,586,437,120 while df says only 20,094,939,136 is actually used — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at 13.00%, i.e. 12 pp BELOW the floor → Schedulable=False, which is exactly the failure that just took SSD1 out. Recommendation: do the move only together with setting storageReserved on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right. Mechanism, if it goes ahead: containerd's root in /etc/containerd/config.toml (the key is present but commented out) — not Docker's data-root, which would move only 0.62 GB. Guard: RequiresMountsFor=/mnt/ssd_2 on containerd.service and docker.service, remembering that RequiresMountsFor RE-MOUNTS rather than refusing (ep0-datastore-volume-move-2026-07-27) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a reboot. Full pre-analysis: audits/SPIKE-dooplex-buildcache-2026-08-05.md §P6 operator + CC
R-210 Process & tooling P4 Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by docker pull WAITING-ON-OPERATOR — NEW 2026-08-05 an operator ruling Nothing was deleted; this is a list, not an action. The registry was queried directly: felhom-controller has 76 tags in Gitea vs 207 locally, felhom-hub 45 vs 107. The 131 + 62 local-only tags are all OLD — controller 0.39.0–0.135.0 plus v0.35.0–v0.39.0, hub 0.9.0–0.57.0 plus v0.7.2–v0.13.0 — while everything from controller 0.136.0 and hub 0.58.0 upward IS in the registry and therefore re-pullable. Size the prize honestly before spending a decision on it: per-tag sizes sum to 139.29 GB, but that double-counts shared layers — docker system df puts the real dedup'd image footprint at 31.02 GB with 27.02 GB reclaimable, i.e. an order of magnitude less than the build cache P3 already returned. docker image prune -a would remove 343 of 345 (only redis:7-alpine and postgres:16-alpine are held by running containers). CC's view: not worth doing for the space — it buys ~27 GB against 199 GB now free, and its only real benefit is dropping unrecoverable clutter operator
R-230 Process & tooling P4 Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation. (a) A ruling is owed on auto-written staleness. The hand-written CLAUDE.md files are now clean of version literals and expired blocks — the gate enforces it — but MEMORY.md, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries 21 lines with component version literals, 5 with bare host addresses, and an entry still reading "demo boxes REMOTE till ~08-02" — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED: the three statements that were actively false were corrected — R-193 decision open (closed 2026-08-05), demo boxes REMOTE till ~08-02 (the box answers on the home LAN), OPEN R-25b (shipped 2026-07-21) — and gate check 6 now WARNs on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. The remaining 32 version literals and 4 host addresses were deliberately left for that loop. What is still owed is the bulk-correction ruling. Correcting the premise: the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) CLOSED 2026-08-06 (close-out) — the workspace-root CLAUDE.md is now a relative symlink to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong), for two files byte-identity as before, so a clone elsewhere is unaffected. Proven, not assumed: three fresh sessions logged session_start for the link path, and a fourth with no tools at all quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) The spec-as-failing-test pilot, approved in principle and not started (was R-229(d)). READY — owner Viktor — — operator
R-288 Process & tooling P4 The capability map is too long to be read, and that is why it stops being true. architecture/00-capability-map.md is 134 642 bytes / 19 456 words across 99 table rows in only 159 lines — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is 3 024 words, the offsite-password-recovery row 1 087, the unattended-restore-proof row 971, the app/guest-network-failure row 904. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads 2026-07-16 against evidence corpus @ felhom.eu tip 4b18cc5 (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. This is the project's memory, so restructuring it is surgery and wants daylight — filed, deliberately not attempted in the 2026-08-09 session. What the shape should probably be: one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from READY (M) — NEW 2026-08-09 — Do not fold this into another session; it needs its own. SECOND CONCRETE COST, 2026-08-10 — and it is a different failure mode from the first. An execution record — 33 package deletions on the operator's rule — was undiscoverable for two days because it lives inside the row about the Configuration page being slow. Two sessions searched for it: one reported "no register row records a package prune", the other exhausted the Gitea logs, the activity feed and the schema before concluding it might be unestablishable. It was in OPEN-ITEMS.md the whole time. The first cost (2026-08-09) was two records that looked contradictory and were not; this one is a record that could not be found at all. Illegibility now has two measured costs and they are different in kind: prose rows make claims ambiguous, and rows-about-other-things make facts unfindable. The rule this earns is in CONTEXT.md: a record that lives inside a row about something else has not been recorded Viktor
R-290 Process & tooling P4 Most capability-map rows that back a green dot cite no evidence document at all — measured, 20 of 28 probed. The page's Walked means "done end to end on real hardware, evidence on file". Extracting the evidence column for the 28 rows behind the page's claims found a tests/ or audits/ path in 8; the other 20 carry prose only. Consequence, applied this session: of 32 claims the page drew as Walked, 12 were downgraded to Built because no walk document exists for them — install.installer-by-tag, use.lifecycle, drives.enrol, drives.migrate, backup.tier1, backup.whole-machine, backup.restore-proof, fault.selfheal, fault.operator-email, fail.drive-filling, fail.lost-recovery-code, fail.hub-down. This is not a claim that those twelve are false — several are near-certainly fine — it is a claim that nothing on file distinguishes them from an opinion, which is exactly what the status word promises. The gate now enforces it going forward: scripts/check_stands.py fails on status: walked with no evidence: source. What is owed: either a walk document per row, or an honest demotion in the map itself (the map is the source; the dataset only follows it) READY (M) — NEW 2026-08-09 R-288 The dataset was corrected; the capability map itself still says PROVEN-LIVE for these rows and is the thing to fix Viktor
R-315 Process & tooling P4 The wire-contract gate's positive control FAILS on the new wire: it checks name-presence, not decodability. R-311 declared hub -> agent (GET /escrow/retained) as a fourth ROOT, and the gate's tag count rose 174 → 182, so the fields ARE inspected. But renaming the agent-side superseded_at json tag to superseded_at_RENAMED still passed — because the string superseded_at also occurs as a map key in the agent's local-API response, and the check is a repo-wide name search. The gate documents this ("name-reachability is not use"), so it is a known limit rather than a regression — but it means declaring this wire bought documentation, not enforcement, and a report that claimed coverage would have been wrong. The mutation was asserted to have applied before the run READY (M) — NEW 2026-08-12, RANK 3 R-311 Make the check resolve the RECEIVER'S mirror type and compare field-by-field, or state per-root which kind of check it got. A gate whose positive control fails is an instrument nobody has calibrated CC
R-325 Process & tooling P4 The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other. customer_copy_vocab.py is the single list; hub_copy_gate.py imports it. felhom-controller/controller/scripts/retrieval_promise_gate.py still carries its own STEMS literal, because the session that created the shared module was under a hard end-state requirement to leave felhom-controller untouched — its target box was being re-deployed the same evening. Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. So the gap is instrumented rather than left open: hub_copy_gate.py READS the controller gate's STEMS and FAILS if the two disagree — single-source semantics tonight without a cross-repo edit. Watched failing: removing one stem from the shared list produced "the shared vocabulary is no longer shared" with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is INCONCLUSIVE (exit 2), never a pass — the G-1 lesson. This is a scaffold, not the destination READY (S) — NEW 2026-08-13, RANK 3 R-299, R-324 Make retrieval_promise_gate.py import felhom.eu/scripts/customer_copy_vocab.py and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads CC
R-327 Process & tooling P4 The standing picture still describes a defect that has been fixed twice over. Found by the first run of unproven.py (R-326), which is the argument for having built it. where-felhom-stands.yaml's claim.code-naming is status: partial and its title reads "The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show" — both halves of which are now false. The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new reenroll mail kind), and the third near-homograph on 2026-08-13 (R-323). NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule: "A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it." Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met READY (S) — NEW 2026-08-13, RANK 4 R-295, R-323, R-326 Decide the capability-map status for the naming arc, then let the dataset follow it. Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet, so walked would be an over-claim; built is probably right, and the title needs rewriting either way because it describes a defect rather than a capability operator + CC
R-346 Process & tooling P4 ActiveEnterTimestamp answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m. Found while taking R-341's first dated check. ep0's proxmox-backup-proxy has MainPID=551655 started 2026-08-18 09:51:04Z (ps -o lstart=), but systemctl show -p ActiveEnterTimestamp reads 03:54:54Z and NRestarts reads 0 — because the 4.2.5-1 upgrade re-exec'd the daemon rather than restarting the unit, so systemd never observed a stop. Anyone anchoring "when did this proxy generation start" on ActiveEnterTimestamp would divide 388 descriptors by 52.1 h instead of 46.2 h and report 178/day instead of 201.6/day — ~12% low — while every field consulted looks healthy and consistent. This is the workspace rule's own case, in a new place: ask of a timestamp what exactly must have happened for this to be set? Here the answer is "the unit entered active", which is not "this process started". NRestarts=0 is the tell, and it reads like reassurance. READY (XS) — NEW 2026-08-20 — R-341's command already uses ps -o lstart= -p $MainPID and is correct; the risk is a future reader "improving" it to a systemd property. Add the reason as a comment beside that command in the R-341 row (done), and check whether any other slope or uptime check in the repo or in scripts/felhom-tenantsync.sh anchors on a systemd timestamp where it means a process start. CC
R-364 Process & tooling P4 Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times. (1) 2026-07-20, ssh → pct exec → bash -c, nearly a wrong "banner cleared" claim (felhom-controller/.claude/rules/ui-hungarian.md:19-22). (2) 2026-08-13, kubectl exec … sh -c grep returned 0 for three strings that were present, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, tar -tf rendered őszibarack.md as \305\221szibarack.md; recording the fixture's name bytes from that listing would have been wrong. NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record. OPEN — LOW — PROPOSED, NOT BUILT: a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. CC
R-374 Process & tooling P4 Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement. audits/CAMPAIGN-12-class-sweep-2026-08-08.md:123: "'Names a route' is a judgement, not a predicate — two readers could disagree on the borderline cases, and three of the 19 were called borderline and left unfiled." The disclosure is honest and is exactly the right thing to write; what is missing is WHICH three. An unnamed borderline case cannot be re-judged by a second reader, which is the only remedy a judgement call has. Age when filed: 14 days. OPEN — LOW — Name the three in that document, or file them as one row listing them. No code. CC
R-377 Process & tooling P4 CONTEXT.md's standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them. Measured 2026-08-22: CONTEXT.md is 217 KB, of which 187,913 bytes — 86% — is a single ## Standing rulings section carrying 39 S- ids and 153 bullets under one heading. This session deliberately did NOT compress or split it, and the reason is the ruling itself: PROMPT-TEMPLATE.md §3.4 and this file's own contract say the decision log is dated, never edited afterwards, and it is the only place that answers "has this been proposed before, and why did we say no?" — compressing it destroys exactly that. With no per-ruling delimiter, any mechanical split risks cutting a live ruling from its reason, which is the failure this whole arc is correcting. So the problem is navigational, not volumetric, and the fix is structural: give each ruling a sub-heading with its S- id and date. Then it can be linked, cited and found without a single word being edited. 13 mentions of SUPERSEDED already sit inside that blob and cannot be separated from live text safely today. OPEN — LOW R-369 Add per-ruling sub-headings only. Do not compress, do not reorder, do not edit any ruling's text. CC
R-392 Process & tooling P4 No architecture document covers the two-AI workflow. documentation/architecture/ holds eight documents and all eight cover the product — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and .claude/rules/ are scoped, or why. The absence was found by trying to fill the template field, not by a survey: the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. OPEN — LOW — Write one architecture document for the agent-tooling layer: which AI owns which artifact class (TASK-*.md, RUNBOOK-*.md, validation), how skills are scoped and installed, what belongs in a CLAUDE.md versus a skill versus a rules file, and the reasoning for each boundary. The rules themselves already exist in skills/felhom-doc-authoring/SKILL.md; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. CC
R-393 Process & tooling P4 A decision-log skill for unattended runs was considered and deliberately deferred. Filed 2026-08-25 by the session that added the five process skills, so the deferral is a decision on the record rather than a thing that was dropped. The gap it would close: an overnight or unattended run makes dozens of decisions and the operator can only reconstruct them by reading the whole transcript, which is exactly what nobody does. The proposal is an appended row per decision — what was chosen, why, the evidence pointer, and the result — so a long run is reconstructable in a page. Why it was NOT built with the other five: the other five are text files that need nothing but the existing installer glob. This one needs a helper script to append rows and a storage convention for where the log lives and when it is rotated, which makes it an implementation task with its own acceptance criteria, not a skill file. OPEN — LOW — Decide the storage convention FIRST — most likely a per-session file beside the session's evidence directory, never REPORT.md, which is overwritten every session (the R-341 shape). Then the skill, then the helper. Check it does not duplicate felhom-handoff, which already owns the end-of-session note; a decision log is the during-the-run half and the two must point at each other rather than overlap. CC
R-394 Process & tooling P4 felhom-build-deploy/SKILL.md is 179 lines, over the 150-line limit its own repo now enforces. Found 2026-08-25 by scripts/check_skills.py on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. It is not edited and not trimmed here: the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. It is a named single-entry exception in GRANDFATHERED in scripts/check_skills.py, printed as a WARN on every run, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. The rationale for the limit — attention thins across the excess, so the lines that matter are not the ones that survive — is in skills/felhom-doc-authoring/SKILL.md §5. OPEN — LOW — Trim felhom-build-deploy/SKILL.md under 150 lines in a session that can VERIFY the commands it keeps, then delete its GRANDFATHERED entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. CC
R-420 Process & tooling P4 controller_gates.py could not express a NON-BLOCKING gate before 2026-09-01 — every registered gate's non-zero exit failed the run, so the only way to add a check was to give it the power to refuse a push. That is the wrong trade for a notice that must fire at the moment a release is committed, when the golden legitimately cannot exist yet. The capability was added rather than the notice compromised (a fifth blocking field, False for exactly one gate; the felhom.eu runner already had the shape from its --scope work). Recorded because the ABSENCE was invisible: nobody had wanted a non-blocking gate before, so nothing said it was impossible. felhom.eu/scripts/repo_gates.py still has no blocking field — it has exemptible, which is a different idea (scope-dependent, not permanent). If a permanently-advisory gate is ever wanted there, it needs the same addition. OPEN — noted, not needed yet — — CC
R-421 Process & tooling P4 THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident. R-410 (a mkdir turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked. The 2026-09-01 decoy sweep read all 29 scripts and fooled 16. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. The shapes, so the next one is cheap to recognise: (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. The single largest cause was mundane: eight gates set their SCOPE with os.listdir (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. decoy_coverage_gate.py now refuses a new gate that ships without a decoy. OPEN — the class row; it stays open as the place the next instance is recorded — — CC
R-422 Process & tooling P4 reuse_refs_check.py only checks citations whose extension is one of go py html css yml yaml sh. A cited .md path that does not exist is invisible — MEASURED 2026-09-01: documentation/architecture/99-does-not-exist.md added to REUSE.md passed, while the .go control was correctly convicted. REUSE.md and the CLAUDE.md files cite .md paths routinely, so this is the common case, not an exotic one. Fix: widen PATH_RE, then walk the false positives it produces across all four repos — that pass is the work, not the regex. The decoy is kept in scripts/test_gate_decoys.py asserting TODAY's behaviour, so the day this is fixed the test fails and is updated deliberately. OPEN — — CC
R-423 Process & tooling P4 site_gates.py checks a hardcoded PAGES list of seven files; a new page is not scanned at all. MEASURED 2026-09-01: a new website/decoy-page.html carrying an emoji and no nav or analytics passed. felhom.eu/CLAUDE.md already tells the author to add new pages to the list by hand — which is the R-410 shape written down as a procedure. Fix: glob website/*.html and rethink the per-page exemptions (ANALYTICS_EXEMPT and friends) so the list becomes a list of EXCEPTIONS rather than a list of what is checked. OPEN — — CC
R-424 Process & tooling P4 one_register_gate.py: a real defect parked under the roadmap state idea is invisible to it. MEASURED 2026-09-01 with a correctly-shaped 5-column row. This is declared in the gate's own docstring as residual hole 1 — "the state column is a human judgement, and a defect written under idea looks exactly like a proposal to this gate" — so it is honest, not hidden. Recorded here because a hole declared only in a docstring is not in the register, which is this project's own standing rule. No cheap fix: distinguishing a defect from a proposal mechanically is the thing the gate cannot do. OPEN — declared, not hidden; recorded so it is not re-derived — — CC
R-425 Process & tooling P4 offbox_rename_gate.py scans a fixed three-entry FILES list. MEASURED 2026-09-01: NAS-mentés in a new backups_offbox_extra.html passed. The scope was correct when written and silently narrows every time the feature grows a file. Fix: scan the offbox feature's files by pattern, or assert the FILES list against a discovered set so a new file fails until it is classified. OPEN — — CC
R-426 Process & tooling P4 The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named. scripts/decoy_coverage_gate.py's EXEMPT map is debt, and this row owns it so it lives in the register and not only in a Python literal. Four kinds: (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — hub-copy, instructions, docker-v, image-pins; (b) blocked by an open hole and therefore un-assertable as rejecting — site (R-423), one-register (R-424), offbox-rename (R-425); (c) shared scripts whose decoy lives in felhom.eu and is counted there — reuse-refs, instructions, observations in the controller and agent runners; (d) no plausible decoy constructed yet — hostinstall, wire-contract, due-checks, published, image-resolvable, volume-persistence. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. The list is green today and shrinks; a NEW gate with no decoy fails immediately. OPEN — 20 names; group (d) is six untested gates — — CC
R-454 Process & tooling P4 [P3-LOW] Five internal/web test files have been gofmt-unclean for an unknown length of time, and nothing notices. MEASURED 2026-09-02: gofmt -l controller/internal/web/ reports backups_split_test.go, claim_code_naming_test.go, disk_health_test.go, r400_debug_routes_test.go, recovery_test.go — at the baseline commit 960d29b0612c, i.e. not introduced by v0.233.0 (both files added that day are clean). go vet does not check formatting and controller_gates.py has no formatting gate, so the only thing that would ever surface this is someone running gofmt -l by hand, which is how it was found. Not reformatted in the same session, deliberately — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. Small, and the cost of NOT having the instrument is the row: the count can only grow, and every future gofmt -l run produces noise that hides a real one. Fix is two lines: a gofmt -l gate in controller_gates.py plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: CC. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: code-formatting tooling only; no household meets it. — — CC
R-457 Process & tooling P4 [P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named. MEASURED 2026-09-03: TestGroupD_BadgeRendersOnBothSurfaces (shipped the previous day in v0.233.0) pinned a fixture catalog_since: "2026-07-18" and asserted the rendered string "Frissítés elérhető — 46 napja". The pure badge tests inject a clock; the RENDER test does not and cannot — it goes through the production templates, which call the funcmap entry updateBadge, which reads time.Now(). The suite was green on 2026-09-02 and FAILED on 2026-09-03 with "the behind badge is missing" on both surfaces, because the true answer had become 47. Fixed by DERIVING the fixture — catalog_since is computed as today minus 46 days, so the test asserts the real number through the real clock and cannot rot. THE CLASS, which is why this is a row and not just a fix: a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED — six other test files contain both a 20xx-xx-xx literal and time.Now(): internal/backup/offbox_test.go, internal/web/handler_export_upload_test.go, internal/web/r103_tier2_action_test.go, internal/web/dashboard_backup_card_test.go, internal/web/async_restore_test.go, internal/stacks/installed_test.go. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. The instrument that would end the class: run the suite once under a faked future date in CI and see what turns red. Owner: CC. felhom-controller v0.234.0 CHANGELOG READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: test-suite hygiene; no household meets it. — — CC
R-488 Process & tooling P4 [P3-LOW] go test ./internal/backup takes 5½ minutes: 89 off-site tests wait on real clocks. MEASURED 2026-09-13 (-v timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — TestOffbox*, TestOffbox3a*, TestOffboxRun*, TestR4xx* reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. Fix shape: the waits are waitForHealthy-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: developer speed only. — — CC
R-492 Process & tooling P4 [P3-LOW] cfg.Paths.HDDPath is empty on every box and still has readers; delete it. R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, systemInfo, the same fallback. The global now carries no information on any box and its deletion was deferred twice. Fix shape: remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: dead-code cleanup; every reader already falls back correctly. — — CC
R-502 Process & tooling P4 [P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner. MEASURED 2026-09-14: grep -rn bootstrap-modes scripts/*.py .gitea/workflows → nothing; the harness is run by hand in felhom-iso-assistant:trixie. While adding the R-496 checks, its fake hub's register reply turned out to carry no pairing_code, so print_pairing_banner returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. Fix shape: register the harness in repo_gates.py behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: test harness coverage; no household meets it directly. — — CC
R-507 Process & tooling P4 [P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen. MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): qm sendkey 332 tab did not move focus (both password copies landed in one field), mouse_move 1237 772 + mouse_button 1 did not move the cursor or press Next, while alt-n did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. Fix shape: measure QEMU input-send-event with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: release-test tooling; an operator click-through covers it. — — CC
R-551 Process & tooling P4 [P3-LOW] No Tier-0 box can put the escrow ceremony in the state R-546 fixes — paused AND connected to its agent — so the readiness branches are proven only by tests. FOUND 2026-09-17 while live-validating controller v0.246.0 (R-546). The branches (reminder bar held back while the agent's preflight is not ok; the waiting card on /backup/escrow; POST /api/escrow/start refused 409 before staging) need a box whose off-site tier is configured, whose escrow is NOT done, and whose controller reaches the agent. Measured on demo-hp: 9201 reaches the agent but is escrowed (the bar is off by design there; making it paused would be hand-set state on the standing demo box, and a real ceremony would supersede its live escrow — the hub keeps ONE host_escrow row per host, host_id PRIMARY KEY); 9202 is paused-capable but has no local-API token at all (its bootstrap.json holds only schema, customer.id, disposition), so its readiness is always UNKNOWN and the bar always shows. The fresh-bind window where this state occurs naturally (~17 min, chaos night Phase 0) needs a fresh install. What IS proven: five tests driving the real pages and handler through ServeHTTP with a fake agent, each red-proofed; and chaos night measured live that the agent's preflight is red for ~17 minutes after a bind and turns green by itself (evidence-chaos-night-2026-09-17/phase0-escrow-*). Fix shape: give the scratch guest a local-API token through the agent's own provisioning path, or walk R-546 on the next fresh install. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: test-coverage comfort; behaviour is unit-tested and no household meets it. — — CC
R-555 Process & tooling P4 [P3-LOW] The wire-contract gate counts a field as received when its name appears in a Go COMMENT on the receiving side. FOUND 2026-09-17 by CC adding the report's language field (controller v0.247.0): scripts/wire_contract_gate.py passed WITHOUT an allowlist entry, because receiver_tokens() tokenises whole files and the word „language" occurs hub-side only in a comment (hub/internal/web/configs.go:558, „this page's existing language"). The shape is the gate's own named failure class (name-for-fact, R-421): any English tag name that also appears in hub prose passes unread. The field was allowlisted by hand with this row named. Fix shape: strip // and /* */ comments (and template {{/* */}}) before tokenising; add the decoy „a tag whose name appears only in a receiver comment must convict"; expect a handful of currently-passing tags to surface — each is a finding, not noise. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: gate precision; tooling only. — — CC
R-564 Process & tooling P4 [P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them. FOUND 2026-09-17 by slice 1 release B (R-556): after the gate learnt English (EN_PATTERNS), seven English retrieval phrases on backups_remote, backups_escrow, backups_restore and backups_restore_wizard had NO Hungarian registration, because their Hungarian carries the verb particle after the verb („A távoli mentések csak akkor állíthatók vissza …", „a távoli mentések CSAK ezzel a kóddal állíthatók vissza", „csak a hiányzó fájlokat hozod vissza"). The stems (visszaállíthat, visszaszerezhet, visszahozhat, visszanyit) match only the joined form. The seven were registered in English with reasons (none is a false promise: two are preconditions, five describe the action on the same page). Fix shape: add split-form patterns to the Hungarian scan (állíthatók? vissza, `(hoz szerez nyit)\w* vissza`), register the Hungarian occurrences found, decoy with a planted split-verb promise. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: copy gate precision; the seven found sentences were reviewed and are true. —
R-569 Process & tooling P4 [P3-LOW] Four more API handlers pick their status code by matching ENGLISH words in an error — the same shape as R-553, one language over. FOUND 2026-09-17 while fixing R-553: controller/internal/api/router.go matches "protected", "not found", "not deployed", "still running", "not orphaned" in err.Error() at the stop/start, remove and orphan-cleanup handlers (three separate blocks). These strings are internal English, so localisation does not move them — the risk is a reworded internal error, not a translation, which is why this is P3 and was NOT folded into R-553's release. Fix shape: the same util.KindErrorf sentinels in internal/stacks (ErrProtectedStack, ErrStackNotFound, ErrStillRunning, …), a statusFor helper per handler family, and one table test per family passing a reworded message. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: internal robustness; only a future rewording would break it. — — CC
R-571 Process & tooling P4 [P3-LOW] The off-site failure classifier and the dashboard's alert-placement rules are described in no architecture document. FOUND 2026-09-17 while fixing R-553: 07-backup-architecture.md and 02-controller-module-map.md grep clean for ClassifyOffsiteFailure, Inline and PageOnly, so the six failure classes (quota / orphaned / no-repo / no-units / transport / unknown), the head lines they pick and the rule that one warning renders inline under the storage bars while every other renders in the top banner exist only in code. 10-localisation.md §9 now names the SIGNALS each decision reads; the behaviour itself still has no home. Fix shape: a short section in 07-backup-architecture.md for the classifier (its classes, what each means for the customer, and that restic/ssh text signatures are external) and one in 02-controller-module-map.md for alert placement. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: documentation only. — — CC
R-574 Process & tooling P4 [P3-LOW] web/handler_debug.go mixes page copy with JSON payload, so neither half could be converted safely. FOUND 2026-09-18 by localisation slice 2 release A (R-557): the file holds 39 Hungarian literals and the inventory classifies them by STATEMENT, not by data flow (I18N-INVENTORY-2026-09-17.md §4), so which are section headings the debug page renders and which are values inside a diagnostic dump the operator copies out is not established. Converting a dump value would change what an operator pastes into a report; leaving a heading Hungarian leaves a half-English page. Fix shape: walk the file once and label every literal page-copy or payload in the same table slice 2 release A used, then convert only the page-copy half. Belongs to slice 2 release B or C. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: debug page is operator-facing. — — CC
R-576 Process & tooling P4 [P3-LOW] i18n_go_parity.py cannot see a call site that LOST text; it only checks that the text a key carries is real. FOUND 2026-09-18 by localisation slice 2 release B, and found the hard way: the bulk converter silently dropped the continuation of a multi-line concatenation (fmt.Errorf("a: "+ "b: %s", x) kept only "a: "), damaging 7 producers — and the gate stayed GREEN throughout, because every surviving fragment WAS a byte-equal base-commit literal. Its question ("is this text real?") was answered yes while the CALL had lost half its sentence and its arguments. Two behaviour tests caught it (TestR356_ScenarioC_UndeployedAppIsStillRefused, TestR379_ScenarioA_RollbackSucceeds_AppComesBack), because they assert the sentence a customer READS. Fix shape: the gate learns a second question — for every util.MsgError("key", …) call site, the count of its arguments must equal the count of printf verbs in the key's Hungarian value, and no key-naming literal may be adjacent to a +. Both are cheap and would have convicted all 7. The general lesson, worth keeping whatever is built: a structural gate over the TEXT cannot see a defect in the CALL. READY - rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3->P4: gate improvement; tooling only. — — CC
R-579 Process & tooling P4 [P3-LOW] Five page shells loaded style.css with NO cache-buster, so a browser holding an older copy kept being served CSS that did not know about the newest UI. FOUND 2026-09-18 by the operator's screenshot of the v0.254.0 language globe: it rendered as a bare, unstyled <details> — a stray triangle and two plain words outside the card — because login.html, claim.html, recovery.html, launcher_shared.html and launcher_share_password.html requested /static/style.css with no ?v=, while layout.html has used ?v={{.Version}} since v0.166.0. .Version was also absent from three of those five data maps. FIXED AND CLOSED in the same session (controller v0.255.0): the parameter on all five, and Version set once in executeTemplateLang so a new shell cannot miss it; TestGlobeOnAnonymousShells now refuses an absent or EMPTY ?v=. The general form, which is the part worth keeping: a template that loads a versioned asset WITHOUT its version is invisible to every test that reads markup — the markup is correct and the browser fetches the wrong file. A gate over "every stylesheet/script link in a template carries ?v=" would catch the class; not built, because 8 first-boot-wizard templates would fail it and R-554 deletes them. DEFERRED (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the gate "every stylesheet/script link carries ?v=" was deliberately not built until R-554 deletes the old wizard; nothing else owns it) — CLOSED 2026-09-18 - controller v0.255.0 — — CC
R-591 Process & tooling P4 [P3-LOW] Stack.Copy() is a deep copy with one shallow field, and the field is new. FOUND 2026-09-20 while adding the catalog's English overlay: controller/internal/stacks/manager.go Copy() deep-copies Meta.DeployFields (with nested Options), Meta.OptionalConfig (with nested Fields), Meta.Integrations, Meta.HealthCheck and Meta.InitialCreds — and does NOT copy the new Meta.I18n map, which the struct assignment leaves shared between the original and the "copy". It is safe TODAY and that is exactly the shape worth filing: Metadata.For reads the overlay and never writes to it (pinned by TestForDoesNotMutateTheReceiver), so nothing can observe the sharing yet. The hole is in the CONTRACT — a function whose whole purpose is "a snapshot the caller may mutate" now has a field that is not one, and the next person to write through an overlay will find a bug with no failing test in front of it. Fix shape: deep-copy I18n in Copy() and pin it with a test that mutates the copy's overlay and asserts the original is unchanged. Alternatively state in Copy()'s comment that I18n is deliberately shared and immutable, and pin THAT with a test. Either is fine; silence is not. READY - rank P3-LOW; owner: CC (controller) Re-ranked 2026-10-03: P3->P4: latent contract gap, safe today. — — CC
R-594 Process & tooling P4 [P3-LOW] The catalog copy gate can CONVICT a retrieval promise but has no way to REGISTER a true one. FOUND 2026-09-20 translating batch 3 (R-560 slice 5). Vaultwarden's invite step and its sign-up setting both ended „…can open an account", and the English retrieval-promise pattern reads can … open as the claim that sealed backups can be opened. The conviction was a FALSE POSITIVE — opening an account is not opening a backup — and the two sentences were reworded to „can sign up", which is also the better copy, so nothing is blocked today. The gap is structural. The shared vocabulary this gate copies (scripts/customer_copy_vocab.py) states the design explicitly: these stems are NOT banned, because each carries a claim that is sometimes TRUE, and "an occurrence must be REGISTERED with a reason in the consuming gate's allowlist". The hub gate has ALLOWLIST_EN; app-catalog-felhom.eu/scripts/check-copy-i18n.py has none, so the only ways past it are to reword or to bypass the gate — and a catalog app whose English genuinely says a file can be restored (a backup app, a versioned document store) has no honest third option. Fix shape: an ALLOWLIST_EN of (app, path, reason) in the gate, a decoy proving a REGISTERED occurrence passes and an unregistered one still convicts, and a check that every entry still matches something (a stale allowlist entry was R-299's shape, and the retrieval gate has gone red on stale entries before). READY - rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3->P4: gate tooling; nothing blocked today. — — CC
R-603 Process & tooling P4 [P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a strings.Contains assertion reads exactly like a missing sentence. FOUND 2026-09-21 while writing the R-598 render tests. backup.target.absent was first written as "The system backup's drive cannot be reached…"; html/template escapes ' to &#39;, so the page carried the sentence and every assertion for it failed. The failure mode is the expensive part: the test said "the English absent-drive copy never reached the page", which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for ' " < > & (zero). The Hungarian bundle has never hit this because Hungarian copy uses „quotes" and few apostrophes; English copy will hit it again. Fix shape: either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against html.EscapeString(want) so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. READY — rank P3-LOW; owner: CC (controller) Re-ranked 2026-10-03: P3->P4: test tooling. — — CC
R-605 Process & tooling P4 [P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened. FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 check-image-resolvable and check-volume-persistence both returned INCONCLUSIVE and the drawn update action was replaced with use (audits/DRILL-chaos-night-2026-09-17.md:692-695). Neither script is defective — they behaved exactly as designed, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (check-image-resolvable.py cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). What is missing is the DISTINCTION. check-volume-persistence.py's self_test refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while classify() returns a per-app UNDETERMINED for an app that wrote nothing; check-image-resolvable.py likewise separates a harness-level canary failure (exit 2 at check() L180-183) from a per-pin throttle (L121-128). catalog_gates.py's VERDICT map collapses all of them into one INCONCLUSIVE label, so the operator-facing summary cannot say whether the gate ran at all. The cost is real and already paid: no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. Fix shape: have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have catalog_gates.py print the two differently. Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa. Small. READY — rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3->P4: tooling report clarity. — — CC
R-624 Process & tooling P4 [P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down. FOUND 2026-09-21 while widening R-462 from 3 apps to . vaultwarden and zipline close self-registration ON PURPOSE — vaultwarden by SIGNUPS_ALLOWED=false (R-512, „a stranger who guesses vault. must not be able to register"), zipline by answering E1037: User registration is disabled — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and that is correct: the harness must not be the reason a customer-facing app accepts strangers. gitea is a different and fixable case: the template sets no INSTALL_LOCK, so a fresh instance sits in its web-installer state and gitea admin user create refuses (MustInstalled() [F] Unable to load config file for a installed Gitea instance); POSTing the installer form first would work and was simply not written tonight. Why this is a row rather than three notes: 09 §3 decision 6 says the upgrade test goes to all apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — apps whose only account-creating route the catalog deliberately closes — for which the honest maximum is inconclusive unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. Needs: the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json. — NIGHT 2026-09-23: gitea stayed inconclusive on both venues (its installer); vaultwarden/zipline not attempted (closed sign-up, by design); code-server, outline, rallly have no front-door seed route (listed, not moved); bentopdf, glance, crafty-controller, wger, wanderer (meilisearch) and uptime-kuma have no fixture tonight (listed, not moved). -- 2026-09-30: the ceiling was WRONG for three of the six named apps. outline has a front-door first-run route (POST /api/installation.create — its self-hosted setup, refused once a team exists), rallly signs up through its own better-auth API with the e-mail code read from its own table in place of a mailbox, and zipline 4's first-run route is POST /api/setup (the fixture had tried /api/auth/register and /api/auth/setup, which are not it). Fixtures for outline and rallly are in upgrade_fixtures_box.py and both apps moved on both venues; zipline's fixture now tries /api/setup first (measured on the bench: a SUPERADMIN made, the login works). What remains in the class: vaultwarden (closed sign-up by design, no CLI), code-server (a browser IDE). audits/pg-last-six-2026-09-30/. -- 2026-09-30 (evening): gitea's fixable case is DONE — the fixture posts gitea's own first-run installer form (the form's own defaults + an admin; no CSRF on that form), measured on the bench and through the setup gate on 9202; gitea 1.27.0 → 1.27.3 published d7ba60c. What remains in the class: vaultwarden, code-server. audits/more-night-apps-2026-09-30/ READY — rank P3-LOW; owner: CC (catalog harness) Re-ranked 2026-10-03: P3→P4: an upgrade-test coverage limit; no household meets it. — — CC
R-652 Process & tooling P4 [P3-LOW] The memory watch counted the kernel's file cache as the app's memory. MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read 100 % of its 1 GiB and immich's PostgreSQL 100 %, each with 0 kernel oom_kills — memory.peak includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked memory_tight, and the gate would demand a raised mem_limit — a customer-box capacity figure — for cache. Done the same night (09 §3 decision 22, CC — operator may reverse): the watch samples the app's own memory (anon of memory.stat) every 15 s; the mark and the ladder's memory_peak_pct read it where measured; the cgroup peak stays beside it (memory_cgroup_peak_pct). Open: romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: audits/night-2026-09-23/apps/nextcloud/bench-1024M/, apps/immich/bench-noanon/. READY — P3; owner: CC (catalog harness) Re-ranked 2026-10-03: P3→P4: the main fix is done; what is left is harness tuning, no household meets it. — — CC
R-693 Process & tooling P4 [P3-LOW] The memory watch marks a Node app memory_tight at any limit — its heap sizes itself from the limit. Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (anon) peaked at 349 MB of 384 MB (90.9 %), then, with the limit raised to 512 MB, at 431 MB of 512 MB (80.4 %) — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). Needs: a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app memory_scales_with_limit fact. audits/night-2026-09-26/C/bench-run1/, …/bench-run2/ OPEN — P3; owner: CC Re-ranked 2026-10-03: P3→P4: a harness judgement problem; no household meets it. — — CC
R-731 Process & tooling P4 [P3-LOW] The catalog-currency comparison cannot see an upstream that changes its TAG SHAPE, and two newest tags need a release check. MEASURED 2026-09-30 (audits/catalog-currency-2026-09-30.md §2 item 5): the same-shape rule read three apps as up to date that were not — gramps-web (v25.6.0 → upstream dropped the v, at 26.9.1), jellyfin (10.11.11 → two-part 12.1), kimai (apache-2.57.0 → plain 2.67.0; the plain tag's digest equals apache's, so kimai moved on 2026-09-30). A control pass over every shape caught them. Not checked: whether mariadb:13.0 and gitea/gitea:28.0.0 (pushed 2026-09-30 00:15 UTC) are general releases. Needs: the currency script's shape-switch control made standing (it is in the audit's tools today), and the two release checks before either is walked. -- 2026-09-30: the two release checks answered (audits/immich-first-start-2026-09-30/D/D2-release-checks.txt): gitea v28.0.0 is a general release (GitHub: not prerelease, not draft, 2026-09-29) — upstream renumbered 1.27.x → 28; mariadb 13.0 is Stable but a short-term Rolling line (no EOL date), while 12.3 and 11.8 are the LTS lines — so a move of any of the four MariaDB apps to 13.0 would leave LTS. Not done: the shape-switch control made standing. NARROWED 2026-09-30 — only the standing shape-switch control is left; owner: CC Re-ranked 2026-10-03: P3→P4: only a standing check in a tooling script is left. — — CC
R-739 Process & tooling P4 [P3-LOW] The test bench cannot run wanderer at all: its web server calls the database at the PUBLIC name https://<SUBDOMAIN_DB>.<domain>, which the bench has no name or TLS for. MEASURED 2026-09-30 (more-night-apps brief, Part B): on bench 9401 the catalog template (v0.20.0) came up wanderer-db healthy, wanderer-search healthy, wanderer unhealthy for 12 min, every page 500 („Error 0: Something went wrong"); PUBLIC_POCKETBASE_URL=https://hike-db.gate.invalid. So the harness can only ever answer inconclusive at FROM for wanderer, never about an update. Its step v0.20.0 → v0.21.0 (web + db) and meilisearch v1.36 → v1.54 stay untested; the brief's question — does the search index survive meilisearch's move or is it rebuilt — is NOT measured. Needs: a bench venue that gives the stack the DB name (an extra_hosts + plain-http override in the bench's render only, stated in the verdict), or wanderer proven on the box venue alone with the operator's word. audits/more-night-apps-2026-09-30/B/wanderer-probe.txt -- 2026-09-30 late: the bench CAN run wanderer now — upgrade-test.py BENCH_ENV_OVERRIDES points PUBLIC_POCKETBASE_URL at the database container on the bench only (the only address the web image reads, measured in v0.20.0 and v0.21.0), recorded in every verdict: web 200, sign-up (PUT /api/v1/user) 200, login 200. The meilisearch question, answered: v1.54.2 REFUSES v1.36's database („Your database version (1.36.0) is incompatible”); with MEILI_UPGRADE_DB=true it upgrades in place and all three indexes (actors, lists, trails) are there — so wanderer's step needs that switch in the template first. Not done: a fixture (creating a list answered PocketBase's „Failed to create record” — likely a rule for unverified users), whether a household's record survives in the index, the step itself. audits/night-rulings-2026-09-30/ NARROWED — the bench runs it; the step needs a fixture and MEILI_UPGRADE_DB; owner: CC Re-ranked 2026-10-03: P3→P4: test bench coverage; the template switch and fixture are tooling work. — — CC
R-759 Process & tooling P4 [P3-LOW] wger's onboarding record (the checklist pilot) keeps rows open that no other row owns. 2026-10-01, app-catalog-felhom.eu/onboarding/wger.md: 2.5 no backup → remove → restore → read back of wger exists — the box has no per-app backup press outside an Update (R-648) and wger has no newer step to carry one; 3.7 changing the password and adding a family member not measured, and the template has no add_people text; 6.3 no forced-fail undo for wger; 8.2 the app page not read on 9202 this session; 9.1 the runtime volume-persistence gate not re-run (last CLEAN 2026-08-02). The other open rows have their own: 1.5 (R-755), 1.7/2.8 (R-762), 3.4 (R-763), 7.1 (R-764). wger is exempt from the onboarding gate (published before the checklist), so nothing blocks; this row is what keeps the record honest. Needs: the five measured on 9202 — 2.5 and 6.3 ride wger's next ladder step (the update's backing-up phase is the per-app backup). audits/new-app-checklist-2026-10-01/ READY — rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3→P4: record-keeping for a hidden app; nothing blocks. — — CC
R-781 Process & tooling P4 [P3-LOW] The catalog's scripts/test_gate_decoys.py fails 4 of its own "genuine" onboarding cases — on the untouched tree. Measured 2026-10-01 (same 4 FAIL lines before and after this session's checklist edit): the harness copies the catalog into a temp tree with a fake sibling felhom.eu, and the REAL published records (radicale, karakeep, dawarich) name evidence that the fake sibling lacks, so a complete record reads incomplete. The gate itself (onboarding) is green on the real tree. Needs: the genuine cases build their own records, or the fake sibling mirrors the cited evidence. READY — rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3→P4: a test-harness self-check; the real gate is green. — — CC
R-786 Process & tooling P4 [P3-LOW] SparkyFitness's onboarding record has six open rows (app-catalog-felhom.eu/onboarding/sparkyfitness.md): 0.5 runtime internet (food search providers), 0.7 the phone app's sign-in route through traefik, 1.6 the env names the server reads, 1.7 the entrypoint read, 5.4 a second memory watch at another limit, 8.3 no logo/screenshots on felhom.eu (404). Everything else measured this session (bench + 9202). Needs: each row measured, or n/a with a reason. READY — rank P3-LOW; owner: CC (catalog) Re-ranked 2026-10-03: P3→P4: onboarding record completeness; no household meets it directly. — — CC
R-797 Process & tooling P4 [P3-LOW] check-family-gate.py rule 3 (a family_gate template needs a baked golden ≥ 0.287.0) is checked only where the felhom.eu sibling exists — CI's single clone cannot. MEASURED 2026-10-02 (audits/family-gate-2026-10-02/C-metube-bench/C3b-family-gate-with-sibling.txt): without the sibling the gate printed NOT CHECKED and its summary read OK. The summary line now says "rule 3 … NOT CHECKED here". The pre-push hook (with the sibling) is where it bites. Needs: nothing more unless CI gets the sibling. WATCHING — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: a CI coverage gap; the pre-push hook still checks it. — — CC
R-805 Process & tooling P4 [P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount. MEASURED 2026-10-02 (audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md): Grimmory's /app/data bind held 0 files and read CLEAN (its state is in MariaDB); komga, paperless-ngx, radarr and sonarr also carry empty binds that are not judged. Needs: decide whether an empty bind after the exercise is UNDETERMINED like a named volume, with a decoy. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: a test-gate rule; no household meets it. — — CC
R-806 Process & tooling P4 [P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000. MEASURED 2026-10-02 (audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md): both reached data only through their fixture seed. Needs: read the loadbalancer.server.scheme label (upgrade_boxport already does); find why gramps-web's first GET got no answer. READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: a test-gate defect; no household meets it. — — CC
R-807 Process & tooling P4 [P3-LOW] 13 apps stay UNDETERMINED because their seeds write only to the database, so their upload / media / cache / redis volumes stay empty. MEASURED 2026-10-02 (audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md): claper, crafty-controller, dawarich (redis), docmost, gramps-web, immich (ML cache), outline (+redis), sparkyfitness, tandoor, vikunja, wger, wishlist, zipline — plus plex (no seed route: a plex.tv claim token) and wanderer (no fixture; unhealthy under the gate). None is BROKEN: nothing was written outside a preserved folder. Needs: per app, a seed that uploads one file (or the volume named n/a with a reason: a redis cache, an ML model cache). READY — rank P3-LOW; owner: CC Re-ranked 2026-10-03: P3→P4: test coverage; nothing was found broken. — — CC
R-819 Process & tooling P4 scripts/check_stands.py is red and runs in no runner. Measured 2026-10-03 on 9e2786c (before the triage): it convicts where-felhom-stands.yaml for citing R-273 and R-356, which are in neither OPEN-ITEMS.md nor anywhere it reads. After the triage it also convicts R-281, R-198 and R-201, because its rule 3 reads only OPEN-ITEMS.md and those rows are closed. It is in neither repo_gates.py nor CI, so nobody saw it — the R-29 shape. READY — filed 2026-10-03 (triage); owner: CC. Let rule 3 accept an id in CLOSED-ITEMS.md (and check the stand's status agrees), fix the two dangling ids, then register it in repo_gates.py with a decoy. — — CC
R-857 Process & tooling P4 Baking a golden twice under the SAME version leaves two stale facts. 2026-10-04 (golden 0.292.0 re-baked with live-restore): (1) golden_currency_gate.py keeps reporting the FIRST bake directory's sha (golden-0.292.0-2026-10-04, d6cf8b33…) — it picks one of two directories with the same version, not the newest; (2) the publish replaces the package in place (pre-delete + upload), so until the operator re-vouches, the hub vouches a sha the registry no longer holds — a fresh install in that window fails its sha check (fail-closed; none ran; re-vouched 17:48). Fix direction: the gate prefers the newest bake directory (or refuses two for one version); the runbook says "re-vouch at once after a same-version re-bake". documentation/tests/golden-0.292.0-2026-10-04-rebake/ READY — owner: CC — — CC
item due (UTC) what to measure
R-872 2026-10-06 the first live 05:00 deadline run judges a down box on the longer lines (Tester 2, if still off): hub log + the two events (detail in the R-872 row)