Files
felhom.eu/documentation/backlog/CLOSED-ITEMS.md
T

292 KiB
Raw Blame History

CLOSED-ITEMS — finished work, compressed

What this is. Every register row that reached a terminal state, compressed to its title, the version it shipped in, its evidence paths, and any sentence that states a RULE rather than a narrative. Nothing was deleted: each entry names the commit that holds its full original text, and git show <commit>:documentation/backlog/OPEN-ITEMS.md returns it verbatim.

Why it exists (operator ruling, 2026-08-22). OPEN-ITEMS.md had grown to 672 KB across 286 entries, over half of it finished work, with one single entry at 16 KB. A file that cannot be read is a file that cannot be checked — and this project has already paid for that twice: a record nobody could find because it sat inside an entry about something else, and a finding rediscovered because nobody could see it. The register now holds open work only, so its size tracks the work rather than the project's age.

A sibling rather than the bottom of the register, deliberately: appending to the same file keeps the byte count and the scroll, which is the thing being fixed.

Load-bearing reasoning was NOT compressed away. Where a closed row states a rule, a fence or a deliberate refusal, that sentence is carried here verbatim under Reasoning kept. Rules that outlive their work item also live in their proper homes — workspace-CLAUDE.md standing rules, felhom.eu/CLAUDE.md, CONTEXT.md, and the architecture folder — and this file is not their primary record.

This file is not the register. Nothing here is open. OPEN-ITEMS.md remains the single source of truth for open work; scripts/one_register_gate.py enforces that against ROADMAP.md.


2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124)

The full text of every row below: git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md (R-880 was opened and closed in this session).

Row What Closed Evidence
R-135 A cookie-less POST skipped the hub's CSRF gate, so a browser with cached Basic credentials could be made to POST cross-site. hub v0.135.0: without a session a state change needs Basic credentials AND the header X-Felhom-Operator (decision 120); the gate sits before the route switch. 39 paths through RequireAuth→ServeHTTP; red-proof: the old shape lets all 39 through. Live: Basic + no header → 403 (also with Origin: evil, also on an unknown path); with the header → passes; header without credentials → 401. Reasoning kept: a browser cannot add a custom header cross-site without a CORS preflight, which the hub never answers. CLOSED 2026-10-05 — FIXED hub v0.135.0 audits/hub-safety-2026-10-05/partA/; web/r135_csrf_test.go
R-133 Every box's break-glass console password was plaintext in hub.db. hub v0.135.0: sealed with the off-site seal and key (decision 121); legacy rows sealed at start-up — live: 4 rows sealed, 0 left plain; the demo-hp reveal still returned a password that minted a PVE ticket (HTTP 200; a wrong one 401); a wrong key → 500, nothing in the body or the log, no event. Reasoning kept: the running hub still holds the key — this closes the database-copy route only; a database backup without OFFSITE_SECRET_KEY cannot open the console passwords (R-173). CLOSED 2026-10-05 — FIXED hub v0.135.0 audits/hub-safety-2026-10-05/partB/; store/r133_recovery_seal_test.go, web/r133_reveal_wrongkey_test.go
R-604 A per-customer controller floor silently kept a box out of every global raise (demo-hp missed four). hub v0.135.0: a global raise logs one line per customer whose own LOWER floor wins and sends ONE operator mail naming them (floor_raise_skipped); a per-customer floor records when it was set; the System page's "Version floors" table lists every per-customer floor with its age and which ones the global cannot move. 2 red-proofs. Live: the table shows the three per-customer floors (age "unknown" — set before v0.135.0). The mail was not exercised live (it needs a global raise below an override). CLOSED 2026-10-05 — FIXED hub v0.135.0 audits/hub-safety-2026-10-05/partD/; web/r604_floor_held_back_test.go
R-530 Nothing listed which boxes still run an old agent (agents update only by a per-box signed job). hub v0.135.0: the System page's Agent cell (box → vouched, how far, since when; red after the wait) and agent_behind after 7 days (decision 119). Live: Tester 2 reads 0.142.0 → 0.146.1. Signing stays per box (the 2026-09-16 ruling: CC may sign until the first paying customer). CLOSED 2026-10-05 — FIXED hub v0.135.0 audits/hub-safety-2026-10-05/partD/; osupdates/r530_agent_alarm_test.go
R-508 Customer tester-1 had no registered e-mail, and the page did not say so. The address has been set since 2026-09-14 (the connect mails reach it — Gmail-read 2026-10-05); hub v0.135.0 adds the page warning: a configured customer with no box and no e-mail shows a red line (three branches tested, red-proof). CLOSED 2026-10-05 — FIXED hub v0.135.0 audits/hub-safety-2026-10-05/partG/; web/r508_no_email_banner_test.go
R-509 A box installed for an existing customer never got the connect e-mail. Fixed in hub v0.114.0; the owed real-mail proof: three mails from the automatic "host delete" trigger, each within 1 s of the hub's own send line (2026-09-16 12:22:59 and 18:17:46, 2026-09-30 07:23:03 UTC), read through the Gmail connector (metadata only). The "e-mail set" trigger shares the send core and is unit-proven. CLOSED 2026-10-05 — VERIFIED audits/hub-safety-2026-10-05/partG/r509-real-mails.txt
R-880 An installed felhom-os-apply refuses a bundle naming a path it does not know (R16), so a release whose bundle ADDS a path cannot reach any box on an older bundle (found 2026-10-05 before delivering v0.146.1, which adds four). Fixed by a step: felhom-agent/scripts/build-step-bundle.py — the box's current bundle with ONLY felhom-os-apply replaced (same paths), published as 0.146.1-step1; then the release's bundle. Tests StepBundle (the R16 refusal reproduced; the step accepted; exactly one file changed). Live: demo-hp, demo-felhom and Tester 1 each took step1 (written=1 same=20) then 0.146.1 (written=3 same=22), self-check ok. Reasoning kept: every future bundle that adds a path needs this step (decision 124); the step package stays published while any box may still be on the old bundle (Tester 2). CLOSED 2026-10-05 — FIXED (tooling, agent e4b5cf9) audits/hub-safety-2026-10-05/part{F,H}/; memory bundle-adding-a-path-needs-step-bundle

2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118)

Row What Closed Evidence
R-871 No architecture covered a box that is not always on, and a missed night was never made up. Decision 109 (option A) built: controller v0.295.0 internal/nightchain — a ledger of when each backup leg ran to its end; on a start or a host resume ONE catch-up 15 min later, backup legs only, never the update leg; a late daily timer after a suspend is skipped; the whole-guest backup and the catch-up wait for each other; decision 110's banner. Design 07 §6.1.1. Live: 9202 (dump made 15 min after the start; after a crash mid-wait, all three legs at the next start), demo-felhom (dump 15 min after the start; the household's timeline line reached the hub); banner served on 9202 and closed by its real route. 13 red-proofs. Reasoning kept: the ledger records that a leg RAN, and is never read as evidence that a backup EXISTS. Full text: git show 1b0678fa:documentation/backlog/OPEN-ITEMS.md. CLOSED 2026-10-05 — FIXED audits/catchup-2026-10-05/partA/, partB/
R-873 A household whose box is off every night was mailed "cannot be reached" every night. hub v0.134.0: at most once per 7 days to the household (persisted), the operator every edge, the recovery mail stays paired (decision 116). Proven by test through the real dispatcher (red-proof); no live occurrence in the session (Tester 2 stayed off). CLOSED 2026-10-05 — FIXED audits/catchup-2026-10-05/partC/r873-red-proof.txt
R-874 A restore-test never ran on a box with short power-on sessions. agent v0.145.0: first due-check 30 min after start (decision 117). Live on demo-felhom: start 07:38:46 UTC → restore-test first evaluation after start (R-874) at 08:08:46 → a due tier restored and passed in 29 s. CLOSED 2026-10-05 — FIXED audits/catchup-2026-10-05/partC/r874-*
R-875 A kept report's reason said "the agent stopped mid-pass" for a hub-away pass. agent v0.145.0: "sent late — kept on the box until the hub could take it". Test + red-proof. CLOSED 2026-10-05 — FIXED audits/catchup-2026-10-05/partC/r875-red-proof.txt
R-876 After a power cut mid-update every later pass failed until a person ran dpkg --configure -a. agent v0.145.0: dpkg's state = --audit AND the update journal in one call; repair on either; belt: repair + retry once when apt says "interrupted" (decision 118). Live (operator's go): crash at 07:56:03 UTC mid-unpack → back by itself → next pass REPAIR configured=0 journal=1 → DONE rc=0 upgraded=12, no person, no mail; package list identical. Reasoning kept: a check that reads one of two places dpkg keeps its state is a check that misses the other. CLOSED 2026-10-05 — FIXED audits/catchup-2026-10-05/partD/
R-877 The Tester 1 VM on demo-hp had no start-on-boot: the morning's demo-hp crash (06:14 UTC) left it off for 1 h 17 min, unnoticed (the night-fixes report called every box healthy). Found 07:31 UTC; qm set 341 --onboot 1, started; the afternoon crash then brought it back by itself. Filed and closed in the same commit. CLOSED 2026-10-05 — FIXED audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt

2026-10-05 (day) — the night's fixes: off-site clean-up guard, first-install image race, R8 download, a killed pass's report (controller v0.294.0, agent v0.144.0 + v0.144.1, golden 0.294.0; rulings 100–103, CC decisions 104–108)

Row What Closed Evidence
R-867 The off-site clean-up guard refused honest 7-day retention and mailed an error every window. Controller v0.294.0: the guard's "young" line is keep-daily CALENDAR days, built from the same constants as the policy (decision 104); every other refusal kept. Tests run restic 0.14.0's policy itself (proven identical to the binary over 92 snapshots), 5 red-proofs. Live by hand: demo-felhom window 5 16 → 14, demo-hp window 6 145 → 127 — exactly the predicted snapshots; hub rows pruned, no event, no mail, key files clean. Reasoning kept: a line that can sit inside the keep window is a line that refuses the honest case — derive it from the policy, never pick an age. Full text: git show 7221ee5c:documentation/backlog/OPEN-ITEMS.md. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partA/
R-95 The box could delete its own off-site history. Append-only key since 2026-10-03 (decisions 68–69); the last open item — a clean-up window that actually removes snapshots — was observed 2026-10-05 on both demo boxes (R-867's live windows, opened by the operator's one-shot grant; the dated check's four conditions all hold). The unattended weekly window is the same code path and is due ~2026-10-11/12; it is not separately re-checked (the DUE-CHECKS entry is removed with this row). Residual: R-822 (an add-only attacker steering older keeps). Full text: git show 7221ee5c:documentation/backlog/OPEN-ITEMS.md. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partA/; audits/offsite-lock-build-2026-10-03/
R-863 A new box's first app install could fail: the one-time image clean-up deleted the image compose had just pulled. Controller v0.294.0: every compose command that can pull holds a shared lock (dockerexec.BeginImageWork); a clean-up pass takes it exclusively without waiting and otherwise does not run; the one-time pass is retried every 2 min and writes its marker only after a pass that ran (decision 105). Live on 9202: the clean-up fired at minute 3 (05:31:36 UTC) inside BookStack's 35.5 s install, skipped itself, BookStack installed first time; the retry at 05:33:36 ran and kept everything; teardown through the product. The update path was already guarded (Updating); restore/undo were exposed in principle and now hold the same lock. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partB/
R-864 A failed compose logged the head of stderr (pull progress) and cut the reason. Controller v0.294.0 tailStr (rune-safe), pinned with the night's real first line. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partB/red-proofs.txt
R-869 The move-aside log line printed an empty destination. Controller v0.294.0: assigned before the log line; test + red-proof. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partD/r869-red-proof.txt
R-865 R8 measured every download as 0 B (--print-uris with -s prints no URIs). Agent v0.144.0: no -s; the fake answers like real apt (verbatim 9202 output), 2 tests, red-proof. Live: the installed wrapper's download_bytes on demo-hp read 12 802 456 B for 13 pending upgrades (0 before), nothing installed. A live R8 REFUSAL line was not produced: it needs < 500 MB free on demo-hp's guest, i.e. 28.5 GB written into a thin pool with 18.7 GB free — it would have stopped every guest. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partC/
R-866 The debug OS pass could not run with the hub away. Agent v0.144.0: the daemon saves the hub's block; the selftest falls back to it and says block=SAVED(<time>; hub unreachable: …) in its header (decision 107). Live on demo-felhom with the hub blackholed: pass ran (nothing ×3); its reports, kept on disk, reached the hub at the next agent start. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partD/r866-*
R-868 An OS pass whose agent was killed installed but never reported. Agent v0.144.0: the wrapper keeps its report beside the plan; the agent sends kept copies and deletes them (decision 106). Measured live NOT to work (05:45 UTC, demo-hp): the wrapper died on a broken stderr pipe before saving. Agent v0.144.1 (decision 108): the wrapper survives a dead reader; the agent looks again every 5 min. Live again (06:04 UTC, demo-hp, the A5 shape): wrapper DONE upgraded=13, copy kept, the hub got ONE applied report (13 packages) at 06:09:05 by the 5-minute look. Reasoning kept: a root helper whose reader can die must never write unguarded to that reader — the log is not the outcome. CLOSED 2026-10-05 — FIXED audits/night-fixes-2026-10-05/partD/

2026-10-04 (night) — the config bundle, test approvals, the Docker-socket self-heal (agent v0.143.0, hub v0.133.0, controller v0.293.0, installer 1.31.0; rulings 96–99)

Row What Closed Evidence
R-840 A new root wrapper or sudoers line could not reach an installed box — no product route. Built (decision 96): the config bundle — every root-owned file the installer writes, one reproducible file beside the binary; a signed agent_config_update the box's felhom-os-apply verifies itself (signature against the root-owned signers file, sha, the 22-path table, every content check before the first write, self-check, undo); the trust root is never a bundle path; the installer installs the same bundle. Live: both demo boxes (bootstrap, then 0 written of 22, probe 71/71), a deliberate change and its undo on demo-hp, a wrong sha and a replay refused, the installer path on demo-felhom. Reasoning kept: the route cannot reach a root side that predates it — the first bundle on such a box is one by-hand step, and that must be said, not hidden. Tester 2's step → R-862. CLOSED 2026-10-04 — BUILT audits/r840-config-bundle-2026-10-04/partB/
R-859 An OS approval made under a TEST wait stayed in force after the test, and a real ring-1 box installed it (Tester 2, 2026-10-04: 49 guest + 106 host packages). Hub v0.133.0: every approval under an override is marked; a start without the override cancels every unsuperseded test approval (no ring-1 plan, ring-1 boxes bumped, one operator event each); amber on the System page; one-time backfill of the automatic early approvals. Live: 4 cancelled at the 18:20 UTC start; the operator's Docker approval stays (CC decision, operator may reverse). CLOSED 2026-10-04 — FIXED audits/r840-config-bundle-2026-10-04/partD/; runbooks/os-updates-test-waits.md
R-860 A re-created Docker socket file left the controller and traefik blind until a person acted (generalises R-858). Measured: only a docker.socket restart re-creates the file; a dockerd crash or systemctl restart docker does not. Controller v0.293.0 (internal/sockheal): 60 s of refusals → the controller exits and Docker restarts it on the current socket; then it restarts any other socket user on an older inode (traefik). Live: 9202 healed in 104 s, demo-hp 9201 in 120 s, every container id unchanged. Reasoning kept: count only refusals, never timeouts, and only after Docker answered once — or a slow daemon restarts the controller, and a broken socket loops it. CLOSED 2026-10-04 — FIXED audits/r840-config-bundle-2026-10-04/partE/

2026-10-04 (~18:30) — R-858 incident (agent v0.142.1, ruling 95)

Row What Closed Evidence
R-858 A Docker engine step left felhom-controller and traefik on the OLD docker socket (found by the operator: demo-felhom DOWN 14:18–15:57 UTC). live-restore kept them running with the deleted, bind-mounted socket file; their own health stayed "healthy", so the step passed. Repaired by restarting the two; agent v0.142.1 restarts ONLY the socket-mounting containers after a step and fails the health rule when the controller cannot reach Docker. Proven live twice on demo-hp (signed undo and forward): both restarted, guest / controller / traefik on the same socket inode, applied, healthy. Reasoning kept: a container that bind-mounts a socket FILE does not follow a restarted daemon — live-restore makes this worse, not better. Check the consequence (the controller reaches Docker), not the mechanism (the container still runs). CLOSED 2026-10-04 — FIXED audits/os-docker-crash-2026-10-04/partE-incident/

2026-10-04 (evening) — System page, Docker slow lane, crash guard (agent v0.142.0, hub v0.132.0, installer 1.30.0)

Evidence: audits/os-docker-crash-2026-10-04/.

Row What Closed Evidence
R-852 The operator could see no box's Debian, Proxmox, kernel or Docker version anywhere. The box reports them (system stanza: Proxmox + kernel from the API, the wrapper's read-only facts); hub v0.132.0 shows them on the new System page (with the ring / switch / approve buttons) and a Proxmox / kernel column on Hosts. Reasoning kept: a value the box could not read says unknown, amber, with the reason — never empty, never guessed. CLOSED 2026-10-04 — FIXED (09 decision 89) partA/
R-835 Turning Docker's live-restore OFF by a restart stops every container and starts none. Resolved by design (decision 87): live-restore is ON everywhere — the golden bakes it, installed boxes get it once by a RELOAD (measured: same ids on 9202 6/6, demo-hp 24/24, demo-felhom 5/5), and the wrapper refuses a Docker step while it is off (R15). Rule kept: never turned off by a plain restart. CLOSED 2026-10-04 — RESOLVED BY DESIGN partB/b1..b4
R-848 A held host package was invisible to the hub. The facts report held; the System page shows it amber. CLOSED 2026-10-04 — FIXED partA/
R-849 The guest's "restart needed since" never cleared. The wrapper scans the guest on every pass too (agent v0.142.0). CLOSED 2026-10-04 — FIXED partB/agent-redproofs.txt
R-851 A host that panicked stayed stopped (kernel.panic = 0). The crash guard (decision 88): kernel.panic = 10; the 3rd unclean stop within 60 minutes leaves the box off; re-arms after 24 h or by felhom-crash-guard rearm. MEASURED on demo-hp: crash 1 and 2 restarted by themselves (54 s, 53 s), the guard tripped, crash 3 stayed off until the operator switched it on; the hub mailed the trip. Reasoning kept: a crash, a power cut and a hard reset cannot be told apart on these boxes (pstore saved nothing for a real panic) — the guard counts every unclean stop. CLOSED 2026-10-04 — FIXED partC/
R-854 The first Docker steps were stored as GUEST reports — they ran (agent rc) before hub v0.132.0 was live, and the older hub maps an unknown layer to guest; the guest candidate briefly read 0 packages. One-time; a pass with the released agent re-reported both layers. Nothing to fix. CLOSED 2026-10-04 — ONE-TIME, NO ACTION partB/b5-*

2026-10-04 (afternoon) — OS updates, host fast lane + fleet view + alarms (agent v0.141.0/v0.141.1, hub v0.131.0/v0.131.1, controller v0.292.0)

Evidence: audits/os-host-lane-2026-10-04/.

Row What Closed Evidence
R-841 Every box reported its tunnel inactive (the agent asked a host unit that does not exist). Agent v0.141.0 reads the guest's cloudflared container and controller v0.292.0's Docker health check on cloudflared's own /ready (200 only with a connection); three states running / not_running / unknown; hub v0.131.0 alarms tunnel_down after two not_running reports and tunnel_recovered on the next running. LIVE on demo-hp: port 7844 blocked → tunnel_down mailed after the 2nd report; unblocked → tunnel_recovered. Reasoning kept: unknown never alarms; a container state alone says "up" for a dead tunnel (wrong token: running, /ready 503). A plain docker stop is healed by the controller's protected-container check within 5 min — before the 15-min host report sees it. CLOSED 2026-10-04 — FIXED partA/
R-845 The OS leg was slow. One pct exec per package (~0.9 s each) replaced by one call per layer; restart scan only after an install (host: every pass, v0.141.1); repair only when dpkg --audit reports. MEASURED: nothing to install, both layers — 23.3 s (demo-felhom), 31.5 s (demo-hp); before, guest only, nothing to install — 14.0 s; the 174–245 s passes of R-845 were passes WITH an install. A 108-package host install pass: 70 s. CLOSED 2026-10-04 — FIXED partD/, partB/live/
R-846 Host "reboot needed" was wrong in agent v0.141.0 (found live on demo-felhom): the restart scan skipped every cgroup containing lxc, hiding lxc-start (0::/lxc.monitor/<vmid>, 20 deleted maps after libc6); and it scanned only after an install, so a reboot never cleared it (the hub's 14-day alarm would fire on a rebooted host). Agent v0.141.1 (:/lxc/, host scans every pass, reboot_scanned) + hub v0.131.1. Red-proved. CLOSED 2026-10-04 — FIXED partB/live-defects-redproofs.txt
R-850 Hub v0.131.0 put the layer into the release fingerprint, so the unchanged guest set (272 packages) counted as new and waited its 24 h again. One-time; ring 1 kept the previous release meanwhile; approved again under the TEST wait (os-guest-20261004-123933). Nothing to fix. CLOSED 2026-10-04 — ONE-TIME, NO ACTION hub log 2026-10-04

2026-10-04 (~12:20) — operator ruling

Row What Closed Evidence
R-842 The guest OS update had no automatic undo (no snapshot of a guest with host binds is possible). Ruled option A (09 §3 decision 81): the whole-guest backup, minutes old, restored by hand, is the undo. Nothing new built; a home-made LVM-thin snapshot rejected (unmeasured, thin-pool risk). CLOSED 2026-10-04 — RULED (decision 81) audits/os-guest-lane-2026-10-04/partA/README.md

2026-10-04 (day) — OS updates, guest fast lane (agent v0.140.0, hub v0.130.0, controller v0.291.0)

Evidence: audits/os-guest-lane-2026-10-04/.

Row What Closed Evidence
R-837 The guest snapshot undo was unmeasured. MEASURED: impossible — PVE refuses any non-vzdump snapshot of a guest with host-path binds (mp8/mp9), as the agent's token (which has the rights) and as root (PVE/AbstractConfig.pm:755-757). The automatic undo was not built (the brief's stop rule); the decision is R-842. Rule kept: a guest with host binds has no PVE snapshot; the night's backup is its undo. CLOSED 2026-10-04 — MEASURED partA/README.md
R-838 The infrastructure images never moved (cloudflared four months behind). Controller v0.291.0: traefik v3.7.13, cloudflared 2026.9.3, filebrowser 1.5.6-stable (release notes read; nothing we use breaks). A controller release DOES move all three (bring-up + start-up mount sync; measured on 9202 and both demo boxes, public gap ≤ 19.6 s); scripts/check-infra-pins.py + the runbook's "Infrastructure pins" section put them on the monthly re-test. CLOSED 2026-10-04 — FIXED controller v0.291.0 partF/
R-726 A returning household's new box made no off-site copy on night one. Decision 78 built (controller v0.291.0): a claimed box that has never made an off-site copy sets the orphaned old copy aside and starts a new one; nothing deleted; a box that has made copies still asks. Red-proved both ways. CLOSED 2026-10-04 — FIXED controller v0.291.0 partE/r726-redproof.txt
R-843 --selftest=wgtunnel (since S3) and --selftest=os-update were dispatched but refused by the flag's allow-list. Found live (the OS debug action could not run); both accepted in agent v0.140.0; TestSelftestFlag_AcceptsEveryDispatchedMode pins every dispatched mode. Rule kept: a dispatch case without a flag case is dead code — the test reads both. CLOSED 2026-10-04 — FIXED agent v0.140.0 (opened and closed the same day) partC/agent-leg-redproofs.txt

2026-10-04 (day) — the off-site topic closed (hub v0.129.0, agent v0.139.0)

Evidence: audits/backup-close-2026-10-04/.

Row What Closed Evidence
R-834 A whole-guest restore brought back the production config (onboot: 1, the real drive binds). Every route listed: the restore-test was MEASURED safe already (onboot 0, throwaway mp8/mp9, every poll); the agent's DR bring-up now REFUSES beside a live original (source guest present, a drives bind, or an unreadable config — agent v0.139.0, proven live on demo-hp); the hand route is scripts/felhom-restore-beside.sh (onboot 0 at restore, host binds removed, NICs down, never started, read back — proven live); provisioning restores the golden, which has no binds. Rule kept: a restore that does not replace the box's own guest on a replaced host ends with onboot 0 and no host-path bind; DR on a replaced host keeps its binds. No sudoers line added (the restore-test sets onboot 0 through the API; DR refuses rather than degrades). CLOSED 2026-10-04 — FIXED agent v0.139.0 + script partA/; tests TestRunBringUp_DRRefusesBesideALiveOriginal, TestRestoreTest_NoHostPathBindBesideTheOriginal, scripts/test_felhom_restore_beside.py
R-833 After a long gap the clean-up wedged at the default cap. Hub v0.129.0: POST /offsite/window-grant/<id> with max_remove=<n> (operator login only, 1..500) raises ONE window's cap; the guard still applies in full; consumed once; the close check uses the window's own cap; operator event offsite_window_large_grant; a box cannot grant itself. Lab proof on a real restic 0.14.0 repo: 98 snapshots, plan 85, default cap 49 refused, raised 90 → 98 → 13, next window passes the default cap. CLOSED 2026-10-04 — FIXED hub v0.129.0 partB/ (4 red-proofs, lab proof, live refusals)

2026-10-04 — off-site safety finished (hub v0.128.0, controller v0.290.0, decisions 71–74)

Evidence: audits/offsite-finish-2026-10-04/.

Row What Closed Evidence
R-824 The fake-snapshot guard refused every window after a manual run. Fixed controller v0.290.0: a young snapshot superseded the same day in its group is EXCLUDED (removed later, once old); any other young removal, a future date or a plan above the week's cap still refuses. Live: window 2 on demo-hp ran without refusal (outcome nothing, 127→127, the only candidates were young same-day copies). Weekly windows ON. Reasoning kept: a guard that refuses the honest case is switched off within a fortnight — exclude the benign shape, keep refusing the poisoning shape. CLOSED 2026-10-04 — FIXED controller v0.290.0, live window without refusal; a window that removes something not yet observed (R-95) partA-live-window.txt, red-proofs-controller.txt RPC1–2
R-823 The household's set-aside deletion stopped happening on the append-only tier. Built (decision 74): the box hands the due request to the hub; the hub deletes only <repo>.orphaned-* after 7 days unless the household or operator cancels; the page keeps a dated, cancellable deletion. Live on tester-1: the live repo named → refused; a planted set-aside dir deleted after the (test-shortened, logged, reverted) delay, read back absent. Reasoning kept: the only deletion a broken-into box can trigger now waits a week under operator mails. CLOSED 2026-10-04 — BUILT hub v0.128.0 + controller v0.290.0, proven live partE-live-tester1.txt, red-proofs-hub.txt RPH1–2, red-proofs-controller.txt RPC3
R-826 Sub-accounts kept every earlier box's key unpinned. tester-1's 3 lines removed through the hub registrar (decision 72, POST /offsite/remove-unpinned); the daily check reads 0 lines, no alarm. Demo boxes' stale lines went on 2026-10-03. CLOSED 2026-10-04 partC-tester1-keys.txt
R-827 The daily key check created .ssh when absent. Fixed hub v0.128.0: only the write path creates it. CLOSED 2026-10-04 — FIXED hub v0.128.0 (red-proved) red-proofs-hub.txt RPH3
R-828 The ep0 copy grew without bound. Decision 71: prune job prune-ep0-copy keep-weekly 8 (all namespaces) daily 07:30 after the 05:00 pull; GC Sundays 08:30; remove-vanished stays false. Dry-run by reasoning: nothing to remove yet (2 snapshots per group, 2 weeks). CLOSED 2026-10-04 partB-copy-retention.txt; runbooks/ep0-datastore-copy.md
R-830 The restore route from the DooPlex copy had never been walked. Walked on demo-hp with a scratch VMID: list 2 s, restore 186 s, data read via pct mount, all torn down. DooPlex PBS is LAN-only; the options for another network are written. Reasoning kept: a copy that was never restored is not proven — and the restore found a trap (R-834). CLOSED 2026-10-04 — WALKED (route 1); route 2 not walked partD-restore-walk.txt; runbooks/ep0-datastore-copy.md

2026-10-03 (evening) — the off-site lock built (hub v0.127.0, controller v0.289.0/0.289.1, decisions 68–70)

Evidence: audits/offsite-lock-build-2026-10-03/.

Row What Closed Evidence
R-820 A box could obtain its sub-account password at will (the hub's self-heal re-armed it; the box consumed it) — and that password removes any append-only pin. Fixed: the box sends only its PUBLIC key; the hub (key registrar) writes it pinned; consume-password answers 410. Measured live: demo-hp's own API key gets 410 gone, no password. Reasoning kept: a pinned key protects nothing while any route can rewrite authorized_keys — the hub is now that file's only writer, and it reads every file daily. CLOSED 2026-10-03 — FIXED hub v0.127.0 + controller v0.289.0/0.289.1, proven live on both demo boxes partD/no-password-for-box.txt, partB/red-proofs-hub.txt; hub internal/offsitekeys
R-821 The hub DB held every sub-account password in the clear. Fixed: AES-256-GCM at rest under OFFSITE_SECRET_KEY (Secret/offsite-secret-key, out of git); the 4 legacy rows sealed at start-up (raw rows read back enc:v1:); no key → the hub refuses to store or use one. Reasoning kept: the running hub still holds the key and can open the passwords — a hub compromise remains an off-site compromise; this closes the database-copy route only. CLOSED 2026-10-03 — FIXED hub v0.127.0, verified on the live DB partB/hub-rollout.txt; TestOffsiteSecret_*
R-342 ep0's server snapshot never covered /mnt/pbs-datastore. Built (decision 70): nightly PBS pull-sync to DooPlex (ep0-copy), remove-vanished false, weekly verify, failures mailed (test mail received). First pull 201 s / 12 GB / 4 of 4 snapshots, matching ep0. ep0 changed by one read-only token only; reached through an SSH forward from DooPlex (operator ruling). Reasoning kept: Hetzner cannot snapshot a Volume — the copy is the only safeguard; never prune it tighter than ep0. CLOSED 2026-10-03 — BUILT; restore route unwalked (R-830), growth unbounded (R-828) partF/; runbooks/ep0-datastore-copy.md
R-825 Controller v0.289.0 reported 0 off-site snapshots as MEASURED over a store holding 12 (demo-felhom), and the hub mailed offsite_snapshots_dropped 11→0 — a false alarm. Cause: the provider's rclone prints a NOTICE line that restic forwards into the combined output; every --json parse failed. Fixed in v0.289.1 within 15 minutes: the notice is stripped, and an unreadable count is never a measured zero (keeps the last value, stats_known=false). Reasoning kept: a failed measurement must never be written as a measurement (R-331) — the detector that caught it is the one it would have blinded. CLOSED 2026-10-03 — FIXED controller v0.289.1 (red-proved), found and fixed in-session operator mail 2026-10-03 17:17 CEST; partC/red-proofs-controller.txt RPC4

2026-10-03 — off-site append-only, measured on the provider (R-436, R-430)

Spike, no product change. Evidence and design: audits/offsite-append-only-2026-10-03/.

Row What Closed Evidence
R-436 Hetzner's --append-only forced command holds for the key it is pinned to. On u629488-sub4 (tester-1's, operator-ruled venue; scratch repo spike-r436, removed): pinned key command="rclone serve restic --stdio --append-only spike-r436",restrict — init, two backups, snapshots, restore (bytes identical), check OK; forget d807418c --prune, forget --keep-last 1, prune → blob not removed, server response: 403 Forbidden (403), rc=1, count unchanged; control with an unpinned key: 1 / 1 files deleted. The client's path and flags are ignored; no shell, sftp, scp, rsync or port forward (administratively prohibited). Reasoning kept: the pin protects a repository only if no other route can rewrite authorized_keys — and the password can (R-820). CLOSED 2026-10-03 — MEASURED; the due-check (2026-10-06) is cleared by this measurement live/E1-E3-init-backup.txt, live/E4-E6-deletes-and-C1.txt, live/B2-forced-key-misuse.txt, live/TEARDOWN.txt (authorized_keys restored, sha256 identical)
R-430 restic unlock prints successfully removed locks after removing nothing — by design; and unlock --remove-all DOES work through the append-only key. Measured live and in the lab: a crash lock (not yet stale: under 30 min, new hostname) survives plain unlock, which still prints success; --remove-all removes it because the rclone append-only server allows lock deletion. A crash lock blocks check, not backup. So resticStep's self-heal stays valid under R-436's transport; the earlier sticky-directory model does not describe it. CLOSED 2026-10-03 — ANSWERED; not a precondition for the rclone transport live/C2-A5-locks.txt, lab/A5-locks.txt

2026-10-03 — the triage: finished rows moved out of the open register

Every row below sat in OPEN-ITEMS.md with a finished LEADING verdict (or was verified finished against live source on this day, or folded into an older duplicate). Full original text of each: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md. Since this day closed_register_gate.py RULE 3 refuses a finished row left in the open register.

Row What Closed Full text
R-88a Failing backup re-quiesces every 5 min, no backoff Reasoning kept: «breaker 15m→4h, per-tier, never permanent» SHIPPED (controller v0.176.0, 2026-07-27). Live on both boxes; breaker 15m→4h, per-tier, never permanent. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-88b /backup/due cannot say unknown SHIPPED + PROVEN-LIVE (agent v0.105.0 + controller v0.178.0, 2026-07-27). age_state=unknown captured on real hardware during a deliberate ep0 outage; controller deferred, zero app stacks stopped. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-123 R-105 and R-106 were READY in ROADMAP.md with no row on THIS page CLOSED 2026-10-03 (triage, verified) — its last residue — "R-105 still needs a row" — is done: R-105 has its own row; the process gap is R-369's gate full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-131 sess-f is a fourth orphaned scratch customer CLOSED 2026-10-03 (triage, verified) — sess-f was already gone by 2026-08-05 (documentation/tests/campaign11-evidence-2026-08-05/journal.md:24-27) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-176 Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap. (a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated (withdrawn, not deferred; CONTEXT.md:2076). — residue NOT-A-FINDING (2026-10-03): its residue ("re-run (a) once against agent v0.120.0") is moot: every node was reinstalled, none migrated (CONTEXT.md:2076) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/SPIKE-r165-phase0-2026-08-03.md
R-191 Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes. Reasoning kept: «R-89 moved PBS pruning SERVER-SIDE — "boxes set keep_last: 0, ep0 runs prune jobs; box tokens stay write-only, never widen the grant".» CLOSED — SHIPPED 2026-08-04 (installer 1.25.0; both live boxes corrected to offsite keep_last: 0, prune_pbs_allowed=false). ep0 prune jobs verified first (18 tasks OK). hostinstall_gates.py asserts it (red-proved). The 'next weekly run OK' observation: offsite snapshots dated 2026-08-11 recorded in audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md:118. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md
R-201 Nothing in the offsite DR chain has ever been exercised past the ceremony — and the one live proof that exists predates the field it is cited for. Reasoning kept: «The pass condition is unchanged — a byte-identical sentinel sha256, not "the store opened"» «a good snapshot is not durable against a later bad run on the same day.» CLOSED 2026-08-07 — BOTH HALVES PASS on the fifth walk (DATA 2026-08-04 night run; JOURNEY 2026-08-07, zero guest command lines). Does NOT claim a smooth journey (R-252/R-253, both since CLOSED 2026-08-08) nor shape (c)'s positive half. Walls along the way: R-203, R-204, R-241. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DRILL-r201-night-run-2026-08-04.md, audits/RECON-offsite-dr-chain-2026-08-04.md, tests/finalwalk-r201-2026-08-07/journal.md, tests/walk5-r201-2026-08-07/journal.md
R-202 NOW EVIDENCED, NOT ARGUED (2026-08-10). CLOSED 2026-10-03 (triage, verified) — the false recoverability promise was removed (R-294/R-299; controller/internal/web/templates/backups_remote.html:93-108) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-209 Should the containerd store move to SSD2 at all? Reasoning kept: «A TRAP was found while proving the guard, and it is the reusable part: RequiresMountsFor on a path with NO mount unit is a SILENT NO-OP» «-X is load-bearing — overlayfs stacking rides trusted.overlay.*.» EXECUTED 2026-08-05 on operator ruling — reboot validation DEFERRED → R-209a (still open, WATCHING, OPEN-ITEMS.md:339). Zero-loss move verified on four observables; storageReserved on SSD2 0 → 80 GB; pre-move tree moved aside. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/SPIKE-dooplex-buildcache-2026-08-05.md
R-214 The physical console never stops asking to be paired. CLOSED 2026-09-20 - ISO 1.28.0+/R-535, proven on a fresh install (localisation slice 6's walk on ISO 1.29.0: last console paint after self-bind is the bilingual bound banner). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/i18n-slice6-2026-09-20/drill/screens/22-console-after-claim.png
R-229 The instruction-file rightsizing landed for felhom-controller and the workspace root; three pieces were deliberately deferred. CLOSED 2026-10-03 (triage, verified) — legs (a), (b), (c) CLOSED 2026-08-06 in the row's own text; leg (d) moved to R-230 full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-233 The golden bake's acceptance checks were a list of strings the script does not print. Reasoning kept: «The general lesson: a runbook's pass markers must be copied from a captured log, never written from memory» CLOSED 2026-08-06 — fixed in RUNBOOK-manual-build.md (§4.1): markers re-captured from the real log, pre-gate URL corrected to golden.tar.zst, token moved off the command line into an in-VM runner script, positive control on the token-leak grep, vouch step rewritten. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: runbooks/RUNBOOK-manual-build.md
R-245 Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT. Reasoning kept: «The real harm, if it comes, is QUOTA — old history blocking new backups — and that is a condition, not a calendar. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. If this is ever built, build it that way.» DECIDED 2026-08-07 — not built; reopens on quota. Built instead: escalating reminders (1/3/7/14 days) and operator levers --abandon-extend / --abandon-stop. Re-filed 2026-08-08 as a decision taken. Reopens verbatim: "THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups." full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-248 A flag that changes behaviour is visible to nobody who would look for it. FOLDED into R-246 2026-10-03 — the same ruling: give stale_at a visible, evidence-bearing setter, or retire it full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-254 The same render-then-hide pattern R-249 fixed is live in two more places, and one of them carries a real per-install secret. Reasoning kept: «escrow_handlers.go's rule that a secret is revealed by an XHR and never templated server-side into HTML.» «A form must carry what it submits.» CLOSED 2026-08-08 — both sites closed in controller v0.208.0: POST /apps/<slug>/initial-credentials/reveal (re-reads container, no-store, CSRF, logged) and POST /stacks/<name>/auto-field/reveal for deployed pages; pre-deploy hidden input kept deliberately. Gate scripts/secret_in_markup_gate.py; its blind spot → R-255. Measured exposure: none; rotation not indicated. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: scripts/secret_in_markup_gate.py
R-272 RANK 1 — Felhom's --uninstall leaves the exact condition that makes Felhom's own reinstall REFUSE. CLOSED 2026-10-03 (triage, verified) — the uninstall purges the dnsmasq Felhom installed (R-316; scripts/felhom-host-install.sh:858-882, _dnsmasq_purge_owned) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-295 One name per secret — CONTROLLER HALF SHIPPED. CLOSED 2026-10-03 (triage, verified) — hub half shipped in hub v0.104.0 (4d6ec7c, hub CHANGELOG.md:844,905); the controller half shipped v0.211.0 full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-303 markOrphaned has no guard against an active abandon countdown — the co-render is made HARMLESS, not IMPOSSIBLE. Reasoning kept: «the wrong fix (suppressing the orphan card during a countdown) would hide a real second fault.» DECIDED 2026-08-13 — left as it is; reopens on a real-world sighting. Verbatim trigger: "an observation of the combined state occurring OUTSIDE a constructed test." Related: R-302. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-312 There is no in-product route from the recovery screen to a set-aside store, and building one is not wiring — it is new surface. Reasoning kept: «retention is an operator-only capability, and nothing anywhere may promise the customer can perform it themselves» DECIDED 2026-08-13 — deliberately not built; re-evaluate on a real customer request. Verbatim trigger: "a real need appearing — one request from a customer who is not us." Related: R-304, R-311. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-313 demo-felhom's set-aside store is UNRECOVERABLE — 36 snapshots whose key we destroyed ourselves. DECIDED 2026-08-13 — kept as a test fixture; delete when R-312 ships or is abandoned. Verbatim: "when R-312 is built, or when R-312's ruling above is made permanent. On either event, delete it deliberately and record why". Related: R-198, R-307, R-312. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-341 Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO. Reasoning kept: «the anchor must be ps -o lstart= -p $MainPID, NOT systemctl show -p ActiveEnterTimestamp, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, NRestarts is still 0) and would put the rate ~15% low.» CLOSED 2026-08-30 — second check taken (fd back to baseline 17, ESTAB 0) but it cannot answer the question: the leak was REMOVED mid-interval by our own fix (R-344, agent 0.130.0). Question moot; fix confirmed holding at 12 days. First check 2026-08-20: unchanged, 201.6 fd/day. R-336 keeps the request-rate scaling item; ActiveEnterTimestamp trap → R-346. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/SPIKE-ep0-established-connections-2026-08-20.md, audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt, evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt
R-343 The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later. CLOSED 2026-10-03 (triage, verified) — obsolete — its only step was "confirm 0.216.0 is healthy in normal operation"; both demo boxes have run every release since and are on 0.288.0 (STATUS.md, 2026-10-02) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-369 There are TWO registers, only one calls itself the source of truth, and work filed in the other is invisible to every standing rule that says "grep the register". CLOSED 2026-10-03 (triage, verified) — RULED 2026-08-22 (one register) and enforced by scripts/one_register_gate.py (ef6ac6f) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-378 A status word inside a longer verdict fooled the session's own compressor, moving six still-open rows into the closed file. Reasoning kept: «Match the LEADING verdict.» CLOSED 2026-08-22 — corrected in the same session; the six rows (R-123, R-190, R-214, R-264, R-295, R-352) restored verbatim from commit fddfe00ce268. Successor: R-369. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-385 A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green. Reasoning kept: «Membership, not baked > released, deliberately: a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it.» CLOSED — 2026-08-23. v0.221.1 given its own CHANGELOG heading (commit da75603); scripts/golden_currency_gate.py now requires the baked version's own ## vX.Y.Z heading anywhere in the CHANGELOG (membership); INCONCLUSIVE exit 2 preserved; both directions red-proofed. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt
R-387 The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite. Reasoning kept: «The coercion STAYS; only the silence is fixed — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one.» CLOSED — hub v0.107.0, 2026-08-23. Coercion kept; a WARN now names customer, event type, rejected value and consequence. Dispatcher severity branch kept (it is the only guard for monitor-originated events). Proven live with an error control silent. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt
R-398 resticStep is not a seam, so no test can drive any restic-backed path. CLOSED 2026-10-03 (triage, verified) — premise corrected 2026-08-30; the execution test exists (controller/internal/backup/r358_scratch_marker_test.go:200-205, TestR358_MarkerOrderingIsExecuted) and the resticStepFn seam was deliberately NOT built — it would hide the unlock --remove-all escalation the test asserts full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-405 R-87 sat in CLOSED-ITEMS.md for nine days while still open; the ranking paragraph ranked it fourth pointing at nothing. CLOSED 2026-08-31 — corrected + gated in the same session: R-87 restored verbatim from ef6ac6f^; scripts/closed_register_gate.py is the 12th gate; R-398 stub turned to prose. Gate's residual holes in its docstring; hole 4 is R-406. Predecessor R-378. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-413 R-87's proof caught a naturally-produced hollow snapshot, end to end, unattended. CLOSED 2026-08-31 — the claim it upgrades is recorded (opengist volumes_expected_none_captured, one offsite_proof_empty at error; four apps ahead passed). Related R-87, R-412. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-428 The decoy-coverage gate identified a repository by its DIRECTORY NAME. CLOSED 2026-09-01 — fixed, and kept as the class's best example (felhom.eu CI job 490 found it; the gate now identifies a repo by which registered runner FILE exists under the root; verified under a renamed directory). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-432 A customer's own sub-account can REACH the snapshot door and is REFUSED writes — but sees it EMPTY, so per-file recovery is not product-reachable. ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary (777,600 exact snapshot names, zero hits; /.zfs/snapshot is a different filesystem from /home). Per-file recovery unreachable from the box entirely → R-433. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-drill-r95-recovery-2026-09-01/
R-434 The snapshot-drop alarm promises a recovery that cannot be performed. Reasoning kept: «when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep.» «That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite — which is the whole reason it was not blocked.» CLOSED 2026-09-01 — hub v0.111.1; the promise was DELETED (not replaced) with a sentence true under every provider answer; three tests in hub/internal/monitor/offsite_r434_test.go, red-proofed. The "blocked on R-433" verdict above was mine and it was wrong. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-463 The day the catalog moves postgres:16 to 17, ELEVEN apps are affected and the container image will NOT perform the conversion. (P2) Reasoning kept: «A later move of any of the three needs its own two-venue proof: the engine gate enforces it per app.» «PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check.» CLOSED 2026-09-30 — 8 of 11 moved by the box's own conversion (docmost, paperless-ngx, tandoor, claper, calcom, rallly, outline, sparkyfitness; controller v0.273.0+, 09 decisions 16/42/43); 3 stay by decision 42's rule (zipline, adventurelog, immich). A later move of any of the three needs its own two-venue proof: the engine gate enforces it per app. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md, audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md, audits/pg-calcom-claper-2026-09-28/, audits/pg-last-six-2026-09-30/
R-497 No product channel ever gives the customer the „Tulajdonosi jelmondat”, yet the self-bind mail says they received it at setup. (P2) CLOSED — hub v0.113.0 (2026-09-14, ArgoCD Synced): customer page renders the hand-over sentence; tests red first (passphrase_handover_test.go); mail change pinned by test, not read live. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DOORSTEP-walk-1270-2026-09-14.md
R-500 [P3-LOW] The dashboard shows the last backup in UTC while every backup page shows it in local time — two different clock times for one backup. CLOSED 2026-10-03 (triage, verified) — controller v0.283.0 — the dashboard's last-backup time uses fmtTime, local time (dashboard.html:154; controller CHANGELOG.md:222) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-505 A fresh box on customer tester-1 connects its Cloudflare tunnel but receives NO routes, so the dashboard answers 503. (P1) CLOSED 2026-09-29 — proven end to end on a fresh box. Cause: the tester-1 tunnel had no public hostnames; operator added a route 2026-09-14; box-side hop proven by the 2026-09-29 new-household drill. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-506 [P3-LOW] day0-install.md A.1 says "the controller manages per-app hostnames itself via the tunnel" — it does not. CLOSED 2026-10-03 (triage, verified) — the day-0 runbook names the published-route step and says the controller creates no routes or DNS (documentation/runbooks/day0-install.md:47-53) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-511 A rebuilt customer's box keeps its ep0 PBS token, and the DR tier can be neither provisioned nor re-issued. (P2) CLOSED 2026-09-16 — proven live on a fresh box: hub v0.114.0 re-issue ADOPTS the endpoint token (shipped 2026-09-15); made live by R-534's ep0 grant. Token-only release on host delete not built → R-526. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-p1fixes-2026-09-15/
R-520 A power cut during a guarded Update leaves no record that an update was running. (P3) Reasoning kept: «pct mount maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts, and 9202 keeps /var/lib/docker on its own ext4 mount under the mp0 volume, so from the host that path is an EMPTY STUB.» «An empty directory is not evidence of an absent file.» CLOSED 2026-09-21 — measured on a real version change (scratch 9202, controller v0.260.0, uptime-kuma 2.4.0→2.5.0 cut in pulling; box put the pin back itself). The starting-phase cut → R-610. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/update-arc-2026-09-21/
R-524 When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető” — and the offered Update is a downgrade. (P2) Reasoning kept: «Ahead is NARROW on purpose: every differing service must be orderable AND newer, or the verdict falls back to Behind — this gate can BLOCK an update, so it errs towards letting one run.» CLOSED 2026-09-21 — controller v0.260.0: stacks.CatalogOrder four verdicts incl. Ahead; UpdatePreflight refuses downgrade (409); three red-proofs; 09 §3 decision 10 (decided by CC unattended — operator may reverse). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-534 The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks Datastore.Modify. (P1) CLOSED 2026-09-16 — grant given and proven end to end: DatastoreAdmin for the hub's felhom@pbs on /datastore/felhom-offsite only (narrowest role carrying Modify; PBS has no custom roles); re-issue then ADOPTED on a fresh box. Closes R-511. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt, audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt, audits/evidence-backup-promise-2026-09-16/phaseC-reissue.txt
R-535 The box's console still says it waits for pairing long after the box is bound — and promises the screen refreshes itself. (P2) CLOSED 2026-09-16 — shipped in ISO 1.28.0 (published); on-screen effect not photographed (print_bound_banner). Later SEEN on screen by R-214's closure 2026-09-20. Deliberately does not name the dashboard URL nor reflect the later claim. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png
R-536 The hub is told „Alkalmazás telepítve” the moment a deploy is ACCEPTED, so an install that never finishes is recorded as completed. (P2) Reasoning kept: «The accept-time app.yaml is deliberately NOT deleted on failure — it is the crash-safe record with Deployed:false and it holds the settings the customer typed; the state every surface reads is not_deployed.» CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0: app_deploy_started at accept, app_deployed from the async end, app_deploy_failed (warning); both new types in allowedEventTypes and customerMessages; two red-proofs. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-537 The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and shows the data-drive size, but the tier-1 unit contains NO drive-side app data (P1-HIGH) Reasoning kept: «The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured» CLOSED 2026-09-16 — controller v0.244.0, proven live on demo-hp; label computed per tier; re-proven on a fresh box (ISO 1.28.0). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt
R-538 A tier-1 app restore reports plain success, leaves Nextcloud listing files whose bytes were never backed up, and destroys the app's own trash (P1-HIGH) Reasoning kept: «never present a DB-only restore of a class-A app as a complete one.» «A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can.» CLOSED 2026-09-16 — controller v0.244.0, proven live on demo-hp; unit restore refuses when the unit cannot return the drive-side files; re-proven on a fresh box, off-site route returned 5/5 photos sha256-identical. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt
R-543 Off-site ON by default is not off-site WORKING: a fresh box's tier 3 waits on escrow and nothing asks the household. (P1) Reasoning kept: «The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism.» CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live: escrow-pending bar on every authenticated page (via executeTemplate) + tier-1 sentence rendered by tier3State; both red-proofed; live on 9202 (paused) and 9201 (escrowed). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-recovery-code-2026-09-16/
R-557 Localisation slice 2 — Go-side customer strings follow the language. (P3) Reasoning kept: «Plurals are a BUNDLE rule (a key with .one/.other takes its count first), not a call-site flag.» «EXCEPT one producer: "Sikeres — nincs mentésre jelölt alkalmazás" (controller/internal/backup/offbox.go) must stay Hungarian until R-570 closes, because the page's legacy fallback still reads it on boxes that have not run off-site since 0.251.0.» CLOSED 2026-09-18 - controller v0.252.0 + v0.253.0 + v0.254.0 (release A: 226 literals + i18n_go_parity.py; B: all 179 error literals keyed via util.MsgError; C: saved notes in box language, globe switch). Gaps filed: R-570, R-572..R-578; R-566 closed with it. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-558 Localisation slice 3 — the hub's customer e-mails follow the household's language. (P3) Reasoning kept: «The bind page is per-language with expired pinned to Hungarian (the language would otherwise be the oracle the text refuses to be).» CLOSED 2026-09-18 - hub v0.118.1 + controller v0.256.1 (hub v0.118.0/v0.118.1 + controller v0.256.0/v0.256.1; 56 Hungarian mail goldens unchanged; proven live twice). Gaps filed: R-581..R-585; R-555 closed with it. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-559 Localisation slice 4 — the console banner and the download page in English. (P3) CLOSED 2026-09-18 — ISO 1.29.0 published (sha256 dceacae5…e94829), proof installs on both menu entries + reboot; felhom.eu/en/download live; release gate G16 rewritten per ruling 1b. Gaps filed: R-586, R-587, R-588. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-560 Localisation slice 5 — catalog cards, settings and first steps in English. (P3) Reasoning kept: «lists replace WHOLE and every other list is matched by its own key (env_var, option value, match_group, target, path), never by position;» «The blocks are GENERATED from a flat {path: english} map, not hand-written: a mistyped env_var is INERT on the box rather than an error, and the generator can only write paths that exist on the Hungarian side.» CLOSED 2026-09-20 — controller v0.257.0, catalog fully translated (catalog e81d41e + three batches; 1 031/1 032 strings, ceiling 1), fleet floor 0.257.0 (min_agent 0.131.0). Gaps filed: R-589..R-594. — residue NOT-A-FINDING (2026-10-03): its residue (Peti's box and tester-1 taking floor 0.257.0 unattended) is moot: Peti's box was retired 2026-09-25 and tester-1 takes every floor like any box full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-561 Localisation slice 6 — the volunteer guide in English, then a stranger's first hour in English; closes R-516. (P3) CLOSED 2026-09-20 - controller v0.258.0, the English guide, and the walk (runbooks/VOLUNTEER-first-hour.en.md; R-589/R-590/R-573 fixed; R-214 closed as a side effect). Verdict: not yet ready for an English tester (R-596). Gaps filed: R-596..R-599. R-516 does NOT close. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DRILL-first-hour-en-0258-2026-09-20.md, audits/i18n-slice6-2026-09-20/R-516-item-by-item.md, runbooks/VOLUNTEER-first-hour.en.md
R-566 Three page titles built in Go around an app name stay Hungarian in the English browser tab. (P3) CLOSED 2026-09-18 - controller v0.252.0 (four %s keys via data["TitleArgs"]; pinned by TestParameterisedPageTitles and i18n_go_parity.py). Part of R-557. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-572 [P3-LOW] Two copy-producing template helpers have no English form, so an English page renders „vasárnap" and „%d órája" in Hungarian. CLOSED 2026-10-03 (triage, verified) — controller v0.258.0 — pruneLabel/nextPruneLabel deleted as dead code (controller/internal/web/funcmap.go:368-377) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-580 [P3-LOW] curl -w '%{redirect_url}' prints Basic-auth credentials back into the session transcript. FOLDED into R-132 2026-10-03 — the same curl -w %{redirect_url} credential echo, seen again 2026-09-18 full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-582 An English copy guard written from the Hungarian one matches ordinary words instead of the claim. (P3) Reasoning kept: «The general form worth keeping: a guard ported between languages must be re-derived from what the claim IS in the new language, not translated word for word — and a guard that convicts 141 true sentences is worse than no guard, because it earns an allowlist entry per sentence and then nobody reads it.» CLOSED 2026-09-18 - hub v0.118.0 (recorded for the lesson; three decoys incl. an innocent control): patterns carry the modal and admit an adverb; gate scans bundle VALUES only. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-583 The test-notification mail was the one customer mail that did not follow the language. (P3) Reasoning kept: «The general form worth keeping: the surface you would use to CHECK a feature is the one most worth checking first — a broken instrument that reports success is worse than a broken feature.» CLOSED 2026-09-18 - hub v0.118.1 (mail.test.subject/mail.test.body, red-proofed, proven live in both languages 74 s apart on demo-hp). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-589 The update badge is Hungarian on an English app page — and it is the badge, not a corner case. (P3) Reasoning kept: «The lesson, because it cost a reviewer pass and half a task brief: a reviewer who reads ONE producer cannot see a SECOND producer that overrides it — reading updatebadge.go alone gives exactly the wrong answer.» CLOSED 2026-09-21 — shipped in v0.258.0; the row was stale (English built in web.localeFuncs "updateBadge", pinned by TestUpdateBadgeFollowsTheLanguage; verified at controller 19ef0329ab66). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DRILL-first-hour-en-0258-2026-09-20.md
R-590 [P3-LOW] The data-folder card tells an English household, in Hungarian, whether its files are backed up. CLOSED 2026-10-03 (triage, verified) — controller v0.258.0 — consequenceFor renders through message keys in the household's language (controller/internal/web/datapath_card.go:32-50) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-592 Two defects in the new catalog copy gate, each found by its own decoy rather than by reading it. (P3) CLOSED 2026-09-20 — fixed in the same session, decoys added (scripts/test_gate_decoys.py, 33 cases): coverage scoped to named apps, credential and ASCII-stem bare-substring matches. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-595 The catalog's new copy gate could not RUN in CI at all — six pushes red, six alarm mails, while the local hook was green. (P2) Reasoning kept: «The general form, and the reason this is P2 rather than P3: a new gate is written and tested on the machine that has every library, and the runner deliberately has none.» CLOSED 2026-09-20 — degraded mode; CI job 791 green (catalog 18a6d2d: without PyYAML the gate runs the freeze check and prints what it did not check; five decoys with PyYAML shadowed). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-596 The claim page — the FIRST screen an English household touches — is English chrome with HUNGARIAN messages. (P1) Reasoning kept: «It was deleted, not translated: a translated dead field would have read for ever after as evidence that this page's title is decided in the handler.» CLOSED 2026-09-21 — controller v0.259.0, proven live (fourteen sites via s.msg(r, "claim.msg.*"); dead data["Title"] deleted; language chain pinned by a test; live on guest 9201 both languages; red-proofed). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-597 The setup code is three Hungarian words inside an otherwise fully English e-mail to an English household. (P2) Reasoning kept: «The task's proposed "read it over the phone" filter was MEASURED and NOT adopted — it removes 5270 of 7772 words (68%, 12.92 → 11.29 bits/word) and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster;» «The hub does not own that secret and no row was opened for it: a second definition here is the drift backupTargetAbsentText already demonstrates across two repos.» CLOSED 2026-09-21 — hub v0.119.0: RandomPassphraseFor(lang, use); setup code 4 en words (51.7 bits), owner passphrase 6 en (77.5); entropy floor computed from embedded lists, red-proofed. Recovery code was never Hungarian (agent, EFF list). Decision in source, operator may reverse. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-598 The Backup page's two protection warnings are Hungarian on an English dashboard. (P2) Reasoning kept: «The English is asserted to carry the same NEGATION the Hungarian does ("protects against corrupted files, but not against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian.» CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live (tier names proven live on 9201; degradedMessageFor returns a KEY). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-601 demo-hp is unreachable — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE. (P2) Reasoning kept: «The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.» «Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence.» CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed (both ~/.ssh/config entries repointed to 192.168.0.104; nodes.md corrected; tailscale not installed on demo-hp). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-608 The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the proposed auto-update window. (P2) Reasoning kept: «THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH: Stack.Updating is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it.» «stacks never imports selfupdate.» CLOSED 2026-09-21 — controller v0.261.0: Manager.AnyUpdating() → Updater.SetAppUpdatingCheck; UpdatePreflight refuses self_updating; lock does not latch (TestR608_LockReleasesAfterHold); four red-proofs. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-609 An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never" (P3-LOW) Reasoning kept: «busy/updating/deploying/migrating/self_updating are TRANSIENT and held/downgrade are TERMINAL until a person acts.» CLOSED 2026-09-21 — controller v0.261.0: 409 body gains data.reason additively; the held refusal in actionStack (before UpdatePreflight) carries it too; red-proofed. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-611 A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement (P3-LOW) Reasoning kept: «the report's FIRST section is "not done", even when empty.» «Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason.» CLOSED 2026-09-21 — run by the successor session (scenarios F and G); "not done" is now the report's first section. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/update-arc-gaps-2026-09-21/
R-614 A stale update_phase survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated (P3-LOW) Reasoning kept: «The name is the only thing a new install shares with the old one, so the record has to go when the app does.» CLOSED 2026-09-22 — controller v0.262.0: RemoveStack calls ClearUpdateState; red-proofed. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-623 The unattended-update caller turned every SUCCESS into a timeout, then refused to press that app again — the instrument, not the box (P3-LOW) Reasoning kept: «an instrument that can report a success as a timeout is not a measurement» CLOSED 2026-09-21 — fixed in audits/update-arc-gaps-2026-09-21/unattended-caller.py (unwraps data). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/update-arc-gaps-2026-09-21/unattended-caller.py
R-627 Nothing checked that the register is a well-formed table; an append ate two rows' state cells and R-254 was broken for 45 days (P2-MEDIUM) CLOSED 2026-09-22 — gate 14 (scripts/register_shape_gate.py in repo_gates.py), four decoys, register repaired 317→315. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-628 An empty search of a mailbox I do not control was turned into a claim about what a THIRD PARTY had done, and entered the register as fact (P2-MEDIUM) Reasoning kept: «an empty search may be reported as "absent from the place I looked", never as "it did not happen" — and only after a control query that MUST hit has been seen to hit.» CLOSED 2026-09-22 — rule recorded in the Gmail-access memory; R-433 corrected with the real answers. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-629 The drill catalog sent the operator 47 CI-failure alarms in one night; the drill method did not mention CI (P2-MEDIUM) CLOSED 2026-09-22 — has_actions false on admin/app-catalog-drill, 09 §6.5 method updated; the 47 mails are the operator's to clear. — residue NOT-A-FINDING (2026-10-03): its residue (47 unread CI mails) is the operator's inbox housekeeping, not product work full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-630 paperless-ngx's health probe has never run on any box; a successful update stops the app (P2-MEDIUM, raised to P1) Reasoning kept: «A stack with no probe is not healthy and not failing — it is SETTLED ON CONTAINER STATE (09 §3), and never a reason to stop a running app.» CLOSED 2026-09-22 — controller v0.262.0 + catalog: the no-probe wait settles on container state, the target is decidable (explicit container field), the gate refuses an ambiguity. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/the-28-2026-09-22/sidejobs/r630-controller-words.txt, audits/the-28-2026-09-22/sidejobs/r630.json
R-631 Five templates cannot be judged by the probe gate at all, and one is correct only by accident (P3-LOW) Reasoning kept: «Add expect: {status: 200} to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy and every successful update of it starts stopping it.» CLOSED 2026-09-22 — all five read live on guest 9202, all five probes correct; home-assistant's fragility is now a number (401 with no expect). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/the-28-2026-09-22/sidejobs/r631.json
R-632 Twenty-eight of the 53 templates have never been deployed by any update drill (P3-LOW) CLOSED 2026-09-22 — all 28 walked in one night (26 deployed, 6 proven); findings → R-630, R-633, R-634. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/DRILL-the-28-2026-09-22.md, audits/probe-fix-2026-09-22/not-judged.json
R-633 A remove sent while a restore runs reports success, deletes the record, and leaves a container restarting forever with a live public route (P2-MEDIUM) Reasoning kept: «down returning 0 is a request, not a result» «A refusal from the wrong rule is not evidence for the new one.» CLOSED 2026-09-22 — controller v0.262.0 (busy guard + verified teardown) + v0.262.1 (typed RemoveBusyError → HTTP 409); both halves proven live on 9202. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt, audits/the-28-2026-09-22/apps/gokapi/log.txt
R-701 demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it (P3-LOW) CLOSED — option (b), 2026-09-28 (operator, 09 §3 decision 44): restore_storage → nvme-scratch, FelhomAgentStore granted on /storage/nvme-scratch; manual and first scheduled cycle passed. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt, audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt, audits/logins-nvme-2026-09-28/C/
R-702 Every claper install creates an admin admin@claper.co with the public password claper, published on the household's domain (P1-HIGH) CLOSED — claper fixed 2026-09-28 (catalog 9dc8a05, controller v0.279.0 after_install); the class → R-707. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt
R-703 calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it (P2) CLOSED — catalog 9555e73, 2026-09-28: memory 768M → 1536M (anon peak 817 MiB, 53%). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt, audits/pg-calcom-claper-2026-09-28/box/C0-calcom-memory.txt
R-708 grafana falls back to password admin when its admin field is empty (P3-LOW) Reasoning kept: «no default in the compose (${GF_SECURITY_ADMIN_PASSWORD:?} refuses to start instead).» CLOSED — 2026-09-29 (catalog d0e7e2e): ${GF_SECURITY_ADMIN_PASSWORD:?…} refuses to start empty. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/login-gate-2026-09-29/D/D4-grafana-r708.txt
R-709 The deploy page writes the generated admin passwords of installed apps into its HTML (P3-LOW) CLOSED — 2026-09-29, controller v0.280.0: password field renders empty with a reveal eye; red-proofs RP14/RP15; live on 9202. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/login-gate-2026-09-29/D/D6-r709-live.txt
R-710 An app installed before its template gained an after_install: is never warned about its default login (P2-MEDIUM) CLOSED — 2026-09-29, controller v0.280.0 (RP13): absent record is "not run yet" only 30 min after install; „Megváltoztattam" press; live on demo-hp. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/login-gate-2026-09-29/A/A2-page-warning-after.txt, audits/login-gate-2026-09-29/A/A3-demo-hp-changed-it.txt
R-711 About a dozen class-4 apps keep open sign-up after their first admin exists — the setup gate does not close that (P2-MEDIUM) CLOSED — 2026-09-29 (decision 47, controller v0.281.0, catalog 6faf432): per-app signup_block written when the gate opens; wanderer → R-714. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/login-gate-2026-09-29/B/B-VERDICT.md, audits/gate-rollout-2026-09-29/
R-712 wger refused every browser sign-in behind traefik: "CSRF verification failed" (P2-MEDIUM) CLOSED — 2026-09-29 (catalog d0e7e2e): CSRF_TRUSTED_ORIGINS + X_FORWARDED_PROTO_HEADER_SET; proven on a fresh install. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/login-gate-2026-09-29/D/D2-live.txt
R-713 claper's after_install pastes the household's password into Elixir code, and the controller does not refuse a value that would break such code (P3-LOW) Reasoning kept: «a code-bound value holding a quote, backslash, $, {, }, backtick or line break is refused» CLOSED — 2026-09-29, controller v0.281.0 (RP24) + catalog 6faf432: code-bound unsafe values refused, ${NAME¦base64} added; proven live. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/gate-rollout-2026-09-29/
R-714 wanderer cannot be gated: its web part calls its own database host through the public name (P2-MEDIUM) CLOSED — 2026-09-29 (decision 48, controller 0.282.0): resolved without a gate; sign-up closed by "Close sign-up now" block + PUBLIC_DISABLE_SIGNUP; proven on 9202. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/signup-lock-2026-09-29/
R-715 The setup gate's probe reads only an HTTP-200 JSON object, so three apps with a real status get the button (P3-LOW) Reasoning kept: «a catalog gate that refuses a probe without a before/after measurement in its comment.» CLOSED — 2026-09-29, controller v0.282.0 (RP29, RP30): list indexes in field, done_status:; new catalog gate check-probe-measured.py (5 decoys). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/signup-lock-2026-09-29/
R-716 Apps installed before controller 0.281.0 keep their open sign-up — decision 47 closes it only where the box opened the gate (P3-LOW) Reasoning kept: «a catalog change never touches an installed app» CLOSED — 2026-09-29: operator ruled A (decision 49); built in controller v0.282.0 and pressed on demo-hp adventurelog/opengist and demo-felhom opengist. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/gate-rollout-2026-09-29/0/P0-3-demo-boxes-after-push.txt, audits/signup-lock-2026-09-29/
R-720 A new household's apps are not in the off-site copy: every app starts „3. mentés Kikapcsolva" (P2-MEDIUM) CLOSED 2026-09-30 — controller v0.283.0 (decision 50): fresh install on an off-site box switches the app ON; one-press offer for older apps. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-fixes-first-tester-2026-09-30/partA/
R-721 The household presses Stop during a whole-guest backup, and the backup starts the app again (P2-MEDIUM) Reasoning kept: «unquiesce restarts only stacks whose desired state is still running» CLOSED 2026-09-30 — controller v0.283.1, proven live (v0.283.0 was wrong in production; v0.283.1 wires the adapter and pins the production types). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-drill-new-household-2026-09-30/phase1/step10-stop-undone-by-quiesce.log, audits/evidence-fixes-first-tester-2026-09-30/partE/
R-722 The volunteer guide is stale in five places a volunteer reads literally (P2-MEDIUM) CLOSED 2026-09-30 — both guides rewritten as measured (operator: CC rewrites); every changed line dated. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: runbooks/VOLUNTEER-first-hour.md
R-727 The whole-guest restore test picks a PREVIOUS box's archive, fails on its key, and the ✗ is labelled with the wrong tier (P2-MEDIUM) Reasoning kept: «The restore test skips an archive written with another key (logged by name).» CLOSED 2026-09-30 — agent v0.138.0 + decision 51 (restore test skips archives written with another key); ✗ card names the tier in controller v0.283.0; ep0 drill archives in tester-1 removed; delivered by signed jobs to both demo boxes. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/evidence-fixes-first-tester-2026-09-30/partC/
R-730 The published ISO 1.29.0 was built from a tree that git cannot name, so it cannot be tagged (P3-LOW) Reasoning kept: «scripts/iso/build-felhom-iso.sh now runs a clean-tree gate right after argument parsing: any uncommitted or untracked change, or HEAD ≠ origin/main, refuses the build (no bypass flag); the manifest records repo-commit from the gate and iso-version-tag : iso-v<version>.» «the installer-v* tags are the host-install SCRIPT's versions (SCRIPT_VERSION 1.28.0 on main; the website serves /scripts/ from installer-v1.28.0), a different number line from the ISO — an installer-v1.29.0 tag would block the next script release.» CLOSED 2026-09-30 — the gate + test; 1.29.0 untagged by decision (its build tree is not a commit). build-felhom-iso.sh clean-tree gate (no bypass), manifest records iso-v; test scripts/iso/test/clean-tree.sh red-proofed. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/pg-last-six-2026-09-30/E/E4-installer-tag.txt, scripts/iso/test/clean-tree.sh
R-732 immich's FIRST start at the catalog pin could not finish on the bench: its database container was OOM-killed at 512 MiB (P2-MEDIUM) CLOSED 2026-09-30 — catalog 56c4888 (768M + v3.2.4); older step 0b8272068aab36bf re-proven at 768M (catalog 48440ce, 63a96b0); an installed immich gets it with its next guarded Update. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/pg-last-six-2026-09-30/F/immich-first-start-oom.txt, audits/immich-first-start-2026-09-30/A-cause.md, audits/more-night-apps-2026-09-30/
R-735 The test bench generates a password:N:special deploy field WITHOUT a special character, so calibre-web refuses it (P3-LOW) CLOSED 2026-09-30 — catalog 5b1972b (_gen mirrors randomWithSpecial; test seen failing first). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/more-night-apps-2026-09-30/bench/apps/calibre-web/, A/R735-red.txt, A/R735-green.txt
R-736 Removing or updating an app never deletes the image it pulled, so a box's Docker disk fills until an install or update is refused (P2-MEDIUM) Reasoning kept: «Keep set box-wide at delete time; exact id; no pass while an update runs; one summary line per pass; a one-time sweep.» «the Remove button runs RemoveStack, not DeleteStack; docker image ls without -a hides the untagged digest-pulled images» CLOSED 2026-09-30 — controller v0.284.2, floor 0.284.2 (decision 53, option A; 0.284.0/0.284.1 never floored). Old controller images → R-745. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/more-night-apps-2026-09-30/box/T2-images-removed-by-name.txt, audits/night-rulings-2026-09-30/
R-737 wger's app login API answers 500 on a CORRECT password: the template sets no JWT_PRIVATE_KEY (P3-LOW) CLOSED 2026-09-30 — catalog 45d8482 (RSA pair generated once by wger's manage.py generate-jwt-keys, kept 0600 on the data volume; bench-proven, not proven on a box). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/more-night-apps-2026-09-30/B/wger-probe.txt, audits/night-rulings-2026-09-30/
R-740 A same-tag upstream security fix (postgres:18-alpine, redis:7-alpine, mariadb:12.3 …) reaches NO box, because the catalog never records a re-test of a tag at a new digest (P2-MEDIUM) Reasoning kept: «Not a cron job (runbook runbooks/monthly-floating-retest.md says why); who presses it monthly is the operator's word (STATUS).» CLOSED 2026-09-30 — built (catalog 6a3ead9); the monthly run is a standing step. Decision 52 option A: re-test entries + scripts/retest-floating.py; proven end to end on 9202. Exact-tag re-pushes → R-743. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/more-night-apps-2026-09-30/C/C3-floating-repush.txt, audits/more-night-apps-2026-09-30/C/C4-same-tag-retest.txt, runbooks/monthly-floating-retest.md, audits/night-rulings-2026-09-30/
R-741 For a few seconds after a fresh install, an after_install app answers its PUBLIC default login through the front door (P3-LOW) Reasoning kept: «an after_install app is installed HELD behind the setup gate's door until the login is replaced (or the household says it changed it).» CLOSED 2026-09-30 — controller v0.284.2 (install hold: an after_install app is installed held behind the setup gate's door until the login is replaced; red-proofs RP-IH1..4). mealie lock → R-747. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/more-night-apps-2026-09-30/A/A3-calibre-default-login-window.txt, audits/night-rulings-2026-09-30/
R-742 zipline 4.8.0 cannot be reached from 4.6.1 in one step: it refuses to start until the database ran the release before it (P3-LOW) CLOSED 2026-09-30 — catalog a9700e2 + fb87030 (two steps 4.6.1 → 4.7.0 → 4.8.0, each proven on both venues). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/more-night-apps-2026-09-30/bench/apps/zipline/, audits/more-night-apps-2026-09-30/box/zipline/
R-743 Exact version tags are re-pushed under the same name too — not only floating lines (P3-LOW) Reasoning kept: «decisions 54 (a CC session the operator starts with the standing brief runs it monthly) and 55 (every app with a proven ladder; --engines-only a switch).» CLOSED 2026-10-01 — decisions 54, 55; both tags re-tested (nextcloud catalog 3b59dfb, sonarr 1a37032). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/night-rulings-2026-09-30/A/, audits/retest-2026-10/
R-744 outline's fixture cannot seed outline 1.10.1: no csrfToken cookie after installation.create (P3-LOW) CLOSED 2026-10-01 — fixture fixed, proven on 9202 (accepts __Host-csrfToken; catalog 9fc7052). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/rulings-2026-10-01/D/D2-outline-fixture-9202.txt
R-745 Old CONTROLLER images are never deleted: ~50 versions on each demo box (P3-LOW) Reasoning kept: «the agent rolls back to the RUNNING image (what /etc/felhom-controller-image named), never to a previous one the controller hands it.» CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0 (decision 56: keep running + previous, delete older/untagged). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/night-rulings-2026-09-30/E/, audits/rulings-2026-10-01/B/
R-746 image_digest.resolve ignores a @digest in its argument — it answers the TAG's current digest (P3-LOW) CLOSED 2026-10-01 — catalog 804884a (manifest asked by digest; malformed digest refused; scripts/test_image_digest.py red-proofed). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/rulings-2026-10-01/D/D1-r746-digest.txt
R-748 The register-shape gate skipped every row whose id has a letter suffix (R-88a, R-88b, R-209a) (P3-LOW) CLOSED 2026-09-30 — scripts/register_shape_gate.py (R-\d+[a-z]?; decoy suffix-row-eaten-state). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-749 retest-floating.py could never start on a fresh bench: it checked for /opt/upg/upgrade-test.py before the step that copies it (P3-LOW) CLOSED 2026-10-01 — catalog 9e53205. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/rulings-2026-10-01/A/
R-750 The registry no longer holds controller releases older than 0.213.0 — something removed them, and nothing records what (P3-LOW) Reasoning kept: «gitea-image-prune.sh keeps the newest 20 + every version in use (floor, vouched golden, vouched agent, min_agent, the golden's baked images, the running hub), refuses (exit 3) when that list is unreadable» CLOSED 2026-10-01 — decision 62, misc-scripts c9d5ed5 (cause: manual gitea-image-prune.sh --keep 7 run 2026-08-22/23, HM-024; prune now keeps newest 20 + every in-use version, exit 3 when unreadable). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/rulings-2026-10-01/B/, audits/lockouts-2026-10-01/C/C1-registry-read.txt, audits/calibre-name-and-prune-2026-10-01/B/
R-751 The image clean-up after an app update could crash the whole controller: nil stack dereference when the app was gone (P2) CLOSED 2026-10-01 — controller v0.285.0, floor 0.285.0 (TestRetainImagesAfterUpdate_AppGoneDoesNotPanic seen panicking first). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/rulings-2026-10-01/B/B1-red-proofs.txt
R-752 Four more catalog apps let a stranger lock the household out with wrong passwords for a known login name — like mealie (P3-LOW) Reasoning kept: «Installed apps: a settings-only change reaches the stack file at the next sync (images equal, ≤15 min) and the running app at the next compose up -d — Restart/Start (measured: the env changed only at Restart), an Update, or a backup's restart (backup.go:972, read).» CLOSED 2026-10-01 — decisions 58–61 (wger catalog 82fff32; BookStack, Grafana kept; calibre-web generated ADMIN_USER catalog e9f50b5). Installed calibre-web's invented name → R-757. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/lockouts-2026-10-01/B/
R-753 Behind the tunnel every visitor reaches an app with the SAME address, so every per-address guard is an 'everyone' guard (P3-LOW) Reasoning kept: «cloudflared alone on felhom-tunnel at the fixed 172.16.253.2, traefik trusts forwarded headers from it only, an entrypoint middleware removes client-writable host/path/address headers» CLOSED 2026-10-01 — controller v0.286.1 (decision 63; v0.286.0 never floored; catalog 04e9516..50e4fb4) (rows R-776..R-779 carry what is left). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/lockouts-2026-10-01/A/A1-client-address.txt, audits/visitors-2026-10-01/A/
R-754 01-topology-and-trust.md §7 says cloudflared runs on the Proxmox HOST; on every box it runs INSIDE the guest (P3-LOW) Reasoning kept: «the data path is not independent of the guest.» CLOSED 2026-10-01 — document corrected (operator: the build is right). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-765 A new app that builds its login from a type: password value at EVERY start loses the login on a restore after a remove — Radicale (P2-MEDIUM) Reasoning kept: «a restore carries only type: secret values (stacks.PortableSecretEnvVars, the D5 ruling — a type: password is an internet-reachable login and stays out of the drive's backup)» «the login file is written on the FIRST start only and lives on the data volume.» CLOSED 2026-10-01 — Radicale's template, before publishing (login file written on first start only, on the data volume). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/new-apps-2026-10-01/box/radicale-attempt1/restore.txt, audits/new-apps-2026-10-01/box/radicale/restore.txt
R-767 MeTube has no login at all, by design, and the box has no permanent household-only door — so it is not built (P3-LOW) CLOSED 2026-10-02 — controller v0.287.0 + catalog (MeTube published) behind the family gate, decision 64, no exception. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/new-apps-2026-10-01/FIT.md, audits/family-gate-2026-10-02/A/items.txt, audits/family-gate-2026-10-02/B/box/metube-fresh.txt
R-772 A health probe that finds NO container to probe records healthy: true (P3-LOW) Reasoning kept: «a not-run probe records healthy: false, not_checked: true, is looked at on the next tick» CLOSED 2026-10-01 — controller v0.286.1 (not-run probe records healthy false, not_checked true; red-proofs RP-D1, RP-D1b). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/new-apps-2026-10-01/box/karakeep-768M/neg-and-crawl.txt, audits/visitors-2026-10-01/D/
R-773 After a remove + restore, an app's sign-up route block (decision 47) is gone (P3-LOW) Reasoning kept: «a REMOVED app restored from its backup gets the lock record (opened_by: restore) and the block written before anything starts» CLOSED 2026-10-01 — controller v0.286.0 (restored app gets lock record opened_by: restore + block before start; red-proof RP-D2). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/new-apps-2026-10-01/box/karakeep/restore.txt, audits/visitors-2026-10-01/D/r773-live.txt
R-780 A permanent household gate with family accounts — the spike PASSED; the build waits for the operator's go (P2-MEDIUM) Reasoning kept: «Build requirement F1: exceptions must be anchored (PathPrefix(/api/v1/opds) also matched /api/v1/opdsx).» CLOSED 2026-10-02 — controller v0.287.0 (decision 64; internal/family, Család card, family_gate fields, catalog gate check-family-gate.py; Grimmory and MeTube published). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/permanent-gate-2026-10-01/VERDICT.md, audits/family-gate-2026-10-02/A/items.txt, audits/family-gate-2026-10-02/E/floor.txt
R-787 No catalog app's LICENCE has been checked against Felhom being a paid service (P2-MEDIUM) Reasoning kept: «New apps carry their licence in record row 0.1.» CLOSED 2026-10-02 — the read (decisions are R-784, R-789..R-795). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/licences-2026-10-02/TABLE.md
R-788 The volume-persistence gate reads an EMPTY declared volume as CLEAN (P3-LOW) Reasoning kept: «an app that wrote nothing is UNDETERMINED (exit 2), never a pass.» CLOSED 2026-10-02 — catalog (empty declared volume after the exercise is UNDETERMINED, red-proofed; APP_EXERCISE added). Bind mounts → R-805. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/family-gate-2026-10-02/C-metube-bench/C2-volume-persistence-metube.txt, audits/persistence-sweep-2026-10-02/A/RP-R788-empty-volume.txt
R-790 Emby is proprietary: a personal, non-commercial, non-transferable licence (P2-MEDIUM) Reasoning kept: «KEPT — the household is the licensee; Felhom installs the official, unchanged image.» CLOSED 2026-10-02 — operator ruling (kept), decision 66; part of the lawyer's review R-802. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/licences-2026-10-02/TABLE.md
R-791 n8n's Sustainable Use License allows own internal or personal use; distribution only free of charge for non-commercial purposes (P3-LOW) Reasoning kept: «KEPT — the household is the user; the official, unchanged image.» CLOSED 2026-10-02 — operator ruling (kept), decision 66; part of the lawyer's review R-802. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/licences-2026-10-02/TABLE.md
R-792 Plex is proprietary: the HOUSEHOLD is the licensee under its own Plex account (P3-LOW) Reasoning kept: «KEPT — the household is the licensee under its own Plex account.» CLOSED 2026-10-02 — operator ruling (kept), decision 66; part of the lawyer's review R-802. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/licences-2026-10-02/TABLE.md
R-795 recipe-importer, Felhom's own image, has no licence file (P3-LOW) Reasoning kept: «Felhom's own code — no action unless it is shared with anyone.» CLOSED 2026-10-02 — operator ruling (no action unless shared), decision 66. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/licences-2026-10-02/TABLE.md
R-800 [P2-MEDIUM] "Delete your data from the hard drive" keeps the app's files in the household's userdata folder — and neither the dialog nor the result says so. CLOSED 2026-10-03 (triage, verified) — controller v0.288.0 (580b656, decision 67) — the remove dialog and result name the kept userdata folder; live on 9202 (0945332) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
R-801 The volume-persistence gate never sent a request to ANY app: it read the routed port from label VALUES, not the label NAME (P2-MEDIUM) CLOSED 2026-10-02 — catalog (the fixed gate routed_ports() + the re-sweep of all 58: CLEAN 39 · UNDETERMINED 19 · BROKEN 0). Follow-ups R-805..R-807. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/family-gate-2026-10-02/B/volume-persistence-metube-exerciser-diag.txt, audits/family-gate-2026-10-02/B/RP-R801-routed-ports.txt, audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md
R-803 papra cannot start under its own memory limit: a fresh install is OOM-killed in its migration and crash-loops (P1-HIGH) CLOSED 2026-10-02 — catalog (papra 768M, mem_request 256M; measured from birth, 0 kills, persistence CLEAN; no box ran papra). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/persistence-sweep-2026-10-02/A/sweep/papra-oom-diagnosis.txt, audits/persistence-sweep-2026-10-02/A/papra/SUMMARY.md

Rows that never had an R- id (the 2026-07-28 table of campaign findings and watches):

Row What Closed Full text
E-2d Prove E-2 on a fresh VM — a real host-install 1.22.0 run, Case B, a claimable customer, add a drive and unplug it. CLOSED — PARTIALLY PROVEN (2026-07-29). C1, C2 proven; C3, C4 proven live; C5 FAILED → R-116 (the single named open leg; R-116 since CLOSED in CLOSED-ITEMS). Per Session-C runbook §9 a failed claim closes as partially proven, no re-run. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md
— (watch, line 285) Storage Box snapshots on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but 0 taken yet CLOSED 2026-10-03 (triage) — snapshots exist: R-429 recorded seven daily Storage Box snapshots (2026-09-01) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
— (watch, line 289) demo-felhom's next weekly PBS backup (newest is 2026-07-26) CLOSED 2026-10-03 (triage) — folded into R-91, whose trigger is exactly this backup full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
— (watch, line 290) demo-felhom's next restore-test (84 h cadence, last 2026-07-27 06:38 UTC) CLOSED 2026-10-03 (triage) — restore tests have run on demo-felhom on their cadence ever since (R-672, R-689 in CLOSED-ITEMS.md) full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
F-CRIT-2 A failed offsite backup left a phantom snapshot that RESET the tier's freshness clock — 7 days silent Reasoning kept: «NewestArchiveTime now counts only plausibly-complete entries (measured 1 MiB floor; undecidable ⇒ not counted).» SHIPPED + PROVEN-LIVE (agent v0.106.0, 2026-07-28). Campaign fault 2 replayed on demo-hp: phantom rejected + logged once, tier DUE and backed up, no thrash on the inverse. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
F-CRIT-1 An app that fails to restart after a quiesce never alarms on any channel SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28). Live on demo-hp: alarmed 9s after grace expiry; a deliberate user stop stayed silent through 9 dead-app scans. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
F-A1 A restore-test in flight made a healthy backup report as FAILED (HTTP 409 read as a tier failure) Reasoning kept: «409 → contention: tier stays DUE, dropped before anything stops (15m), and BLOCKED alarm if contention outlives the agent's 120m ceiling (3h).» SHIPPED + PROVEN-LIVE (controller v0.179.0, 2026-07-28). Hub DB: 409 → 0 operator emails, real failure → 1. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
C9-F1 Tier-2 „Fájlok visszaállítása" is offered for apps whose copy has no restorable file leg; restores 0 files and reports all files present Reasoning kept: «Tier2RestoreCoverage refuses UP FRONT without stopping the app and NAMES the working action; a run that proceeds claims only what it examined and discloses that the database and volumes are not covered.» SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28). Live on demo-felhom: bookstack refused without restart; paperless re-run byte-identical, 16/16 docs. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
C9-F2 An app in a Docker crash loop never alarms on any channel; StateRestarting is in no down-set Reasoning kept: «StateRestarting deliberately NOT added to IsDownState (that alarms on every deploy fleet-wide)» SHIPPED + PROVEN-LIVE (controller v0.183.0, 2026-07-28). Sustained restarting becomes down after crashLoopAfter=5m; dashboard counter uses the same predicate. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
C9-F3 → R-104 An interrupted offsite run leaves an exclusive restic lock the self-heal cannot reach: resticStep (offbox.go:634-648) has unlock --remove-all, but ensureOffboxRepo's probe fails first, classifyResticProbe (offbox.go:77-93) has no lock case → … FOLDED into R-104 — the same finding; R-104 has its own open row full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
D5 Move app secrets into the LOCAL recovery unit so Tier-1/Tier-2 restore stop needing the guest Reasoning kept: «Operator ruling 2026-07-30: type: secret travels (45 fields), type: password NEVER (7) plus a code register (vaultwarden/ADMIN_TOKEN); plaintext. The exclusion is what LICENSES the plaintext — coupled, not independent.» «the register is code, not a catalog flag (a boundary a catalog push can move is not a boundary — R-97a).» «Precedence: the UNIT WINS over the guest, because the unit's secrets were captured in the same run as the dumps beside them and therefore match the data being restored; pinned both directions.» SHIPPED + PROVEN-LIVE (controller v0.188.0, 2026-07-30). Operator ruling 2026-07-30 on which secrets travel; manifest schema 2; live proof on scratch guest (AdventureLog secrets recovered 2/2, app read seeded row over TCP); 4 red-proofs. Related: R-127. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/D5-drive-alone-restore-2026-07-30.md
F-DIAG Four distinct offsite failure causes collapse into two operator-visible strings Reasoning kept: «Unclassifiable says so rather than being folded into a neighbour.» SHIPPED (controller v0.182.0, 2026-07-28). ClassifyOffsiteFailure → quota / orphaned / no_repo / no_units / transport / unknown; redaction by the target's actual host/user/path. Unit-proven; not yet exercised by a live offsite failure of each class. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
F-OPS A manual pct restore inherits the source guest's bind mounts — during a real DR, on a different host, under pressure Reasoning kept: «Docs only by design — the agent already neutralises binds on its own restore paths, and a second implementation would drift» DOCUMENTED (2026-07-28) in runbooks/RUNBOOK-manual-guest-restore.md (mpN volumes vs binds, the mp9 source-VMID trap, strip-and-re-add, positive pre-start verification). full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: runbooks/RUNBOOK-manual-guest-restore.md
F-REBOOT A guest rebooted during its backup does not come back — 9m47s outage with every alarm silent Reasoning kept: «onboot is the deliberate-stop discriminator (already the stale-lock path's, and what pve-guests consults)» SHIPPED + PROVEN-LIVE (agent v0.107.0, 2026-07-28). 60 s guest-power watchdog, retry 3x/1m-2m-4m then escalate. Live on demo-hp: 120 s unattended vs 587 s; onboot:0 guest left stopped. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
F-LEAK A failed restore-test cannot destroy its own scratch guest (403 VM.Allocate); the 10-slot VMID band shrinks silently Reasoning kept: «4th root-fenced exception, band enforced in sudoers literally (pct destroy 99000[0-9] --purge) + in code + at the caller; API destroy still tried first.» SHIPPED + PROVEN-LIVE (agent v0.110.0 + host-install v1.21.0, 2026-07-28). Third attempt shipped: 4th root-fenced exception; live band permitted, 9201/9100/9999/990010/1 and pct start 990000 refused. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
F-OBS deadapp-check leaves NO positive observable on a default (info-level) box SHIPPED + PROVEN-LIVE (controller v0.180.0 + agent v0.109.0, 2026-07-28). INFO summary every 20th scan; agent v0.109.0 fixes the same shape in the guest-power watchdog. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
E-2 Drive-role machinery around the moved vzdump target CLOSED — PARTIALLY PROVEN (Session C, 2026-07-29). C1/C2 proven in E-2d; C3, C4 proven live (R-114, R-112); C5 FAILED → R-116. Broken legs R-112, R-113, R-114, R-116 all since CLOSED (CLOSED-ITEMS). Arc's definition of done: R-106+R-109, R-108, D5. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md; evidence: audits/SESSION-C-2026-07-29.md, audits/E2D-fresh-vm-2026-07-29.md
E-2a The target move needs a root-fenced wrapper — the agent cannot do it Reasoning kept: «the agent's PVE role was NOT widened.» «has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint.» SHIPPED + PROVEN-LIVE (agent v0.113.0 + host-install v1.22.0, 2026-07-29). felhom-backup-target-apply behind literal FELHOM_BACKUPTARGET sudoers alias; all five laws proven live as root on demo-hp, 0 stray storages. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
E-2b NotifyStorageDisconnected/Reconnected defined and called NOWHERE — a drive going absent emitted no event SHIPPED + PROVEN-LIVE (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29). Seam wired in ReconcileDriveGates; keying bug (guest path vs host MountPath) caught before deploy. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md
E-2c E-1 put the whole-guest backups on a drive POST /disks/eject would eject Reasoning kept: «NOT a role reclassification — RoleForStorage untouched, because on both boxes that drive is ALSO the enrolled user-data drive» SHIPPED + PROVEN-LIVE (agent v0.112.0, 2026-07-29). Eject + decommission refuse 409 on the backup-target mount; refused live on both boxes. full text: git show 9e2786c:documentation/backlog/OPEN-ITEMS.md

2026-09-17 — nothing decides by reading a Hungarian word (controller v0.251.0)

Row What Closed Full text
R-553 Five places decided behaviour by matching their own Hungarian wording (P3). Closed in controller v0.251.0 (c00fed6 + f806baf): each producer attaches a signal and each decision reads it — deploy sentinels (stacks.ErrRequiredField / ErrPathMissing / ErrNotEnoughMemory / ErrAlreadyDeployed → api.deployStatusFor, 400/409), backup.ErrOffsiteQuota in ClassifyOffsiteFailure, monitor.HealthReport.WarningKinds for alert placement, settings.OffboxTarget.LastWarningKind for the stale off-site note. util.KindErrorf produces the same message bytes fmt.Errorf did while carrying the sentinel, so not one Hungarian byte moved (pinned per producer). The hub wire is unchanged: report.HealthReport still carries exactly status, issues, warnings — proven live and by test. Each site red-proofed by restoring its pre-fix predicate. Reasoning kept: a text signature is replaced by a signal only where WE produce the text — restic's and ssh's English output stays matched by text, because we neither write nor translate it. audits/r553-2026-09-17/ CLOSED 2026-09-17 — PROVEN-LIVE (the 409 refusal and the page bytes; the 400 / quota / stale-note paths are red-proofed tests, not provoked on a live box) full text: git show b408b28:documentation/backlog/OPEN-ITEMS.md
R-563 The remote-backup page started its progress poll by reading its own Hungarian word „Fut" (P3). Closed in controller v0.251.0 (f806baf): the status element carries data-status="{{.Offbox.LastStatus}}" and the poll reads v.dataset.status === 'running'. The word became {{T "backups_remote.fut"}} (en „Running…"), so the English page can finally see a running backup — slice 1 had to leave that one word Hungarian for this reason. 12 parity fixtures re-captured; each differs from its predecessor by exactly the attribute and that poll line, the other 94 byte-identical. Red-proofed twice (poll back on the word; word back inline). audits/r553-2026-09-17/site5-fixture-diff.txt CLOSED 2026-09-17 — PROVEN-LIVE (the rendered attribute and poll line on demo-hp) full text: git show b408b28:documentation/backlog/OPEN-ITEMS.md

2026-09-17 — localisation slice 1: the whole dashboard in English (controller v0.248.0–v0.250.0)

Row What Closed Full text
R-556 Localisation slice 1 — the other 31 dashboard templates in English, Hungarian byte-identical (P3). Closed in controller v0.248.0 (ff68b0b, apps and settings), v0.249.0 (11b790e, backups) and v0.250.0 (ac149c2 + be2efe5, storage, sharing, sign-in, claim, guest share, catch-all, debug, and the switch for everyone). Every page's parity fixtures were captured from its UNCONVERTED template before conversion (6be55a0, f8ebc47, 0555091); 106 states. TestI18nParityCoversEveryMarker (every marker rendered by a case); English page mask narrowed to ≥ 2 words / ≥ 12 chars; retrieval-promise gate scans English; executeTemplateLang for the six pages outside the dashboard chrome (no session CSRF, no escrow reminder — TestI18nDirectRenderPagesHaveNoAdminChrome); TestHandlerTitleKeysMatchHungarianTitle. Decision 6 (switch hidden) superseded: the switch is on every dashboard page; the 89 layout fixtures re-captured and each equals its predecessor plus exactly one switch form. Proven live on demo-hp 9201 for each release: Hungarian pages equal their previous-version fetch apart from live numbers (and, in C, the switch form); POST /settings/language made them English and back; the hub stored hu, en, hu. Startup parse: hu ~24–27 ms, en ~20–28 ms. Reasoning kept: a word a page COMPARES stays unconverted until the comparison is language-neutral (R-563); a fixture is never regenerated to make a conversion pass — the two re-captures in this slice (C's case titles, the switch) were captured from the tree that defines the truth and their diffs were proven to be exactly the intended change. Found: R-563, R-564, R-565, R-566, R-567, R-568; R-516 extended. audits/i18n-slice1-2026-09-17/{A,B,C}/ CLOSED 2026-09-17 — PROVEN-LIVE full text: git show bc15153:documentation/backlog/OPEN-ITEMS.md

2026-09-16 — the drill: the tunnel answers from outside

ID Title Shipped Evidence
R-510 [P1-HIGH] tester-1's tunnel now has its route, and still gives a fresh box 502: the route sends traffic to https://traefik WITH certificate checking, and traefik answers the name traefik with its default certificate. operator ticked „No TLS Verify" 2026-09-15; proven with a box connected 2026-09-16 three GETs from DooPlex to https://felhom.enkicsifelhom.hu at 10:04:19/22/25Z → 302 ×3 (server: cloudflare), following it → 200 on the Hungarian claim page „A szerver beállítása … Add meg az e-mailben kapott beállító kódot" (no 502, no 530). audits/evidence-drill-0243-2026-09-16/phase1-tunnel.txt

2026-09-16 — the drill's Phase 0 (hub v0.115.0, golden 0.243.0, the signing ruling)

Two rows closed. Full original text: git show <this commit>^ -- documentation/backlog/OPEN-ITEMS.md.

ID Title Shipped Evidence
R-529 [P3-LOW] The agent-plane host_stale / host_down / host_recovered mails still wait out the one-hour quiet rule. hub v0.115.0 (2026-09-16, operator ruling 2) nodeLivenessEvents + red-proof TestOperatorCooldown_NodeLivenessBypassesQuietHour; 08-alarm-ladder.md §6.2
R-533 [P3-LOW] The operator signing keys were placed on DooPlex world-readable (mode 664) and sit outside any documented location. operator ruling 1, 2026-09-16 keys at /mnt/5_hdd/felhom.eu/felhom-op-{operational,rec-recovery} + felhom_op_ed25519, mode 0600 owner kisfenyo; CONTEXT.md, 04-control-plane-authorization.md §3.1

2026-09-15 — the big night's P1 fixes (agent v0.131.0, controller v0.243.0, hub v0.114.0, catalog templates, ISO 1.27.1 published)

Nine rows closed. Full original text: git show <this commit>^ -- documentation/backlog/OPEN-ITEMS.md. Evidence folder: documentation/audits/evidence-p1fixes-2026-09-15/.

ID Title Shipped Evidence
R-493 [P1-HIGH] There are NO customer-facing install instructions, so a volunteer cannot begin — this blocks inviting anyone. ISO 1.27.1 published + felhom.eu/letoltes live 2026-09-15: round trip over https://iso.felhom.eu sha256 25637007…c053, 1 705 322 496 B; felhom.eu/letoltes 200 naming 1.27.1 (documentation/tests/iso-release-1.27.1-2026-09-14/README.md)
R-495 [P2-MEDIUM] The public installer asks a stranger four questions nothing answers, and REFUSES its own default on one of them. answered by the guide; ISO 1.27.1 published 2026-09-15 same as R-493
R-496 [P2-MEDIUM] The box's console tells a stranger, in English and FIRST, to open the Proxmox admin page — and calls the owner passphrase „a jelszavadat”. ISO 1.27.1 (console Felhom-only) published 2026-09-15 same as R-493; G15 live proof on VM 332
R-512 [P2-MEDIUM] Vaultwarden is installed with open registration, and the one control the page tells the customer to use to close it is read-only. catalog template (no version change), 2026-09-15 spike: stranger 400 / invite 200 / invited 200 (E1-vaultwarden-spike.txt); live on 9202 from the catalog: SIGNUPS_ALLOWED=false, stranger 400 „Registration not allowed" (E1-vaultwarden-9202-live.txt)
R-513 [P1-HIGH — SECURITY] Every box's file manager (FileBrowser, files.<domain>, a launcher tile) accepts the login admin / admin, and on demo-hp that login page is on the public internet. controller v0.243.0 9202 (default login): generated, stored encrypted, reveal 200 (16 chars), revealed=200 admin=401 wrong=401; 9201 (hand-set): recorded operator 08:52:10Z, untouched; public files.enkisfelhom.hu admin → 401 (B1/B4 files)
R-514 [P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut". catalog template: 1 worker × 1 thread, 1280M live on 9202: 20 PDFs at once → 20/20 SUCCESS, memory.peak 772 370 432 B, no OOM (E2-paperless-9202-live.txt). The OOM-visibility half is NOT proven live → R-528
R-515 [P3-LOW] The Paperless-ngx app page tells the customer to log in with admin / admin, and that login does not exist: the deploy form generates the admin password. catalog template, 2026-09-15 default_creds removed; first steps point at „Automatikusan generált értékek" (app-catalog commit e6aa443)
R-517 [P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true. controller v0.243.0 + agent v0.131.0 live on 9201: per-tier rows „Helyi tároló (local) ✓ … Naprakész", „PBS ✓ … Naprakész", remote tick on a real PBS success (C4-backup-page-9201.txt). Found live and fixed after the release: an unknown size printed „0 B" → „–" (controller main d3eacbb, unreleased)
R-523 [P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule. agent v0.131.0 + hub v0.114.0 (+ golden script --restart always, no bake) measured first: kill leaves both unless-stopped and always exited (A1). Live on 9201: idle kill → dashboard 200 in 59 s; parked → stayed dead 100 s with the PARKED line, unpark → 200 in 25 s; kill during a swap → supervisor deferred ×3, swap rolled back itself; the crash-loop guard tripped for real after 3 test restarts and the hub mailed controller_crashloop (A4-* files). NOT measured: restart timing during a deploy (the budget was spent) → R-531

| R-442 | remove_hdd_data: true was INERT — the customer's data stayed on the drive while the API reported success (HTTP 200, hdd_paths_removed: null, 128 MB of Nextcloud left; demo-hp 2026-09-01). Shipped in controller v0.236.0 (2026-09-13). Removal resolved the drive from the GLOBAL cfg.Paths.HDDPath (no default, set on NO box — demo-hp AND demo-felhom both measured 0 hdd_path / 0 FELHOM_PATHS_*, so the fleet shares the shape); it now reads the app's OWN app.yaml HDD_PATH — the 07-backup-architecture.md ~L437 rule that deploy, the start gate and the backup destination already implemented — and a data removal it cannot resolve, or whose drive is absent, is REFUSED (409, exact Hungarian sentence, typed stacks.RemoveRefusedError) BEFORE compose down, app kept. SSD app → hdd_paths_removed: [], never null, plus hdd_note. Missing folders stated in hdd_paths_missing. The backup-half refusal reaches the response (backup_paths_refused) and its base follows the same rule. Reasoning kept: "declares no drive" ≠ "could not resolve the drive" — the first is a fact, the second a refusal; an app gone with its data left behind is unrecoverable from the UI — the customer cannot even re-run the removal; no fallback to the global — that silent fallback is the exact path this closes. Observation carried: 8 of the 13 needs_hdd catalog apps bind ONLY ${USERDATA_PATH} (the shared library) and no ${HDD_PATH} folder, so for them "delete my data" correctly removes nothing on the drive and the modal shows no checkbox. | CLOSED 2026-09-13 — shipped controller v0.236.0, proven live on demo-hp | audits/R442-2026-09-13/ — A: 63 MB written by Nextcloud ITSELF, gone after removal and listed with its size; C: 409 + sentence, all 69 files untouched, app still deployed, [ERROR] … refused logged; D: gokapi [] + note; ASCII controls (llap: C=2 A=0 D=0). 15 tests + two red-proofs in felhom-controller/REPORT.md. Full original text: git show d6837d98ee24:documentation/backlog/OPEN-ITEMS.md. | | R-449 | UPDATE ARC SLICE 5 — an upgrade test that runs again. BUILT AND RUN 2026-09-06: app-catalog-felhom.eu/scripts/upgrade-test.py + upgrade_fixtures.py, 7 edges across 3 apps, evidence in audits/upgrade-spike-2026-09-06/. Reasoning kept — success is an APPLICATION-LEVEL READBACK, never file identity: survive2.py's sha256+inode rule is right for a redeploy and WRONG for an upgrade, because a migration is supposed to rewrite files and that rule would fail every correct upgrade. Reasoning kept — nothing is ever seeded into a volume by hand (R-156); an app with no non-browser route is recorded inconclusive, which is a result and not a licence to plant a file. Reasoning kept — run the negative control FIRST: C3's TO image exits immediately and came back failed; a harness that cannot fail a known-broken upgrade proves nothing with its greens. WHAT IT MEASURED: all five real catalog upgrades kept the customer's data; and whether an upgrade can be undone is a property of the individual APP, not of upgrades — docmost REFUSES ("corrupted migrations: previously executed migration 20260213T085259-notifications is missing"), privatebin does not, which independently reproduces the Nextcloud finding on a second app by a DIFFERENT mechanism and puts two measurements behind §4's ruling that "rollback" is the wrong word. It also found a defect in our own catalog (R-459). What stays open, as its own rows rather than inside this one: R-459 (the skipped MariaDB datadir upgrade), R-460 (bookstack's file half is unprovable headlessly), R-462 (the widening, costed). Full original text: git show 417df06f3529:documentation/backlog/OPEN-ITEMS.md. | CLOSED 2026-09-06 — harness built, run, and proven by a red negative control | audits/SPIKE-upgrade-test-2026-09-06.md; audits/upgrade-spike-2026-09-06/evidence/; catalog 0474ce387e6f | | R-438 | The catalog sync rewrote a DEPLOYED app's docker-compose.yml and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06). Syncer.copyTemplates copied into every stack folder on a 15-minute cycle with no deployed check, so a deployed app's file and its running containers disagreed from that moment, and the next compose up -d from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s with a pull, and a boot reconciliation upgraded it with nobody pressing anything. Reasoning kept — the distinction this row existed to protect: RestartStack's use of up -d to pick up template changes was CHOSEN and written down in its own comment, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. Reasoning kept — one fear was measured SMALLER than stated: a plain power cut does NOT upgrade anything, because Docker's restart: unless-stopped restores the containers on the old image and the reconciler logs no boot-orphaned apps; the unattended upgrade needs the narrower precondition "and the app did not come back". Full original text: git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md. | CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0 | audits/SPIKE-app-update-2026-09-01.md §2, §3, §8; architecture/09-update-architecture.md; tests/VALIDATION-update-slice3-2026-09-06.md | | R-447 | UPDATE ARC SLICE 3 — the live compose file is now DERIVED; the syncer renders instead of copying. SHIPPED controller v0.235.0 (2026-09-06). The pin lives in app.yaml (pinned_images), the definition it came from is stored beside the app as applied-compose.yml, and Syncer.renderSource writes the catalog template verbatim while the catalog still offers the pinned version and the stored definition once it moves past it. Reasoning kept — the ruling, in the operator's own words: while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update. Reasoning kept — why nothing was added to the thirteen compose up -d call sites: most of them are REPAIRS (the boot reconciler, the drive-return gate, the app-stop guard), and a repair path that refuses to repair leaves a customer's app down, which is worse than the problem; they were made safe by removing the reason, not by gating them. Reasoning kept — why the frozen branch writes a WHOLE file and never a substitution: wger 2.6 needs a full DB configuration the older template cannot supply, so an old image under a new template is a third state nobody chose. Reasoning kept — why this is not "skip deployed apps" (option B, rejected): that also stops health-check fixes, memory limits and new deploy fields, and destroys the self-healing measured live in the spike §3. Reasoning kept — pinned_images is INTENT and installed_images is an OBSERVATION; never feed one from the other (the R-166 category error, one field over). Full original text: git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md. | CLOSED 2026-09-06 — SHIPPED controller v0.235.0 | architecture/09-update-architecture.md §3.4, §5; tests/VALIDATION-update-slice3-2026-09-06.md; controller CHANGELOG.md v0.235.0 | | R-441 | The restore path and the catalog sync disagreed about the image, and the SYNC WON within 15 minutes. CLOSED controller v0.235.0 (2026-09-06). stackAdapter.RecreateStackDefinitionFromUnit wrote the recovery unit's captured docker-compose.yml — carrying the OLD pin — into the live stack dir, and Syncer.copyIfChanged overwrote it from the catalog on the next tick, so a restore's image-level recovery had a <=15-minute half-life. The restore now PINS to what the unit captured and stores it as the applied definition, so the render obeys the restored file instead. Reasoning kept: the overwrite half was MEASURED live 2026-09-01 ([INFO] [sync] Updated bentopdf/docker-compose.yml at 18:10:29Z over a locally-modified file); the "the restore writes to that same path" half was READ, and this closure rests on the render behaviour being measured live rather than on the reading. Full original text: git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md. | CLOSED 2026-09-06 — SHIPPED controller v0.235.0 | internal/stacks/pin.go; TestGroupF_RestorePinIsReportedToTheSyncer; tests/VALIDATION-update-slice3-2026-09-06.md | | R-455 | DooPlex had no Docker Hub login, and the unauthenticated ceiling blocked a BUILD rather than only a gate. CLOSED 2026-09-06 — a Docker Hub PAT was added to the credentials file (operator). Measured 2026-09-02: build.sh 0.233.0 --push failed at 429 Too Many Requests resolving debian:bookworm-slim, with neither base image in the local store. Reasoning kept — the workaround and its verification, because the answer outlives the incident: both bases were pulled from Google's official Docker Hub mirror (mirror.gcr.io/library/...) and retagged, and the identity claim was later MEASURED rather than assumed — docker pull docker.io/library/<img> answered Status: Image is up to date for both, i.e. Hub's own manifest resolved to the images already local, and the docker manifest inspect bodies were identical between the registries. So the mirror is a verified-sound fallback if the PAT is ever unavailable. Full original text: git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md. | CLOSED 2026-09-06 — credential added by the operator | felhom-controller/REPORT.md (2026-09-03) §5 | | R-429 | CORRECTED 2026-09-01 — the snapshots ARE being taken; what was broken was that nobody could tell, and my own probe looked for the wrong name. Viktor read the control panel on 2026-09-01: seven automatic daily snapshots on storage-box-pool-1 (plan BX11), six days old to ~9 h old, filesystem ~2.5 GB, per-snapshot 0–14 MB, Display snapshot directory ON. The mitigation works. MY ERROR, NAMED: the 2026-09-01 spike probed for a directory called .snapshots; the vendor documents the path as /.zfs/snapshot. The probe's controls were sound and its subject was wrong, so not found was true and meant nothing. The task that set the spike asserted the mechanism without citing the vendor documentation, and I did not check it — that is how a correct instrument produced a wrong headline. WHAT REMAINS TRUE, and is the actual finding: the row claiming it had no R-number so nothing could cite it; its "confirm tomorrow" went 36 days unanswered; and the DUE-CHECKS block built for that exact class (R-341) was empty. The finding was never the snapshots. It was that nobody could tell. Re-probed at the documented path — see R-432 for what a sub-account can actually reach. Evidence: audits/SPIKE-r95-offsite-delete-2026-09-01.md §Q1 and this row. | CLOSED 2026-09-01 — mitigation CONFIRMED WORKING; the visibility gap is R-432 | panel read 2026-09-01 (Viktor); re-probe at /.zfs/snapshot in audits/SPIKE-r95-offsite-delete-2026-09-01.md | | R-431 | An unexplained fall in a customer's off-site snapshot count is now noticed within a day — SHIPPED hub v0.111.0. Third signal in OffsiteChecker, beside FILL and STALENESS. It lives on the HUB deliberately: the event being detected is a box deleting its own backups, so a detector on that box is one the same event can silence; the hub already receives the count and keeps the history. Threshold: a fall of more than HALF the previous count and at least 5 — reasoned, not invented, because the measurement gave nothing to calibrate against: over 12 898 reports (2026-06-05 → 2026-09-01) every one of the nine decreases lands exactly on ZERO and every one predates stats_known (the R-331 shape), and in the 380-report window where stats_known is true there are zero decreases. Retention keeps 7 daily + 4 weekly + 6 monthly per group, so it cannot halve a total; a mass deletion goes to ~0. Three pre-conditions, each with its scar: StatsKnown (R-331 — a zero is not a zero when unmeasured), the declared State (R-204 — the box names its own situation), and run success (R-100 — presence is not success; incomplete excluded too). An untrustworthy report neither alarms nor moves the baseline. ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms, fixture committed. Severity error, operator-only, no customer template. | CLOSED 2026-09-01 — SHIPPED hub v0.111.0 | hub/internal/monitor/offsite.go; offsite_r431_test.go incl. 9 009-point real-history replay; hub v0.111.0 | | R-419 | observations_gate.py accepted an observation whose body merely CONTAINED the string NOT-A-FINDING, even in prose disclaiming it. Found by accident on 2026-09-01 when a planted test observation reading "it carries no FILED: and no NOT-A-FINDING: marker" was reported OK 1. NOT-A-FINDING and a real push went green over an unfiled finding. The gate's whole job is to force an explicit choice, and a sentence disclaiming the choice counted as making it. FIXED 2026-09-01: a marker must now start a line or follow a sentence boundary, and inline code spans are stripped before matching — a marker inside backticks is being talked about, never used. | CLOSED — FIXED + PINNED (2026-09-01, R-421 sweep) | verified in BOTH directions: the decoy and a backticked mention are convicted; a real **FILED: R-419** and a real **NOT-A-FINDING: ...** still pass. Decoy kept in felhom.eu/scripts/test_gate_decoys.py | | R-404 | DECISION — should a documents-only push be subject to the golden-currency gate? RULED 2026-09-01: NEITHER option as framed. Block the push that can create the debt; notify the push that cannot. The two options on the table were narrow the gate and leave it and build a waiver, and both were wrong for the same reason: they argued about the GATE, and the gate was never the problem. The DIAGNOSIS, measured from live source, is that the check was aimed at the wrong repository. golden_currency_gate.py never looks at the push at all — it compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing, which is correct for a standing invariant and wrong as a push gate. Meanwhile controller_gates.py had NO golden-currency entry, so the repo where a release happens never checked, and the repo that cannot create the debt enforced it on every push. 18 of the last 24 pushes here touched no code — MEASURED, and the classifier agrees exactly — most of them for a structural reason: the controller's code is in one repo and its register, architecture and status live in this one, so every controller change produces a documents-only push here by construction. Six of those 18 were bake records — the push that PAYS the debt is itself documents-only, so the gate was blocking its own cure. WHY NOT THE WAIVER the gate's own docstring prescribes: that clause was written for a release nobody wants a golden for. The case that actually occurred (R-417) was a release we did want a golden for, on a night the runbook forbade baking. A waiver would have recorded a lie. SHIPPED: scripts/push_scope.py (allow-list; every uncertainty answers code), a fifth exemptible field in repo_gates.py + --scope, a new ADVISORY verdict printed in its own block, the pre-push hook reading git's stdin, the same rule in CI from the push event payload, and felhom-controller/controller/scripts/golden_notice.py — a NON-BLOCKING notice at the moment a release is committed. The gate's own logic, exit codes and wording are byte-identical; only the consequence changed. The exemption is ONE gate wide and TestR3/Scenario C pins it, red-proved by widening it. Proven live on the real hook: docs+debt → ADVISORY, pushed; code+debt → refused; docs+debt+a second gate → refused for that gate alone | CLOSED 2026-09-01 — ruled and shipped | full text and the ruling: git show 1e6c387:documentation/backlog/CLOSED-ITEMS.md; the diagnosis is in CONTEXT.md and scripts/CHANGELOG.md; scripts/push_scope.py + test_repo_gates_scope.py | | R-417 | A drill night that forbids baking a golden made golden_currency_gate.py red, so pushing the drill's own evidence needed --no-verify — the very signal CI e-mails about. Measured 2026-09-01: five consecutive felhom.eu CI runs red (jobs 469/470/471/473/476), all mine, all on step 3 Run the gate entry point; job 478 green the moment the golden-0.232.0 evidence was committed. Cause confirmed by isolation — moving that directory aside reproduces exit=1, restoring it gives exit=0. The gate was right every time: 0.231.0 and 0.232.0 were released with no golden carrying them. CAUSE REMOVED, not worked around (R-404): a drill's pushes are documents-only, so the conviction now prints as a loud ADVISORY and the push proceeds — in the hook AND in CI, so a drill night no longer produces red runs indistinguishable from real ones. The expectation is now written where the next drill author reads it (documentation/runbooks/target-selection.md), which is the half I had left out. | CLOSED 2026-09-01 — by R-404 | five red CI jobs 469/470/471/473/476, green at 478; reproduced by isolation (documentation/audits/AUDIT-gate-decoys-2026-09-01.md records the technique) | | R-361 | The pre-restore safety dump overwrote the app's own DB dump, and the comment beside it said it could not. Shipped in controller v0.221.0 (+v0.221.1). Evidence: audits/DRILL-r361-2026-08-22/evidence/. Reasoning kept: DumpOne writes <stack>-<dbtype>.sql — the app's canonical dump, the name the replay loop matches EXACTLY — so nothing else may ever be written to it. The fix is a DESTINATION, not a rename: DumpOneTo takes the final path and derives its own .tmp from it, so neither the destination nor the scratch file can collide with a nightly dump running beside it. DumpOne's signature did not move — it has callers outside this concern. The manifest no longer lists the undo copies: every consumer of Manifest.DBDumps was grepped and named — three, all inside recovery_unit.go, none reading it for recovery. AND THAT CHANGE MADE ANOTHER UNREACHABLE: a stable db_dumps let CaptureRecoveryUnit's already-current early return fire, and the undo-copy prune sat after it — four copies on disk against a cap of three, counted live. The prune now runs ABOVE the check; it is housekeeping on the dump directory and is independent of whether the manifest needs rewriting. PROVEN LIVE the only way it can be: the canonical dump's sha256, unchanged across a restore — docmost 5d35678349bb…, bookstack 7837aa5de295…, both byte-identical before and after. A test asserting merely that the undo copy exists passes just as well when the app's backup was destroyed. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.221.1, 2026-08-23) | full text: git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md | | R-379 | The pre-restore undo copy was valid, was named to the customer, and no product action could apply it. Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: audits/DRILL-r379-rollback-2026-08-22/evidence/. Reasoning kept: R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken. The undo set is matched on THE RUN'S OWN STAMP, never on the pre-restore- prefix (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) and never just the first file (a two-database app would have had one restored and the other left half-written). The rollback RE-DISCOVERS the container — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of waitDBReady against a dead id while its data was recoverable. No unit test saw that: they all inject the import seam and never look at container identity. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-380 | A failed MariaDB replay left a partially-applied database behind an app reporting health=healthy. Shipped in controller v0.220.0. Evidence: audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt. Reasoning kept: no engine flag closes this — --single-transaction was added to the Postgres import and does make it all-or-nothing, but MariaDB's DDL is not transactional, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: bookstack's migrations table back at 102 rows, the exact cell the defect was measured in. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.220.0, 2026-08-22) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-381 | The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface. Shipped in controller v0.220.0. Reasoning kept: the full engine text now goes to the operator log, which never had it before — the diagnostic was ADDED, not removed. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an INSERT INTO migrations VALUES (…) listing); now 257 bytes with no engine tokens. A red-proof for this PASSED and the test was hollow: it injected below ImportDump, so a leak reintroduced inside ImportDump could not fail it. The guard now sits at that layer. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.220.0, 2026-08-22) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-382 | The reconstitution's summary log line omitted the volume count it already held. Shipped in controller v0.220.0. Proven live: 0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed. | CLOSED — SHIPPED (controller v0.220.0, 2026-08-22) | full text: git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md | | R-356 | The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no. Shipped in controller v0.219.0. Evidence: audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/. Reasoning kept: the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (Manager.GetAppDrivePath, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong. An app with no drive is not misconfigured — 01-topology-and-trust.md §8 carries the [DESIGN] marker; between 19 and 22 August that design was called a defect four times. Deployment is asked of ListDeployedStacks() and FAILS CLOSED on a nil provider: "cannot tell" must not become "go ahead" when the caller's next act is a write. Two different failures get two different sentences — installed-but-no-resolvable-data-root has its own refusal and its own route; widening nincs telepítve to cover it would send a customer to reinstall a running app and hide the real fault. Measured, and load-bearing: 53 templates, 13 needs_hdd: true, 40 false (catalogue @ 459766cb1639). The capture side's raw GetStackHDDPath is FENCED and was not changed — capture resolves an app's declared userdata/import file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.219.0, 2026-08-22; privatebin on demo-hp: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md | | R-216 | A correct recovery code was reported to the customer as wrong. Shipped in 0.120.0, v0.125.0. | SHIPPED (controller v0.201.0 + hub v0.97.0/0.97.1) — but see R-223: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-218 | Succeeding at recovery stopped the box asking for what it still needed. Shipped in v0.203.0. Evidence: documentation/tests/part4-rewalk-2026-08-06/journal.md. | CLOSED 2026-08-06 — controller v0.203.0, proven live. (State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.) The over-claim, as it stood: the fix covered the DECLARATION half only. Measured on the R-201 re-walk: the box declared, and offsiteheal re-staged the secret at 11:44:57 saying "the box re-consumes on its next cycle" — the next cycle came and went (host-report 11:55:46, Received report 11:55:54, a full cycle with a positive control that it ran) and the credential was still not consumed. 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on /backups/remote (config, reset, run, toggle) found none that fetches a staged credential, and the only lever is systemctl restart felhom-controller-bootstrap.service inside the guest — which worked in 18 s (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: the only thing missing is anything at all to trigger a retry. This is the FIRST of the two dead ends that keep the recovery journey failing | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-219 | The listing the screen promises could never render on the shape it exists for. | SHIPPED (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-217 | An unreadable store reported as "opened, with unattributable content". | SHIPPED (controller v0.201.0) — opened / empty / unreadable are three distinguishable states | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-222 | Reaching for a RETAINED earlier package read as a wrong code. | SHIPPED (controller v0.201.0 + hub v0.97.0) — the ACK carries superseded_present/superseded_at and the screen names the situation. It states what the hub knows and promises nothing — the read path is still unbuilt (R-199's inventory) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-215 | GET /recovery rendered the recovery story on a box that never had off-site backups. | SHIPPED (controller v0.201.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-220 | After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action. | CLOSED 2026-08-06 — shipped in agent v0.127.0 and PROVEN LIVE on a genuinely rebuilt box. The fix is corroborated, not a widened prefix: a mountpoint outside /mnt/felhom-drives is forgiven only when the SAME device is also mounted under the managed path — a pairing only Felhom's own enrolment produces, so a disk another system is using at /srv/data or even /mnt/someone-elses-disk is still refused (own test + red-proof). Read from /proc/mounts deliberately: the lsblk invocation is pinned verbatim in the sudoers file, so switching to plural MOUNTPOINTS would have shipped a sudoers change with the binary. Fail-safe: an unreadable mount table corroborates nothing. Measured on the Part 4 venue after a real guest purge, with both raw mounts still present on the surviving host: /disks/candidates returned both drives in attach and initialize (before the fix: two empty lists), and both re-attached through the customer endpoint (registered: true). The customer-facing refusal was corrected in controller v0.203.0. | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-221 | A rebuilt box cannot run the escrow ceremony at all. | CLOSED 2026-08-08 — agent v0.128.0. Apply re-asserts the seed BEFORE the idempotent early return; the return itself is kept and pinned by a zero-Proxmox-calls assertion. The writer was established at file:line rather than assumed — see the follow-through section below | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-223 | The Day-0 manifest vouched agent 0.120.0 while the recovery feature needs 0.125.0 — and a reinstall DOWNGRADES a box that was fixed by hand. Shipped in 0.120.0, 0.125.0, 0.192.0. | CLOSED 2026-08-05. Golden 0.201.0 baked in the drill VM (658 165 766 B, sha e730d7cab343eb35…f007654, round-trip verified from Gitea), then manifest set in one save: agent=0.125.0 golden=0.201.0 min_agent=0.125.0. A fresh install now lands on current agent AND current controller | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-224 | Every non-code failure on the unlock path is reported to the customer as a statement about their code. Reasoning kept: And R-216's gate cannot catch it: the box's own ring reads recovery capability gate: offsite_key_recovery=yes (source=version) — the gate discriminates the agent's age, not its reachability, so a dead agent of the right version sails through the guard whose own comment says "An attemp | CLOSED 2026-08-06 — controller v0.202.0 + agent v0.126.0. The discriminator is now a VALUE: escrow.ErrBundleFetch → HTTP 502 at the agent, agentapi.RecoveryRefusal carrying the status at the controller, and ClassifyRecoveryFailure mapping it to one of five classes from the value, never the text. PROVEN LIVE on the venue, same wrong code, only the hub's reachability changed: hub up → 400 "…did not open the sealed bundle" · hub REJECTed → 502 "…could not be fetched — the recovery code was NOT used" · hub restored → 400. Red-proof: deleting the agent case reproduces got 400, want 502 with the wrong-code sentence. Coupled MinAgent 0.126.0 — an older agent answers 400 for both causes, so the reading is withheld and the 400 degrades to NEUTRAL; the gate blocks nothing. The customer-facing messages were NOT re-driven end-to-end: /recovery correctly redirects since F7 set the old data aside, and restoring that state is the reconfiguration §11 forbids — they are covered by handler tests + red-proofs | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-225 | The remote store reports 0 pillanatkép · 0 / 50 GB when the box cannot read it — directly above a card stating the store holds backups. | CLOSED 2026-08-06 — controller v0.202.0. StatsKnown is a named state (the OffsiteInventory.Empty pattern), because zero is what an unread store and an empty one both look like and omitempty makes "absent" and "0" the same bytes. The fill bar renders only when the fill is known — a 0 %-wide bar is a picture of emptiness. PROVEN LIVE both ways: before a run the venue read „a pillanatképek száma még ismeretlen"; after one, „2 pillanatkép … / 50 GB". A measured zero still says zero | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-226 | M1 — the only message that tells a customer to check their typing — is unreachable on any box that has re-escrowed. | CLOSED 2026-08-06 — controller v0.202.0. The retained-package message now names both possibilities and restores the ten-words prompt, because the two are indistinguishable at the engine and saying so is the honest thing. It still does not promise the earlier package can be opened. Red-proof: removing the clause makes the prompt unreachable again | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-228 | After „I do not want the old data", the set-aside history becomes invisible — the box records where it is and shows it to nobody. | CLOSED 2026-08-06 — controller v0.202.0. OrphanedRenamedTo is surfaced as two facts and stops. It does not promise the history can be reopened — it cannot be, by anyone, today (R-199's inventory is unbuilt) — and the set-aside confirmation copy was corrected for the same reason: "a helyreállítási kód nélkül többé nem lesznek megnyithatók" implied that WITH the code they could be. The field's own comment said "recovery-code-recoverable", the same over-promise in the code. PROVEN LIVE: the notice renders on the venue | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-227 | A controller restart mid-unlock returns a raw English Bad Gateway. | CLOSED 2026-08-06 — controller v0.202.0, partially and stated as such. The layer that answers is traefik, whose config this repo generates — but traefik v3 serves no static files, so a branded proxy page needs a new always-up container for every 502 on the box: scoped, not built. Shipped: the unlock posts via fetch and answers a gateway failure in Hungarian in-page. Progressive enhancement — with no JS the plain POST still shows the proxy's error | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-234 | An off-site run reports success while silently omitting an app the customer just switched on. Shipped in v0.205.0. | CLOSED 2026-08-06 — controller v0.205.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-236 | After a guest rebuild the hub never re-stages the off-site credential. | CLOSED 2026-08-06 — not a defect | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-237 | After a successful recovery the customer is shown no backups at all, because the restore surface is keyed on apps that are currently installed and currently marked for future remote backup. | CLOSED 2026-08-06 — controller v0.204.0: the list is now built from OffsiteInventoryList (the repository's own snapshot tags). Installed-ness became a property OF a row, never a filter; an unreadable store renders as UNKNOWN and keeps the action offered; felhom-offbox and _shares are excluded. 7 new tests incl. a rendered-page test for the rebuilt shape, and a red-proof that keys the list back on installed-and-toggled apps. | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-238 | „Teljes visszaállítás előkészítése" accepts the click and does nothing. Shipped in v0.204.0. | CLOSED 2026-08-06 — controller v0.204.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-239 | The fixes are written, tested, pushed — and a machine installed tonight gets none of them. Shipped in 0.127.0, 0.203.0, 0.204.0. Evidence: tests/finalwalk-r201-2026-08-07/journal.md, tests/golden-0.205.0-2026-08-07/. | CLOSED 2026-08-07 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-241 | The credential self-heal, succeeding, locks the customer out of their own recovery. Shipped in v0.206.0, v0.98.0. Evidence: audits/SPIKE-r241-recovery-offer-2026-08-07.md, tests/finalwalk-r201-2026-08-07/journal.md. | FIXED 2026-08-07 — v0.206.0 / hub v0.98.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-247 | The box is being told something false, in its own words, and it recommends the destructive act. Shipped in v0.206.0. | CLOSED 2026-08-08 — controller v0.209.0. The field is received and the box tells the two conditions apart; see the Campaign-12 follow-through section below. The WRONG FLAG itself is R-246 (operator act, hub-side) and the customer-facing card copy is unchanged — both stated rather than folded in | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-249 | The retrieval passphrase ships in the customer page's HTML, so any headless read puts it in a transcript. Shipped in v0.206.0, v0.207.0. Reasoning kept: Severity MEDIUM: it is a live per-customer secret that fetches the whole config (GET /api/v1/config/<id> with X-Retrieval-Password), but the exposure is to someone who can already read the operator page — a defence-in-depth failure, not a boundary crossed. | CLOSED 2026-08-08 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-252 | After a rebuild the restore refuses because the data drives are not registered, and nothing on the recovery path says so. Shipped in v0.207.0. | CLOSED 2026-08-08 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-253 | The restore page promises it will reinstall the app, and the restore then refuses because the app is not installed — in the customer's own language, three lines apart. Shipped in v0.207.0. | CLOSED 2026-08-08 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-212 | The orphaned-ciphertext deletion HALTED: the stores on the storage box do not match this register's record. | CLOSED 2026-08-05 — all three deleted after the operator confirmed the corrected list | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-94 | A hand-synced version constant drifts, and the gate that would catch it is never run Shipped in 1.22.0, 9.9.9. | CLOSED — SHIPPED (hub v0.87.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-110 | main is the installer's publish channel — there is no staging. Shipped in 1.22.0, v1.23.0, v4.4.0. Reasoning kept: E-2a's felhom-backup-target-apply (:2116) is installed 0755 to /usr/local/sbin and root-fenced in sudoers, validated only by bash -n — a root-executed artifact taken from main with no pinned integrity, which is this row's class exactly. Fixing only (i) leaves a tagged installer pulling nine untagged files from main at run time — a staging story that is false in the place it matters most, since one of those nine (felhom-backup-target-apply) is installed 0755 into /usr/local/sbin and root-fenced in sudoers, validated only b | CLOSED — SHIPPED (installer v1.23.0, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-111 | The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent 0.96.0, not 0.113.0. Shipped in 0.113.0, 0.114.0, 0.161.0. Evidence: audits/E2D-fresh-vm-2026-07-29.md. Reasoning kept: The global controller floor was deliberately NOT raised: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. | SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-115 | Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable. Shipped in 0.113.0, 0.114.0, 0.119.0. Reasoning kept: Class: → R-29, one layer up — a control that exists and is never walked; deliberately NOT given its own ID. It calls the existing publish-agent.sh rather than reimplementing it, refuses a dirty or unpushed tree, refuses to re-release an existing version (one version name must never mean two binaries), and deliberately does not vouch — vouching points machines at a version and stays the operator's ac | CLOSED — SHIPPED (release-agent.sh + check-published-versions.py, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-116 | The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC storage_disconnected, return the SPECIFIC backup_target_restored; backup_target_absent never fired at all Shipped in 0.185.1, v0.115.0, v1.25.0. Evidence: audits/R116-v0116-2026-07-30.md, audits/SPIKE-r117-bind-liveness-2026-07-30.md. | SHIPPED + PROVEN-LIVE (agent v0.116.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-120 | The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message Shipped in 0.113.0, 0.116.0, 0.156.0. Evidence: audits/R120-golden-rebake-2026-07-30.md. | CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate (golden 0.186.0 + hub v0.82.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-117 | A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered. Shipped in 0.113.0, 0.117.0. Evidence: audits/R117-v0117-2026-07-30.md. Reasoning kept: No block I/O proven by strace (only /proc/self/mountinfo, 0 statfs) — the Part 1 CLAUDE.md fence applied to its own first consumer. | SHIPPED + PROVEN-LIVE (agent v0.117.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-113 | The drive-absent gate CANNOT FIRE on device loss — E-2b's alarm is wired to an unreachable condition. Shipped in 0.113.0, 0.114.0, v0.185.0. Evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md. | SHIPPED + PROVEN-LIVE (agent v0.114.0, 2026-07-29) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-112 | E-2's degraded banner and offer have NO UI CONSUMER — the endpoint is correct and the customer never sees it. Shipped in v0.185.1. Evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md. | SHIPPED + PROVEN-LIVE (controller v0.186.0, 2026-07-29) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-114 | On target-drive loss the customer is told the wrong story and offered the drive that just vanished. Shipped in 0.113.0. Evidence: audits/E2D-fresh-vm-2026-07-29.md, audits/SESSION-C-2026-07-29.md. | SHIPPED + PROVEN-LIVE (controller v0.186.0, 2026-07-29) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-29 | The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it. Shipped in 1.19.0, 1.22.0, v0.129.0. | CLOSED — both halves shipped (2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-86 | Restore-tests are interval-scheduled, not backup-aligned | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03 (agent v0.121.0, hub v0.91.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-185 | The agent cannot see the host backup tier's archives on demo-felhom — the PVE token has no ACL on /storage/felhom-backup, so the content listing returns EMPTY where root sees three archives. Shipped in v0.123.0. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03 (agent v0.123.0, installer 1.24.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-186 | A released agent binary's sha256 cannot be reproduced from its tag. Shipped in v0.120.1, v0.121.0, v0.121.2. | CLOSED — SHIPPED + MEASURED 2026-08-03 (agent v0.122.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-187 | R-115's one-command release had never actually run its publish leg — the first real use died there. Shipped in v0.121.0. | CLOSED — SHIPPED 2026-08-03 (felhom-agent) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-188 | Every agent release has a ~50 % chance of emailing the operator a CI failure for a release that is correct. Shipped in 0.121.2, v0.121.0, v0.121.1. | CLOSED — SHIPPED 2026-08-03 (agent v0.122.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-189 | A passing restore-test can be invisible to the hub forever — and R-86 made that window a week instead of a day. Shipped in v0.121.1. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03 (agent v0.122.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-195 | A customer with no machine ever bound e-mailed an expected_dbdump_missed ERROR every morning. Shipped in v0.73.0. Reasoning kept: Fail-open on a read error (an unreadable binding must never SUPPRESS a real alarm), and the deferral is LOGGED with its own counter (the v0.73.0 Part-7 precedent: a quiet check must not look like a check that did not run). | SHIPPED (hub v0.92.0, 2026-08-04) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-196 | escrow_stale is wired to the ONE path that does not change the repo password, and absent from the path that does. Shipped in v0.95.0. Evidence: audits/DRILL-r201-night-run-2026-08-04.md, audits/SPIKE-offsite-credential-recovery-2026-08-04.md. | CLOSED 2026-08-05 — hub v0.95.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-197 | The hub holds both halves of the evidence that a box's offsite DATA key changed, and reads neither. Shipped in v0.78.0, v0.93.0. Evidence: audits/SPIKE-offsite-credential-recovery-2026-08-04.md. Reasoning kept: The in-between shapes (a first-ever hash, a hash-less supersession) are LOGGED rather than dropped, so "we chose not to alarm" and "the check did not run" never look identical. | SHIPPED (hub v0.93.0, 2026-08-04) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-198 | The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy. Shipped in v0.92.0, v0.93.0. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md. | SHIPPED (hub v0.93.0, 2026-08-04) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-199 | The hub serves recovery blobs on two endpoints that have no client anywhere in the system. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md. | SHIPPED + PROVEN-LIVE 2026-08-04 (hub v0.94.0, agent v0.125.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-203 | A customer-declared MANDATORY data directory was silently absent from the off-site snapshot while the run reported ok. Evidence: audits/DRILL-r201-offsite-recovery-2026-08-04.md. | SHIPPED + PROVEN-LIVE 2026-08-04 (controller v0.197.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-204 | A rebuilt box can recover its off-site key and still cannot use it: the remedy that reconfigures the tier is the thing that blocks the recovery. Shipped in v0.198.0, v0.199.0, v0.95.0. Evidence: audits/DRILL-r201-night-run-2026-08-04.md. Reasoning kept: Live on demo-felhom 9201, nothing restarted (restarts=0, container older than both mints): the superseded code returned „Hibás vagy lejárt kód" and the current one was accepted first time. Item 2 (a re-issue marks a healthy escrow stale) — CLOSED, → R-196. Test-proven; deliberately NOT fir What is deliberately NOT automated: the escrow ceremony. A credential is replaceable; the recovery code is not. | ALL FOUR ITEMS CLOSED 2026-08-05 (items 1–3 controller v0.198.0 + hub v0.95.0; item 4 controller v0.199.0 + hub v0.96.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-192 | offsite_delivery_stuck tells the operator the opposite of what the detector measured, and the self-heal silently refuses for exactly the reason the message denies. Shipped in 0.187.0, 0.192.0, v0.199.0. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md, audits/SPIKE-offsite-credential-recovery-2026-08-04.md. | CLOSED 2026-08-05 — the guard's scoping half closed BY REPLACEMENT (hub v0.96.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-193 | A guest rebuild silently drops the off-site app-data tier, and nothing restages the credential. Shipped in 0.156.0, 0.187.0, 0.192.0. Evidence: audits/RECON-offsite-dr-chain-2026-08-04.md, audits/SPIKE-offsite-credential-recovery-2026-08-04.md. | CLOSED 2026-08-05 — controller v0.200.0 (credential half v0.199.0/v0.96.0; the recovery SCREEN v0.200.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-90 | ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged | CLOSED — the operator rescaled ep0 to a CX33 on 2026-08-03 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-97 | Whole-guest backup tier had no hub signal; quiesce blamed the apps Shipped in v0.79.0. | SHIPPED (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-100 | ~~A restic offsite tier that fails every night never goes stale on the hub — isStale counted from LastRun Reasoning kept: The real defect is defeated defence in depth: the hub-side pull net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost push (cf. | SHIPPED + PROVEN-LIVE (controller v0.181.0 + hub v0.80.0, 2026-07-28) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-101 | ~~Tier-2 LastRun is written on failure and rendered to the customer as „Legutóbbi másolat" — including in t | SHIPPED + PROVEN-LIVE (controller v0.182.0, 2026-07-28) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-108 | Network storage can host an app's namespace, and FileBrowser binds a network share at its ROOT Evidence: audits/R108-network-app-namespace-2026-07-30.md. | SHIPPED + PROVEN-LIVE (controller v0.187.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-109 | The DR recipe records no backup target Evidence: audits/R106-R109-recipe-completeness-2026-07-30.md. | SHIPPED + PROVEN-LIVE (agent v0.118.1 + hub v0.83.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-106 | The DR recipe records the PBS namespace as "root" on every box | SHIPPED + PROVEN-LIVE (agent v0.118.1, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-122 | ~~AssembleDRRecipe silently DROPPED offsite_restic — the offsite recovery location never reached any reci | SHIPPED (hub v0.83.0, 2026-07-30) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-125 | A "test through the production path" is only true up to the seam it injects at. Shipped in v0.118.0. Evidence: audits/R106-R109-recipe-completeness-2026-07-30.md. | FIXED (agent v0.118.1) — filed for the DOCTRINE point | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-128 | ~~build-felhom-iso.sh:44 comments that ISO_VERSION "aligns with felhom-host-install SCRIPT_VERSION" — a c | CLOSED (iso v1.26.0, 2026-07-31) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-154 | [first-boot] is automated-install-only and nothing in the Felhom tree said so Evidence: audits/SPIKE-universal-iso-3-2026-07-31.md. | CLOSED (iso v1.26.0, 2026-07-31) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-155 | iso-repack.sh refuses any ISO without auto-installer-mode.toml, blocking the no-answer.toml posture | CLOSED (iso v1.26.0, 2026-07-31) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-156 | An app's data is neither persisted nor backed up, and it reports healthy. Shipped in 26.6.1. Reasoning kept: Provenance, stated because it decides the row: the observation is docker ps -a on demo-hp's guest 9201 returning empty, supplied with the 2026-08-02 task; this session did not re-measure (documentation-only, every box fenced). | CLOSED — all three apps fixed (papra template, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-157 | bootrecon's start-ONCE sweep misses the boot orphan it exists to recover — TWO mechanisms. | CLOSED — SHIPPED + PROVEN-LIVE (B: controller v0.189.0; A: v0.190.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-170 | The drive-backed boot gate infers a customer's Stop from a container count. Shipped in v0.190.0. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.190.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-171 | The boot sweep started apps whose data drive was ABSENT — a regression introduced by v0.189.0, now FIXED. Shipped in v0.189.0. Evidence: audits/DIAG-bootrecon-drive-absent-2026-08-02.md. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.190.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-172 | A false host_stale alarm fires when the hub's SQLite refuses two consecutive host reports. Reasoning kept: Retry options (b) and (c) were deliberately NOT taken — with readers no longer blocking writers a surviving SQLITE_BUSY would be a real signal, and a retry would hide it; revisit only on evidence. | CLOSED — SHIPPED + PROVEN-LIVE (hub v0.88.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-174 | The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code. Shipped in v0.189.0. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.191.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-175 | 07-backup-architecture.md §7.5 states ONE box's size bound as if it were the fleet's. Shipped in 0.192.0, 7.5.1. Evidence: audits/SPIKE-r165-mp1-merge-2026-08-02.md. | CLOSED — FIXED 2026-08-03 (same pass as R-165) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-183 | A fresh install fetched the vouched agent BINARY and its sixteen CONFIG files from two different refs, and nothing compared them. Shipped in v0.120.0. Reasoning kept: Why it is a defect and not only untidiness: these files are the agent's own operating surface — its systemd unit, its sudoers, its guarded wrappers — and configs/felhom-backup-target-apply is installed 0755 into /usr/local/sbin and root-fenced in sudoers, validated only by bash -n. | CLOSED — SHIPPED (installer v1.23.0, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-182 | A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour. Shipped in v0.194.0, v0.90.0, v0.90.1. | CLOSED — SHIPPED (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-181 | The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide. Shipped in v0.192.0, v0.193.0, v0.193.1. | CLOSED — SHIPPED (controller v0.193.0 + v0.193.1, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-178 | The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED. Shipped in 0.119.0, 0.120.0, 0.192.0. | CLOSED — BOTH BOXES REINSTALLED AND PROVEN (2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-158 | A local Tier-1 app-data backup failure reaches no hub channel. Shipped in v0.78.0. | CLOSED BY R-167 — SHIPPED + PROVEN-LIVE (controller v0.191.0 + hub v0.89.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-159 | wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy. | SHIPPED (templates/wishlist/docker-compose.yml, 2026-08-02) — filed to record the CLASS | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-160 | gramps-web persisted three paths and wrote to none of them. | SHIPPED (templates/gramps-web/docker-compose.yml, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-163 | mp1 is RETENTION, not staging — and it is sized as if it were neither. Shipped in v0.192.0. | CLOSED by R-165 — the ceiling it describes no longer exists (golden v3.0.0, 2026-08-03) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-165 | Merge mp1 into mp0 — the dedicated 20 G backup partition stops existing. Shipped in 0.192.0, v0.192.0. Evidence: audits/SPIKE-r165-phase0-2026-08-03.md. | SHIPPED — golden build-golden.sh v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. IMPLEMENTED — the LAYOUT is proven live on both boxes (R-178, 2026-08-03); the BULKHEAD'S REPLACEMENT IS NOT (→ R-181) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-166 | App state gets a desired/observed model with its own store. Shipped in v0.189.0. | SHIPPED + PROVEN-LIVE (controller v0.189.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-167 | Storage monitoring and backup alerts. Shipped in v0.191.1, v0.191.2. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-168 | CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK Shipped in 0.1.0. Evidence: audits/SPIKE-ci-runner-2026-08-02.md. | SHIPPED — and the alarm is DEMONSTRATED (2026-08-02) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-205 | RootFsPressureDespiteHousekeeping can never fire Evidence: audits/SPIKE-dooplex-buildcache-2026-08-05.md. | CLOSED — SHIPPED + RED-PROVEN LIVE (homelab-manifests 6808a4b, 2026-08-05) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-258 | C3 — the customer's per-app backup tick is green on the PRESENCE of a restore point, and its only red condition is a GLOBAL one. | CLOSED 2026-08-08 — controller v0.210.0. appDumpVerdict reads THIS app's own dump result; three states, no icon when nothing is known. Recency deliberately not added — see the observation in the follow-through section | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-259 | C4 — a disk read that FAILS renders as „0.0 GB / 0.0 GB (0%)" in the nominal colour, on the dashboard's most-looked-at meter. | CLOSED 2026-08-08 — controller v0.210.0. readDiskUsage reports success; SystemInfo.DiskKnown/HDDKnown; the template draws no figure, no percentage and no meter fill when unknown. The hub leg is deliberately NOT fixed and is now R-266 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-260 | C5 — the agent reports at least eight decision-bearing facts the hub models NOWHERE, and the sharpest one blinds the check that answers „can the operator get into this box". | CLOSED 2026-08-08 — the class is GATED (G-1, scripts/wire_contract_gate.py) and the sharpest instance is fixed (hub v0.99.0). The remaining unconsumed facts are R-264, OPEN — allowlisted with reasons, which is not the same as decided. See the follow-through section below | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-265 | A CI run can fail with NO LOG PERSISTED, and the alarm mail then points the operator at a log that does not exist. | CLOSED 2026-08-08 — timeout-minutes: 5 on the gates job, and the alarm mail now states elapsed seconds and qualifies its own "names itself in the run log" sentence. ⚠ The unknown is NOT closed and must not be read as closed: whether the if: failure() alarm fires for a REAPED job is still unverified. The timeout makes the reap unreachable in practice; it does not answer what happens inside one | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-267 | The Configuration page is 2.6× faster and is still ~10 s, and the remaining cost is ONE Gitea call whose latency swings 20× with load. Shipped in 0.100.2, v0.100.0. | CLOSED 2026-08-08 — hub v0.101.0 + a registry prune. Final: cold 5.4 s, warm 0.14 s (was 26.2 s). Three serialisation legs took it to 9.85 s mean, the 60 s in-memory memo took the warm path to a quarter-second, and the prune halved what remains of the cold path. ⚠ TWO CORRECTIONS TO THIS ROW'S OWN EARLIER TEXT, because both were wrong and both mattered. (1) "Only 50 generic versions exist" WAS NOT A COUNT, IT WAS A PAGE LIMIT. ?type=generic&limit=1000 returns at most 50; the 50 I measured was exactly the cap, and three older agent versions (0.81.0, 0.80.0, 0.79.0) only became visible after the first 30 deletions moved them onto page one. An unpaginated listing is not evidence of a total — this repo's own "an empty listing is not evidence of emptiness" rule, walked into while measuring. (2) THE OPERATOR'S "REDUCE THE NUMBER OF ARTIFACTS" WAS THE BETTER CALL AND MY MEASUREMENT SAID OTHERWISE. I reported it helps "sub-linearly" and "is not the lever". Measured after: trimming to 10+10 took the COLD load from 13.4 s to 5.4 s — a 2.5× improvement on the path the memo cannot help, because the fan-out is per-version. Recorded rather than quietly dropped (the R-96 standing rule). Pruned to the newest 10 per package on the operator's rule, with the live-vouched golden/agent/floor asserted into the KEEP set before a single DELETE was issued; 33 deletions, all HTTP 204, and golden 0.210.0 / agent 0.128.0 / agent 0.127.0 verified still fetchable afterwards. drill-r50 runs agent 0.113.0, now deleted — flagged to the operator first; it is a disposable nested drill VM and only its re-download path is gone | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-268 | A live per-guest local-API token was printed into a session transcript. Shipped in 169.254.253. | CLOSED — ROTATED + PROVEN LIVE 2026-08-09 (rehearsal pre-phase, audits/REHEARSAL-byo-reinstall-2026-08-09.md §3) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-273 | RANK 1 — the hub vouched an agent version that was never git-tagged, and every install fleet-wide now fails at step 5/8. Shipped in 0.127.0, 0.128.0, v0.127.0. | CLOSED 2026-08-09 — tag pushed, install PROVEN | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-278 | demo-felhom's off-site tier has never completed a run and has been stuck for six days. Shipped in 0.200.0, v0.93.0. | CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-280 | RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks". | CLOSED — controller v0.211.0, delivered via golden 0.211.0 (vouched 2026-08-10) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-281 | The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal. | WITHDRAWN 2026-08-09 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-293 | CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population. | CLOSED-INFORMATIONAL 2026-08-10 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-294 | The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented. Evidence: documentation/design/SPEC-orphan-card-copy-2026-08-10.md. | CLOSED — controller v0.211.0; see R-299 for the sentence it missed | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-296 | The orphan card's OTHER sentence makes the same promise, and the spec says it is fine. | CLOSED — shipped in controller v0.212.0 (R-299); verified: the sentence at backups_remote.html:98 was replaced and the stem guard covers it | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-297 | An install took whatever golden was lying around. Shipped in 0.153.0, 0.210.0, 0.213.0. | CLOSED — observed live + PUBLISHED as installer-v1.27.0 (both refs bumped) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-299 | The orphan card's OTHER sentence made the same unevaluable promise, and the spec called it accurate. Shipped in v0.211.0. | CLOSED — controller v0.212.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-300 | Our own uninstall left the thing that makes our own reinstall refuse. Shipped in 0.0.0, 10.0.2, 127.0.0. | CLOSED — observed live + PUBLISHED as installer-v1.27.0 (both refs bumped) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-301 | The abandon countdown banner makes the retired promise a third time, and as a flat statement. Reasoning kept: the customer chose to abandon a recovery offer that exists — which is why it was NOT changed (this session was fenced to the orphan card). | CLOSED — premise CONFIRMED and fixed in controller v0.213.0 (R-302) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-302 | The abandon banner promised retrieval it could not see was still true — fixed by PINNING a fingerprint at the decision. Shipped in v0.213.0. Reasoning kept: Empty is not a match on either side; a countdown started before v0.213.0 carries no pin and takes the cautious branch (deliberately NOT backfilled). | CLOSED — controller v0.213.0 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-305 | The R-300 cleanup fires exactly once per machine, and the second reinstall hits the original wall. Shipped in 0.0.0, v1.27.0. | CLOSED — superseded by R-316 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-307 | demo-felhom carries a LIVE abandon countdown that this drill did not start — and the end state says there should be none. Reasoning kept: The drill's fence forbade starting, shortening or triggering a countdown, and none was; but its required end state was "no abandon countdown anywhere", and one exists. | CLOSED — countdown cancelled 2026-08-12 on the operator's ruling | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-308 | The stored controller password no longer opens demo-felhom — WITHDRAWN 2026-08-12, this was MY BUG, not a defect. | WITHDRAWN — not a defect (my error) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-309 | The day-0 runbook says pushing the installer publishes it. It has not since R-110. Shipped in 1.25.0, 1.27.0. Evidence: documentation/runbooks/day0-install.md. | grep -m1 '^SCRIPT_VERSION'. **Measured while writing it: served 1.28.0, main 1.28.0, both pins installer-v1.28.0— the three agreeing is the observation; any one alone is not.** **The claim was copied elsewhere and the copy was hunted:**audits/SPIKE-universal-iso-3-2026-07-31.md:184said the same thing and **citedday0-install.mdas its source**, which is how it spread. It was **true on the day it was written** (R-110 shipped 2026-08-03), so the dated finding is kept verbatim and carries a SUPERSEDED note rather than being rewritten — falsifying a dated record to tidy it is its own defect. Two other hits are correct in context:hostinstall_gates.py:198states the consequence of the manifest LOSING its tag, and the 2026-08-12 drill record already names the sentence as false | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-311** | **A correct recovery code for a retained package stopped being reported as wrong.** Shipped in 0.126.0, 0.128.0, 0.129.0. | **CLOSED — shipped + delivered: hub v0.103.0 + agent v0.129.0 + controller v0.214.0** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-316** | **The removal now genuinely reverses the installation — R-305's once-per-machine defect closed.** Shipped in 0.0.0, v1.27.0, v1.28.0. | **CLOSED — shipped + published, observed on the cycle that actually fails** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-318** | **No honest marker exists that says Felhom installed dnsmasq on a machine already in the field, and none can be invented.** Shipped in v1.27.0. **Reasoning kept:**/var/log/dpkg.logdoes record the install — and is a **timestamp**, which the standing rule refuses as a heuristic dressed as a fact. | **CLOSED — established, no action possible for existing boxes** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-319** | **The guest-network watchdog finally has a reader — the first of R-264's twenty-one, and it is the repair COUNT that matters, not the state.** Shipped in 0.92.0, v0.92.0. | **CLOSED — shipped hub-side 2026-08-13** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-320** | **Evidence has been destroyed twice in three days, in the same place, by the same act.** Shipped in v1.28.0. Evidence:audits/DRILL-retained-key-2026-08-12.md, audits/REPORT-r316-installer-v1.28.0-2026-08-13.md. **Reasoning kept:** **The rule, now standing rule 5 in workspace-CLAUDE.md(so it loads in every session) and repeated where a session actually meets it —runbooks/target-selection.md, RUNBOOK-rehearsal-v3.md, and the PROMPT-TEMPLATE.mdreport section: evidence is copied off the machine at the end of the phase | **CLOSED — rule written, four homes** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-321** | **A box on which reporting is deliberately switched off still alarms as stale, then down.** Shipped in v0.105.0. | **CLOSED — shipped hub v0.105.0, both doors** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-322** | **The claim guard has never scanned the hub, and the hub sends the customer's first sentence.** **Reasoning kept:** **Recommended shape, and the reason it is not one line:** the gate is invoked bycontroller_gates.py, so pointing it at a sibling repo makes a controller gate fail on a felhom.eu edit — the cross-repo lesson from G-1 (a gate needing a sibling passes locally and exits INCONCLUSIVE in CI, and must n | **CLOSED 2026-08-13 by R-324** — scripts/hub_copy_gate.py, registered in repo_gates.py, scanning 95 hub files for retired names and four declared customer surfaces for retrieval stems, with a plant→convict→remove→pass selftest that caught a defect in its own instrument on the first run. The stem list IS shared (scripts/customer_copy_vocab.py) and no controller gate was made to depend on a felhom.eu clone; the controller gate's adoption of the shared list is R-325, and until it happens the two are drift-checked rather than left to diverge | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-323** | **The third near-homograph — the five-word phrase is „Tulajdonosi jelmondat” now.** | **CLOSED — shipped hub v0.105.0** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-324** | **The hub's customer copy is under a guard for the first time — and the guard has been watched catching, ignoring and releasing.** | **CLOSED — shipped, selftest green** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-326** | **"Which claims are unproven?" is a question a machine can answer now — and the number everyone was repeating answered a different question.** | **CLOSED — shipped** | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-328** | **The disk alert was emailed to nobody, and one word is the whole reason.** Evidence:audits/DIAG-smart-passed-trap-2026-08-14.md. | **CLOSED — controller v0.215.0, PROVEN LIVE 2026-08-14.** Now "warning", and DiskAlertKind.Severity()is exported so the contract is assertable from any package rather than duplicated as a literal. **The proof is a side-by-side pair pushed through the REAL hub event endpoint** from demo-hp's controller: severity"warning"→ storedwarning, notification_log**id 689, channeloperator, status sent**; the identical push at "warn" → stored **info**, and **no notification_logrow exists at all**. Pinned byTestNotifyDiskHealthDegraded_SeverityRoutes, which asserts membership of the hub's accepted set (not just the literal) and names both hub locations; its red-proof — restoring "warn"— fails all three assertions | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-334** | **CLOSED 2026-08-18 — golden 0.216.0 baked, published and VOUCHED; CI green by run id.** Shipped in 0.214.0, 0.215.0, 0.216.0. Evidence:documentation/tests/golden-, documentation/tests/golden-0.214.0-2026-08-12. | **CLOSED 2026-08-18.** Baked from RUNBOOK-manual-build.md §4.0+§4.1 in the DooPlex drill VM and published: **GOLDEN_VERSION=0.216.0**, **GOLDEN_SHA256=ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b**, 656,970,239 bytes at …/generic/felhom-golden/0.216.0/golden.tar.zst. **The published bytes were verified, not just the script's print** — the artifact was downloaded back out of Gitea and hashed, and it matches. **Vouched by the operator, all THREE fields together**, confirmed by reading the hub's own store rather than the save: artifact_golden_version=0.216.0, artifact_agent_version=0.129.0, artifact_min_agent=0.129.0(2026-08-18 11:00:59–11:01:00), and the hub's recorded sha256 matches the downloaded artifact. The R-216 shape was checked on the machine:MinAgent 0.129.0 is **equal to**, not above, the newest **published** agent. **golden_currency_gate.pyrc=0 andrepo_gates.py --fastrc=0 — all nine gates — and CI is GREEN BY RUN ID: run **353**,head_sha 7d81681d6, conclusion success** (the two prior runs 351/352 on this same afternoon were red on exactly this row, which is the contrast). That push needed **no --no-verify** — the first of the day that did not. Evidence: documentation/tests/golden-0.216.0-2026-08-18/, report REPORT-golden-0.216.0.md. **Closed with the run id quoted deliberately**: this row was re-confirmed once and widened once, and closing it on a local green a third time would have left the same ambiguity | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Shipped in v0.215.0. **Reasoning kept:** **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** — — CC | **CLOSED — controller v0.216.0, 2026-08-14.** EachdiskKeyis evaluated once per run; both entries stay markedseenso neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned byTestDiskCheck_SameDiskTwiceIsEvaluatedOnce; companion red-proof run and reverted — deleting the guard makes the first sighting emit Kind:2(Hiba-from-sectors) at 8 sectors | full text:git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | **R-344** | **felhom-agentleaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Shipped in 0.129.0, 0.130.0. Evidence:audits/SPIKE-ep0-established-connections-2026-08-20.md. **Reasoning kept:** **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that e **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, pro | **CLOSED 2026-08-20 — fixed, proven live on both boxes, published and vouched** | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md| | **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Shipped in 0.129.0, 0.130.0, 0.216.0. Evidence:documentation/runbooks/publish-train-rules.md. | **CLOSED 2026-08-20 — published, vouched, and the fleet reconciled onto the published bytes** | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** | \.NamespaceRoot\b' --include=*.go found no non-test reader anywhere — the reconstitution opened the manifest (offbox_reconstitute.go:235) purely for the coherence stamp and resolved its destination from the LIVE app instead. A restore into a destination different from the recorded one therefore succeeded silently, under a green message. (b) The second press. All seven restore handlers gated on backupMgr.IsRunning() — the CONCURRENCY flag, acquired inside the goroutine (offbox_reconstitute.go:180) after the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and overwrote the first restore's op/stack. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. (c) The banner gated its terminal result on a page-local sawRunning, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-354 | The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed. Shipped in 0.217.0, 0.218.0. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22 (controller v0.218.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-355 | paperless-ngx's PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database. Shipped in 0.217.0, 0.218.0. | CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22 (controller v0.218.0) | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-339 | The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it. Reasoning kept: That is correct for a fill signal — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were indistinguishable on the operator channel. | SHIPPED — hub v0.106.0, 2026-08-18. Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, pbsdr_box_unreachable / offsite_box_unreachable (severity warning) past a default 3 windows (≈30–45 min), with paired *_recovered all-clears wired into recoveredPairedDownTypes — necessary because both recoveries are severity info and severityNotifies drops info. Threshold tunable via alerting.box_unreachable_windows. The fill logic is untouched: no threshold, throttle, band or escalate-once behaviour changed. Evidence: internal/monitor/box_reachability_test.go (Scenarios A–F) + internal/notify/dispatcher_box_reachability_test.go (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-370 | PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never documentation/architecture/. Evidence: documentation/architecture/. Reasoning kept: R-352 (re-framed), R-369 The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | CLOSED — corrected 2026-08-22 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-96 | Two standing rules were agreed in chat and never committed Evidence: documentation/runbooks/workspace-CLAUDE.md:48-70. Reasoning kept: Two standing rules were agreed in chat and never committed MIGRATED FROM ROADMAP.md 2026-08-22 (R-369) — originally filed 2026-07-27, size XS, roadmap state idea — found 2026-07-27. Moved verbatim; nothing added or reinterpreted. | CLOSED — migrated from ROADMAP 2026-08-22 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-107 | No offsite action unpacks the named-volume tars Tier-3 captures on every run. Shipped in v0.218.0. | CLOSED — migrated from ROADMAP 2026-08-22 | full text: git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md | | R-383 | The double-failure message told the customer their previous state was saved, and named a file that was not there. Shipped in controller v0.222.0. Evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/. Reasoning kept: One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed. Do NOT simply drop the filename: an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. A zero-length dump counts as MISSING, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is os.Stat and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | CLOSED — SHIPPED (controller v0.222.0, 2026-08-23; undoCopyPhrase, four cases, plus an AST seam test that the message is still wired to the builder) | full text: git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md | | R-384 | An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first. Shipped in controller v0.222.0. Evidence: audits/DRILL-r384-dead-db-alarm-2026-08-23/. Reasoning kept: The defect was the ORDER of two questions, not the unhealthy exclusion. "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end unhealthy, so the symptom the fault causes was what suppressed the alarm for it. IsDownState is byte-identical and unhealthy stays excluded — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; no new state was minted, StateDegraded already means this. Two things had to move and either alone leaves the defect standing: the hoist, AND widening "some members are up" from running > 0 to any member not in the down bucket — the old guard made the R-51 block unreachable in precisely the case it was written for. The register's own suggested fix was WRONG and is recorded as such: it proposed a sustained-unhealthy threshold on the crashLoopAfter model; the actual defect needed no threshold at all. PROVEN LIVE the only way it can be — the same fixture that printed 0 currently down on 2026-08-22 printed 1 currently down on 2026-08-23, with app_start_failed 7 s after the stop and the banner reading „…nem fut: BookStack (degraded)". Scenario D measured 0 alarms across 9 scans through a full stop→start cycle. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.222.0, 2026-08-23) | full text: git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md | | R-329 | app_start_failed was emitted with severity "warn", so every one of them was delivered to nobody. Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: audits/DRILL-r329-r386-2026-08-23/. Reasoning kept: The vocabulary is EXACT and it is the HUB's, not ours — {info, warning, error, critical}; anything else is coerced to info at ingest and dropped by severityNotifies before BOTH legs. This was the SECOND occurrence (DiskAlertKind.Severity until v0.215.0), and its comment had recorded the lesson — a comment is not a guard, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and an unlisted limit is not a limit, it is a hole. The register's own framing was that the DECISION was the work — should a stopped app mail the customer at all? Answered: operator always, customer OFF by default, because processOperator never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. Deliberately NOT added to operatorOnlyEvents — that would make the new toggle visible, flickable and structurally incapable of delivering. Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, warning/sent, after it. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md | | R-386 | A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact. Shipped in controller v0.223.0. Evidence: audits/DRILL-r329-r386-2026-08-23/. Reasoning kept: the state test was guessing at something the product already knows. DesiredState records the customer's intent, has exactly one writer, and is tri-state; StateExited never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: Stopped → no alarm, Running → alarm, absent → UNKNOWN, keep today's behaviour AND announce it. Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about. A rule without a mechanism is a wish: every such suppression sets IntentUnknown and the names are logged at INFO, so an operator can answer "how many apps am I blind to?". failedRestart must still lift a Stopped intent or F-CRIT-1 re-opens. Fenced act: adding a DesiredState WRITER — twelve of StopStack's fourteen callers are machines. Proven live: alarm 24 s after an out-of-band docker compose stop, heartbeat 1 currently down against the previous day's 0; and with intent removed, suppressed plus the log line naming the app. 0 of 8 deployed apps on demo-hp carry an absent intent. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.223.0, 2026-08-23) | full text: git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md | | R-389 | Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app. Shipped in hub v0.108.0. Evidence: audits/DRILL-cooldown-grain-2026-08-23/. Reasoning kept: the fix is a THIRD SIBLING of cooldownTierSuffix/cooldownRunSuffix, separate for the reason the second one's docstring already gives — the existing two keep byte-identical semantics for every type that uses them. cooldownStackSuffix takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property: tier and run_id appear only on types that want that grain, stack_name does not. perAppCooldownEvents is a named allow-list with app_start_failed and nothing else — the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app, and this is not hypothetical: crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a different struct, so a payload-shape rule would have split it silently. The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown. app_start_failed qualifies precisely because it has NO digest — there is no apps_down_run the way backup_run_failures summarises a run. The hour is unchanged; the grain was the complaint. Fail-soft: absent or malformed details degrade to the old key and the mail still goes. PROVEN LIVE 2026-08-23: two apps four minutes apart gave 2 sent / 0 suppressed where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (…:opengist, …:calibre-web) against the previous day's shared key=demo-hp:app_start_failed; and crossdrive_failed for two different apps stayed coarse under key=demo-hp:crossdrive_failed, byte-identical to the derived v0.107.0 value. AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | CLOSED — SHIPPED + PROVEN-LIVE (hub v0.108.0, 2026-08-23) | full text: git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md |

2026-08-30 — the off-site store gets checked (controller v0.227.0/v0.227.1)

Two rows closed, one CORRECTED and deliberately left open. Full original text: git show <this commit> -- documentation/backlog/OPEN-ITEMS.md.

ID Title Shipped Evidence
R-359 The off-site restic store was never verified by anything, ever controller v0.227.0/v0.227.1 documentation/tests/r359-integrity-2026-08-30/
R-397 NotifyIntegrityOK/NotifyIntegrityFailed had no caller, and the product advertised a weekly check that did not exist controller v0.227.0 documentation/tests/r359-integrity-2026-08-30/

R-398 is deliberately NOT a row here. It was proposed for closure in the same pass and was CORRECTED instead — the premise was wrong, the seam already existed — so it stays in OPEN-ITEMS.md as the record. It is written as prose rather than a table row because a register row in both files is exactly what closed_register_gate.py convicts on.

The rules these leave behind:

  • The integrity check TAKES the single-writer flag and SKIPS rather than waits. resticStep escalates to unlock --remove-all on a lock error and is only safe while every caller holds that mutex; a check without it can strip a LIVE prune's lock. Never remove that guard.
  • Due-ness, not a weekday. R-341 is the other shape: a dated check quietly missed and never caught up. No Weekly primitive was added; a daily job that asks "is it due?" catches up after downtime.
  • "I could not look" is not "I looked and it is broken". Skipped / unreachable / failed are three facts. A timeout is unreachable, never damage. A failure advances due-ness; a skip does not.
  • A success that mails nobody is a design choice, not a gap. backup_integrity_ok is severity info and is dropped before both delivery legs. A weekly success e-mail is how alerts stop being read.
  • ⚠ THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION. Measured: a pack corrupted without a size change returned no errors were found, exit 0. Only --read-data* caught it. An ok at the shipped depth means the index and the snapshot graph are sound — narrower than the word suggests. That is R-399, open.
  • A damage classifier must match PHRASES, not words. "pack ", "tree " and "snapshot " all appear in restic's ORDINARY progress output; the first draft would have called a healthy run corrupt. The negative control caught it — which is why a control that has only seen the failing case is worth nothing.
  • The restic exec seam has always existed (SetOffboxRunner). R-398 said otherwise and was wrong.

2026-08-30 — the restore tells the truth (controller v0.226.0) + R-395

Six rows closed. Full original text: git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md. Compressed here to title, shipping version, evidence, and the sentences that state a RULE.

ID Title Shipped Evidence
R-353 A local unit restore reported a bare completion whether it returned an entire dataset or nothing controller v0.226.0 documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt — live sentence A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.
R-357 The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths controller v0.226.0 seam tests only — NOT live-validated, by design
R-358 OffboxFullScratchReady asked "non-empty directory", which is what a failed restic run leaves controller v0.226.0 documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt — marker {"schema":1,…,"full":false}, gate logged place-to-live closed
R-360 The verification-copy delete gated on the concurrency flag, which a verification restore never holds controller v0.226.0 documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt — refused in the live flag state; planted canary survived
R-396 A unit-only verification restore unlocked the DESTRUCTIVE full restore controller v0.226.0 same evidence; found while answering R-358's open question
R-395 STATUS.md contradicted itself about the controller version doc fix, same session STATUS.md at e027b5d9

The rules these leave behind — the reason the rows are kept rather than deleted:

  • A restore outcome is a claim about THE BACKUP, never about the app. R-355 extended to the Tier-1 path. The off-site twin has SafetyDump as an honest discriminator; the local path has none, so no claim about the app is available to it at all. Not merely unproven — unprovable from a manifest: §6.3 records that an absent dump has causes that say nothing about the app.
  • Zero-replayed has two causes and they are opposite news. "The backup held no data" and "the backup listed data that did not come back" must never share a sentence.
  • The destructive reconstitute uses NO headroom margin, matching PlaceOffsiteRestore. The ×1.1 elsewhere exists because that gate PREDICTS a download; this one measures a tree that already exists.
  • Fail closed when a probe reads ≤ 0. free < need with need == 0 is FALSE, so an unmeasurable input sails through — a gate present and inert, which is worse than no gate because it reads as protection.
  • A hidden button is not a guard. Template enable-flags control a button; the handler must refuse.
  • One boolean must not drive three intents (R-396): ScratchReady answered "is there a scratch" while being consumed as "may we place" and "may we destructively restore".
  • Never restate a version in a second place on the same page (R-395). Live versions belong in the hub, never in a doc.
  • A doc comment claiming a guard exists is why nobody looks for the missing guard (R-360). Correct such a sentence in place; do not delete it. | R-399 | How deep should the off-site integrity check go — Viktor ruled full depth. Shipped in controller v0.228.0, 2026-08-31. Evidence: felhom-controller/REPORT.md (v0.228.0) — restic argv observed from the guest at both depths on demo-hp. Reasoning kept: the structure check does not detect a size-preserving pack corruption — measured 2026-08-30, plain restic check reported no errors were found and exited 0 over a damaged pack that every read-data form caught. That is the reason for the default and it is what should stop anyone turning it back down to save four seconds. An empty value means "not configured", therefore the default; off is the off token, because a setting with no off switch is not a setting. A malformed value falls back to the DEFAULT, never to structure — falling back to structure would silently remove the protection on a typo, which is R-357's shape. Superseded by R-401 for anything about a large store. Original text: git show 300d7e8:documentation/backlog/OPEN-ITEMS.md | CLOSED — SHIPPED (controller v0.228.0, 2026-08-30) | controller v0.228.0; documentation/tests/r359-integrity-2026-08-30/ | | R-400 | A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD. Shipped in controller v0.228.0, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. backup/crossdrive implemented (proven live: real Tier-2 copies for three apps); backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status and both storage/simulate-* deleted with their panels and JavaScript. Reasoning kept: implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push. Keep handleDebugAPI's exact-match switch with its NotFound default; a prefix match would have made the defect invisible instead of merely silent. A panel left behind renders nothing forever, which is how this class hides. A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded. Enforced by controller/scripts/debug_route_gate.py, both directions, red-proofed. Original text: git show 300d7e8:documentation/backlog/OPEN-ITEMS.md | CLOSED — SHIPPED (controller v0.228.0, 2026-08-30) | controller v0.228.0; controller/scripts/debug_route_gate.py + its decoy in test_gate_decoys.py | | R-102 (was C9-F4) | Tier-2 wrote a full recovery-unit/ mirror on every run and no code path read it - RecoveryUnitPath joined a hard-coded backups/primary/, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller v0.229.0: four unit-directory-relative path primitives in appbackup, RestoreFromRecoveryUnitAt(stack, unitDir), RestoreTier2Unit. Evidence: audits/DRILL-r102-tier2-unit-2026-08-31/. Reasoning kept: THE SOURCE MOVES; THE DESTINATION DOES NOT - unitDir changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by GetAppDrivePath exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: a directory that exists is not a package - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside (07 §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md | | R-103 (was C9-F1b) | The Tier-2 no-coverage refusal named the working action but did not route to it - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller v0.229.0: POST /backup/tier2/unit-restore and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: audits/DRILL-r102-tier2-unit-2026-08-31/. Reasoning kept: a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: two questions, two predicates - CanRestore() was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: tier2UnitNotCoveredMsg was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE (the refusal now carries tier2UnitAvailableMsg, verified at the endpoint) | full text: git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md | | R-87 | The restic tier was never restore-tested — RE-SCOPED by its own spike to "prove the off-site snapshot still CONTAINS a recoverable unit". Shipped in controller v0.231.0 + hub v0.110.0. Evidence: tests/r87-offsite-proof-2026-08-31/; reasoning: audits/SPIKE-restic-restore-test-2026-08-31.md. Reasoning kept: The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing. The acceptance rule has TWO parts and the obvious one is a trap — "everything declared is present" passes a hollow unit, which is the shape it exists to catch. The expectation comes from INSIDE the unit, never the live box: the snapshot may predate the app's shape. The volume half is an EXISTENCE check and not a name match — the naming held on all eight real units, but "held on eight" is not "derivable" (R-355), and half a rule that is true beats a whole rule that is invented. THREE outcomes: pass, fail, and cannot-judge — collapsing the third hides a gap in one direction and alarms on our own blind spot in the other. It proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app — §8 matrix row 4 was deliberately NOT moved. The proof's scratch is a SEPARATE root because the job deletes on every path, and sharing the customer's root would mean a nightly job deleting a copy the customer is looking at. | CLOSED 2026-08-31 — SHIPPED + PROVEN-LIVE (controller v0.231.0, hub v0.110.0) | full text: git show 303129e:documentation/backlog/OPEN-ITEMS.md | | R-406 | Two unrelated findings shared the identifier R-133. Resolved 2026-09-01 by renumbering the hub-uniqueness finding to R-415. Reasoning kept: citations were MEASURED before choosing — 3 for hub-uniqueness, 5 for the plaintext break-glass credential — and the FEWER-cited one moved; this is the opposite of the task's literal instruction, whose stated ground ("the older number has the longer reference trail") the measurement contradicts; the principle was followed and the letter was not; the within-register duplicate rule was deliberately NOT added in the commit that removed its only subject — a guard whose red-proof can only be a planted fixture is not this project's standard (R-416). | CLOSED — RENUMBERED (2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-410 | golden_currency_gate.py was satisfied by a DIRECTORY NAME — mkdir turned it green with no bake behind it. Shipped in felhom.eu, 2026-09-01. Reasoning kept: a directory name is a label; GOLDEN_SHA256=<64 hex> is a fact only a completed publish produces; the self-test ships a POSITIVE CONTROL, without which "it fails on an empty directory" would be satisfied by a gate that fails on everything; directories that look right and hold nothing are printed by name rather than silently ignored, so a half-finished bake is visible. | CLOSED — SHIPPED (felhom.eu, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-414 | The nightly off-site proof was INERT on a box with no registered data drive, every night, with only a WARN. Shipped in controller v0.232.0. Evidence: audits/R411-R414-2026-09-01/, determination in 00-part2.1-determination.md. Reasoning kept: the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded — its own tests say "the scratch still resolves … only the DESTINATION moves"; the fallback is SCOPED because the two callers ask different questions, and one predicate answering both is the R-356 defect itself — unit-only may fall back (§7: a driveless app's unit already lives on the system data path, "intended, not a defect"), a full restore may not (§2.2: state-only tier); absence on last_proof_result already means "controller too old", so a second meaning on one field is the StatsKnown trap one level up; a cannot_run is recorded but does NOT advance per-snapshot due-ness, or the app would never be retried once a drive is registered. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-407 | restic check DOES write a lock file, and the comment above it said it never writes. Corrected in controller v0.232.0. Reasoning kept: corrected in place, not deleted — R-360's rule is that a comment claiming a guard is why nobody looks for the missing one; the same paragraph now carries the fact that restic stats also takes a lock. | CLOSED — CORRECTED IN PLACE (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-408 | RestoreOffboxScratch took no single-writer flag while a comment asserted every off-site operation did. Shipped in controller v0.232.0. Reasoning kept: the real deliverable is the WALK, not the acquire — the sentence was false for months and nothing checked it, the ninth instance of this project's most-repeated class; it is an AST pass and not strings.Contains, because a commented-out call still contains the string; adding a line to offsiteExempt is a deliberate act and belongs in the commit that adds it; the R-87 proof's exemption is kept HONEST by a second test that fails if that path ever gains unlockStale, routes through resticStep, or loses --no-lock. | CLOSED — SHIPPED (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-411 | A background job deleted the lock of a live customer restore and logged it as a crash that did not happen. Shipped in controller v0.232.0. Evidence: audits/DRILL-soak-2026-08-31/phase1-lock-collision/, audits/R411-R414-2026-09-01/. Reasoning kept: restic stats TAKES a repository lock — the fact nobody had, and the one that made the chain reachable; restic check takes one too, restic snapshots and restic list do not; the fix was wider than the row — FOUR entry points were unflagged, three of them found by R-408's walk rather than by the report; the escalation in resticStep was NOT removed — real stale locks exist and it clears them; the defect was that a sibling could be live. | CLOSED — SHIPPED + PROVEN-LIVE (controller v0.232.0, 2026-09-01) | full text: git show 22e1c95:documentation/backlog/OPEN-ITEMS.md | | R-403 | A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with --delete. Shipped in controller v0.230.0. MEASURED before it was fixed — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: audits/DRILL-r403-tier2-delete-2026-08-31/. Reasoning kept: hollowness is a MANIFEST question, never a size question - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. The guard fences ONE shape and not shrinking - 07 §8 row 5's derived-copy rebuild is a DESIGN DECISION, --delete stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). The rehydrate happens INSIDE the restore - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and the capture is deliberately NOT guarded, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. A warning that fires on everything costs the same as the comforting lie it replaces - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md | | R-459 | The skipped MariaDB conversion is STABLE but never self-resolving; converting costs 7 s and keeps the abort — RULED YES and shipped 2026-09-13: MARIADB_AUTO_UPGRADE=1 on bookstack-db, kimai-db, nextcloud-db, romm-db (catalog eec1228/bd32830/3525e35; no image moved, catalog_since untouched). Measured audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md; proven audits/r459-close-2026-09-13/ — harness E3/E3b proven with engine_state_after = already upgraded to 12.3.3-MariaDB [exit=1], skipped due to $MARIADB_AUTO_UPGRADE 0 lines, C3 still failed; landed on demo-hp via the real 15-min cycle with both container IDs unchanged, one deliberate restart → MariaDB upgrade not required, /login 200. Reasoning kept: ask the engine, not the log — the entrypoint prints MariaDB upgrade not required on an unsupported downgrade too (R-464); mariadb-upgrade --check-if-upgrade-is-needed exit 0 = needed, 1 = not. Not established, unchanged: whether any MariaDB feature misbehaves on an UNCONVERTED datadir. The precaution that keeps the setting inert until Slice 4: the engine-major rule + gate, removal tracked as R-469. | CLOSED 2026-09-13 — shipped in the catalog, PROVEN by harness and live | full text: git show ae59c31:documentation/backlog/OPEN-ITEMS.md | | R-467 | Controller v0.236.0 owed a golden — PAID 2026-09-13: golden 0.236.0 baked (GOLDEN_SHA256=58a3cc24…958bf, 654 115 664 B), round-tripped from the DOWNLOADED bytes, ./etc/felhom-controller-image says felhom-controller:0.236.0, hub dropdown agreed, three-field vouch re-read (0.236.0 / agent 0.130.0 / min_agent 0.129.0, not the R-216 shape), floor raised 0.232.0 → 0.236.0. Evidence documentation/tests/golden-0.236.0-2026-09-13/. Reasoning kept: this bake carried FOUR unbaked releases and is the LAST per-release bake — goldens are weekly and before any install from today (R-468); the MinAgent line was missing from four headers (R-470). | CLOSED 2026-09-13 — baked, vouched, floor raised | full text: git show ae59c31:documentation/backlog/OPEN-ITEMS.md | | R-448 | UPDATE ARC SLICE 4 — a guarded update: verified-backup precondition, abort path, truth at the moment of action. Shipped controller v0.237.0 (job) + v0.238.0 (page) + v0.238.1 (nightly legs skip an app mid-update, found live). Proven live on demo-hp 2026-09-13, scenarios A/B/E/F/H and the restore walk. Evidence: audits/slice4-2026-09-13/. Reasoning kept: the precondition is the existing verified backup, not a new copy (ruling 2026-09-02) — backup.Tier2UnitRestorePoint, extracted from the backups page, not copied; age a copy by its last SUCCESSFUL Tier-2 copy, never the manifest created_at (measured: the manifest moves only on definition changes); no automatic rollback — measured per-app, the route back is the restore; anything that writes a restore point skips an app that is held OR updating. Open consequences: R-472, R-475, R-476. | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show HEAD~1:documentation/backlog/OPEN-ITEMS.md | | R-443 | The Update button reported success over an app it had just broken. Closed by slice 4 (v0.237.0): 202 accepted, not completed; the outcome exists only as update_phase after health. Pinned by TestR443_UpdateIsNeverReportedCompleteSynchronously. Evidence: audits/slice4-2026-09-13/. Reasoning kept: a compose exit code is never a success signal (spike §4: HTTP 200 over a crash loop). | CLOSED 2026-09-13 | full text: git show HEAD~1:documentation/backlog/OPEN-ITEMS.md | | R-439 | The restore hold was not honoured by the update path. Closed by slice 4 (v0.237.0): update joined the router's hold check; live Scenario H refused start/restart/update and the boot sweep. Evidence: audits/slice4-2026-09-13/. Reasoning kept: a hold that only one path honours is not a hold — the audit found the drive-return gate (restart + boot recreate) and the nightly volume dump ignoring any hold; all three fixed and red-proofed. | CLOSED 2026-09-13 | full text: git show HEAD~1:documentation/backlog/OPEN-ITEMS.md | | R-470 | Four controller CHANGELOG headers (v0.233.0–v0.236.0) carried no MinAgent: line, while the vouch and now the declared floor read it from the header. Closed 2026-09-13 (felhom-controller f946b0d): the four headers backfilled with **MinAgent: 0.129.0** (unchanged) — v0.232.0's value, proven unchanged (no commit under internal/agentapi since 2026-09-01; highest featureMinAgent 0.129.0) — and controller/scripts/minagent_header_gate.py (fast, blocking) refuses a newest header without the line; a prose or code-span mention does not count (decoy; red-proof F). | CLOSED 2026-09-13 — GATED | full text: git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md | | R-472 | The golden cadence ruling and the hub's floor rule contradicted each other: a floor above the vouched golden delivered nothing. Operator ruling 2026-09-13, hub v0.112.0 (f181efd): a floor saved with the release's declared MinAgent is served above the golden under the same agent comparison; an undeclared one is still held and both forms refuse it (floor_needs_min_agent). Proven live: controller 0.239.0 reached demo-hp in 14 s and demo-felhom in 15 s from the save, hub managed floor SERVED … from declared. Evidence: audits/rulings-r472-r475-2026-09-13/ 02, 03. Reasoning kept: the manifest leads the floor inside the golden; above it, the release's own declared MinAgent does (publish-train rule 1); a declaration binds to its exact floor. | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md | | R-475 | The update precondition was Tier-2-only, so an app with no second-drive copy could not be updated. Operator ruling 2026-09-13, controller v0.239.0 (b93c154): the first fresh copy in the order Tier 2, Tier 1, Tier 3 (bounded, unreachable = absent + WARN); backup_max_age applies to the chosen tier; nothing anywhere → back up first; refused only when no copy and no backup can be taken; the hold names the tier; a Tier-2 failure in the pre-backup is a WARN. Proven live on demo-hp: nothing anywhere → backed up first (04), Tier 1 alone (05), held naming „saját meghajtó” (07), restored from „helyi” (08). Red-proof M (age only on Tier 2) fails. Reasoning kept: first FRESH copy, not first copy — a stale mirror must not force a backup while the own unit is minutes old. Follow-ups: R-477..R-480. | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md | | R-473 | The glance catalog template crash-looped on every fresh install — the image ships no default config. Closed 2026-09-13 (catalog 50ad286): an entrypoint wrapper seeds a small Hungarian start page on first boot only when /app/config/glance.yml is absent, then execs the image's own command (read from the image, not guessed); the seed validated with the image's config:validate. Proven live on demo-hp the same evening: a fresh throwaway install healthy in 21 s, restart count 0, front door 200, seed present in the volume (audits/nightly-2026-09-13-adventurelog/08-R473-glance-fresh-install.txt). No version moved. | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-471 | observations_gate.py read only the FIRST observations section, so an appended second section with an unmarked item passed — and the R-419 decoy had been reading LIVE HOLE at HEAD. Closed 2026-09-13 (felhom.eu scripts/observations_gate.py): observation_sections collects every observations heading and observation_items pools their items; the heading shown is all of them joined. Cause established: first-heading-wins (the parser breaks on the first match). Red-proof: the old parser exits 0 on the decoy shape, the new one 1 (audits/v0240-2026-09-13/rp-R471.txt); all 12 felhom.eu decoys behave; the felhom.eu and controller REPORTs still pass through the shared script. The "run by hand" half stands: the decoy suite is still not in any runner (R-426's exemptions) — that is a separate row if wanted. | CLOSED 2026-09-13 — GATED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-486 | Removing an app with its backups KEPT forgot its Tier-2 record, so the second-drive restore was refused over an intact mirror. Closed in controller v0.240.0 (bdcbd50): the record goes only with remove_backups. Proven live on demo-hp: remove keeping backups → cross_drive record kept → „Teljes visszaállítás" restored 2 volumes and the database, data identical. Red-proof: an unconditional SetCrossDriveConfig(nil) fails the wiring test. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-484 | PostGIS was not a database. Closed in v0.240.0: dbTypeForImage matches postgis, pgvector, timescaledb as Postgres. Proven live: adventurelog's unit now carries db-dumps/adventurelog-postgres.sql (8 MB) and the unit restore reports „2 adatkötet és az adatbázis visszaállítva". Red-proof: dropping postgis fails the table test. Immich's own image already matched. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-485 | The backup card read two dead paths and said has_backups:false over 484 MB. Closed in v0.240.0: it sizes the recovery unit and the app's Tier-2 mirror(s). Proven live: 244M + 244M, has_backups:true. Red-proof: the old paths fail TestR485. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-480 | The card kept a held update's „leállítva marad" sentence after a successful restore and after removal. Closed in v0.240.0: the stack remembers its last update ended held; fillHoldReason hides that outcome once the hold is gone or the app is not deployed; a pull failure keeps its sentence. Red-proof: disabling the block fails three assertions. Live: the card after a restore reads clean. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-477 | The update's Tier-3 lookup paid the full 15 s bound and blamed another app. Closed in v0.240.0: OffsiteSnapshotTimes — one snapshots --json, no per-app stats; the inventory page and the update share offsiteNewestPerTag (registered read-only for R-408). Red-proof: routing the lookup through the inventory fails TestR477. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — GATED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-478 | A unit left by a removed install counted as the reinstall's fresh copy. Closed in v0.240.0: usableRestorePoint refuses a copy older than the app's deployed_at; R-474's fix removes such units anyway. Red-proof: dropping the check fails TestR478. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — GATED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-474 | „Delete backups" deleted only db-dumps; the unit and the Tier-2 mirror survived; prefs stayed. Closed in v0.240.0 for the backups half: the whole unit, every mirror (Tier2MirrorDirsForApp / RemoveTier2Mirrors, exact-path guarded) and the prefs go; backup_paths_removed lists them. Proven live: 244M + 243.6 MB removed, nothing left on either drive, prefs None. The volumes_removed: null half is re-filed as R-489. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-466 | Removal with „Mentési adatok törlése" left the unit's compose/ + manifest.json. Subsumed by R-474's fix in v0.240.0 (the whole unit goes). Decided by the same change: the button means the WHOLE unit. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-) | CLOSED 2026-09-13 — GATED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-482 | adventurelog ran Django with DEBUG=True on the public origin. Closed in catalog ed2c018: DEBUG=False on the backend service. Proven live after the sync: the same CSRF failure renders the 300-byte production page, no debug text. wger, tandoor, paperless-ngx are recorded as unmeasured. audits/v0240-2026-09-13/ (10-v0240-validation.txt; red-proofs rp-v240-*) | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-453 | ~/.config/credentials values are single-quoted, and a half-applied strip produced a confidently wrong "password is stale" verdict — twice. Closed 2026-09-13: the instrument the row asked for exists — felhom.eu/scripts/read_credential.py KEY <0600-file> (66156c6, 2026-08-31) — and the whole 2026-09-13 night used it for every controller and hub password with zero quoting incidents; the memory credentials-file-values-are-quoted now names it. The instrumentation lesson (a discriminator that rules out one alternative does not rule in the rest) stays in the memory. | CLOSED 2026-09-13 — INSTRUMENTED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-461 | runbooks/target-selection.md named a venue that does not exist and fenced a fixture that is gone. Closed 2026-09-13, both halves checked against both boxes first: (a) no /mnt/nvme-1tb on demo-hp or demo-felhom; demo-hp's NVMe is nvme0n1 at /mnt/hdd_1 (demo-felhom's /mnt/hdd_1 is sdb) — the runbook now names /mnt/hdd_1 and says it is the same disk as the data drive; (b) qm list is empty on BOTH boxes — drill-r50 (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | CLOSED 2026-09-13 — DOCUMENTED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-452 | Nothing enforced catalog_since, so the badge's one number could silently under-report. Closed 2026-09-13 (catalog): scripts/check-catalog-since.py, the fifth gate in catalog_gates.py — a --range A..B gate in the engine-major shape: an app whose per-service image: lines differ across the range must carry a catalog_since on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (audits/v0240-2026-09-13/rp-R452.txt). | CLOSED 2026-09-13 — GATED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-465 | cfg.Paths.HDDPath — empty on every box — still had six readers; were any inert? AUDITED 2026-09-13 on demo-hp (registered drive /mnt/felhom-drives/hdd_1, hdd_path absent, no FELHOM_PATHS_HDD_PATH). Five of six fall back before the value matters: report/builder.go:69 and monitor/healthcheck.go:35 take storagePaths[0]; web/server.go:740 (primaryHDDPath) takes the default storage path; main.go:511 (metrics) takes the default storage path; main.go:347 passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. One is inert AND unreachable: api/router.go:981 (systemInfo, GET /api/system/info) reads the empty value with no fallback (hdd_configured:false forever) — and the endpoint itself is shadowed: the web layer's ServeSystemAPI claims /api/system/* and answers 404 „ismeretlen végpont" for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as R-490. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | CLOSED 2026-09-13 — AUDITED | full text: git show 681c3d6:documentation/backlog/OPEN-ITEMS.md | | R-479 | For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so. Operator ruling 2026-09-13; closed in controller v0.241.0 (3e81330): an app with classified binds walks second drive → off-site → own unit (UpdateTierOrderFor), and the hold sentence ends with what the chosen copy holds (RestoreHold.CopyHolds). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). audits/v0241-2026-09-13/ | CLOSED 2026-09-13 — PROVEN-LIVE | full text: git show 8914ab0:documentation/backlog/OPEN-ITEMS.md | | R-483 | adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog). Closed 2026-09-13: the k3s ingress routes /media, /static, /admin, /accounts to the backend service on port 80 — the nginx inside the backend image that serves Django's X-Accel-Redirect media; the catalog routed everything to the frontend. Two catalog cuts (3172258 router, ed62cfd port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (GET /media/…webp with a session → 200 image/webp, RIFF/WEBP) and confirmed by the operator in a browser at 21:49 (two photos render). Scripted multipart uploads through the frontend's /api proxy still 500 (upstream RequestContentLengthMismatchError); browser uploads work — not a template matter. audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt | CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed | full text: git show 8914ab0:documentation/backlog/OPEN-ITEMS.md | | R-456 | A partly-dead stack is not a boot orphan, written down nowhere (P3). Closed in controller v0.242.0 (d698ce3) by pinning the rule: an absent member does not make a stack degraded, a present-but-dead member does (internal/bootrecon/r456_partly_dead_test.go). Design unchanged. | CLOSED 2026-09-14 — PINNED | full text: git show 72ee053:documentation/backlog/OPEN-ITEMS.md | | R-476 | The Mentések page dated a Tier-2 copy from the unit manifest, which moves only with the definition (P3). Closed in controller v0.242.0 (d698ce3): Tier2Coverage.UnitDataDate (newest dump in the mirrored unit) is what a refreshed leg names; a PRESERVED package keeps the manifest date (R-403). Unit-proven with the measured shape (manifest 09-12, dump 09-13); live on 9202 two captures under one definition dated the copy by the second capture's dump. Red-proof: the manifest-only date fails. audits/v0242-2026-09-14/ | CLOSED 2026-09-14 — PROVEN | full text: git show 72ee053:documentation/backlog/OPEN-ITEMS.md | | R-481 | No scratch guest on demo-hp for the nightly rotation (P2). Operator ruling 2026-09-13, option 1. BUILT the same night: LXC 9202 demo-hp-scratch on demo-hp, restored from the vouched golden 0.236.0 onto a dir storage re-added at /mnt/hdd_1 (nvme-scratch), sized like 9201 (7 cores / 25 898 MB / 32 G + 70 G, unprivileged), the demo-hp customer seeded with hub OFF, tunnel OFF, agent OFF, off-site OFF, self-update OFF; image set by hand (allowed only there); a claimed settings.json with the demo password and a scratch second drive. It persists on purpose. Disposition in all three layers: the guest (/etc/felhom-scratch-disposition), the host (pct description, tags scratch,r481, the bootstrap file) and the hub side (operations/nodes.md, this row — the hub has no customer notes field and the guest is outside the felhom pool). Two traps written into nodes.md: pct restore wants the golden as a backup volume; directories made on the raw volume from the host must be chowned to 100000. The rotation restarted from bentopdf on it. audits/nightly-2026-09-13b-bentopdf/14-scratch-guest-built.txt | CLOSED 2026-09-13 — BUILT, persists | full text: git show 72ee053:documentation/backlog/OPEN-ITEMS.md | | R-487 | A removed app whose backups were kept was listed on neither backup page (P2). Closed in controller v0.242.0 (d698ce3): the local lists are keyed on the DRIVES the way R-237 keyed the off-site list on the store — ListRemovedAppUnits walks backups/primary/ on the system path and every connected registered drive; the Mentések page lists the unit after the deployed rows („Eltávolítva — visszaállítható", one action), the Visszaállítás picker lists it in its own group, GET /api/backup/snapshots answers for it, and the restore opens the unit where it sits (primaryUnitDirFor — a unit kept on a data drive was unreachable before, the fallback named the system path). Proven live on the scratch guest 9202 with an opengist throwaway: removed with data, backups kept → row + picker + API answered; „Visszaállítás a mentésből" reinstalled it running. Red-proofs: lister inert, picker 404, wrong unit dir, row not built, row not rendered, picker not rendered — all fail. audits/v0242-2026-09-14/ | CLOSED 2026-09-14 — PROVEN-LIVE | full text: git show 72ee053:documentation/backlog/OPEN-ITEMS.md | | R-490 | The monitoring page's memory-distribution card never rendered — /api/system/info was 404 (P3). Closed in controller v0.242.0 (d698ce3): an exact-path mount ahead of the web layer's /api/system/ prefix, and systemInfo reads the default storage path like every other reader of the empty global. Live on 9202: 200 with the drive figures. Red-proofs: mount removed, fallback removed — both fail. The global's deletion stays deferred → R-492. audits/v0242-2026-09-14/ | CLOSED 2026-09-14 — PROVEN-LIVE | full text: git show 72ee053:documentation/backlog/OPEN-ITEMS.md | | R-491 | Removing an app left its update hold in the store, so a reinstall started held (P2). Closed in controller v0.242.0 (d698ce3): removeStack clears an UPDATE hold (Settings.ClearUpdateHold, never an R-379 restore hold), logged. Proven live on 9202: a held opengist removed → the store no longer carries the hold, the app redeployed without refusal. Red-proof: the removal without the clear fails the wiring test. audits/v0242-2026-09-14/ | CLOSED 2026-09-14 — PROVEN-LIVE | full text: git show 72ee053:documentation/backlog/OPEN-ITEMS.md | | R-546 | The first-hour guide and the reminder bar sent the household to create their recovery code before the box could (P2). Closed in controller v0.246.0 (0fe315b): the bar consults the agent's OWN preflight ok (all five blocking items, not a copy of pbs_storage_id), cached 60 s, probed only while paused, held back while not ready; /backup/escrow shows a waiting card that polls and reloads; POST /api/escrow/start refuses 409 before staging (the direct path that produced the raw -storage stderr); unknown readiness keeps the bar. The guide moves the step after the first apps: „amikor a sárga sáv megjelenik”. Red-proofs: bar held back, waiting card, start refusal; controls ready and unknown. Proven by tests through ServeHTTP, NOT live — no Tier-0 box is paused and agent-connected (R-551); chaos night measured live the ~17-minute red window and its self-heal. | CLOSED 2026-09-17 — PROVEN (tests); live walk owed by R-551 | full text: git show 06334e1:documentation/backlog/OPEN-ITEMS.md | | R-549 | The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2). Closed by operator ruling A (2026-09-17): alerting.stale_threshold 30 m → 45 m, node_down/host_down at 90 m (manifests/hub.yaml, commit 06334e1). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub v0.117.0 (37ae31f) makes controllerStatus read the same value — red-proof TestControllerStatus_FollowsConfiguredThreshold (report 40m old: status warn, want ok). Proven live: the running hub printed node_stale after 45m0s, node_down after 1h30m0s and host_stale after 45m0s, host_down after 1h30m0s at 08:22Z. Reasoning kept: the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt | CLOSED 2026-09-17 — PROVEN-LIVE | full text: git show 06334e1:documentation/backlog/OPEN-ITEMS.md | | R-550 | The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2). Closed in controller v0.246.0 (0fe315b) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): restore-status.json in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on /backups/restore until that app's next restore, raised once as restore_interrupted (hub v0.117.0, household). Red-proofs: record across restart (StartedAt:0001-01-01), main() wiring (AST), startup helper, page card. Proven live on demo-hp 9201: a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read ok:false … megszakadt … interrupted:true, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. Known gap filed: R-552 (a removed app keeps its notice). audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt | CLOSED 2026-09-17 — PROVEN-LIVE | full text: git show 06334e1:documentation/backlog/OPEN-ITEMS.md | | R-539 | The restart brake caught a FAST crash loop and was blind to a SLOW one (P3). Closed by operator ruling 3 of 2026-09-16 in agent v0.132.0 (18d03bd, tag v0.132.0, sha256 4afe8157…) + hub v0.117.0 (37ae31f): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth slow_crashloop_since moves (at most once per 24 h) and the hub mints controller_slow_crashloop (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed agent_update to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging slow_crashloop_max=5 slow_crashloop_window=24h0m0s. Proven live with the PRODUCTION window (no test-only interval): five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 SLOW CRASH-LOOP … restarts_24h=5; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). Reasoning kept: the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. audits/evidence-chaos-fixes-2026-09-17/partC-*.txt | CLOSED 2026-09-17 — PROVEN-LIVE | full text: git show 3c1882a:documentation/backlog/OPEN-ITEMS.md | | R-637 | Build the undo (09 §3 decision 15) — the box puts a failed update back by itself (P2). Closed in controller v0.263.2 (8fc2b4a v0.263.0, 5d38573 v0.263.1, 2cd6666 v0.263.2). Copy method chosen by the bake-off (decision 19): a folder copy of every NAMED volume, cp -a into <vol>.pre-update-<stamp> after the pull, finished-marker last; bind folders never touched. Undo in failAndHold: copies validated first, volumes refilled, definition + pin + the pinned version's .felhom.yml (applied-meta/) put back from the job's own copies, old probe, undone or a HOLD whose sentence says the undo failed and the data state. Proven live on 9202 (0.263.2): docmost, romm, vikunja undone by the product with seeds before the backup, after it, and seconds before the press all read back, ledgers equal, page line hu/en; a cut-off copy → HOLD untouched; a power cut (pct stop) during undoing → resumed and undone; a manual press after an undo → done, note cleared; removal deleted kept copies. Two defects the live proof found and the unit tests could not: the undo's probe was gated on running while the current probe held the app unhealthy (v0.263.1), and the "old" .felhom.yml was already the new one because it flows in on every sync (v0.263.2). Evidence: audits/undo-bakeoff-2026-09-23/, audits/undo-live-2026-09-23/. Reasoning kept: a copy counts only with its finished-marker, judged by the helper container's own exit — killing docker run does not stop the copy; the undo reads its own copies, never the recovery unit. | CLOSED 2026-09-23 — PROVEN-LIVE (controller v0.263.2) | full text: git show 4c92bea:documentation/backlog/OPEN-ITEMS.md | | R-639 | After a held update the previous definition survived only in the recovery unit (P3). Closed in controller v0.263.0/v0.263.2: the pre-update compose, applied definition, pin and the pinned version's .felhom.yml are kept until the undo is over; the undo never reads the unit (R-645). Evidence: audits/undo-bakeoff-2026-09-23/, audits/undo-live-2026-09-23/. | CLOSED 2026-09-23 — SHIPPED (controller v0.263.2) | full text: git show 4c92bea:documentation/backlog/OPEN-ITEMS.md | | R-641 | An app with no database server had no last-second copy (P2). Closed in controller v0.263.0: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: audits/undo-bakeoff-2026-09-23/, audits/undo-live-2026-09-23/. | CLOSED 2026-09-23 — PROVEN-LIVE (controller v0.263.0) | full text: git show 4c92bea:documentation/backlog/OPEN-ITEMS.md | | R-642 | Start answered 200 "start completed" over a crash loop (P3). Closed in controller v0.263.0: start/restart answer requested — state now: <state>, never "completed" (startAnswer, pinned by TestR642_*, red-proofed); live on 9202: Stack romm start requested — state now: running. | CLOSED 2026-09-23 — PROVEN-LIVE (controller v0.263.0) | full text: git show 4c92bea:documentation/backlog/OPEN-ITEMS.md |

2026-09-23 — the household is told when an update is undone or held (controller v0.264.0, hub v0.120.0)

Row What Closed Full text
R-606 Every sentence the update path shows a household was Hungarian-only (P2). Closed in controller v0.264.0 (bc27894): UpdateError is stored as a key + args; phase labels, refusals, failure lines, the undone line, the hold sentence and its prefix and UpdateCopyHolds render in the reader's language on both pages and in GET /api/stacks/<n>; the stored Hungarian is byte-identical (parity gate green). Live: 9202 and 9201, both pages, both languages. Three leftovers → R-647. CLOSED 2026-09-23 — PROVEN-LIVE (endpoint-level; audits/undo-fleet-2026-09-23/22-*, 23-*, 35-*) git show HEAD~1:documentation/backlog/OPEN-ITEMS.md
R-620 A disabled notifier dropped every event with no local trace (P3). Closed in controller v0.264.0: one WARN per event type per process naming the dropped event, then DEBUG. Live on 9202: six types each WARNed once; app_start_failed WARN at 12:02:30Z then DEBUG at 12:03:00Z. CLOSED 2026-09-23 — PROVEN-LIVE (24-9202-r620-warn-then-debug.txt) as above
R-646 An app pinned before v0.263.2 had no record of its own .felhom.yml (P3). Closed in controller v0.264.0: a startup pass records applied-meta/ for every deployed, pinned app CURRENT with the catalog; a behind app is skipped by name (its file is gone — not backfillable by design). Live: 9202 recorded 3, 9201 recorded 10, skipped 0. CLOSED 2026-09-23 — PROVEN-LIVE (28-9202-controller-log.txt, 30-9201-before-and-upgrade.txt) as above

2026-09-23 (evening) — clean-up: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0)

Row What Closed Full text
R-634 An app could RUN while the controller recorded it as not deployed (P1). The unremovable half closed in v0.262.0; the mechanism is now diagnosed and fixed in v0.265.0 (0054d4b): the whole-box backup's app list read the in-memory Deployed flag, true from the moment a deploy is accepted, so the volume leg stopped a DEPLOYING app (compose down), dumped half-made volumes and ran a second compose up -d beside the deploy's own — both failed and the deploy recorded „not deployed" (reproduced on demand on 9202, audits/cleanup-2026-09-23/12-*, 13-*). Fix: deploying apps leave every backup list; the volume leg re-asks before the stop; StopStack/StartStack refuse a deploying stack for every caller. sparkyfitness did not reproduce alone. The deploy's own-failure question → R-649. CLOSED 2026-09-23 — PROVEN-LIVE (32-*, 33-*: backup across a live deploy stopped three other apps and never outline; deploy deployed) git show HEAD~1:documentation/backlog/OPEN-ITEMS.md
R-625 A held app invited an update its button refused (P2). v0.265.0: badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed" (tag-error, title = the hold's first sentence per reader), no Update button, 409 held unchanged. CLOSED 2026-09-23 — PROVEN-LIVE (44-*, all four box/reader language pairs) as above
R-636 Six hours of OOM kills sent the same single warning as one hiccup (P2). v0.265.0 + hub v0.121.0: the kernel oom_kill counter (the OOMKilled flag is sticky and cannot count); ≥ 20 kills in 30 min of one container run → ONE app_oom_storm (error, operator-only, per-app cooldown). CLOSED 2026-09-23 — PROVEN-LIVE on the controller (45-*, 46-*: RomM at 320M, storm at 21 kills, still one at 49); hub side by unit tests as above
R-647 Three leftovers of the update mail (P3). v0.265.0: a held update's error is the key update.error.held, rendered per reader on both pages and the API; copy_holds travels as its key; the two log wordings fixed. CLOSED 2026-09-23 — PROVEN-LIVE for (1) (43-*, 44-*); (2)(3) by red-proofed tests as above
R-648 The drill's „Mentés most" was whole-box (P3). No per-app backup endpoint exists; the harness (audits/cleanup-2026-09-23/walk.py backup_now) now presses nothing and the guarded update's own backing-up phase backs up the throwaway app alone. CLOSED 2026-09-23 — PROVEN-LIVE (43-*: phase backing-up for vikunja only) as above
R-649 A failed install could leave containers under „not deployed" (P2). Operator ruling 2026-09-23 (option a): the controller cleans up. Closed in v0.266.0 (964ae75): runComposeDeploy's failure branch runs compose down (volumes kept) before the record reads not-deployed. CLOSED 2026-09-23 — PROVEN-LIVE (audits/r649-2026-09-23/: outline with a never-healthy redis — 0 containers left, 3 volumes kept) git show HEAD~1:documentation/backlog/OPEN-ITEMS.md
R-650 A controller unit test could act on DooPlex's production Docker (P3 → raised to P2 by the night brief). Closed in v0.267.0 (80e6ad8c4772): internal/dockerexec — every docker exec goes through it; under go test a real docker is refused with an error naming the command (opt-in FELHOM_TEST_REAL_DOCKER=1; a stub under the temp dir allowed). The sweep found 8 api tests and the web/stacks/backup/appexport/system fixtures reaching the real daemon (one docker-compose down); api/stacks/web now run under a silent stub. TestR650_NoBareDockerExec pins the invariant. CLOSED 2026-09-23 — red-proofed twice (guard off → the decoy runs; a bare call → the sweep names it) audits/night-2026-09-23/A1-*
R-640 A truncated PostgreSQL copy loaded with rc 0 into an empty database (P2, narrowed to the restore paths). Closed in v0.267.0: appbackup.CheckDumpComplete (the engine's end marker); the unit restore and the off-site restore refuse before the first mutation, every replay checks again before any load. Rule: a dump is judged by its END — the header and a CREATE TABLE say nothing about whether it finished. CLOSED 2026-09-23 — three red-proofs + PROVEN-LIVE on 9202 (the household's restore button refused a half-length docmost copy; containers untouched; the whole copy then restored) audits/night-2026-09-23/A2-*, E1-r640-live.*
R-499 Every driveless app was told its data was „already in the full system backup (PBS)" (P2). Closed in v0.267.0: the Tier-2 page's sentence has four branches from the box's own whole-system backup target (own drive / same disk / drive gone / cannot ask); „(PBS)" and „nincs külön teendő" only where true. CLOSED 2026-09-23 — two red-proofs; live on 9202 (the unknown branch, hu + en, matching /api/storage/backup-target) audits/night-2026-09-23/A4-*, A7-*
R-626 A removed app came back (P2). Measured on v0.266.0, NOT reproduced: navidrome removed through the product, 390 s of docker events (the remove's destroy seen, no create), a controller restart at +150 s, a guest reboot after → no container, no volume. The two known creators (the restore/remove race R-633, the backup/deploy race R-634) are fixed. Rule kept from the row: a check that runs once, immediately, cannot see a thing created just after it — watch a window. Leftover found and filed: R-651. CLOSED 2026-09-23 — by measurement (a positive observable: the destroy events) audits/night-2026-09-23/A3-*

2026-09-24 — the undo after a restore, the held app's page, the ladder (controller v0.268.0, hub v0.122.0, catalog 5ed599c)

Row What Closed Full text
R-658 After a restore, the undo copied NOTHING (P1). v0.268.0 (206b035): the undo selects volumes from the rendered compose file (DeclaredVolumeNames: name: else <project>_<key>, each checked to exist); the label is a logged cross-check. The unit restore creates volumes WITH compose's project/volume/version labels (never a guessed config-hash); the remove counts unlabelled declared volumes. Live on 9202: restored under v0.267.0 → labels null; on v0.268.0 a failing update copied both by name and seeds A and B (B written after the restore) read back; a v0.268.0 restore → labels present, no compose warning. Rule: never select an app's volumes by the compose label. v0.268.0, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-659 A held app's page named a way back the restore refused (P1). Operator ruling 2026-09-24 (09 §3 decision 25, option A). v0.268.0 + hub v0.122.0: the hold names the newest WHOLE copy (WholeOnTier asks the refusal's own predicate); with none, hold.update.no_whole_copy in the household's language, no Mentések button, and app_hold_no_whole_copy (critical, operator-only). Live on 9202: round 11 reproduced (nextcloud, cut-off undo copy) — both languages, no button on either page, the event's R-620 line. Rule: a sentence may name a copy as a way back only when the restore for that copy would accept it. Left open beside it: R-661, R-666. v0.268.0 / hub v0.122.0, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-660 A held app also raised app_start_failed (P3). v0.268.0: a fourth suppression set at classifyRunStates — the update-held apps (backup.UpdateHeldStacks; a RESTORE hold is not in it). Live: after the hold no app_start_failed in ~12 min of scans, while a throwaway stopped out of band raised exactly one. v0.268.0, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-651 A removed app left applied-compose.yml and applied-meta/ (P3). v0.268.0: remove deletes both (the catalog mirror and hold-logs/ stay — the latter is evidence). Live: present before, gone after, on nextcloud and vikunja. v0.268.0, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-653 The memory watch wrote proven over a watch whose load never reached the app (P3). Catalog 5ed599c: load_verdict — reached only when at least half the requests got an HTTP answer, else the edge is inconclusive. Unit-tested and red-proofed; no bench run this session. catalog, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-656 The bench re-used an app's scratch drive folder (P3). Catalog 5ed599c: clear_scratch_folders removes the app's own ${HDD_PATH}/${USERDATA_PATH}/${IMPORT_PATH} bind folders before FROM and says so; never a bare root, never outside. Unit-tested and red-proofed; no bench run this session. catalog, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-40 The update path could not express a multi-hop upgrade (P2). Superseded by 09 §3 decisions 13–14 and shipped as §6.4 part 5 (v0.268.0): the catalog records every tested step with its own definition, and one press climbs one. Rule kept: a hop exists for a box only when the catalog holds it as a TESTED step — a >1-major catalog move without its intermediate steps is still a jump for a box below it. v0.268.0, 2026-09-24 git show 500488cad673:documentation/backlog/OPEN-ITEMS.md
R-663 Two catalog test suites had been red since the night of 2026-09-23 (P3). Filed and fixed the same session: test_gate_decoys.py (kimai-db 11.6 → 11.8) and test_ladder_writer.py (navidrome 0.64.0 → 0.64.1) typed the live pins as literals; they now READ them. 84 decoy cases OK. Rule: a fixture that names a live pin reads it; the catalog moves under it every night. catalog 5ed599c, 2026-09-24 this entry
R-661 A file app could not come back WHOLE from the second drive (decision 26). v0.269.0 backup.RestoreTier2Whole: the mirror's files by four rules (never delete; never write an older file over a newer one; bring back every missing one; an older, different live file is replaced and KEPT beside as <name>.felhom-<UTC>), then the unit from the mirror; refuses before anything moves without a proven, openable mirror with file legs or 2 GB free. Live on 9202: restored 1, kept-newer 1, unchanged 137, 3/3 volumes + 1/1 db in 38 s; again from a hold naming the second drive; again after a power cut mid-update (chaos round 3) and with the disk 1 GB above the floor (round 11). Rule: a restore never deletes a user file and never writes an older file over a newer one; rsyncMirror (--delete) is never reused for a restore. v0.269.0, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-24/A1/, E/round-03.json, E/round-11.json
R-662 A dead „database and settings only" second step lingered in the second drive's restore (P3). Removed in v0.269.0 (decision 25 ruled that restore out); accept_missing_files is read by no handler now — only RestoreTier2Whole passes it internally. v0.269.0, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md
R-664 A ladder step had no .felhom.yml of its own (P3). Catalog cf7cf84: steps/<StepKey>.felhom.yml beside each step's compose, written by --write-ladder, refused absent by check-test-record.py rule 4b, 8 backfilled; controller v0.269.0 reads the step's own file for the probe, the memory request and the applied record. v0.269.0 + catalog cf7cf84, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-24/redproofs/A4-*
R-665 The update judged the new version with the stack dir's .felhom.yml, which a restore rewrites (P3). v0.269.0 journals the new version's own file (entry.NewMeta) and verifyAndConclude loads exactly it. Live (Part C): wishlist, navidrome and romm steps judged by their own files. v0.269.0, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-24/redproofs/A4-r665-r664.txt
R-666 While support is informed, Remove offered to delete the data (decision 27). v0.269.0: the dialog reads keep_data_only and offers only „remove the app, keep my data"; the API refuses data or backup deletion with 409 (hu + en); the no-whole-copy sentence is informal. Live on 9202 (a one-drive hold). v0.269.0, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-24/A2/10-held-no-copy-keep-data.*
R-667 A crash loop never reached the alarm, and nothing stopped it (decision 28). v0.269.0 + hub v0.123.0: ≥ 6 restarts in 10 min (RestartCount, not the resettable restarting_since) or an OOM storm → the box stops the app, holds it (unhealthy_stop), tells household + operator (app_stopped_unhealthy); Start = one more try; a repeat in 24 h says support is informed. Live: gokapi trip 1 and 2; chaos rounds 1, 7, 8, 12. Rule: Docker's back-off caps a steady loop at ~1 restart/min, so a threshold must be below 10 per 10 min. v0.269.0 / hub v0.123.0, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-24/A3/, 08 §6.2
R-668 The Tier-2 copy chose a registered path that no longer existed, on the app's own disk (P2). v0.269.0: the same-disk check fails CLOSED. Live: the next copy went to the SSD. Residual: an older same-disk record counts until the next Tier-2 run replaces it. v0.269.0, 2026-09-24 git show 4502af6bb109:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-24/A1/01-find-mirror.txt
R-669 After a failed update ended in a restore, the box kept the FAILED step's health check as the pinned version's, and the next undo held the app for nothing (P2). v0.270.0: the recovery unit captures the pinned version's .felhom.yml (applied-meta), not the stack dir's file the sync may already have replaced; a restore makes the restored file the applied record. Live on 9202: the sync wrote the bad probe at 14:57:46, the unit captured at 14:57:55 kept the good one; a second-drive restore turned stack 8999 / applied 3000 into 3000 / 3000, the applied record rewritten at the restore. Rule: a copy of an app's definition carries the PINNED version's health check, never the catalog's newest. v0.270.0, 2026-09-24 git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md; audits/r672-2026-09-24/D/30-33*
R-674 The ladder log called a pin at the head "older than the ladder" (P3). v0.270.0: it says "AT THE HEAD". Unit test + red-proof only — no product path reaches it since R-679 refuses a current app first. v0.270.0, 2026-09-24 git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md; audits/r672-2026-09-24/redproofs/D-r674.txt
R-679 An Update on an app already at the head ran the whole guarded update — dump, pull, restart (P2). v0.270.0: 409 already_current before anything moves, hu + en; a re-tested digest of a floating tag still updates. Live on 9202 (privatebin): both languages refused, no backup, no pull. A test comment had called a same-version Update "the repair path" — Restart is. v0.270.0, 2026-09-24 git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md; audits/r672-2026-09-24/D/10-r679-live.txt
R-681 An install cut off by a controller restart was lost silently (P2). v0.270.0: an install marker before the compose-up, removed when the install ends; a marker at start → compose down (volumes kept), stale pin records cleared, app_deploy_failed with the reason, and the apps page says the install was interrupted (hu + en) until the next install; a finished install only loses the marker. Live on 9202: mealie killed 0 s into its pull → reported, page sentence, nothing left running; reinstall cleared the sentence; actualbudget (killed after it finished) left alone. Rule: a long-running act the customer started is journaled so a restart finishes or reports it. v0.270.0, 2026-09-24 git show 54bff69f88ab:documentation/backlog/OPEN-ITEMS.md; audits/r672-2026-09-24/D/20-26*
R-672 The scheduled restore-test filled the production thin pool and turned a customer guest's disks read-only (P1). Agent v0.133.0: space preflight on the UNCOMPRESSED size, off the tested guest's pool, unknown refuses. Delivered 2026-09-24 night to both demo boxes by CC-signed agent_update jobs (ruling 1, 2026-09-16); restore test back ON; live: demo-felhom PASS 85 s (pool 3.13 → 5.68 → 3.13 %), demo-hp REFUSED for space (needs 30.3 GiB, has 22.1). Peti's box has not received it. agent v0.133.0, delivered 2026-09-24 git show 75ff264:documentation/backlog/OPEN-ITEMS.md; audits/r672-2026-09-24/; audits/night-2026-09-25/A/
R-673 9201's whole-box backups failed on a stale snapshot-delete lock (P2). Agent v0.133.0: the stale-lock sweep every 10 min under the one-heavy-operation gate. Delivered to both demo boxes 2026-09-24 night; demo-hp 9201's next whole-box backup ran clean (no lock, 8.18 GB). agent v0.133.0, delivered 2026-09-24 git show 75ff264:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-25/A/A4-*
R-684 demo-hp's whole-box backups could not fit on its root-disk target (P2). Operator ruling 2026-09-24 evening, option A: retention 3 → 1 (local_backup_retention), the two oldest archives pruned by PVE's own prune, and one backup by the product's trigger fitted (8.18 GB; root 89 % → 60 %). The product lesson — warn BEFORE the night — is R-685 (agent v0.134.0 skips with a reason). ruling 2026-09-24; applied 2026-09-24 night git show 75ff264:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-25/A/A4-hp-backup-space.txt, A4-hp-backup-run.txt
R-680 The box did not remember a failed update step (P2). Controller v0.271.0: an undone or held step is recorded in app.yaml (failed_update_step, tied to the ladder's print); the automatic leg skips it until the catalog's ladder changes; a person can still press. Live on 9202: vikunja undone night 1, skipped failed_before night 2, re-tried after the catalog re-tested it. v0.271.0, 2026-09-25 git show 75ff264:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-25/C/; B/redproofs/R680-*
R-678 After a step ended done, steps-left and the badge stayed stale (P3). Controller v0.271.0: the update re-reads the app's catalog fields BEFORE it says done (and after an undo). Live on 9202: every automatic step's page read current at the leg's end. v0.271.0, 2026-09-25 git show 75ff264:documentation/backlog/OPEN-ITEMS.md; B/redproofs/R678-*
R-643 The ruled chain left the automatic update leg at most 15 minutes a night (P2). Decision 20, built in controller v0.271.0: the full-system backup's gate defers while the leg runs, until W+5h (then only for a step in flight, cap W+5h30m — decision 31); the leg starts no step at or after W+5h; one shared constant. Unit + red-proof (TestD20_GateWaitsForTheLeg); live on the demo boxes: see the night record Part D. v0.271.0, 2026-09-25 git show 75ff264:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-25/B/redproofs/D20-*
PETI peti-felhom deliberately not migrated; parked until the tester reinstalls. RETIRED 2026-09-25 (operator ruling): the tester wiped his server and the box will not return. Removed through the hub's customer delete (journal #20): the customer record and 1,519 residue rows, Storage Box sub-account u629488-sub2 (id 269130) — which held only one 81-byte authorized_keys, never a repository (all 482 reports: 0 off-site snapshots, 0 bytes); ep0 held nothing of it (no PBS namespace, no WireGuard peer). The audit trail stays by design. retired 2026-09-25 git show 6b2176e:documentation/backlog/OPEN-ITEMS.md; audits/RETIRE-peti-2026-09-25.md
R-686 The automatic update leg is not resumed after a controller restart during the night. RULED 2026-09-25 (operator, option B): the apps the leg had not reached wait for the next night; nothing is built — 09 §3 decision 34. The page-line side effect (a resumed step's last_auto_update not written) stays as measured. ruled 2026-09-25 git show 6b2176e:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-25/C/night3-kill/, night4-power/
R-685 A whole-box backup that cannot fit must say so before it fails (P2). Agent v0.134.0: a named skip before any vzdump (free space from GET /nodes/<n>/storage), proven live with safe builds on demo-hp. Controller v0.272.0: the backup page's tier row says „A teljes rendszermentés nem fér el: %s kell, %s szabad…” / "The full system backup does not fit…" (render test through the real template, both languages). The page line has no live proof: 9202 has no agent and no demo box is short of space now; the operator event rides the existing whole_guest_backup_failed. agent v0.134.0 + controller v0.272.0, 2026-09-25 git show eb1c56a:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-25/F/; audits/retire-peti-2026-09-25/B/
R-671 The undo copies kept by a hold survived the hold's clearing by a restore (P3). v0.272.0: the restore that lifts an update hold removes that hold's undo copies (never a restore hold, never mid-update). Live on 9202: navidrome held with 1 copy kept → the backup page's restore → hold CLEARED, "removed 1 undo cop(y/ies)", 0 copies left, data read back. v0.272.0, 2026-09-25 git show eb1c56a:documentation/backlog/OPEN-ITEMS.md; audits/retire-peti-2026-09-25/B/live/
R-670 Every undo logged a false backup block rejected … docker-compose.yml unreadable (P3). v0.272.0: probe-only copies load with stacks.LoadProbeMetadata. Live on 9202: an undo of navidrome (a backup block in its .felhom.yml) logged 0 such lines. v0.272.0, 2026-09-25 git show eb1c56a:documentation/backlog/OPEN-ITEMS.md; audits/retire-peti-2026-09-25/B/live/r670-verdict.txt
R-677 A re-tested floating tag's badge age read the tag's date (P3). v0.272.0: stacks.BehindSinceAge — a digest-only move counts from tested_at; both producers. Live on 9202: „Frissítés elérhető — ma” / "Update available — today" with catalog_since two days old. v0.272.0, 2026-09-25 git show eb1c56a:documentation/backlog/OPEN-ITEMS.md; audits/retire-peti-2026-09-25/B/live/live272.json

2026-09-25 (evening) — the box converts a PostgreSQL major; kept data (controller v0.273.0 + v0.274.0, hub v0.125.0)

id what closed closed where the full text is
R-657 A reinstall over kept data ran silently into the old files (nextcloud never installed) (P2). Operator ruling 2026-09-25 (09 §3 decision 36), built in v0.274.0: the install asks „use my kept data" / „start fresh" (409 kept_data_choice until chosen); start fresh renames into <drive>/kept/<app>/<date>/; the „Megőrzött adatok" page lists, loads, deletes (typed). Proven live on 9202 in both languages (audits/night-2026-09-26/E/E5-*). CLOSED 2026-09-25 — controller v0.274.0 git show 3386041e6215:documentation/backlog/OPEN-ITEMS.md; audits/night-2026-09-26/E/
R-690 The removed-app restore (R-487) never found a unit kept on a DATA drive (P1). v0.274.0: isStackDeployed instead of GetStackComposePath (true for every catalog app); the R-487 test's fake had answered it for deployed apps only. Proven live: „use my kept data" loaded from scratch_hdd/backups/primary/nextcloud. CLOSED 2026-09-25 — controller v0.274.0 filed and closed this session; audits/night-2026-09-26/E/E1-README.md
R-692 The Kept-data list named two leftovers "Filebrowser" (P2). Found live on 9202 (0.274.0-rc1, 2026-09-25): the read-only view makes the file browser's compose bind every kept folder by absolute path, and the owner lookup took it. No live folder was at risk. Fixed before release in v0.274.0: an owner binds the folder through ${HDD_PATH} (the folder or one inside it), never a protected stack. TestKept_OwnerIsNeverTheFileBrowser, red-proofed. audits/night-2026-09-26/E/ CLOSED 2026-09-25 — fixed in v0.274.0 before it shipped filed and closed this session; audits/night-2026-09-26/E/E5-redproof-owner.txt

2026-09-26/27 — a backup's data and its version travel together (controller v0.275.0, agent v0.135.0, catalog f1a7d6c)

Full original text of each row: git show <the commit that added this section>^:documentation/backlog/OPEN-ITEMS.md.

id what closed closed where the full text is
R-696 Tier 1's "proven at" was the manifest's refresh time; the kept pre-conversion copy was released on a pre-conversion backup (P2). Widened by A1: the refresh re-captured the unit's DEFINITION over the old data, and a restore in that window left a PostgreSQL app down. v0.275.0: every data file stamped with the versions that wrote it; the unit keeps the definition its data belongs to; a restore never mixes versions (unit restores refuse a mismatch, the off-site restore writes the snapshot's definition); every tier's time is its data's; the release needs a dump on the new major. Rule: 07 §6.6. Proven live on 9202 in the box's own night (00:30 data, 02:15 conversion, 02:16 refresh kept the 16 definition; restore back whole at 16; climb again). CLOSED 2026-09-27 — controller v0.275.0 audits/version-travel-2026-09-26/A1/, A5-redproofs/, A5-live/
R-699 The update's precondition accepted a unit holding no data (a just-installed app's) as a fresh copy (P2). Found live on 9202 this session (tandoor). v0.275.0: such a unit stays listed, never a precondition copy on Tier 1/2; the update backs up first (seen live: adventurelog's fresh install got backing-up). CLOSED 2026-09-27 — controller v0.275.0 filed and closed this session; A5-redproofs/RP6-*
R-695 Two kept-data Deletes in one second could leave a self-perpetuating empty kept folder (P3). v0.275.0: the file-browser sync is single-flight; an empty dated kept folder is never listed or bound. Red-proofed. CLOSED 2026-09-27 — controller v0.275.0 audits/version-travel-2026-09-26/D2/
R-694 A load/restore regenerated a withheld login and the page showed it as the password (P3). Measured per app from each entrypoint: 6 of 7 keep the login in their data (code-server is the exception). v0.275.0: restored_logins; the page shows no value and says to use the password valid at the backup. CLOSED 2026-09-27 — controller v0.275.0 audits/version-travel-2026-09-26/D4/
R-689 demo-hp's restore test picked the golden template in local:backup/ and failed every 6 h (P3). Agent v0.135.0 + v0.136.0 + v0.137.0: only vzdump-<type>-<vmid> files or PBS ct/… or vm/… snapshots with a reported vmid, of a guest that still EXISTS on the node (v0.136.0 — the first half, read live with -selftest=restore-test-due, fell to a leftover archive of a guest deleted in August; v0.137.0 — PVE answers 403, not "does not exist", for a guest outside the agent's pool, which 0.136.0 read as a lookup failure). Delivered to both demo hosts by signed agent_update. CLOSED 2026-09-27 — agent v0.137.0 audits/version-travel-2026-09-26/D1/
R-655 adventurelog v0.13.0 could not become healthy in the catalog's template (P2). Operator ruling decision 41 (keep the download). Catalog 06ea7da: the frontend override dropped; a cut-off world-data file set aside before start (a pending import crash-looped 10× without it). Proven on the bench and on 9202 (204 s under the default 5-min wait). CLOSED 2026-09-27 — catalog 06ea7da audits/version-travel-2026-09-26/C/
R-697 A restore dropped the conversion_copy record but not the kept pre-conversion volume — the copy was never released (P3). Seen 2026-09-26 on 9202 (A1). v0.276.0: the restore's write carries the app's life records from the app.yaml it replaces (carryLifeRecords: conversion copies, desired_state, update history; not the pin, not installed_images); a second conversion after a restore to the old major keeps the first copy in earlier_conversion_copies, released by the same rule. Rule: a write that is not a new install is load-then-save, or it names every field it drops. Red-proofed RP1, RP2. CLOSED 2026-09-27 — controller v0.276.0 audits/records-carried-2026-09-27/