5b6ade033f
gates / gates (push) Successful in 23s
- audits/i18n-slice1-2026-09-17/C/: live before/after/en/back on demo-hp (CSRF redacted), hub report hu/en/hu, red-proofs, the switch fixture diff (89 of 89), green gate. - 10-localisation.md: §2.2 executeTemplateLang facts, §2.3 what stays Hungarian + the English test's ASCII blind spot, §3 decision 6 SUPERSEDED 2026-09-17, §5 formal ceiling 16 and its under-count, English retrieval stems; §10 slice 1 done; §11 decision 6 struck. - Register: R-556 closed to CLOSED-ITEMS; R-516 extended; R-565 (ASCII-only Hungarian invisible to the English page test), R-566 (three app-name page titles), R-567 (wizard nav highlight), R-568 (disk rows reorder). 260 -> 263 open. - Capability map row, STATUS (needs you: the floor delivers the switch). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
331 lines
146 KiB
Markdown
331 lines
146 KiB
Markdown
# CLOSED-ITEMS — finished work, compressed
|
||
|
||
> **What this is.** Every register row that reached a terminal state, compressed to its title, the
|
||
> version it shipped in, its evidence paths, and any sentence that states a RULE rather than a
|
||
> narrative. **Nothing was deleted:** each entry names the commit that holds its full original text,
|
||
> and `git show <commit>:documentation/backlog/OPEN-ITEMS.md` returns it verbatim.
|
||
>
|
||
> **Why it exists (operator ruling, 2026-08-22).** `OPEN-ITEMS.md` had grown to 672 KB across 286
|
||
> entries, over half of it finished work, with one single entry at 16 KB. A file that cannot be read
|
||
> is a file that cannot be checked — and this project has already paid for that twice: a record
|
||
> nobody could find because it sat inside an entry about something else, and a finding rediscovered
|
||
> because nobody could see it. The register now holds **open work only**, so its size tracks the work
|
||
> rather than the project's age.
|
||
>
|
||
> **A sibling rather than the bottom of the register**, deliberately: appending to the same file keeps
|
||
> the byte count and the scroll, which is the thing being fixed.
|
||
>
|
||
> **Load-bearing reasoning was NOT compressed away.** Where a closed row states a rule, a fence or a
|
||
> deliberate refusal, that sentence is carried here verbatim under **Reasoning kept**. Rules that
|
||
> outlive their work item also live in their proper homes — `workspace-CLAUDE.md` standing rules,
|
||
> `felhom.eu/CLAUDE.md`, `CONTEXT.md`, and the architecture folder — and this file is not their
|
||
> primary record.
|
||
>
|
||
> **This file is not the register.** Nothing here is open. `OPEN-ITEMS.md` remains the single source
|
||
> of truth for open work; `scripts/one_register_gate.py` enforces that against `ROADMAP.md`.
|
||
|
||
---
|
||
|
||
## 2026-09-17 — localisation slice 1: the whole dashboard in English (controller v0.248.0–v0.250.0)
|
||
|
||
| Row | What | Closed | Full text |
|
||
|---|---|---|---|
|
||
| **R-556** | **Localisation slice 1 — the other 31 dashboard templates in English, Hungarian byte-identical (P3).** Closed in controller **v0.248.0** (`ff68b0b`, apps and settings), **v0.249.0** (`11b790e`, backups) and **v0.250.0** (`ac149c2` + `be2efe5`, storage, sharing, sign-in, claim, guest share, catch-all, debug, and the switch for everyone). Every page's parity fixtures were captured from its UNCONVERTED template before conversion (`6be55a0`, `f8ebc47`, `0555091`); 106 states. `TestI18nParityCoversEveryMarker` (every marker rendered by a case); English page mask narrowed to ≥ 2 words / ≥ 12 chars; retrieval-promise gate scans English; `executeTemplateLang` for the six pages outside the dashboard chrome (no session CSRF, no escrow reminder — `TestI18nDirectRenderPagesHaveNoAdminChrome`); `TestHandlerTitleKeysMatchHungarianTitle`. Decision 6 (switch hidden) superseded: the switch is on every dashboard page; the 89 layout fixtures re-captured and each equals its predecessor plus exactly one switch form. **Proven live on demo-hp 9201** for each release: Hungarian pages equal their previous-version fetch apart from live numbers (and, in C, the switch form); `POST /settings/language` made them English and back; the hub stored `hu`, `en`, `hu`. Startup parse: hu ~24–27 ms, en ~20–28 ms. **Reasoning kept:** a word a page COMPARES stays unconverted until the comparison is language-neutral (R-563); a fixture is never regenerated to make a conversion pass — the two re-captures in this slice (C's case titles, the switch) were captured from the tree that defines the truth and their diffs were proven to be exactly the intended change. Found: R-563, R-564, R-565, R-566, R-567, R-568; R-516 extended. `audits/i18n-slice1-2026-09-17/{A,B,C}/` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show bc15153:documentation/backlog/OPEN-ITEMS.md` |
|
||
|
||
## 2026-09-16 — the drill: the tunnel answers from outside
|
||
|
||
| ID | Title | Shipped | Evidence |
|
||
|---|---|---|---|
|
||
| **R-510** | [P1-HIGH] `tester-1`'s tunnel now has its route, and still gives a fresh box 502: the route sends traffic to `https://traefik` WITH certificate checking, and traefik answers the name `traefik` with its default certificate. | operator ticked „No TLS Verify" 2026-09-15; proven with a box connected 2026-09-16 | three GETs from DooPlex to `https://felhom.enkicsifelhom.hu` at 10:04:19/22/25Z → **302 ×3 (server: cloudflare)**, following it → **200** on the Hungarian claim page „A szerver beállítása … Add meg az e-mailben kapott beállító kódot" (no 502, no 530). `audits/evidence-drill-0243-2026-09-16/phase1-tunnel.txt` |
|
||
|
||
## 2026-09-16 — the drill's Phase 0 (hub v0.115.0, golden 0.243.0, the signing ruling)
|
||
|
||
Two rows closed. **Full original text: `git show <this commit>^ -- documentation/backlog/OPEN-ITEMS.md`.**
|
||
|
||
| ID | Title | Shipped | Evidence |
|
||
|---|---|---|---|
|
||
| **R-529** | [P3-LOW] The agent-plane `host_stale` / `host_down` / `host_recovered` mails still wait out the one-hour quiet rule. | hub v0.115.0 (2026-09-16, operator ruling 2) | `nodeLivenessEvents` + red-proof `TestOperatorCooldown_NodeLivenessBypassesQuietHour`; `08-alarm-ladder.md` §6.2 |
|
||
| **R-533** | [P3-LOW] The operator signing keys were placed on DooPlex world-readable (mode 664) and sit outside any documented location. | operator ruling 1, 2026-09-16 | keys at `/mnt/5_hdd/felhom.eu/felhom-op-{operational,rec-recovery}` + `felhom_op_ed25519`, mode 0600 owner `kisfenyo`; `CONTEXT.md`, `04-control-plane-authorization.md` §3.1 |
|
||
|
||
## 2026-09-15 — the big night's P1 fixes (agent v0.131.0, controller v0.243.0, hub v0.114.0, catalog templates, ISO 1.27.1 published)
|
||
|
||
Nine rows closed. **Full original text: `git show <this commit>^ -- documentation/backlog/OPEN-ITEMS.md`.** Evidence folder: `documentation/audits/evidence-p1fixes-2026-09-15/`.
|
||
|
||
| ID | Title | Shipped | Evidence |
|
||
|---|---|---|---|
|
||
| **R-493** | [P1-HIGH] There are NO customer-facing install instructions, so a volunteer cannot begin — this blocks inviting anyone. | ISO 1.27.1 published + `felhom.eu/letoltes` live | 2026-09-15: round trip over `https://iso.felhom.eu` sha256 `25637007…c053`, 1 705 322 496 B; `felhom.eu/letoltes` 200 naming 1.27.1 (`documentation/tests/iso-release-1.27.1-2026-09-14/README.md`) |
|
||
| **R-495** | [P2-MEDIUM] The public installer asks a stranger four questions nothing answers, and REFUSES its own default on one of them. | answered by the guide; ISO 1.27.1 published 2026-09-15 | same as R-493 |
|
||
| **R-496** | [P2-MEDIUM] The box's console tells a stranger, in English and FIRST, to open the Proxmox admin page — and calls the owner passphrase „a jelszavadat”. | ISO 1.27.1 (console Felhom-only) published 2026-09-15 | same as R-493; G15 live proof on VM 332 |
|
||
| **R-512** | [P2-MEDIUM] Vaultwarden is installed with open registration, and the one control the page tells the customer to use to close it is read-only. | catalog template (no version change), 2026-09-15 | spike: stranger 400 / invite 200 / invited 200 (E1-vaultwarden-spike.txt); live on 9202 from the catalog: `SIGNUPS_ALLOWED=false`, stranger 400 „Registration not allowed" (E1-vaultwarden-9202-live.txt) |
|
||
| **R-513** | [P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.<domain>`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet. | controller v0.243.0 | 9202 (default login): generated, stored encrypted, reveal 200 (16 chars), revealed=200 admin=401 wrong=401; 9201 (hand-set): recorded `operator` 08:52:10Z, untouched; public `files.enkisfelhom.hu` admin → 401 (B1/B4 files) |
|
||
| **R-514** | [P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut". | catalog template: 1 worker × 1 thread, 1280M | live on 9202: 20 PDFs at once → 20/20 SUCCESS, memory.peak 772 370 432 B, no OOM (E2-paperless-9202-live.txt). The OOM-visibility half is NOT proven live → R-528 |
|
||
| **R-515** | [P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password. | catalog template, 2026-09-15 | `default_creds` removed; first steps point at „Automatikusan generált értékek" (app-catalog commit e6aa443) |
|
||
| **R-517** | [P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true. | controller v0.243.0 + agent v0.131.0 | live on 9201: per-tier rows „Helyi tároló (local) ✓ … Naprakész", „PBS ✓ … Naprakész", remote tick on a real PBS success (C4-backup-page-9201.txt). Found live and fixed after the release: an unknown size printed „0 B" → „–" (controller main d3eacbb, unreleased) |
|
||
| **R-523** | [P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule. | agent v0.131.0 + hub v0.114.0 (+ golden script `--restart always`, no bake) | measured first: kill leaves both `unless-stopped` and `always` exited (A1). Live on 9201: idle kill → dashboard 200 in 59 s; parked → stayed dead 100 s with the PARKED line, unpark → 200 in 25 s; kill during a swap → supervisor deferred ×3, swap rolled back itself; the crash-loop guard tripped for real after 3 test restarts and the hub mailed `controller_crashloop` (A4-* files). NOT measured: restart timing during a deploy (the budget was spent) → R-531 |
|
||
|
||
| **R-442** | **`remove_hdd_data: true` was INERT — the customer's data stayed on the drive while the API reported success (HTTP 200, `hdd_paths_removed: null`, 128 MB of Nextcloud left; demo-hp 2026-09-01).** Shipped in controller **v0.236.0** (2026-09-13). Removal resolved the drive from the GLOBAL `cfg.Paths.HDDPath` (no default, set on NO box — demo-hp AND demo-felhom both measured 0 `hdd_path` / 0 `FELHOM_PATHS_*`, so the fleet shares the shape); it now reads the app's OWN `app.yaml` `HDD_PATH` — the `07-backup-architecture.md` ~L437 rule that deploy, the start gate and the backup destination already implemented — and a data removal it cannot resolve, or whose drive is absent, is REFUSED (409, exact Hungarian sentence, typed `stacks.RemoveRefusedError`) BEFORE `compose down`, app kept. SSD app → `hdd_paths_removed: []`, never `null`, plus `hdd_note`. Missing folders stated in `hdd_paths_missing`. The backup-half refusal reaches the response (`backup_paths_refused`) and its base follows the same rule. **Reasoning kept:** *"declares no drive" ≠ "could not resolve the drive" — the first is a fact, the second a refusal*; *an app gone with its data left behind is unrecoverable from the UI — the customer cannot even re-run the removal*; *no fallback to the global — that silent fallback is the exact path this closes*. **Observation carried:** 8 of the 13 `needs_hdd` catalog apps bind ONLY `${USERDATA_PATH}` (the shared library) and no `${HDD_PATH}` folder, so for them "delete my data" correctly removes nothing on the drive and the modal shows no checkbox. | **CLOSED 2026-09-13 — shipped controller v0.236.0, proven live on demo-hp** | `audits/R442-2026-09-13/` — A: 63 MB written by Nextcloud ITSELF, gone after removal and listed with its size; C: 409 + sentence, all 69 files untouched, app still deployed, `[ERROR] … refused` logged; D: gokapi `[]` + note; ASCII controls (`llap`: C=2 A=0 D=0). 15 tests + two red-proofs in `felhom-controller/REPORT.md`. Full original text: `git show d6837d98ee24:documentation/backlog/OPEN-ITEMS.md`. |
|
||
| **R-449** | **UPDATE ARC SLICE 5 — an upgrade test that runs again. BUILT AND RUN 2026-09-06:** `app-catalog-felhom.eu/scripts/upgrade-test.py` + `upgrade_fixtures.py`, 7 edges across 3 apps, evidence in `audits/upgrade-spike-2026-09-06/`. **Reasoning kept — success is an APPLICATION-LEVEL READBACK, never file identity:** `survive2.py`'s sha256+inode rule is right for a redeploy and WRONG for an upgrade, because a migration is supposed to rewrite files and that rule would fail every correct upgrade. **Reasoning kept — nothing is ever seeded into a volume by hand (R-156);** an app with no non-browser route is recorded `inconclusive`, which is a result and not a licence to plant a file. **Reasoning kept — run the negative control FIRST:** C3's TO image exits immediately and came back `failed`; a harness that cannot fail a known-broken upgrade proves nothing with its greens. **WHAT IT MEASURED:** all five real catalog upgrades kept the customer's data; and **whether an upgrade can be undone is a property of the individual APP, not of upgrades** — docmost REFUSES (*"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*), privatebin does not, which independently reproduces the Nextcloud finding on a second app by a DIFFERENT mechanism and puts two measurements behind §4's ruling that "rollback" is the wrong word. **It also found a defect in our own catalog (R-459).** **What stays open, as its own rows rather than inside this one:** R-459 (the skipped MariaDB datadir upgrade), R-460 (bookstack's file half is unprovable headlessly), R-462 (the widening, costed). Full original text: `git show 417df06f3529:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — harness built, run, and proven by a red negative control** | `audits/SPIKE-upgrade-test-2026-09-06.md`; `audits/upgrade-spike-2026-09-06/evidence/`; catalog `0474ce387e6f` |
|
||
| **R-438** | **The catalog sync rewrote a DEPLOYED app's `docker-compose.yml` and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06).** `Syncer.copyTemplates` copied into every stack folder on a 15-minute cycle with **no deployed check**, so a deployed app's file and its running containers disagreed from that moment, and the next `compose up -d` from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s **with a pull**, and a boot reconciliation upgraded it **with nobody pressing anything**. **Reasoning kept — the distinction this row existed to protect:** `RestartStack`'s use of `up -d` to pick up template changes was **CHOSEN and written down in its own comment**, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. **Reasoning kept — one fear was measured SMALLER than stated:** a plain power cut does NOT upgrade anything, because Docker's `restart: unless-stopped` restores the containers on the old image and the reconciler logs `no boot-orphaned apps`; the unattended upgrade needs the narrower precondition *"and the app did not come back"*. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0** | `audits/SPIKE-app-update-2026-09-01.md` §2, §3, §8; `architecture/09-update-architecture.md`; `tests/VALIDATION-update-slice3-2026-09-06.md` |
|
||
| **R-447** | **UPDATE ARC SLICE 3 — the live compose file is now DERIVED; the syncer renders instead of copying. SHIPPED controller v0.235.0 (2026-09-06).** The pin lives in `app.yaml` (`pinned_images`), the definition it came from is stored beside the app as `applied-compose.yml`, and `Syncer.renderSource` writes the catalog template verbatim while the catalog still offers the pinned version and the stored definition once it moves past it. **Reasoning kept — the ruling, in the operator's own words:** *while the catalog is offering the same version you are running, its fixes flow to you; the moment it moves to a newer version, you are frozen at what you have until you choose to update.* **Reasoning kept — why nothing was added to the thirteen `compose up -d` call sites:** most of them are REPAIRS (the boot reconciler, the drive-return gate, the app-stop guard), and **a repair path that refuses to repair leaves a customer's app down, which is worse than the problem**; they were made safe by removing the reason, not by gating them. **Reasoning kept — why the frozen branch writes a WHOLE file and never a substitution:** `wger 2.6` needs a full DB configuration the older template cannot supply, so an old image under a new template is a third state nobody chose. **Reasoning kept — why this is not "skip deployed apps" (option B, rejected):** that also stops health-check fixes, memory limits and new deploy fields, and destroys the self-healing measured live in the spike §3. **Reasoning kept — `pinned_images` is INTENT and `installed_images` is an OBSERVATION; never feed one from the other** (the R-166 category error, one field over). Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — SHIPPED controller v0.235.0** | `architecture/09-update-architecture.md` §3.4, §5; `tests/VALIDATION-update-slice3-2026-09-06.md`; controller `CHANGELOG.md` v0.235.0 |
|
||
| **R-441** | **The restore path and the catalog sync disagreed about the image, and the SYNC WON within 15 minutes. CLOSED controller v0.235.0 (2026-09-06).** `stackAdapter.RecreateStackDefinitionFromUnit` wrote the recovery unit's captured `docker-compose.yml` — carrying the OLD pin — into the live stack dir, and `Syncer.copyIfChanged` overwrote it from the catalog on the next tick, so a restore's image-level recovery had a **<=15-minute half-life**. The restore now PINS to what the unit captured and stores it as the applied definition, so the render obeys the restored file instead. **Reasoning kept:** the overwrite half was MEASURED live 2026-09-01 (`[INFO] [sync] Updated bentopdf/docker-compose.yml` at 18:10:29Z over a locally-modified file); the "the restore writes to that same path" half was READ, and this closure rests on the render behaviour being measured live rather than on the reading. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — SHIPPED controller v0.235.0** | `internal/stacks/pin.go`; `TestGroupF_RestorePinIsReportedToTheSyncer`; `tests/VALIDATION-update-slice3-2026-09-06.md` |
|
||
| **R-455** | **DooPlex had no Docker Hub login, and the unauthenticated ceiling blocked a BUILD rather than only a gate. CLOSED 2026-09-06 — a Docker Hub PAT was added to the credentials file (operator).** Measured 2026-09-02: `build.sh 0.233.0 --push` failed at `429 Too Many Requests` resolving `debian:bookworm-slim`, with neither base image in the local store. **Reasoning kept — the workaround and its verification, because the answer outlives the incident:** both bases were pulled from Google's official Docker Hub mirror (`mirror.gcr.io/library/...`) and retagged, and the identity claim was later MEASURED rather than assumed — `docker pull docker.io/library/<img>` answered `Status: Image is up to date` for both, i.e. Hub's own manifest resolved to the images already local, and the `docker manifest inspect` bodies were identical between the registries. **So the mirror is a verified-sound fallback if the PAT is ever unavailable.** Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — credential added by the operator** | `felhom-controller/REPORT.md` (2026-09-03) §5 |
|
||
| **R-429** | **CORRECTED 2026-09-01 — the snapshots ARE being taken; what was broken was that nobody could tell, and my own probe looked for the wrong name.** Viktor read the control panel on 2026-09-01: **seven automatic daily snapshots** on `storage-box-pool-1` (plan BX11), six days old to ~9 h old, filesystem ~2.5 GB, per-snapshot 0–14 MB, *Display snapshot directory* ON. **The mitigation works.** **MY ERROR, NAMED:** the 2026-09-01 spike probed for a directory called `.snapshots`; the vendor documents the path as **`/.zfs/snapshot`**. The probe's controls were sound and its subject was wrong, so `not found` was true and meant nothing. **The task that set the spike asserted the mechanism without citing the vendor documentation, and I did not check it** — that is how a correct instrument produced a wrong headline. **WHAT REMAINS TRUE, and is the actual finding:** the row claiming it had no R-number so nothing could cite it; its "confirm tomorrow" went **36 days** unanswered; and the `DUE-CHECKS` block built for that exact class (R-341) was **empty**. **The finding was never the snapshots. It was that nobody could tell.** Re-probed at the documented path — see R-432 for what a sub-account can actually reach. Evidence: `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1 and this row. | **CLOSED 2026-09-01 — mitigation CONFIRMED WORKING; the visibility gap is R-432** | panel read 2026-09-01 (Viktor); re-probe at `/.zfs/snapshot` in `audits/SPIKE-r95-offsite-delete-2026-09-01.md` |
|
||
| **R-431** | **An unexplained fall in a customer's off-site snapshot count is now noticed within a day — SHIPPED hub v0.111.0.** Third signal in `OffsiteChecker`, beside FILL and STALENESS. **It lives on the HUB deliberately:** the event being detected is a box deleting its own backups, so a detector on that box is one the same event can silence; the hub already receives the count and keeps the history. **Threshold: a fall of more than HALF the previous count and at least 5** — reasoned, not invented, because the measurement gave nothing to calibrate against: over 12 898 reports (2026-06-05 → 2026-09-01) every one of the nine decreases lands exactly on ZERO and every one predates `stats_known` (the R-331 shape), and in the 380-report window where `stats_known` is true there are **zero** decreases. Retention keeps 7 daily + 4 weekly + 6 monthly per group, so it **cannot halve a total**; a mass deletion goes to ~0. **Three pre-conditions, each with its scar:** `StatsKnown` (R-331 — a zero is not a zero when unmeasured), the declared `State` (R-204 — the box names its own situation), and run success (R-100 — presence is not success; `incomplete` excluded too). An untrustworthy report neither alarms nor moves the baseline. **ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms**, fixture committed. Severity `error`, operator-only, no customer template. | **CLOSED 2026-09-01 — SHIPPED hub v0.111.0** | `hub/internal/monitor/offsite.go`; `offsite_r431_test.go` incl. 9 009-point real-history replay; hub v0.111.0 |
|
||
| **R-419** | **`observations_gate.py` accepted an observation whose body merely CONTAINED the string `NOT-A-FINDING`, even in prose disclaiming it.** Found by accident on 2026-09-01 when a planted test observation reading *"it carries no `FILED:` and no `NOT-A-FINDING:` marker"* was reported `OK 1. NOT-A-FINDING` and a real push went green over an unfiled finding. The gate's whole job is to force an explicit choice, and a sentence disclaiming the choice counted as making it. **FIXED 2026-09-01:** a marker must now start a line or follow a sentence boundary, and inline code spans are stripped before matching — a marker inside backticks is being talked about, never used. | **CLOSED — FIXED + PINNED** (2026-09-01, R-421 sweep) | verified in BOTH directions: the decoy and a backticked mention are convicted; a real `**FILED: R-419**` and a real `**NOT-A-FINDING: ...**` still pass. Decoy kept in `felhom.eu/scripts/test_gate_decoys.py` |
|
||
| **R-404** | **DECISION — should a documents-only push be subject to the golden-currency gate? RULED 2026-09-01: NEITHER option as framed. Block the push that can create the debt; notify the push that cannot.** The two options on the table were *narrow the gate* and *leave it and build a waiver*, and both were wrong for the same reason: they argued about the GATE, and the gate was never the problem. **The DIAGNOSIS, measured from live source, is that the check was aimed at the wrong repository.** `golden_currency_gate.py` never looks at the push at all — it compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing, which is correct for a standing invariant and wrong as a push gate. Meanwhile `controller_gates.py` had NO golden-currency entry, so **the repo where a release happens never checked, and the repo that cannot create the debt enforced it on every push.** 18 of the last 24 pushes here touched no code — MEASURED, and the classifier agrees exactly — most of them for a structural reason: the controller's code is in one repo and its register, architecture and status live in this one, so **every controller change produces a documents-only push here by construction.** Six of those 18 were bake records — **the push that PAYS the debt is itself documents-only, so the gate was blocking its own cure.** **WHY NOT THE WAIVER** the gate's own docstring prescribes: that clause was written for *a release nobody wants a golden for*. The case that actually occurred (R-417) was *a release we did want a golden for, on a night the runbook forbade baking*. A waiver would have recorded a lie. **SHIPPED:** `scripts/push_scope.py` (allow-list; every uncertainty answers `code`), a fifth `exemptible` field in `repo_gates.py` + `--scope`, a new **ADVISORY** verdict printed in its own block, the pre-push hook reading git's stdin, the same rule in CI from the push event payload, and `felhom-controller/controller/scripts/golden_notice.py` — a NON-BLOCKING notice at the moment a release is committed. **The gate's own logic, exit codes and wording are byte-identical**; only the consequence changed. The exemption is ONE gate wide and `TestR3`/Scenario C pins it, red-proved by widening it. Proven live on the real hook: docs+debt → ADVISORY, pushed; code+debt → refused; docs+debt+a second gate → refused for that gate alone | **CLOSED 2026-09-01** — ruled and shipped | full text and the ruling: `git show 1e6c387:documentation/backlog/CLOSED-ITEMS.md`; the diagnosis is in `CONTEXT.md` and `scripts/CHANGELOG.md`; `scripts/push_scope.py` + `test_repo_gates_scope.py` |
|
||
| **R-417** | **A drill night that forbids baking a golden made `golden_currency_gate.py` red, so pushing the drill's own evidence needed `--no-verify` — the very signal CI e-mails about.** Measured 2026-09-01: five consecutive felhom.eu CI runs red (jobs 469/470/471/473/476), all mine, all on step 3 `Run the gate entry point`; job 478 green the moment the golden-0.232.0 evidence was committed. Cause confirmed by isolation — moving that directory aside reproduces exit=1, restoring it gives exit=0. **The gate was right every time**: 0.231.0 and 0.232.0 were released with no golden carrying them. **CAUSE REMOVED, not worked around** (R-404): a drill's pushes are documents-only, so the conviction now prints as a loud ADVISORY and the push proceeds — in the hook AND in CI, so a drill night no longer produces red runs indistinguishable from real ones. The expectation is now written where the next drill author reads it (`documentation/runbooks/target-selection.md`), which is the half I had left out. | **CLOSED 2026-09-01** — by R-404 | five red CI jobs 469/470/471/473/476, green at 478; reproduced by isolation (`documentation/audits/AUDIT-gate-decoys-2026-09-01.md` records the technique) |
|
||
| **R-361** | **The pre-restore safety dump overwrote the app's own DB dump, and the comment beside it said it could not.** Shipped in controller v0.221.0 (+v0.221.1). Evidence: `audits/DRILL-r361-2026-08-22/evidence/`. **Reasoning kept:** *`DumpOne` writes `<stack>-<dbtype>.sql` — the app's canonical dump, the name the replay loop matches EXACTLY — so nothing else may ever be written to it.* The fix is a DESTINATION, not a rename: `DumpOneTo` takes the final path and derives its own `.tmp` from it, so neither the destination nor the scratch file can collide with a nightly dump running beside it. **`DumpOne`'s signature did not move** — it has callers outside this concern. **The manifest no longer lists the undo copies:** every consumer of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`, none reading it for recovery. **AND THAT CHANGE MADE ANOTHER UNREACHABLE:** a stable `db_dumps` let `CaptureRecoveryUnit`'s already-current early return fire, and the undo-copy prune sat after it — four copies on disk against a cap of three, counted live. The prune now runs ABOVE the check; it is housekeeping on the dump directory and is independent of whether the manifest needs rewriting. **PROVEN LIVE the only way it can be:** the canonical dump's sha256, unchanged across a restore — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`, both byte-identical before and after. A test asserting merely that the undo copy exists passes just as well when the app's backup was destroyed. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.221.1, 2026-08-23) | full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-379** | **The pre-restore undo copy was valid, was named to the customer, and no product action could apply it.** Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/`. **Reasoning kept:** *R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken.* **The undo set is matched on THE RUN'S OWN STAMP, never on the `pre-restore-` prefix** (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) **and never just the first file** (a two-database app would have had one restored and the other left half-written). **The rollback RE-DISCOVERS the container** — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of `waitDBReady` against a dead id while its data was recoverable. **No unit test saw that: they all inject the import seam and never look at container identity.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-380** | **A failed MariaDB replay left a partially-applied database behind an app reporting `health=healthy`.** Shipped in controller v0.220.0. Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt`. **Reasoning kept:** **no engine flag closes this** — `--single-transaction` was added to the Postgres import and does make it all-or-nothing, but **MariaDB's DDL is not transactional**, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: `bookstack`'s `migrations` table back at **102 rows**, the exact cell the defect was measured in. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-381** | **The restore-failure message pasted raw engine stderr — including rows out of the customer's own database — into the Hungarian customer surface.** Shipped in controller v0.220.0. **Reasoning kept:** the full engine text now goes to the operator log, **which never had it before — the diagnostic was ADDED, not removed**. Measured: 407 bytes (Postgres) and 615 (MariaDB, whose middle was an `INSERT INTO migrations VALUES (…)` listing); now 257 bytes with no engine tokens. **A red-proof for this PASSED and the test was hollow**: it injected below `ImportDump`, so a leak reintroduced inside `ImportDump` could not fail it. The guard now sits at that layer. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-382** | **The reconstitution's summary log line omitted the volume count it already held.** Shipped in controller v0.220.0. Proven live: `0 file(s) placed, 3 volume(s) replayed, 1 DB dump(s) replayed`. | **CLOSED — SHIPPED** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-356** | **The off-site restore refused every app that has no data drive — it asked "does this app have an HDD path?" to answer "is this app installed?", and for 40 of 53 catalogue apps the honest answer to the first is permanently no.** Shipped in controller v0.219.0. Evidence: `audits/DRILL-r356-hot-only-restore-2026-08-22/evidence/`. **Reasoning kept:** *the restore destination is resolved by the SAME rule as the capture destination — the drive if the app has one, the system data path otherwise (`Manager.GetAppDrivePath`, one expression). The refusal that protects a drive app from being restored onto the wrong disk applies to apps that HAVE a drive to get wrong.* **An app with no drive is not misconfigured** — `01-topology-and-trust.md` §8 carries the `[DESIGN]` marker; between 19 and 22 August that design was called a defect four times. **Deployment is asked of `ListDeployedStacks()` and FAILS CLOSED on a nil provider:** "cannot tell" must not become "go ahead" when the caller's next act is a write. **Two different failures get two different sentences** — installed-but-no-resolvable-data-root has its own refusal and its own route; widening `nincs telepítve` to cover it would send a customer to reinstall a running app and hide the real fault. **Measured, and load-bearing: 53 templates, 13 `needs_hdd: true`, 40 `false`** (catalogue @ `459766cb1639`). **The capture side's raw `GetStackHDDPath` is FENCED and was not changed** — capture resolves an app's declared `userdata`/`import` file legs against that value, and a system-data fallback there would write a snapshot claiming to hold files it does not. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.219.0, 2026-08-22; `privatebin` on `demo-hp`: planted, backed up, deleted, restored, 15/15 files byte-identical including two Hungarian accented names) | full text: `git show e18668f9e19f:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-216** | **A correct recovery code was reported to the customer as wrong.** Shipped in 0.120.0, v0.125.0. | **SHIPPED** (controller v0.201.0 + hub v0.97.0/0.97.1) — **but see R-223**: the feature does not work on a NEW box until the manifest vouches agent 0.125.0. Until then such a box is correctly HELD, not lied to | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-218** | **Succeeding at recovery stopped the box asking for what it still needed.** Shipped in v0.203.0. Evidence: `documentation/tests/part4-rewalk-2026-08-06/journal.md`. | **CLOSED 2026-08-06 — controller v0.203.0, proven live.** *(State corrected 2026-08-06: this field read REOPENED while the body below already recorded the fix shipped and proven. The history of the over-claim is kept deliberately — it is why the row is worded as it is.)* **The over-claim, as it stood: the fix covered the DECLARATION half only.** Measured on the R-201 re-walk: the box declared, and **`offsiteheal` re-staged the secret at 11:44:57** saying *"the box re-consumes on its next cycle"* — **the next cycle came and went** (`host-report` 11:55:46, `Received report` 11:55:54, a full cycle **with a positive control that it ran**) **and the credential was still not consumed.** 23 minutes after the re-stage the box's last off-site-apply attempt was still the pre-re-stage one. A census of the customer-reachable actions on `/backups/remote` (`config`, `reset`, `run`, `toggle`) found **none that fetches a staged credential**, and the only lever is `systemctl restart felhom-controller-bootstrap.service` **inside the guest** — which worked in **18 s** (Campaign 11 measured 17), confirming nothing was wrong with the credential, the target or the key: **the only thing missing is anything at all to trigger a retry.** **This is the FIRST of the two dead ends that keep the recovery journey failing** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-219** | **The listing the screen promises could never render on the shape it exists for.** | **SHIPPED** (controller v0.201.0) — the unlock now places the key, brings the tier up, then lists | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-217** | **An unreadable store reported as "opened, with unattributable content".** | **SHIPPED** (controller v0.201.0) — opened / empty / unreadable are three distinguishable states | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-222** | **Reaching for a RETAINED earlier package read as a wrong code.** | **SHIPPED** (controller v0.201.0 + hub v0.97.0) — the ACK carries `superseded_present`/`superseded_at` and the screen names the situation. **It states what the hub knows and promises nothing** — the read path is still unbuilt (R-199's inventory) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-215** | **`GET /recovery` rendered the recovery story on a box that never had off-site backups.** | **SHIPPED** (controller v0.201.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-220** | **After a rebuild the customer's drives cannot be re-enrolled, and the refusal names an impossible action.** | **CLOSED 2026-08-06 — shipped in agent v0.127.0 and PROVEN LIVE on a genuinely rebuilt box.** The fix is **corroborated, not a widened prefix**: a mountpoint outside `/mnt/felhom-drives` is forgiven only when the SAME device is also mounted under the managed path — a pairing only Felhom's own enrolment produces, so a disk another system is using at `/srv/data` or even `/mnt/someone-elses-disk` is still refused (own test + red-proof). Read from `/proc/mounts` deliberately: the lsblk invocation is pinned verbatim in the sudoers file, so switching to plural `MOUNTPOINTS` would have shipped a sudoers change with the binary. Fail-safe: an unreadable mount table corroborates nothing. **Measured on the Part 4 venue after a real guest purge, with both raw mounts still present on the surviving host:** `/disks/candidates` returned both drives in `attach` and `initialize` (before the fix: two empty lists), and both **re-attached through the customer endpoint** (`registered: true`). The customer-facing refusal was corrected in controller v0.203.0. | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-221** | **A rebuilt box cannot run the escrow ceremony at all.** | **CLOSED 2026-08-08 — agent v0.128.0.** `Apply` re-asserts the seed BEFORE the idempotent early return; the return itself is kept and pinned by a zero-Proxmox-calls assertion. **The writer was established at `file:line` rather than assumed** — see the follow-through section below | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-223** | **The Day-0 manifest vouched agent 0.120.0 while the recovery feature needs 0.125.0 — and a reinstall DOWNGRADES a box that was fixed by hand.** Shipped in 0.120.0, 0.125.0, 0.192.0. | **CLOSED 2026-08-05.** Golden **0.201.0** baked in the drill VM (658 165 766 B, sha `e730d7cab343eb35…f007654`, **round-trip verified from Gitea**), then manifest set in one save: `agent=0.125.0 golden=0.201.0 min_agent=0.125.0`. A fresh install now lands on current agent AND current controller | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-224** | **Every non-code failure on the unlock path is reported to the customer as a statement about their code.** **Reasoning kept:** **And R-216's gate cannot catch it**: the box's own ring reads `recovery capability gate: offsite_key_recovery=yes (source=version)` — the gate discriminates the agent's **age**, not its **reachability**, so a dead agent of the right version sails through the guard whose own comment says *"An attemp | **CLOSED 2026-08-06 — controller v0.202.0 + agent v0.126.0.** The discriminator is now a VALUE: `escrow.ErrBundleFetch` → **HTTP 502** at the agent, `agentapi.RecoveryRefusal` carrying the status at the controller, and `ClassifyRecoveryFailure` mapping it to one of five classes **from the value, never the text**. **PROVEN LIVE on the venue**, same wrong code, only the hub's reachability changed: `hub up → 400 "…did not open the sealed bundle"` · `hub REJECTed → 502 "…could not be fetched — the recovery code was NOT used"` · `hub restored → 400`. Red-proof: deleting the agent case reproduces `got 400, want 502` with the wrong-code sentence. **Coupled `MinAgent 0.126.0`** — an older agent answers 400 for both causes, so the reading is withheld and the 400 degrades to NEUTRAL; the gate blocks nothing. **The customer-facing messages were NOT re-driven end-to-end**: `/recovery` correctly redirects since F7 set the old data aside, and restoring that state is the reconfiguration §11 forbids — they are covered by handler tests + red-proofs | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-225** | **The remote store reports `0 pillanatkép · 0 / 50 GB` when the box cannot read it — directly above a card stating the store holds backups.** | **CLOSED 2026-08-06 — controller v0.202.0.** `StatsKnown` is a **named** state (the `OffsiteInventory.Empty` pattern), because zero is what an unread store and an empty one both look like and `omitempty` makes "absent" and "0" the same bytes. The fill bar renders only when the fill is known — a 0 %-wide bar is a picture of emptiness. **PROVEN LIVE both ways**: before a run the venue read „a pillanatképek száma még ismeretlen"; after one, „2 pillanatkép … / 50 GB". A measured zero still says zero | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-226** | **M1 — the only message that tells a customer to check their typing — is unreachable on any box that has re-escrowed.** | **CLOSED 2026-08-06 — controller v0.202.0.** The retained-package message now names **both** possibilities and restores the ten-words prompt, because the two are indistinguishable at the engine and saying so is the honest thing. It still does not promise the earlier package can be opened. Red-proof: removing the clause makes the prompt unreachable again | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-228** | **After „I do not want the old data", the set-aside history becomes invisible — the box records where it is and shows it to nobody.** | **CLOSED 2026-08-06 — controller v0.202.0.** `OrphanedRenamedTo` is surfaced as two facts and stops. **It does not promise the history can be reopened** — it cannot be, by anyone, today (R-199's inventory is unbuilt) — and the set-aside **confirmation copy was corrected** for the same reason: *"a helyreállítási kód nélkül többé nem lesznek megnyithatók"* implied that WITH the code they could be. The field's own comment said "recovery-code-recoverable", the same over-promise in the code. **PROVEN LIVE**: the notice renders on the venue | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-227** | **A controller restart mid-unlock returns a raw English `Bad Gateway`.** | **CLOSED 2026-08-06 — controller v0.202.0, partially and stated as such.** **The layer that answers is traefik**, whose config this repo generates — but traefik v3 serves no static files, so a branded proxy page needs a **new always-up container** for every 502 on the box: **scoped, not built**. Shipped: the unlock posts via `fetch` and answers a gateway failure in Hungarian in-page. **Progressive enhancement — with no JS the plain POST still shows the proxy's error** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-234** | **An off-site run reports success while silently omitting an app the customer just switched on.** Shipped in v0.205.0. | **CLOSED 2026-08-06** — controller v0.205.0 | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-236** | **After a guest rebuild the hub never re-stages the off-site credential.** | **CLOSED 2026-08-06 — not a defect** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-237** | **After a successful recovery the customer is shown no backups at all, because the restore surface is keyed on apps that are currently installed and currently marked for future remote backup.** | **CLOSED 2026-08-06** — controller v0.204.0: the list is now built from `OffsiteInventoryList` (the repository's own snapshot tags). Installed-ness became a property OF a row, never a filter; an unreadable store renders as UNKNOWN **and keeps the action offered**; `felhom-offbox` and `_shares` are excluded. 7 new tests incl. a rendered-page test for the rebuilt shape, and a red-proof that keys the list back on installed-and-toggled apps. | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-238** | **„Teljes visszaállítás előkészítése" accepts the click and does nothing.** Shipped in v0.204.0. | **CLOSED 2026-08-06** — controller v0.204.0 | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-239** | **The fixes are written, tested, pushed — and a machine installed tonight gets none of them.** Shipped in 0.127.0, 0.203.0, 0.204.0. Evidence: `tests/finalwalk-r201-2026-08-07/journal.md`, `tests/golden-0.205.0-2026-08-07/`. | **CLOSED 2026-08-07** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-241** | **The credential self-heal, succeeding, locks the customer out of their own recovery.** Shipped in v0.206.0, v0.98.0. Evidence: `audits/SPIKE-r241-recovery-offer-2026-08-07.md`, `tests/finalwalk-r201-2026-08-07/journal.md`. | **FIXED 2026-08-07 — v0.206.0 / hub v0.98.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-247** | **The box is being told something false, in its own words, and it recommends the destructive act.** Shipped in v0.206.0. | **CLOSED 2026-08-08** — controller v0.209.0. The field is received and the box tells the two conditions apart; see the Campaign-12 follow-through section below. The WRONG FLAG itself is R-246 (operator act, hub-side) and the customer-facing card copy is unchanged — both stated rather than folded in | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-249** | **The retrieval passphrase ships in the customer page's HTML, so any headless read puts it in a transcript.** Shipped in v0.206.0, v0.207.0. **Reasoning kept:** **Severity MEDIUM:** it is a live per-customer secret that fetches the whole config (`GET /api/v1/config/<id>` with `X-Retrieval-Password`), but the exposure is to someone who can already read the operator page — a defence-in-depth failure, not a boundary crossed. | **CLOSED 2026-08-08** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-252** | **After a rebuild the restore refuses because the data drives are not registered, and nothing on the recovery path says so.** Shipped in v0.207.0. | **CLOSED 2026-08-08** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-253** | **The restore page promises it will reinstall the app, and the restore then refuses because the app is not installed — in the customer's own language, three lines apart.** Shipped in v0.207.0. | **CLOSED 2026-08-08** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-212** | **The orphaned-ciphertext deletion HALTED: the stores on the storage box do not match this register's record.** | **CLOSED 2026-08-05 — all three deleted after the operator confirmed the corrected list** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-94** | **~~A hand-synced version constant drifts, and the gate that would catch it is never run~~** Shipped in 1.22.0, 9.9.9. | **CLOSED — SHIPPED** (hub v0.87.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-110** | **`main` is the installer's publish channel — there is no staging.** Shipped in 1.22.0, v1.23.0, v4.4.0. **Reasoning kept:** E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly. Fixing only (i) leaves a tagged installer pulling nine untagged files from `main` at run time — a staging story that is false in the place it matters most, since one of those nine (`felhom-backup-target-apply`) is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only b | **CLOSED — SHIPPED** (installer v1.23.0, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-111** | **The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.** Shipped in 0.113.0, 0.114.0, 0.161.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`. **Reasoning kept:** **The global controller floor was deliberately NOT raised**: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. | **SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** Shipped in 0.113.0, 0.114.0, 0.119.0. **Reasoning kept:** **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. It calls the existing `publish-agent.sh` rather than reimplementing it, refuses a dirty or unpushed tree, refuses to re-release an existing version (one version name must never mean two binaries), and **deliberately does not vouch** — vouching points machines at a version and stays the operator's ac | **CLOSED — SHIPPED** (`release-agent.sh` + `check-published-versions.py`, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-116** | **The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC `storage_disconnected`, return the SPECIFIC `backup_target_restored`; `backup_target_absent` never fired at all** Shipped in 0.185.1, v0.115.0, v1.25.0. Evidence: `audits/R116-v0116-2026-07-30.md`, `audits/SPIKE-r117-bind-liveness-2026-07-30.md`. | **SHIPPED + PROVEN-LIVE** (agent v0.116.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-120** | **The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message** Shipped in 0.113.0, 0.116.0, 0.156.0. Evidence: `audits/R120-golden-rebake-2026-07-30.md`. | **CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate** (golden 0.186.0 + hub v0.82.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-117** | **A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered.** Shipped in 0.113.0, 0.117.0. Evidence: `audits/R117-v0117-2026-07-30.md`. **Reasoning kept:** **No block I/O proven by strace** (only `/proc/self/mountinfo`, **0** statfs) — the Part 1 `CLAUDE.md` fence applied to its own first consumer. | **SHIPPED + PROVEN-LIVE** (agent **v0.117.0**, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-113** | **The drive-absent gate CANNOT FIRE on device loss — E-2b's alarm is wired to an unreachable condition.** Shipped in 0.113.0, 0.114.0, v0.185.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (agent v0.114.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-112** | **E-2's degraded banner and offer have NO UI CONSUMER — the endpoint is correct and the customer never sees it.** Shipped in v0.185.1. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that just vanished.** Shipped in 0.113.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** Shipped in 1.19.0, 1.22.0, v0.129.0. | **CLOSED — both halves shipped** (2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-86** | **Restore-tests are interval-scheduled, not backup-aligned** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.121.0**, hub **v0.91.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-185** | **The agent cannot see the host backup tier's archives on demo-felhom — the PVE token has no ACL on `/storage/felhom-backup`, so the content listing returns EMPTY where root sees three archives.** Shipped in v0.123.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.123.0**, installer **1.24.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-186** | **A released agent binary's sha256 cannot be reproduced from its tag.** Shipped in v0.120.1, v0.121.0, v0.121.2. | **CLOSED — SHIPPED + MEASURED 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-187** | **R-115's one-command release had never actually run its publish leg — the first real use died there.** Shipped in v0.121.0. | **CLOSED — SHIPPED 2026-08-03** (`felhom-agent`) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-188** | **Every agent release has a ~50 % chance of emailing the operator a CI failure for a release that is correct.** Shipped in 0.121.2, v0.121.0, v0.121.1. | **CLOSED — SHIPPED 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-189** | **A passing restore-test can be invisible to the hub forever — and R-86 made that window a week instead of a day.** Shipped in v0.121.1. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-195** | **A customer with no machine ever bound e-mailed an `expected_dbdump_missed` ERROR every morning.** Shipped in v0.73.0. **Reasoning kept:** Fail-**open** on a read error (an unreadable binding must never SUPPRESS a real alarm), and the deferral is LOGGED with its own counter (the v0.73.0 Part-7 precedent: a quiet check must not look like a check that did not run). | **SHIPPED** (hub **v0.92.0**, 2026-08-04) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-196** | **`escrow_stale` is wired to the ONE path that does not change the repo password, and absent from the path that does.** Shipped in v0.95.0. Evidence: `audits/DRILL-r201-night-run-2026-08-04.md`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. | **CLOSED 2026-08-05 — hub v0.95.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-197** | **The hub holds both halves of the evidence that a box's offsite DATA key changed, and reads neither.** Shipped in v0.78.0, v0.93.0. Evidence: `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. **Reasoning kept:** The in-between shapes (a first-ever hash, a hash-less supersession) are LOGGED rather than dropped, so *"we chose not to alarm"* and *"the check did not run"* never look identical. | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-198** | **The hub's superseded-escrow retention does NOT retain the offsite repository password — and the ceremony the system tells the customer to run is what destroys the last copy.** Shipped in v0.92.0, v0.93.0. Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`. | **SHIPPED** (hub **v0.93.0**, 2026-08-04) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-199** | **The hub serves recovery blobs on two endpoints that have no client anywhere in the system.** Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`. | **SHIPPED + PROVEN-LIVE 2026-08-04** (hub **v0.94.0**, agent **v0.125.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-203** | **A customer-declared MANDATORY data directory was silently absent from the off-site snapshot while the run reported `ok`.** Evidence: `audits/DRILL-r201-offsite-recovery-2026-08-04.md`. | **SHIPPED + PROVEN-LIVE 2026-08-04** (controller **v0.197.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-204** | **A rebuilt box can recover its off-site key and still cannot use it: the remedy that reconfigures the tier is the thing that blocks the recovery.** Shipped in v0.198.0, v0.199.0, v0.95.0. Evidence: `audits/DRILL-r201-night-run-2026-08-04.md`. **Reasoning kept:** **Live on demo-felhom 9201, nothing restarted (`restarts=0`, container older than both mints): the superseded code returned „Hibás vagy lejárt kód" and the current one was accepted first time.** **Item 2 (a re-issue marks a healthy escrow stale) — CLOSED, → R-196.** Test-proven; deliberately NOT fir **What is deliberately NOT automated: the escrow ceremony.** A credential is replaceable; the recovery code is not. | **ALL FOUR ITEMS CLOSED 2026-08-05** (items 1–3 controller v0.198.0 + hub v0.95.0; item 4 controller v0.199.0 + hub v0.96.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-192** | **`offsite_delivery_stuck` tells the operator the opposite of what the detector measured, and the self-heal silently refuses for exactly the reason the message denies.** Shipped in 0.187.0, 0.192.0, v0.199.0. Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. | **CLOSED 2026-08-05 — the guard's scoping half closed BY REPLACEMENT** (hub v0.96.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-193** | **A guest rebuild silently drops the off-site app-data tier, and nothing restages the credential.** Shipped in 0.156.0, 0.187.0, 0.192.0. Evidence: `audits/RECON-offsite-dr-chain-2026-08-04.md`, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`. | **CLOSED 2026-08-05 — controller v0.200.0** (credential half v0.199.0/v0.96.0; the recovery SCREEN v0.200.0) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-90** | **~~ep0 RAM headroom — 4 GiB swap survived its first reboot 2026-07-27; 3.8 GB RAM unchanged~~** | **CLOSED — the operator rescaled ep0 to a CX33 on 2026-08-03** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-97** | **~~Whole-guest backup tier had no hub signal; quiesce blamed the apps~~** Shipped in v0.79.0. | **SHIPPED** (controller v0.177.0 + hub v0.78.0/v0.79.0, 2026-07-27) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-100** | **~~A restic offsite tier that fails every night never goes stale on the hub — `isStale` counted from `LastRun`** **Reasoning kept:** The real defect is **defeated defence in depth**: the hub-side *pull* net was anchored on a field the failing controller keeps refreshing, so it could not compensate for a lost *push* (cf. | **SHIPPED + PROVEN-LIVE** (controller v0.181.0 + hub v0.80.0, 2026-07-28) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-101** | **~~Tier-2 `LastRun` is written on failure and rendered to the customer as „Legutóbbi másolat" — including in t** | **SHIPPED + PROVEN-LIVE** (controller v0.182.0, 2026-07-28) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-108** | **~~Network storage can host an app's namespace, and FileBrowser binds a network share at its ROOT~~** Evidence: `audits/R108-network-app-namespace-2026-07-30.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.187.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-109** | **~~The DR recipe records no backup target~~** Evidence: `audits/R106-R109-recipe-completeness-2026-07-30.md`. | **SHIPPED + PROVEN-LIVE** (agent v0.118.1 + hub v0.83.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-106** | **~~The DR recipe records the PBS namespace as `"root"` on every box~~** | **SHIPPED + PROVEN-LIVE** (agent v0.118.1, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-122** | **~~`AssembleDRRecipe` silently DROPPED `offsite_restic` — the offsite recovery location never reached any reci** | **SHIPPED** (hub v0.83.0, 2026-07-30) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-125** | **A "test through the production path" is only true up to the seam it injects at.** Shipped in v0.118.0. Evidence: `audits/R106-R109-recipe-completeness-2026-07-30.md`. | **FIXED** (agent v0.118.1) — filed for the DOCTRINE point | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-128** | **~~`build-felhom-iso.sh:44` comments that `ISO_VERSION` "aligns with felhom-host-install SCRIPT_VERSION" — a c** | **CLOSED** (iso v1.26.0, 2026-07-31) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-154** | **~~`[first-boot]` is automated-install-only and nothing in the Felhom tree said so~~** Evidence: `audits/SPIKE-universal-iso-3-2026-07-31.md`. | **CLOSED** (iso v1.26.0, 2026-07-31) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-155** | **~~`iso-repack.sh` refuses any ISO without `auto-installer-mode.toml`, blocking the no-`answer.toml` posture~~** | **CLOSED** (iso v1.26.0, 2026-07-31) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-156** | **An app's data is neither persisted nor backed up, and it reports healthy.** Shipped in 26.6.1. **Reasoning kept:** **Provenance, stated because it decides the row:** the observation is `docker ps -a` on demo-hp's **guest 9201** returning empty, supplied with the 2026-08-02 task; **this session did not re-measure** (documentation-only, every box fenced). | **CLOSED — all three apps fixed** (papra template, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-157** | **`bootrecon`'s start-ONCE sweep misses the boot orphan it exists to recover — TWO mechanisms.** | **CLOSED — SHIPPED + PROVEN-LIVE** (B: controller v0.189.0; A: v0.190.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-170** | **The drive-backed boot gate infers a customer's Stop from a container count.** Shipped in v0.190.0. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-171** | **The boot sweep started apps whose data drive was ABSENT — a regression introduced by v0.189.0, now FIXED.** Shipped in v0.189.0. Evidence: `audits/DIAG-bootrecon-drive-absent-2026-08-02.md`. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-172** | **A false `host_stale` alarm fires when the hub's SQLite refuses two consecutive host reports.** **Reasoning kept:** **Retry options (b) and (c) were deliberately NOT taken** — with readers no longer blocking writers a surviving `SQLITE_BUSY` would be a real signal, and a retry would hide it; revisit only on evidence. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.88.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-174** | **The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code.** Shipped in v0.189.0. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-175** | **`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.** Shipped in 0.192.0, 7.5.1. Evidence: `audits/SPIKE-r165-mp1-merge-2026-08-02.md`. | **CLOSED — FIXED 2026-08-03** (same pass as R-165) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-183** | **A fresh install fetched the vouched agent BINARY and its sixteen CONFIG files from two different refs, and nothing compared them.** Shipped in v0.120.0. **Reasoning kept:** **Why it is a defect and not only untidiness:** these files are the agent's own operating surface — its systemd unit, its sudoers, its guarded wrappers — and `configs/felhom-backup-target-apply` is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n`. | **CLOSED — SHIPPED** (installer v1.23.0, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-182** | **A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour.** Shipped in v0.194.0, v0.90.0, v0.90.1. | **CLOSED — SHIPPED** (controller v0.194.0 + hub v0.90.0/.1, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-181** | **The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide.** Shipped in v0.192.0, v0.193.0, v0.193.1. | **CLOSED — SHIPPED** (controller v0.193.0 + v0.193.1, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-178** | **The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED.** Shipped in 0.119.0, 0.120.0, 0.192.0. | **CLOSED — BOTH BOXES REINSTALLED AND PROVEN (2026-08-03)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-158** | **A local Tier-1 app-data backup failure reaches no hub channel.** Shipped in v0.78.0. | **CLOSED BY R-167 — SHIPPED + PROVEN-LIVE** (controller v0.191.0 + hub v0.89.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-159** | **wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy.** | **SHIPPED** (`templates/wishlist/docker-compose.yml`, 2026-08-02) — filed to record the CLASS | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-160** | **gramps-web persisted three paths and wrote to none of them.** | **SHIPPED** (`templates/gramps-web/docker-compose.yml`, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-163** | **`mp1` is RETENTION, not staging — and it is sized as if it were neither.** Shipped in v0.192.0. | **CLOSED by R-165 — the ceiling it describes no longer exists** (golden v3.0.0, 2026-08-03) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-165** | **Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.** Shipped in 0.192.0, v0.192.0. Evidence: `audits/SPIKE-r165-phase0-2026-08-03.md`. | **SHIPPED — golden `build-golden.sh` v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. IMPLEMENTED — the LAYOUT is proven live on both boxes (R-178, 2026-08-03); the BULKHEAD'S REPLACEMENT IS NOT (→ R-181)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-166** | **App state gets a desired/observed model with its own store.** Shipped in v0.189.0. | **SHIPPED + PROVEN-LIVE** (controller v0.189.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-167** | **Storage monitoring and backup alerts.** Shipped in v0.191.1, v0.191.2. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-168** | **~~CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK~~** Shipped in 0.1.0. Evidence: `audits/SPIKE-ci-runner-2026-08-02.md`. | **SHIPPED — and the alarm is DEMONSTRATED** (2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-205** | **`RootFsPressureDespiteHousekeeping` can never fire** Evidence: `audits/SPIKE-dooplex-buildcache-2026-08-05.md`. | **CLOSED — SHIPPED + RED-PROVEN LIVE** (`homelab-manifests` `6808a4b`, 2026-08-05) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-258** | **C3 — the customer's per-app backup tick is green on the PRESENCE of a restore point, and its only red condition is a GLOBAL one.** | **CLOSED 2026-08-08 — controller v0.210.0.** `appDumpVerdict` reads THIS app's own dump result; three states, no icon when nothing is known. **Recency deliberately not added** — see the observation in the follow-through section | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-259** | **C4 — a disk read that FAILS renders as „0.0 GB / 0.0 GB (0%)" in the nominal colour, on the dashboard's most-looked-at meter.** | **CLOSED 2026-08-08 — controller v0.210.0.** `readDiskUsage` reports success; `SystemInfo.DiskKnown`/`HDDKnown`; the template draws no figure, no percentage and no meter fill when unknown. **The hub leg is deliberately NOT fixed and is now R-266** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-260** | **C5 — the agent reports at least eight decision-bearing facts the hub models NOWHERE, and the sharpest one blinds the check that answers „can the operator get into this box".** | **CLOSED 2026-08-08** — the class is GATED (G-1, `scripts/wire_contract_gate.py`) and the sharpest instance is fixed (hub v0.99.0). The remaining unconsumed facts are **R-264, OPEN** — allowlisted with reasons, which is not the same as decided. See the follow-through section below | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-265** | **A CI run can fail with NO LOG PERSISTED, and the alarm mail then points the operator at a log that does not exist.** | **CLOSED 2026-08-08 — `timeout-minutes: 5` on the gates job, and the alarm mail now states elapsed seconds and qualifies its own "names itself in the run log" sentence.** ⚠ **The unknown is NOT closed and must not be read as closed:** whether the `if: failure()` alarm fires for a REAPED job is still unverified. The timeout makes the reap unreachable in practice; it does not answer what happens inside one | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-267** | **The Configuration page is 2.6× faster and is still ~10 s, and the remaining cost is ONE Gitea call whose latency swings 20× with load.** Shipped in 0.100.2, v0.100.0. | **CLOSED 2026-08-08 — hub v0.101.0 + a registry prune.** **Final: cold 5.4 s, warm 0.14 s** (was 26.2 s). Three serialisation legs took it to 9.85 s mean, the 60 s in-memory memo took the warm path to a quarter-second, and the prune halved what remains of the cold path. **⚠ TWO CORRECTIONS TO THIS ROW'S OWN EARLIER TEXT, because both were wrong and both mattered.** **(1) "Only 50 generic versions exist" WAS NOT A COUNT, IT WAS A PAGE LIMIT.** `?type=generic&limit=1000` returns at most 50; the 50 I measured was exactly the cap, and three older agent versions (0.81.0, 0.80.0, 0.79.0) only became visible after the first 30 deletions moved them onto page one. **An unpaginated listing is not evidence of a total** — this repo's own "an empty listing is not evidence of emptiness" rule, walked into while measuring. **(2) THE OPERATOR'S "REDUCE THE NUMBER OF ARTIFACTS" WAS THE BETTER CALL AND MY MEASUREMENT SAID OTHERWISE.** I reported it helps "sub-linearly" and "is not the lever". Measured after: trimming to 10+10 took the COLD load from 13.4 s to 5.4 s — a 2.5× improvement on the path the memo cannot help, because the fan-out is per-version. Recorded rather than quietly dropped (the R-96 standing rule). **Pruned to the newest 10 per package on the operator's rule**, with the live-vouched golden/agent/floor asserted into the KEEP set before a single DELETE was issued; 33 deletions, all HTTP 204, and golden 0.210.0 / agent 0.128.0 / agent 0.127.0 verified still fetchable afterwards. `drill-r50` runs agent 0.113.0, now deleted — flagged to the operator first; it is a disposable nested drill VM and only its re-download path is gone | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-268** | **A live per-guest local-API token was printed into a session transcript.** Shipped in 169.254.253. | **CLOSED — ROTATED + PROVEN LIVE 2026-08-09** (rehearsal pre-phase, `audits/REHEARSAL-byo-reinstall-2026-08-09.md` §3) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-273** | **RANK 1 — the hub vouched an agent version that was never git-tagged, and every install fleet-wide now fails at step 5/8.** Shipped in 0.127.0, 0.128.0, v0.127.0. | **CLOSED 2026-08-09 — tag pushed, install PROVEN** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-278** | **demo-felhom's off-site tier has never completed a run and has been stuck for six days.** Shipped in 0.200.0, v0.93.0. | **CLOSED 2026-08-10 — protection RESTORED, and the recovery it waited for could never have worked** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-280** | **RANK 1 — after a reinstall the data drive cannot be re-attached through ANY dashboard route, and the restore page promises it is "two clicks".** | **CLOSED — controller v0.211.0, delivered via golden 0.211.0 (vouched 2026-08-10)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-281** | **The hub said NOTHING through an entire reinstall — and the tripwire for a sealed-backup unseal did not fire on a real unseal.** | **WITHDRAWN 2026-08-09** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-293** | **CENSUS, 2026-08-10 — no machine that is not ours can be in the state that cost demo-felhom its history, and here is the whole population.** | **CLOSED-INFORMATIONAL 2026-08-10** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-294** | **The orphan card promises restorability that the box rendering it cannot evaluate — specified, not implemented.** Evidence: `documentation/design/SPEC-orphan-card-copy-2026-08-10.md`. | **CLOSED — controller v0.211.0; see R-299 for the sentence it missed** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-296** | **The orphan card's OTHER sentence makes the same promise, and the spec says it is fine.** | **CLOSED — shipped in controller v0.212.0 (R-299); verified: the sentence at backups_remote.html:98 was replaced and the stem guard covers it** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-297** | **An install took whatever golden was lying around.** Shipped in 0.153.0, 0.210.0, 0.213.0. | **CLOSED — observed live + PUBLISHED as `installer-v1.27.0` (both refs bumped)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-299** | **The orphan card's OTHER sentence made the same unevaluable promise, and the spec called it accurate.** Shipped in v0.211.0. | **CLOSED — controller v0.212.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-300** | **Our own uninstall left the thing that makes our own reinstall refuse.** Shipped in 0.0.0, 10.0.2, 127.0.0. | **CLOSED — observed live + PUBLISHED as `installer-v1.27.0` (both refs bumped)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-301** | **The abandon countdown banner makes the retired promise a third time, and as a flat statement.** **Reasoning kept:** the customer chose to abandon a recovery offer that exists — which is why it was NOT changed (this session was fenced to the orphan card). | **CLOSED — premise CONFIRMED and fixed in controller v0.213.0 (R-302)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-302** | **The abandon banner promised retrieval it could not see was still true — fixed by PINNING a fingerprint at the decision.** Shipped in v0.213.0. **Reasoning kept:** Empty is not a match on either side; a countdown started before v0.213.0 carries no pin and takes the cautious branch (deliberately NOT backfilled). | **CLOSED — controller v0.213.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-305** | **The R-300 cleanup fires exactly once per machine, and the second reinstall hits the original wall.** Shipped in 0.0.0, v1.27.0. | **CLOSED — superseded by R-316** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-307** | **`demo-felhom` carries a LIVE abandon countdown that this drill did not start — and the end state says there should be none.** **Reasoning kept:** The drill's fence forbade starting, shortening or triggering a countdown, and none was; but its required end state was *"no abandon countdown anywhere"*, and one exists. | **CLOSED — countdown cancelled 2026-08-12 on the operator's ruling** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-308** | **~~The stored controller password no longer opens `demo-felhom`~~ — WITHDRAWN 2026-08-12, this was MY BUG, not a defect.** | **WITHDRAWN — not a defect (my error)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-309** | **The day-0 runbook says pushing the installer publishes it. It has not since R-110.** Shipped in 1.25.0, 1.27.0. Evidence: `documentation/runbooks/day0-install.md`. | grep -m1 '^SCRIPT_VERSION'`. **Measured while writing it: served `1.28.0`, `main` `1.28.0`, both pins `installer-v1.28.0` — the three agreeing is the observation; any one alone is not.** **The claim was copied elsewhere and the copy was hunted:** `audits/SPIKE-universal-iso-3-2026-07-31.md:184` said the same thing and **cited `day0-install.md` as its source**, which is how it spread. It was **true on the day it was written** (R-110 shipped 2026-08-03), so the dated finding is kept verbatim and carries a SUPERSEDED note rather than being rewritten — falsifying a dated record to tidy it is its own defect. Two other hits are correct in context: `hostinstall_gates.py:198` states the consequence of the manifest LOSING its tag, and the 2026-08-12 drill record already names the sentence as false | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-311** | **A correct recovery code for a retained package stopped being reported as wrong.** Shipped in 0.126.0, 0.128.0, 0.129.0. | **CLOSED — shipped + delivered: hub v0.103.0 + agent v0.129.0 + controller v0.214.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-316** | **The removal now genuinely reverses the installation — R-305's once-per-machine defect closed.** Shipped in 0.0.0, v1.27.0, v1.28.0. | **CLOSED — shipped + published, observed on the cycle that actually fails** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-318** | **No honest marker exists that says Felhom installed dnsmasq on a machine already in the field, and none can be invented.** Shipped in v1.27.0. **Reasoning kept:** `/var/log/dpkg.log` does record the install — and is a **timestamp**, which the standing rule refuses as a heuristic dressed as a fact. | **CLOSED — established, no action possible for existing boxes** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-319** | **The guest-network watchdog finally has a reader — the first of R-264's twenty-one, and it is the repair COUNT that matters, not the state.** Shipped in 0.92.0, v0.92.0. | **CLOSED — shipped hub-side 2026-08-13** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-320** | **Evidence has been destroyed twice in three days, in the same place, by the same act.** Shipped in v1.28.0. Evidence: `audits/DRILL-retained-key-2026-08-12.md`, `audits/REPORT-r316-installer-v1.28.0-2026-08-13.md`. **Reasoning kept:** **The rule, now standing rule 5 in `workspace-CLAUDE.md` (so it loads in every session) and repeated where a session actually meets it — `runbooks/target-selection.md`, `RUNBOOK-rehearsal-v3.md`, and the `PROMPT-TEMPLATE.md` report section: evidence is copied off the machine at the end of the phase | **CLOSED — rule written, four homes** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-321** | **A box on which reporting is deliberately switched off still alarms as stale, then down.** Shipped in v0.105.0. | **CLOSED — shipped hub v0.105.0, both doors** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-322** | **The claim guard has never scanned the hub, and the hub sends the customer's first sentence.** **Reasoning kept:** **Recommended shape, and the reason it is not one line:** the gate is invoked by `controller_gates.py`, so pointing it at a sibling repo makes a controller gate fail on a felhom.eu edit — the cross-repo lesson from G-1 (a gate needing a sibling passes locally and exits INCONCLUSIVE in CI, and must n | **CLOSED 2026-08-13 by R-324** — `scripts/hub_copy_gate.py`, registered in `repo_gates.py`, scanning 95 hub files for retired names and four declared customer surfaces for retrieval stems, with a plant→convict→remove→pass selftest that caught a defect in its own instrument on the first run. The stem list IS shared (`scripts/customer_copy_vocab.py`) and no controller gate was made to depend on a felhom.eu clone; the controller gate's adoption of the shared list is R-325, and until it happens the two are drift-checked rather than left to diverge | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-323** | **The third near-homograph — the five-word phrase is „Tulajdonosi jelmondat” now.** | **CLOSED — shipped hub v0.105.0** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-324** | **The hub's customer copy is under a guard for the first time — and the guard has been watched catching, ignoring and releasing.** | **CLOSED — shipped, selftest green** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-326** | **"Which claims are unproven?" is a question a machine can answer now — and the number everyone was repeating answered a different question.** | **CLOSED — shipped** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-328** | **The disk alert was emailed to nobody, and one word is the whole reason.** Evidence: `audits/DIAG-smart-passed-trap-2026-08-14.md`. | **CLOSED — controller v0.215.0, PROVEN LIVE 2026-08-14.** Now `"warning"`, and `DiskAlertKind.Severity()` is exported so the contract is assertable from any package rather than duplicated as a literal. **The proof is a side-by-side pair pushed through the REAL hub event endpoint** from demo-hp's controller: severity `"warning"` → stored `warning`, `notification_log` **id 689, channel `operator`, status `sent`**; the identical push at `"warn"` → stored **`info`**, and **no `notification_log` row exists at all**. Pinned by `TestNotifyDiskHealthDegraded_SeverityRoutes`, which asserts membership of the hub's accepted set (not just the literal) and names both hub locations; its red-proof — restoring `"warn"` — fails all three assertions | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-334** | **CLOSED 2026-08-18 — golden 0.216.0 baked, published and VOUCHED; CI green by run id.** Shipped in 0.214.0, 0.215.0, 0.216.0. Evidence: `documentation/tests/golden-*`, `documentation/tests/golden-0.214.0-2026-08-12`. | **CLOSED 2026-08-18.** Baked from `RUNBOOK-manual-build.md` §4.0+§4.1 in the DooPlex drill VM and published: **`GOLDEN_VERSION=0.216.0`**, **`GOLDEN_SHA256=ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b`**, 656,970,239 bytes at `…/generic/felhom-golden/0.216.0/golden.tar.zst`. **The published bytes were verified, not just the script's print** — the artifact was downloaded back out of Gitea and hashed, and it matches. **Vouched by the operator, all THREE fields together**, confirmed by reading the hub's own store rather than the save: `artifact_golden_version=0.216.0`, `artifact_agent_version=0.129.0`, `artifact_min_agent=0.129.0` (2026-08-18 11:00:59–11:01:00), and the hub's recorded sha256 matches the downloaded artifact. The R-216 shape was checked on the machine: `MinAgent` 0.129.0 is **equal to**, not above, the newest **published** agent. **`golden_currency_gate.py` rc=0 and `repo_gates.py --fast` rc=0 — all nine gates — and CI is GREEN BY RUN ID: run **353**, `head_sha 7d81681d6`, conclusion `success`** (the two prior runs 351/352 on this same afternoon were red on exactly this row, which is the contrast). That push needed **no `--no-verify`** — the first of the day that did not. Evidence: `documentation/tests/golden-0.216.0-2026-08-18/`, report `REPORT-golden-0.216.0.md`. **Closed with the run id quoted deliberately**: this row was re-confirmed once and widened once, and closing it on a local green a third time would have left the same ambiguity | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Shipped in v0.215.0. **Reasoning kept:** **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** — — CC | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-344** | **`felhom-agent` leaks one TCP connection to PBS per poll cycle, forever, on both sides — and it is the whole of the ep0 descriptor leak.** Shipped in 0.129.0, 0.130.0. Evidence: `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **Reasoning kept:** **The proof obligation is the fd count, not the diff:** per standing rule 3 the positive observable is ep0's ESTAB count going FLAT between proxy restarts, measured over a window long enough to matter — a green test suite proves nothing here, and a 30-minute window proves nothing here either (that e **control 4 cycles -> 4 leaks; fixed 4 cycles -> 0 leaks.** **Positive observable per standing rule 3** (a zero leak is equally consistent with "the agent stopped working"): the fixed box's four poll cycles are in ep0's log, and the boxes' other traffic is near-identical (libwww-perl 924 vs 926, pro | **CLOSED 2026-08-20 — fixed, proven live on both boxes, published and vouched** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-347** | **The R-344 fix exists on two demo boxes by hand and NOWHERE ELSE — a box installed from the current image still ships the leaking agent.** Shipped in 0.129.0, 0.130.0, 0.216.0. Evidence: `documentation/runbooks/publish-train-rules.md`. | **CLOSED 2026-08-20 — published, vouched, and the fleet reconciled onto the published bytes** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-351** | **The restore never read back where the backup said the data lived, and a second press started a second restore.** | \.NamespaceRoot\b' --include=*.go` found **no non-test reader anywhere** — the reconstitution opened the manifest (`offbox_reconstitute.go:235`) purely for the coherence stamp and resolved its destination from the LIVE app instead. **A restore into a destination different from the recorded one therefore succeeded silently, under a green message.** **(b) The second press.** All seven restore handlers gated on `backupMgr.IsRunning()` — the CONCURRENCY flag, acquired *inside* the goroutine (`offbox_reconstitute.go:180`) **after** the handler returned. Established with a test before any change: both the reconstitute and place handlers answered „…elindult" and **overwrote the first restore's op/stack**. The wizard had read the correct flag since v0.154.0 and said so in a comment; the handlers were never moved over. **(c)** The banner gated its terminal result on a page-local `sawRunning`, so a restore that finished before the page opened — the 8.666 s OpenGist restore — was shown to nobody. | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-354** | **The off-site full restore has NO named-volume leg — the tar is in the unit, in the snapshot and in the checking folder, and is never replayed.** Shipped in 0.217.0, 0.218.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-355** | **`paperless-ngx`'s PostgreSQL is dumped into a directory for a stack that does not exist, so its unit has never contained a database dump — and the destructive restore therefore takes no safety dump and tells the customer the app has no database.** Shipped in 0.217.0, 0.218.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-22** (controller **v0.218.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** **Reasoning kept:** That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-370** | **PROCESS: between 2026-08-19 and 2026-08-22 the reviewing side called a documented architectural decision a defect, in four places, because it read the register and live source and never `documentation/architecture/`.** Evidence: `documentation/architecture/`. **Reasoning kept:** R-352 (re-framed), R-369 The record is corrected in place with the framing marked rather than deleted, per the standing rule that a document which quietly changes its mind teaches nobody. | **CLOSED — corrected 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-96** | **Two standing rules were agreed in chat and never committed** Evidence: `documentation/runbooks/workspace-CLAUDE.md:48-70`. **Reasoning kept:** **Two standing rules were agreed in chat and never committed** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-27, size XS, roadmap state `idea — found 2026-07-27`.** Moved verbatim; nothing added or reinterpreted. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-107** | **No offsite action unpacks the named-volume tars Tier-3 captures on every run.** Shipped in v0.218.0. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-383** | **The double-failure message told the customer their previous state was saved, and named a file that was not there.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed.* **Do NOT simply drop the filename:** an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. **A zero-length dump counts as MISSING**, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is `os.Stat` and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | **CLOSED — SHIPPED** (controller v0.222.0, 2026-08-23; `undoCopyPhrase`, four cases, plus an AST seam test that the message is still wired to the builder) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-329** | **`app_start_failed` was emitted with severity `"warn"`, so every one of them was delivered to nobody.** Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *The vocabulary is EXACT and it is the HUB's, not ours* — `{info, warning, error, critical}`; anything else is coerced to `info` at ingest and dropped by `severityNotifies` before BOTH legs. **This was the SECOND occurrence** (`DiskAlertKind.Severity` until v0.215.0), and its comment had recorded the lesson — **a comment is not a guard**, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and *an unlisted limit is not a limit, it is a hole*. **The register's own framing was that the DECISION was the work** — should a stopped app mail the customer at all? Answered: **operator always, customer OFF by default**, because `processOperator` never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. **Deliberately NOT added to `operatorOnlyEvents`** — that would make the new toggle visible, flickable and structurally incapable of delivering. **Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, `warning`/`sent`, after it.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-386** | **A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact.** Shipped in controller v0.223.0. Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *the state test was guessing at something the product already knows.* `DesiredState` records the customer's intent, has **exactly one writer**, and is tri-state; `StateExited` never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: `Stopped` → no alarm, `Running` → **alarm**, **absent → UNKNOWN, keep today's behaviour AND announce it**. *Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about.* **A rule without a mechanism is a wish:** every such suppression sets `IntentUnknown` and the names are logged at INFO, so an operator can answer *"how many apps am I blind to?"*. **`failedRestart` must still lift a `Stopped` intent or F-CRIT-1 re-opens.** **Fenced act: adding a `DesiredState` WRITER** — twelve of `StopStack`'s fourteen callers are machines. Proven live: alarm 24 s after an out-of-band `docker compose stop`, heartbeat `1 currently down` against the previous day's `0`; and with intent removed, suppressed *plus* the log line naming the app. **0 of 8 deployed apps on `demo-hp` carry an absent intent.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-389** | **Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app.** Shipped in hub v0.108.0. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. **Reasoning kept:** the fix is a THIRD SIBLING of `cooldownTierSuffix`/`cooldownRunSuffix`, separate for the reason the second one's docstring already gives — *the existing two keep byte-identical semantics for every type that uses them.* **`cooldownStackSuffix` takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property:** `tier` and `run_id` appear only on types that want that grain, `stack_name` does not. **`perAppCooldownEvents` is a named allow-list with `app_start_failed` and nothing else** — *the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app*, and this is not hypothetical: **`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` through a different struct**, so a payload-shape rule would have split it silently. **The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown.** `app_start_failed` qualifies precisely because it has NO digest — there is no `apps_down_run` the way `backup_run_failures` summarises a run. **The hour is unchanged; the grain was the complaint.** Fail-soft: absent or malformed details degrade to the old key and the mail still goes. **PROVEN LIVE 2026-08-23:** two apps four minutes apart gave **2 sent / 0 suppressed** where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (`…:opengist`, `…:calibre-web`) against the previous day's shared `key=demo-hp:app_start_failed`; and `crossdrive_failed` for two different apps stayed **coarse** under `key=demo-hp:crossdrive_failed`, byte-identical to the derived v0.107.0 value. **AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED** — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.108.0, 2026-08-23) | full text: `git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md` |
|
||
|
||
|
||
## 2026-08-30 — the off-site store gets checked (controller v0.227.0/v0.227.1)
|
||
|
||
Two rows closed, one CORRECTED and deliberately left open.
|
||
**Full original text: `git show <this commit> -- documentation/backlog/OPEN-ITEMS.md`.**
|
||
|
||
| ID | Title | Shipped | Evidence |
|
||
|---|---|---|---|
|
||
| **R-359** | The off-site restic store was never verified by anything, ever | controller v0.227.0/v0.227.1 | `documentation/tests/r359-integrity-2026-08-30/` |
|
||
| **R-397** | `NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller, and the product advertised a weekly check that did not exist | controller v0.227.0 | `documentation/tests/r359-integrity-2026-08-30/` |
|
||
|
||
**R-398 is deliberately NOT a row here.** It was proposed for closure in the same pass and was **CORRECTED instead** — the premise was wrong, the seam already existed — so it stays in `OPEN-ITEMS.md` as the record. It is written as prose rather than a table row because a register row in both files is exactly what `closed_register_gate.py` convicts on.
|
||
|
||
**The rules these leave behind:**
|
||
|
||
- **The integrity check TAKES the single-writer flag and SKIPS rather than waits.** `resticStep`
|
||
escalates to `unlock --remove-all` on a lock error and is only safe while every caller holds that
|
||
mutex; a check without it can strip a LIVE prune's lock. Never remove that guard.
|
||
- **Due-ness, not a weekday.** R-341 is the other shape: a dated check quietly missed and never caught
|
||
up. No `Weekly` primitive was added; a daily job that asks "is it due?" catches up after downtime.
|
||
- **"I could not look" is not "I looked and it is broken".** Skipped / unreachable / failed are three
|
||
facts. A timeout is unreachable, never damage. A failure advances due-ness; a skip does not.
|
||
- **A success that mails nobody is a design choice, not a gap.** `backup_integrity_ok` is severity
|
||
`info` and is dropped before both delivery legs. A weekly success e-mail is how alerts stop being read.
|
||
- **⚠ THE STRUCTURE CHECK DOES NOT CATCH SILENT CORRUPTION.** Measured: a pack corrupted without a size
|
||
change returned `no errors were found`, exit 0. Only `--read-data*` caught it. An `ok` at the shipped
|
||
depth means the index and the snapshot graph are sound — narrower than the word suggests. That is
|
||
**R-399**, open.
|
||
- **A damage classifier must match PHRASES, not words.** `"pack "`, `"tree "` and `"snapshot "` all
|
||
appear in restic's ORDINARY progress output; the first draft would have called a healthy run corrupt.
|
||
The negative control caught it — which is why a control that has only seen the failing case is worth
|
||
nothing.
|
||
- **The restic exec seam has always existed** (`SetOffboxRunner`). R-398 said otherwise and was wrong.
|
||
|
||
## 2026-08-30 — the restore tells the truth (controller v0.226.0) + R-395
|
||
|
||
Six rows closed. **Full original text: `git show e027b5d9 -- documentation/backlog/OPEN-ITEMS.md`.**
|
||
Compressed here to title, shipping version, evidence, and the sentences that state a RULE.
|
||
|
||
| ID | Title | Shipped | Evidence |
|
||
|---|---|---|---|
|
||
| **R-353** | A local unit restore reported a bare completion whether it returned an entire dataset or nothing | controller v0.226.0 | `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` — live sentence `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` |
|
||
| **R-357** | The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths | controller v0.226.0 | seam tests only — **NOT live-validated**, by design |
|
||
| **R-358** | `OffboxFullScratchReady` asked "non-empty directory", which is what a failed restic run leaves | controller v0.226.0 | `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` — marker `{"schema":1,…,"full":false}`, gate logged place-to-live closed |
|
||
| **R-360** | The verification-copy delete gated on the concurrency flag, which a verification restore never holds | controller v0.226.0 | `documentation/audits/evidence-r353-r360-live-2026-08-30/live-validation.txt` — refused in the live flag state; planted canary survived |
|
||
| **R-396** | A unit-only verification restore unlocked the DESTRUCTIVE full restore | controller v0.226.0 | same evidence; found while answering R-358's open question |
|
||
| **R-395** | `STATUS.md` contradicted itself about the controller version | doc fix, same session | `STATUS.md` at `e027b5d9` |
|
||
|
||
**The rules these leave behind — the reason the rows are kept rather than deleted:**
|
||
|
||
- **A restore outcome is a claim about THE BACKUP, never about the app.** R-355 extended to the Tier-1
|
||
path. The off-site twin has `SafetyDump` as an honest discriminator; the local path has none, so no
|
||
claim about the app is available to it at all. Not merely unproven — unprovable from a manifest:
|
||
§6.3 records that an absent dump has causes that say nothing about the app.
|
||
- **Zero-replayed has two causes and they are opposite news.** "The backup held no data" and "the
|
||
backup listed data that did not come back" must never share a sentence.
|
||
- **The destructive reconstitute uses NO headroom margin**, matching `PlaceOffsiteRestore`. The ×1.1
|
||
elsewhere exists because that gate PREDICTS a download; this one measures a tree that already exists.
|
||
- **Fail closed when a probe reads ≤ 0.** `free < need` with `need == 0` is FALSE, so an unmeasurable
|
||
input sails through — a gate present and inert, which is worse than no gate because it reads as
|
||
protection.
|
||
- **A hidden button is not a guard.** Template enable-flags control a button; the handler must refuse.
|
||
- **One boolean must not drive three intents** (R-396): `ScratchReady` answered "is there a scratch"
|
||
while being consumed as "may we place" and "may we destructively restore".
|
||
- **Never restate a version in a second place on the same page** (R-395). Live versions belong in the
|
||
hub, never in a doc.
|
||
- **A doc comment claiming a guard exists is why nobody looks for the missing guard** (R-360). Correct
|
||
such a sentence in place; do not delete it.
|
||
| **R-399** | **How deep should the off-site integrity check go — Viktor ruled full depth.** Shipped in controller **v0.228.0**, 2026-08-31. Evidence: `felhom-controller/REPORT.md` (v0.228.0) — restic argv observed from the guest at both depths on `demo-hp`. **Reasoning kept:** *the structure check does not detect a size-preserving pack corruption — measured 2026-08-30, plain `restic check` reported `no errors were found` and exited 0 over a damaged pack that every read-data form caught. That is the reason for the default and it is what should stop anyone turning it back down to save four seconds.* *An empty value means "not configured", therefore the default; `off` is the off token, because a setting with no off switch is not a setting.* *A malformed value falls back to the DEFAULT, never to structure — falling back to structure would silently remove the protection on a typo, which is R-357's shape.* **Superseded by R-401** for anything about a large store. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` | **CLOSED — SHIPPED** (controller v0.228.0, 2026-08-30) | controller v0.228.0; `documentation/tests/r359-integrity-2026-08-30/` |
|
||
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` | **CLOSED — SHIPPED** (controller v0.228.0, 2026-08-30) | controller v0.228.0; `controller/scripts/debug_route_gate.py` + its decoy in `test_gate_decoys.py` |
|
||
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-87** | **The restic tier was never restore-tested — RE-SCOPED by its own spike to "prove the off-site snapshot still CONTAINS a recoverable unit".** Shipped in controller **v0.231.0** + hub **v0.110.0**. Evidence: `tests/r87-offsite-proof-2026-08-31/`; reasoning: `audits/SPIKE-restic-restore-test-2026-08-31.md`. **Reasoning kept:** *The weekly check proves the stored bytes are the bytes we stored; it cannot tell us we stored the WRONG thing.* *The acceptance rule has TWO parts and the obvious one is a trap — "everything declared is present" passes a hollow unit, which is the shape it exists to catch.* *The expectation comes from INSIDE the unit, never the live box: the snapshot may predate the app's shape.* *The volume half is an EXISTENCE check and not a name match — the naming held on all eight real units, but "held on eight" is not "derivable" (R-355), and half a rule that is true beats a whole rule that is invented.* *THREE outcomes: pass, fail, and cannot-judge — collapsing the third hides a gap in one direction and alarms on our own blind spot in the other.* *It proves the snapshot CONTAINS a recoverable unit; it does NOT prove a restore puts data back into a running app — §8 matrix row 4 was deliberately NOT moved.* *The proof's scratch is a SEPARATE root because the job deletes on every path, and sharing the customer's root would mean a nightly job deleting a copy the customer is looking at.* | **CLOSED 2026-08-31 — SHIPPED + PROVEN-LIVE** (controller v0.231.0, hub v0.110.0) | full text: `git show 303129e:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-406** | **Two unrelated findings shared the identifier R-133.** Resolved 2026-09-01 by renumbering the hub-uniqueness finding to **R-415**. **Reasoning kept:** *citations were MEASURED before choosing — 3 for hub-uniqueness, 5 for the plaintext break-glass credential — and the FEWER-cited one moved*; *this is the opposite of the task's literal instruction, whose stated ground ("the older number has the longer reference trail") the measurement contradicts; the principle was followed and the letter was not*; *the within-register duplicate rule was deliberately NOT added in the commit that removed its only subject — a guard whose red-proof can only be a planted fixture is not this project's standard (R-416)*. | **CLOSED — RENUMBERED** (2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-410** | **`golden_currency_gate.py` was satisfied by a DIRECTORY NAME — `mkdir` turned it green with no bake behind it.** Shipped in `felhom.eu`, 2026-09-01. **Reasoning kept:** *a directory name is a label; `GOLDEN_SHA256=<64 hex>` is a fact only a completed publish produces*; *the self-test ships a POSITIVE CONTROL, without which "it fails on an empty directory" would be satisfied by a gate that fails on everything*; *directories that look right and hold nothing are printed by name rather than silently ignored, so a half-finished bake is visible*. | **CLOSED — SHIPPED** (`felhom.eu`, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-414** | **The nightly off-site proof was INERT on a box with no registered data drive, every night, with only a WARN.** Shipped in controller **v0.232.0**. Evidence: `audits/R411-R414-2026-09-01/`, determination in `00-part2.1-determination.md`. **Reasoning kept:** *the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded — its own tests say "the scratch still resolves … only the DESTINATION moves"*; *the fallback is SCOPED because the two callers ask different questions, and one predicate answering both is the R-356 defect itself* — unit-only may fall back (§7: a driveless app's unit already lives on the system data path, *"intended, not a defect"*), a full restore may not (§2.2: state-only tier); *absence on `last_proof_result` already means "controller too old", so a second meaning on one field is the `StatsKnown` trap one level up*; *a `cannot_run` is recorded but does NOT advance per-snapshot due-ness, or the app would never be retried once a drive is registered*. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-407** | **`restic check` DOES write a lock file, and the comment above it said it never writes.** Corrected in controller **v0.232.0**. **Reasoning kept:** *corrected in place, not deleted — R-360's rule is that a comment claiming a guard is why nobody looks for the missing one*; *the same paragraph now carries the fact that `restic stats` also takes a lock*. | **CLOSED — CORRECTED IN PLACE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-408** | **`RestoreOffboxScratch` took no single-writer flag while a comment asserted every off-site operation did.** Shipped in controller **v0.232.0**. **Reasoning kept:** *the real deliverable is the WALK, not the acquire* — the sentence was false for months and nothing checked it, the ninth instance of this project's most-repeated class; *it is an AST pass and not `strings.Contains`, because a commented-out call still contains the string*; *adding a line to `offsiteExempt` is a deliberate act and belongs in the commit that adds it*; *the R-87 proof's exemption is kept HONEST by a second test that fails if that path ever gains `unlockStale`, routes through `resticStep`, or loses `--no-lock`*. | **CLOSED — SHIPPED** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-411** | **A background job deleted the lock of a live customer restore and logged it as a crash that did not happen.** Shipped in controller **v0.232.0**. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`, `audits/R411-R414-2026-09-01/`. **Reasoning kept:** *`restic stats` TAKES a repository lock* — the fact nobody had, and the one that made the chain reachable; *`restic check` takes one too, `restic snapshots` and `restic list` do not*; *the fix was wider than the row — FOUR entry points were unflagged, three of them found by R-408's walk rather than by the report*; *the escalation in `resticStep` was NOT removed — real stale locks exist and it clears them; the defect was that a sibling could be live*. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-459** | **The skipped MariaDB conversion is STABLE but never self-resolving; converting costs 7 s and keeps the abort — RULED YES and shipped 2026-09-13: `MARIADB_AUTO_UPGRADE=1` on `bookstack-db`, `kimai-db`, `nextcloud-db`, `romm-db`** (catalog `eec1228`/`bd32830`/`3525e35`; no image moved, `catalog_since` untouched). Measured `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md`; proven `audits/r459-close-2026-09-13/` — harness E3/E3b `proven` with `engine_state_after` = `already upgraded to 12.3.3-MariaDB [exit=1]`, `skipped due to $MARIADB_AUTO_UPGRADE` 0 lines, C3 still `failed`; landed on demo-hp via the real 15-min cycle with both container IDs unchanged, one deliberate restart → `MariaDB upgrade not required`, `/login` 200. **Reasoning kept:** *ask the engine, not the log* — the entrypoint prints `MariaDB upgrade not required` on an unsupported downgrade too (R-464); `mariadb-upgrade --check-if-upgrade-is-needed` exit 0 = needed, 1 = not. *Not established, unchanged:* whether any MariaDB feature misbehaves on an UNCONVERTED datadir. *The precaution that keeps the setting inert until Slice 4:* the engine-major rule + gate, removal tracked as R-469. | **CLOSED 2026-09-13 — shipped in the catalog, PROVEN by harness and live** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-467** | **Controller v0.236.0 owed a golden — PAID 2026-09-13:** golden `0.236.0` baked (`GOLDEN_SHA256=58a3cc24…958bf`, 654 115 664 B), round-tripped from the DOWNLOADED bytes, `./etc/felhom-controller-image` says `felhom-controller:0.236.0`, hub dropdown agreed, three-field vouch re-read (`0.236.0` / `agent 0.130.0` / `min_agent 0.129.0`, not the R-216 shape), floor raised 0.232.0 → 0.236.0. Evidence `documentation/tests/golden-0.236.0-2026-09-13/`. **Reasoning kept:** *this bake carried FOUR unbaked releases and is the LAST per-release bake* — goldens are weekly and before any install from today (R-468); *the MinAgent line was missing from four headers* (R-470). | **CLOSED 2026-09-13 — baked, vouched, floor raised** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-448** | **UPDATE ARC SLICE 4 — a guarded update: verified-backup precondition, abort path, truth at the moment of action.** Shipped controller **v0.237.0** (job) + **v0.238.0** (page) + **v0.238.1** (nightly legs skip an app mid-update, found live). Proven live on demo-hp 2026-09-13, scenarios A/B/E/F/H and the restore walk. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *the precondition is the existing verified backup, not a new copy* (ruling 2026-09-02) — `backup.Tier2UnitRestorePoint`, extracted from the backups page, not copied; *age a copy by its last SUCCESSFUL Tier-2 copy, never the manifest `created_at`* (measured: the manifest moves only on definition changes); *no automatic rollback — measured per-app, the route back is the restore*; *anything that writes a restore point skips an app that is held OR updating*. Open consequences: R-472, R-475, R-476. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-443** | **The Update button reported success over an app it had just broken.** Closed by slice 4 (v0.237.0): 202 `accepted, not completed`; the outcome exists only as `update_phase` after health. Pinned by `TestR443_UpdateIsNeverReportedCompleteSynchronously`. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a compose exit code is never a success signal* (spike §4: HTTP 200 over a crash loop). | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-439** | **The restore hold was not honoured by the update path.** Closed by slice 4 (v0.237.0): `update` joined the router's hold check; live Scenario H refused start/restart/update and the boot sweep. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a hold that only one path honours is not a hold* — the audit found the drive-return gate (restart + boot recreate) and the nightly volume dump ignoring any hold; all three fixed and red-proofed. | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-470** | **Four controller CHANGELOG headers (v0.233.0–v0.236.0) carried no `MinAgent:` line, while the vouch and now the declared floor read it from the header.** Closed 2026-09-13 (`felhom-controller` `f946b0d`): the four headers backfilled with `**MinAgent: 0.129.0** (unchanged)` — v0.232.0's value, proven unchanged (no commit under `internal/agentapi` since 2026-09-01; highest `featureMinAgent` 0.129.0) — and `controller/scripts/minagent_header_gate.py` (fast, blocking) refuses a newest header without the line; a prose or code-span mention does not count (decoy; red-proof F). | **CLOSED 2026-09-13 — GATED** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-472** | **The golden cadence ruling and the hub's floor rule contradicted each other: a floor above the vouched golden delivered nothing.** Operator ruling 2026-09-13, hub **v0.112.0** (`f181efd`): a floor saved with the release's declared MinAgent is served above the golden under the same agent comparison; an undeclared one is still held and both forms refuse it (`floor_needs_min_agent`). Proven live: controller 0.239.0 reached demo-hp in 14 s and demo-felhom in 15 s from the save, hub `managed floor SERVED … from declared`. Evidence: `audits/rulings-r472-r475-2026-09-13/` 02, 03. **Reasoning kept:** *the manifest leads the floor inside the golden; above it, the release's own declared MinAgent does* (publish-train rule 1); a declaration binds to its exact floor. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-475** | **The update precondition was Tier-2-only, so an app with no second-drive copy could not be updated.** Operator ruling 2026-09-13, controller **v0.239.0** (`b93c154`): the first fresh copy in the order Tier 2, Tier 1, Tier 3 (bounded, unreachable = absent + WARN); `backup_max_age` applies to the chosen tier; nothing anywhere → back up first; refused only when no copy and no backup can be taken; the hold names the tier; a Tier-2 failure in the pre-backup is a WARN. Proven live on demo-hp: nothing anywhere → backed up first (04), Tier 1 alone (05), held naming „saját meghajtó” (07), restored from „helyi” (08). Red-proof M (age only on Tier 2) fails. **Reasoning kept:** *first FRESH copy, not first copy* — a stale mirror must not force a backup while the own unit is minutes old. Follow-ups: R-477..R-480. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-473** | **The `glance` catalog template crash-looped on every fresh install — the image ships no default config.** Closed 2026-09-13 (catalog `50ad286`): an `entrypoint` wrapper seeds a small Hungarian start page on first boot only when `/app/config/glance.yml` is absent, then execs the image's own command (read from the image, not guessed); the seed validated with the image's `config:validate`. Proven live on demo-hp the same evening: a fresh throwaway install healthy in 21 s, restart count 0, front door 200, seed present in the volume (`audits/nightly-2026-09-13-adventurelog/08-R473-glance-fresh-install.txt`). No version moved. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-471** | **`observations_gate.py` read only the FIRST observations section, so an appended second section with an unmarked item passed — and the R-419 decoy had been reading LIVE HOLE at HEAD.** Closed 2026-09-13 (felhom.eu `scripts/observations_gate.py`): `observation_sections` collects every observations heading and `observation_items` pools their items; the heading shown is all of them joined. Cause established: first-heading-wins (the parser `break`s on the first match). Red-proof: the old parser exits 0 on the decoy shape, the new one 1 (`audits/v0240-2026-09-13/rp-R471.txt`); all 12 felhom.eu decoys behave; the felhom.eu and controller REPORTs still pass through the shared script. The "run by hand" half stands: the decoy suite is still not in any runner (R-426's exemptions) — that is a separate row if wanted. | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-486** | **Removing an app with its backups KEPT forgot its Tier-2 record, so the second-drive restore was refused over an intact mirror.** Closed in controller **v0.240.0** (`bdcbd50`): the record goes only with `remove_backups`. Proven live on demo-hp: remove keeping backups → `cross_drive` record kept → „Teljes visszaállítás" restored 2 volumes and the database, data identical. Red-proof: an unconditional `SetCrossDriveConfig(nil)` fails the wiring test. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-484** | **PostGIS was not a database.** Closed in v0.240.0: `dbTypeForImage` matches `postgis`, `pgvector`, `timescaledb` as Postgres. Proven live: adventurelog's unit now carries `db-dumps/adventurelog-postgres.sql` (8 MB) and the unit restore reports „2 adatkötet és az adatbázis visszaállítva". Red-proof: dropping `postgis` fails the table test. Immich's own image already matched. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-485** | **The backup card read two dead paths and said `has_backups:false` over 484 MB.** Closed in v0.240.0: it sizes the recovery unit and the app's Tier-2 mirror(s). Proven live: `244M` + `244M`, `has_backups:true`. Red-proof: the old paths fail `TestR485`. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-480** | **The card kept a held update's „leállítva marad" sentence after a successful restore and after removal.** Closed in v0.240.0: the stack remembers its last update ended held; `fillHoldReason` hides that outcome once the hold is gone or the app is not deployed; a pull failure keeps its sentence. Red-proof: disabling the block fails three assertions. Live: the card after a restore reads clean. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-477** | **The update's Tier-3 lookup paid the full 15 s bound and blamed another app.** Closed in v0.240.0: `OffsiteSnapshotTimes` — one `snapshots --json`, no per-app `stats`; the inventory page and the update share `offsiteNewestPerTag` (registered read-only for R-408). Red-proof: routing the lookup through the inventory fails `TestR477`. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-478** | **A unit left by a removed install counted as the reinstall's fresh copy.** Closed in v0.240.0: `usableRestorePoint` refuses a copy older than the app's `deployed_at`; R-474's fix removes such units anyway. Red-proof: dropping the check fails `TestR478`. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-474** | **„Delete backups" deleted only `db-dumps`; the unit and the Tier-2 mirror survived; prefs stayed.** Closed in v0.240.0 for the backups half: the whole unit, every mirror (`Tier2MirrorDirsForApp` / `RemoveTier2Mirrors`, exact-path guarded) and the prefs go; `backup_paths_removed` lists them. Proven live: 244M + 243.6 MB removed, nothing left on either drive, prefs `None`. **The `volumes_removed: null` half is re-filed as R-489.** `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-466** | **Removal with „Mentési adatok törlése" left the unit's `compose/` + `manifest.json`.** Subsumed by R-474's fix in v0.240.0 (the whole unit goes). Decided by the same change: the button means the WHOLE unit. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-482** | **`adventurelog` ran Django with DEBUG=True on the public origin.** Closed in catalog `ed2c018`: `DEBUG=False` on the backend service. Proven live after the sync: the same CSRF failure renders the 300-byte production page, no debug text. `wger`, `tandoor`, `paperless-ngx` are recorded as unmeasured. `audits/v0240-2026-09-13/` (10-v0240-validation.txt; red-proofs rp-v240-*) | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-453** | **`~/.config/credentials` values are single-quoted, and a half-applied strip produced a confidently wrong "password is stale" verdict — twice.** Closed 2026-09-13: the instrument the row asked for exists — `felhom.eu/scripts/read_credential.py KEY <0600-file>` (`66156c6`, 2026-08-31) — and the whole 2026-09-13 night used it for every controller and hub password with zero quoting incidents; the memory `credentials-file-values-are-quoted` now names it. The instrumentation lesson (a discriminator that rules out one alternative does not rule in the rest) stays in the memory. | **CLOSED 2026-09-13 — INSTRUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-461** | **`runbooks/target-selection.md` named a venue that does not exist and fenced a fixture that is gone.** Closed 2026-09-13, both halves checked against both boxes first: (a) no `/mnt/nvme-1tb` on demo-hp or demo-felhom; demo-hp's NVMe is `nvme0n1` at `/mnt/hdd_1` (demo-felhom's `/mnt/hdd_1` is `sdb`) — the runbook now names `/mnt/hdd_1` and says it is the same disk as the data drive; (b) `qm list` is empty on BOTH boxes — `drill-r50` (VM 300) exists nowhere; the fence text stays with the measured absence written beside it, and R-93 carries the fact. | **CLOSED 2026-09-13 — DOCUMENTED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-452** | **Nothing enforced `catalog_since`, so the badge's one number could silently under-report.** Closed 2026-09-13 (catalog): `scripts/check-catalog-since.py`, the fifth gate in `catalog_gates.py` — a `--range A..B` gate in the engine-major shape: an app whose per-service `image:` lines differ across the range must carry a `catalog_since` on or after the moving commit's day and not in the future; comments, README and CHANGELOG mentions are not the fact. Hook-enforced; the shallow CI clone skips it out loud (the CI-shape half the row named stays as is, by the same reasoning engine-major uses). Five decoy cases; red-proof: dropping the date comparison lets the untouched-date fact through (`audits/v0240-2026-09-13/rp-R452.txt`). | **CLOSED 2026-09-13 — GATED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-465** | **`cfg.Paths.HDDPath` — empty on every box — still had six readers; were any inert?** AUDITED 2026-09-13 on demo-hp (registered drive `/mnt/felhom-drives/hdd_1`, `hdd_path` absent, no `FELHOM_PATHS_HDD_PATH`). Five of six fall back before the value matters: `report/builder.go:69` and `monitor/healthcheck.go:35` take `storagePaths[0]`; `web/server.go:740` (`primaryHDDPath`) takes the default storage path; `main.go:511` (metrics) takes the default storage path; `main.go:347` passes it only as the auto-discovery FALLBACK, and discovery seeds from the apps. **One is inert AND unreachable:** `api/router.go:981` (`systemInfo`, `GET /api/system/info`) reads the empty value with no fallback (`hdd_configured:false` forever) — and the endpoint itself is shadowed: the web layer's `ServeSystemAPI` claims `/api/system/*` and answers **404 „ismeretlen végpont"** for everything but the two memory routes (measured live). Its only consumer is the monitoring page's memory-distribution card, which therefore never renders — split out as **R-490**. Conclusion: the global can be deleted once R-490 is fixed; no report field, health check or metric depends on it. | **CLOSED 2026-09-13 — AUDITED** | full text: `git show 681c3d6:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-479** | **For a bind-data app the Tier-1 route back restored settings only, and the hold did not say so.** Operator ruling 2026-09-13; closed in controller **v0.241.0** (`3e81330`): an app with classified binds walks second drive → off-site → own unit (`UpdateTierOrderFor`), and the hold sentence ends with what the chosen copy holds (`RestoreHold.CopyHolds`). Delivered by the floor (16 s / 18 s). Proven live on demo-hp with a nextcloud throwaway on the registered drive, Tier 2 off: the failed update held it naming „saját meghajtó, … — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem." Red-proof: a layout-blind order fails the bind case. The 2→3→1 preference itself is unit-proven (a live off-site run touches the standing apps' leg and was not done). `audits/v0241-2026-09-13/` | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-483** | **adventurelog photos uploaded but rendered as broken „Uploaded content" on demo-hp (P2, catalog).** Closed 2026-09-13: the k3s ingress routes `/media`, `/static`, `/admin`, `/accounts` to the backend service on port **80** — the nginx inside the backend image that serves Django's `X-Accel-Redirect` media; the catalog routed everything to the frontend. Two catalog cuts (`3172258` router, `ed62cfd` port 80 — the first cut hit gunicorn and returned empty 200s). Applied to the operator's instance through the guarded Update; proven headless (`GET /media/…webp` with a session → 200 `image/webp`, RIFF/WEBP) and **confirmed by the operator in a browser at 21:49** (two photos render). Scripted multipart uploads through the frontend's `/api` proxy still 500 (upstream `RequestContentLengthMismatchError`); browser uploads work — not a template matter. `audits/nightly-2026-09-13-adventurelog/11-R483-repro.txt` | **CLOSED 2026-09-13 — PROVEN-LIVE, operator-confirmed** | full text: `git show 8914ab0:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-456** | **A partly-dead stack is not a boot orphan, written down nowhere (P3).** Closed in controller **v0.242.0** (`d698ce3`) by pinning the rule: an absent member does not make a stack degraded, a present-but-dead member does (`internal/bootrecon/r456_partly_dead_test.go`). Design unchanged. | **CLOSED 2026-09-14 — PINNED** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-476** | **The Mentések page dated a Tier-2 copy from the unit manifest, which moves only with the definition (P3).** Closed in controller **v0.242.0** (`d698ce3`): `Tier2Coverage.UnitDataDate` (newest dump in the mirrored unit) is what a refreshed leg names; a PRESERVED package keeps the manifest date (R-403). Unit-proven with the measured shape (manifest 09-12, dump 09-13); live on 9202 two captures under one definition dated the copy by the second capture's dump. Red-proof: the manifest-only date fails. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-481** | **No scratch guest on demo-hp for the nightly rotation (P2).** Operator ruling 2026-09-13, option 1. **BUILT the same night:** LXC **9202** `demo-hp-scratch` on demo-hp, restored from the vouched golden 0.236.0 onto a `dir` storage re-added at `/mnt/hdd_1` (`nvme-scratch`), sized like 9201 (7 cores / 25 898 MB / 32 G + 70 G, unprivileged), the demo-hp customer seeded with hub OFF, tunnel OFF, agent OFF, off-site OFF, self-update OFF; image set by hand (allowed only there); a claimed `settings.json` with the demo password and a scratch second drive. **It persists on purpose.** Disposition in all three layers: the guest (`/etc/felhom-scratch-disposition`), the host (`pct` description, tags `scratch,r481`, the bootstrap file) and the hub side (`operations/nodes.md`, this row — the hub has no customer notes field and the guest is outside the felhom pool). Two traps written into nodes.md: `pct restore` wants the golden as a `backup` volume; directories made on the raw volume from the host must be chowned to 100000. The rotation restarted from bentopdf on it. `audits/nightly-2026-09-13b-bentopdf/14-scratch-guest-built.txt` | **CLOSED 2026-09-13 — BUILT, persists** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-487** | **A removed app whose backups were kept was listed on neither backup page (P2).** Closed in controller **v0.242.0** (`d698ce3`): the local lists are keyed on the DRIVES the way R-237 keyed the off-site list on the store — `ListRemovedAppUnits` walks `backups/primary/` on the system path and every connected registered drive; the Mentések page lists the unit after the deployed rows („Eltávolítva — visszaállítható", one action), the Visszaállítás picker lists it in its own group, `GET /api/backup/snapshots` answers for it, and the restore opens the unit where it sits (`primaryUnitDirFor` — a unit kept on a data drive was unreachable before, the fallback named the system path). Proven live on the scratch guest 9202 with an opengist throwaway: removed with data, backups kept → row + picker + API answered; „Visszaállítás a mentésből" reinstalled it running. Red-proofs: lister inert, picker 404, wrong unit dir, row not built, row not rendered, picker not rendered — all fail. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-490** | **The monitoring page's memory-distribution card never rendered — `/api/system/info` was 404 (P3).** Closed in controller **v0.242.0** (`d698ce3`): an exact-path mount ahead of the web layer's `/api/system/` prefix, and `systemInfo` reads the default storage path like every other reader of the empty global. Live on 9202: 200 with the drive figures. Red-proofs: mount removed, fallback removed — both fail. **The global's deletion stays deferred → R-492.** `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-491** | **Removing an app left its update hold in the store, so a reinstall started held (P2).** Closed in controller **v0.242.0** (`d698ce3`): `removeStack` clears an UPDATE hold (`Settings.ClearUpdateHold`, never an R-379 restore hold), logged. Proven live on 9202: a held opengist removed → the store no longer carries the hold, the app redeployed without refusal. Red-proof: the removal without the clear fails the wiring test. `audits/v0242-2026-09-14/` | **CLOSED 2026-09-14 — PROVEN-LIVE** | full text: `git show 72ee053:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-546** | **The first-hour guide and the reminder bar sent the household to create their recovery code before the box could (P2).** Closed in controller **v0.246.0** (`0fe315b`): the bar consults the agent's OWN preflight `ok` (all five blocking items, not a copy of `pbs_storage_id`), cached 60 s, probed only while paused, held back while not ready; `/backup/escrow` shows a waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging (the direct path that produced the raw `-storage` stderr); unknown readiness keeps the bar. The guide moves the step after the first apps: „amikor a sárga sáv megjelenik”. Red-proofs: bar held back, waiting card, start refusal; controls ready and unknown. **Proven by tests through ServeHTTP, NOT live** — no Tier-0 box is paused and agent-connected (**R-551**); chaos night measured live the ~17-minute red window and its self-heal. | **CLOSED 2026-09-17 — PROVEN (tests); live walk owed by R-551** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||
| **R-539** | **The restart brake caught a FAST crash loop and was blind to a SLOW one (P3).** Closed by operator ruling 3 of 2026-09-16 in agent **v0.132.0** (`18d03bd`, tag `v0.132.0`, sha256 `4afe8157…`) + hub **v0.117.0** (`37ae31f`): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth `slow_crashloop_since` moves (at most once per 24 h) and the hub mints `controller_slow_crashloop` (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed `agent_update` to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`. **Proven live with the PRODUCTION window (no test-only interval):** five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 `SLOW CRASH-LOOP … restarts_24h=5`; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). **Reasoning kept:** the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. `audits/evidence-chaos-fixes-2026-09-17/partC-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 3c1882a:documentation/backlog/OPEN-ITEMS.md` |
|