diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 50f36677..1a5f2495 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -921,6 +921,25 @@ flavour, a tag simply **absent** from the store, gives the pull-failure leg. colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by running it**, four positive cases and a negative control, 2026-09-21. +### Creating the drill repo — and the step that mails the operator if you skip it + +`POST /api/v1/repos/migrate` is the route that works (the project's Gitea tokens carry +`write:repository` but not `write:user`, so `POST /user/repos` answers 403). **A migrated repo +inherits `has_actions: true`, and the catalog's CI workflow comes with it.** Every drill push then +runs that workflow, it fails — the drill repo is deliberately gate-less — and each failure **mails +`admin@felhom.eu`**. The 2026-09-21 night sent **47 such alarms overnight**, into the same mailbox +that was carrying real off-site alarms at the time (**R-629**). + +So creating the drill repo is two acts, not one: + +``` +POST /api/v1/repos/migrate {"repo_name":"app-catalog-drill","private":true, …} +PATCH /api/v1/repos/admin/app-catalog-drill {"has_actions": false} +``` + +then read the repo back and quote `has_actions: False` — an alarm channel trained to be ignored is +worse than no alarm channel, and R-168 made CI mail the thing that notices a bypassed gate. + ### Pointing a box at the drill catalog — the step that is NOT obvious **`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`; diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 5a1d9a41..6ef4d19e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -210,25 +210,17 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour | **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | | **R-231** | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | | **R-233** | **The golden bake's acceptance checks were a list of strings the script does not print** — found 2026-08-06 while baking golden 0.203.0 by following `runbooks/RUNBOOK-manual-build.md` §4.1 verbatim. Two of the three named pass markers **cannot ever match**: `overlay2 OK` is not in `build-golden.sh` at all (the line it means is ` docker OK (overlay2; data-root /var/lib/docker)`), and `including mount point … mp1` refers to a volume that stopped existing in `build-golden.sh` **v3.0.0**, when R-165 collapsed the two data volumes into one. The 404 pre-gate's URL was also wrong — the published filename is `golden.tar.zst`, not `felhom-golden-.tar.zst`, so the pre-gate would 404 for the wrong reason and pass **even when the version already existed**. This is the *"an instrument that can silently drop results is not a measurement"* class landing on the bake's own acceptance check: a grep for an impossible string reads `0` forever, and `0` is indistinguishable from failure. **The bake was never actually unguarded** — the script's own `[ "$drv" = "overlay2" ] || { echo FATAL; exit 1; }` is fail-closed and the run exited 0 with no `FATAL`. The **document** was the broken part, which is why nothing had ever gone wrong and nobody had noticed. **FIXED in the same session:** all markers re-captured from the real log rather than paraphrased, the corrected pre-gate URL, the token handling moved off the command line into an in-VM runner script (the old `--setenv=GITEA_TOKEN=$GT` form put the value where `systemctl show` prints it), a required **positive control** on the token-leak grep, and the vouch step rewritten as the three-field change it actually is. **The general lesson:** a runbook's pass markers must be **copied from a captured log, never written from memory** — §4.0 of that same file already learned this for the qemu launch line and says so; §4.1 had not. | **CLOSED 2026-08-06** — fixed in `RUNBOOK-manual-build.md` | - | **R-235** | **The appliance console keeps telling an already-paired box to go and pair itself.** Measured 2026-08-06 on VM 323: **25 minutes after** the operator bind, with the guest provisioned, the controller reporting 0.203.0 and the agent ONLINE, the physical console still displayed „Felhom — a doboz készen áll, és a **párosításra vár**" together with the now-spent pairing code `US3-6GP` — and, in the same panel, the promise **„Ez a képernyő magától frissül — nincs teendő a doboznál"**. It does not refresh. A customer looking at their screen is told the setup has not happened, and is told the screen would have updated if it had. Cosmetic in mechanism, not in effect: it invites the customer to re-pair a working box, or to call for help about a box that is already fine. Same family as R-234 — a surface asserting a state that stopped being true. | **READY** — owner Viktor | - - | **R-240** | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | - - | **R-242** | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | | **R-243** | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | - | **R-244** | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. | **READY** — owner Viktor | - | **R-245** | **Should a customer who never decides be auto-abandoned after 30 days? RECORDED, NOT BUILT — and the reasoning against it is recorded with it so the decision can be revisited properly.** **The operator's proposal (2026-08-07):** a box that has been offered recovery for 30 days without the customer deciding is auto-abandoned, entering the 14-day grace, so an undecided box does not sit for ever holding history nobody has claimed. **What was built instead:** the escalating reminders (1/3/7/14 days) and the operator levers `--abandon-extend` / `--abandon-stop`. **The reasoning, as settled with the operator the same day:** (1) **nobody is absent** — a box does not reinstall itself, so whoever rebuilt it was standing there and met the recovery question; the "customer away for months" case does not arise from this situation, because **a reinstall implies a person**. (2) **A customer who cannot find their code will get in touch**, which is the moment to extend or disarm by hand — so the automation would be firing at people we are already talking to, which is why the levers were the thing worth building. (3) **The cost is theirs**: the old history sits in the customer's own storage allowance, and if they are paying to keep something they have not decided about, that is their call and they feel it before we do. (4) **The real harm, if it comes, is QUOTA** — old history blocking new backups — and **that is a condition, not a calendar**. An automatic ending should trigger on the harm, with a dated warning, never on a date alone. **If this is ever built, build it that way.** **✅ RE-FILED 2026-08-08 AS A DECISION TAKEN, not a question pending.** It sat in the operator's queue as `WAITING-ON-OPERATOR` for a day, and **nothing was actually pending** — the operator and the reviewer settled it on 2026-08-07: it is **not built**, the levers were built instead, and the whole reasoning above is the record of why. A settled decision parked in a queue is a queue nobody trusts, and an audit of every `WAITING-ON-OPERATOR` row the same day found this was the ONLY one — so the drift was caught while it was still a single row. **THE CONDITION THAT REOPENS IT, which the reasoning already names: QUOTA — old set-aside history blocking new backups.** Not a calendar. If a customer's retained history ever refuses a new backup, revisit this with a dated warning that triggers on the refusal; until then it stays decided. | **DECIDED 2026-08-07 — not built; reopens on quota** | - | **R-246** | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor | | **R-248** | **A flag that changes behaviour is visible to nobody who would look for it.** Q4 of the 2026-08-08 spike, answered plainly. **The customer** sees only a derived card stating a false reason (R-247). **The box** cannot see it at all (R-247's dropped field). **The operator** can see it on exactly ONE page — the **PBS-DR** view (`hub/internal/web/pbsdr.go:487`, `v.EscrowStale = escrow.StaleAt != ""`) — which is the wrong tier for this symptom: an operator investigating an OFF-SITE problem has no reason to open a PBS-DR page. **No alert, no report field, no off-site surface.** The one-shot `escrow_stale` event fired on 2026-08-04 and **was never notified** (a full census of `notification_log` for that customer that day returns 8 rows, none of them this one); it has fired twice ever, both on 4 August. **So the practical answer is: only a database read.** That is a finding in its own right — a flag that silently changes behaviour and cannot be seen is the shape this fortnight has been about (R-241's discarded comparison, R-228's unread field, R-243's unobserved state). **What it needs:** surface `stale_at` on the off-site operator surface and in the host report, or stop using a field nobody can observe to change what a customer is told. | **READY** — owner Viktor | | **R-250** | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | | **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags `felhom-offbox,calibre-web`; the listing renders **two rows** — `calibre-web · 2026-08-07 14:57 · 12.8 MB` and `felhom-offbox · 2026-08-07 14:57 · 12.8 MB`. `felhom-offbox` is the tier's own marker tag, not an application. **The screen's whole job is to let the customer check that what is in the store is what they expect** (*"Nézd át, hogy tényleg azt találod-e itt, amire számítasz"*), and it shows them a stranger's name beside their own data and a total that is double the truth. **Cosmetic, not a data defect** — the restore page correctly offers only `calibre-web`. **Fix:** filter the marker tag out of the listing, or key the rows on the app tag. | **READY** — owner Viktor | -| **R-254** | **The same render-then-hide pattern R-249 fixed is live in two more places, and one of them carries a real per-install secret.** Found by the §7.1 census that R-249's fix required — *a pattern found once is worth a census*, and it was. **(1) `app_info.html:185` — the serious one.** An app's auto-generated first-login password is rendered into `` beside a „Megjelenítés" button. `hidden` is the same class of control as R-249's `display:none`: it stops a browser drawing the value and leaves it in the response body, so a fetch of the app page returns it. **The value is a REAL per-install secret** — `ReadInitialCredentials` reads it live out of the deployed container (`internal/stacks/initialcreds.go`), it is not the catalog's published `default_creds`. **(2) `deploy.html:482` — the weaker one.** An auto-generated `type: secret` deploy field renders into a `` with a „Megjelenítés" toggle. On the PRE-deploy form this is close to unavoidable — the form must post the value, and it does, in a sibling hidden input — but on an **already-deployed** app's page (`$isDeployed`) the hidden input is correctly omitted while the readonly input still carries the value, and there the exposure is gratuitous. **Not fixed here, deliberately:** this session's scope was R-249/R-252/R-253, and each of these needs its own reveal endpoint and its own body-asserting test rather than a shared quick edit. **The fix shape already exists** — `POST /settings/retrieval-password/reveal` (v0.207.0) and `escrow_handlers.go`'s rule that a secret is revealed by an XHR and never templated server-side into HTML. **Severity MEDIUM for (1)** — a real credential in a page any logged-in customer opens, with no audit event, reaching caches, history and screen-shares; **LOW for (2)**. **Recommended next**, because R-249 proved the pattern is not theoretical: it was found by the value landing in a session transcript. **✅ BOTH SITES CLOSED — controller v0.208.0, 2026-08-08, and they turned out to be two different problems.** +| **R-254** | **The same render-then-hide pattern R-249 fixed is live in two more places, and one of them carries a real per-install secret.** Found by the §7.1 census that R-249's fix required — *a pattern found once is worth a census*, and it was. **(1) `app_info.html:185` — the serious one.** An app's auto-generated first-login password is rendered into `` beside a „Megjelenítés" button. `hidden` is the same class of control as R-249's `display:none`: it stops a browser drawing the value and leaves it in the response body, so a fetch of the app page returns it. **The value is a REAL per-install secret** — `ReadInitialCredentials` reads it live out of the deployed container (`internal/stacks/initialcreds.go`), it is not the catalog's published `default_creds`. **(2) `deploy.html:482` — the weaker one.** An auto-generated `type: secret` deploy field renders into a `` with a „Megjelenítés" toggle. On the PRE-deploy form this is close to unavoidable — the form must post the value, and it does, in a sibling hidden input — but on an **already-deployed** app's page (`$isDeployed`) the hidden input is correctly omitted while the readonly input still carries the value, and there the exposure is gratuitous. **Not fixed here, deliberately:** this session's scope was R-249/R-252/R-253, and each of these needs its own reveal endpoint and its own body-asserting test rather than a shared quick edit. **The fix shape already exists** — `POST /settings/retrieval-password/reveal` (v0.207.0) and `escrow_handlers.go`'s rule that a secret is revealed by an XHR and never templated server-side into HTML. **Severity MEDIUM for (1)** — a real credential in a page any logged-in customer opens, with no audit event, reaching caches, history and screen-shares; **LOW for (2)**. **Recommended next**, because R-249 proved the pattern is not theoretical: it was found by the value landing in a session transcript. | **✅ BOTH SITES CLOSED — controller v0.208.0, 2026-08-08, and they turned out to be two different problems.** | | **R-255** | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | **SITE ONE — the same defect, fixed the same way.** `app_info.html` no longer renders the value; the page carries the username, the note and a boolean, and the password comes from **`POST /apps//initial-credentials/reveal`**, which **re-reads the running container** rather than serving a cached copy (caching it in the handler would have put it straight back into the body one layer in). `no-store`, CSRF-covered, and **logged as an act**. Both controls — Megjelenítés *and* Másolás — go through it, and a reveal that cannot read the value **says why** instead of returning an empty string that would render as a blank password. @@ -284,7 +276,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -468,8 +460,6 @@ applied.** The one that matters: Scenario A **fails against today's tree** with | ID | What | State | |---|---|---| | **R-266** | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | - - | **R-269** | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** `localapi.TokenStore.Mint` documents *"last-write wins — any previous token for this guest is revoked"*. Across processes that is FALSE until something unrelated forces a reload: the long-lived agent serves `Lookup` from an in-memory index and re-reads the store **only on a miss** (the B3 reload-on-miss optimisation, `tokenstore.go`), so a superseded token is a direct map **hit** and returns `(vmid, true)`. **Red-proved twice.** (a) A unit probe — `TestTokenStore_ReloadOnMiss_RemintCoherence` with the two lookups swapped, i.e. present the rotated-out token FIRST — **fails**; the shipped test passes only because it looks up the NEW token first, and that miss is what evicts the old hash. (b) On hardware, 2026-08-09: after the on-disk rotation the old token returned **HTTP 200**, then 401 only once a new-token lookup had forced the reload, and reliably 401 after `systemctl restart felhom-agent`. **This is the `CLAUDE.md` invariant-comment case exactly** — the comment reads as settled and the test that looks like its pin is order-dependent. **Fix options:** pin the reversed order with a test, or make eviction not depend on an unrelated miss. Until then, **an operator rotating a leaked token MUST restart the agent** — the runbook step is not optional | **READY (S) — NEW 2026-08-09** | — | Found by doing R-268's rotation rather than reading about it | CC | | **R-270** | **R-268's stated rotation recipe is incomplete: the controller never re-reads `bootstrap.json`'s `local_api`, so a rotation leaves the agent channel dead across restarts.** `bootstrap.ensureLocalAPI` returns early when `cfg.LocalAPI.Endpoint != ""` — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own `controller.yaml`, not in the mount. Proved live 2026-08-09: two controller restarts after a correct `bootstrap.json` rotation, still `HTTP 401`; the channel came up only once `local_api.token` was written into `controller.yaml`. The neighbouring `DetectEndpointDrift` compares the ENDPOINT and deliberately does not compare the token (*"a token mismatch is a different failure"*), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is **R-78** | **READY (S) — NEW 2026-08-09** | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC | | **R-271** | **The `agent_channel_unauthorized` alarm can never be closed, because its own prescribed remedy is what silences the recovery.** `channelhealth.Checker.Check`'s UP branch notifies only when `prev != "" && prev != "up"`; a controller restart resets `state` to `""`, so an unseeded→up transition is silent by construction. The alert text says *"token stale/rotated (**re-bootstrap**)"* — i.e. restart the controller — so **following the instruction guarantees no recovery event.** Observed live 2026-08-09: two `agent_channel_unauthorized` errors on the hub (one `sent`, one `suppressed` by the 1 h operator cooldown) and **nothing afterwards**, though the channel came up 3 minutes later and stayed up. The down side is deliberately asymmetric (F2: a born-down channel alerts on cycle 1); the up side never got the matching treatment. Customer dashboard is fine — `SetDashboard` reflects current state every cycle. It is the OPERATOR's trail that ends on "down" | **READY (S) — NEW 2026-08-09** | — | Notify on unseeded→up when the previous *persisted* state was down, or seed from the hub's last event | CC | @@ -595,7 +585,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | | **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **CLOSED 2026-08-30 — SECOND CHECK TAKEN, AND IT CANNOT ANSWER THE QUESTION: the leak this row measured was REMOVED mid-interval by our own fix (R-344, agent 0.130.0). Question moot; the fix is confirmed holding at 12 days** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again **FIRST CHECK — TAKEN 2026-08-20 08:02:13Z, and it was taken at +46.2 h, NOT +24 h.** No session ran between 18 and 20 August, so the 2026-08-19 date passed as pure elapsed time; the delay is stated here rather than backfilled. **The longer window is a BETTER measurement, not a degraded one** — the leaked count is **388** against the 4 and 5 descriptors the original answer rested on, roughly **97x**, so the uncertainty falls from about +/-50% to about +/-5%. **Precondition PASSED:** PID still **551655**, `ps -o lstart` 2026-08-18 09:51:04, `NRestarts=0` on both units. **Result:** fd **17 -> 405** over **166,251 s** = **201.6 fd/day**, Poisson +/-10.2/day (1s), 2s band **181.2-222.1**. Pre-registered range was **370-450** (confirm-band 330-490); **observed 388**, near the centre. **VERDICT: unchanged — the EXPECTED result, and not a failed upgrade.** **Composition:** ESTAB 0 -> 388, **CLOSE-WAIT 0 — absent from the histogram entirely**, so 100% of the growth is established connections and `CLOSE-WAIT` is not merely the minority half. Runway from fd 405 at 201.6/day to the 65536 ceiling: **~323 days (~2027-07-09)**. Full working: `audits/SPIKE-ep0-established-connections-2026-08-20.md` + `evidence-ep0-established-connections-2026-08-20/step1-slope-computation.txt`. **MEASUREMENT TRAP for the 2026-08-25 check, found on this run — see R-346:** the anchor must be `ps -o lstart= -p $MainPID`, NOT `systemctl show -p ActiveEnterTimestamp`, which reads 03:54:54Z for this generation (the upgrade re-exec'd the proxy; systemd never saw a stop, `NRestarts` is still 0) and would put the rate ~15% low. **PERTURBATION NOTE for the 2026-08-25 check:** this spike's Phase C quietens `pvestatd` on demo-hp for one overnight window inside that 7-day interval. Quantified so nobody reads the shortfall as a change in the leak: even if the leak were fully proportional to the request rate (which Parts 1-2 predict it is NOT), a 10 h half-rate window inside 168 h shifts the 7-day slope by **~3%**, far inside the +/-10% Poisson band on a ~1,400-descriptor count. **The +7 d reading remains usable.** **SECOND CHECK — TAKEN 2026-08-30 16:40Z, five days late, and its own premise did not survive the interval.** Evidence: `audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. **Precondition PASSED and this is what makes the reading interpretable at all:** PID still **551655**, `ps -o lstart=` still **2026-08-18 09:51:04**, `NRestarts=0` — the SAME proxy generation as t0, so nothing restarted and re-based the count. **Result: fd = 17.** Not 17 more; **seventeen total** — exactly the documented baseline, against **405** at the first check. The socket histogram contains **one LISTEN and nothing else**: ESTAB **0**, CLOSE-WAIT **0**. **THE VERDICT IS "UNANSWERABLE", NOT "THE UPGRADE FIXED IT", AND THE DIFFERENCE MATTERS.** This row asks whether the PBS 4.2.5-1 upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** — R-344 / agent 0.130.0 fixed the agent's unclosed keep-alive transport, and both boxes were confirmed on 0.130.0 on 2026-08-30. A slope of ~0 after that is a measurement of OUR fix, not of the upgrade, and reading it as "the upgrade worked" would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation quantified in advance for this window was Phase C at ~3%; **the actual perturbation was the removal of the entire phenomenon**, which no pre-registration anticipated. **The question is now MOOT rather than open** — there is no leak left whose slope could differ. **What this check DOES establish, and it is worth more than the original question:** twelve days after the R-344 fix, on the same proxy generation and with no restart to hide behind, ep0 sits at its baseline with **zero** established connections. The 388-descriptor accumulation has not returned. R-336's runway concern (~323 days to the 65536 ceiling) is retired with it; what remains of R-336 is the REQUEST-RATE scaling item, which this check does not touch. | CC on both dates | - | **R-342** | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-343** | **The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later.** Raised by the operator 2026-08-18 **12:36:58Z**, in a **separate save after** the artifact vouch. **Read back from the store, not the form** (`hub_settings.min_controller_version`, `GetGlobalMinControllerVersion` — `hub/internal/store/store.go:1751`): `min_controller_version = 0.216.0`, updated 12:36:58. **WHY IT HAD BEEN BEHIND — this was NOT drift, and describing it as "two releases behind" without this context reads as a defect it was not.** `publish-train-rules.md` rule 1 is *manifest before floor*, and rule 2 requires the floor field to be filled **LAST, in a separate save**, because the DB row overrides the env floor and **acts immediately on the next report cycle**. That rule was earned: on the 2026-07-11 publish train the floor was saved together with the manifest, acted at once, and pushed controller 0.113.0 onto Peti's box **~9 minutes ahead of agent 0.81.0** — the exact forbidden skew, benign only because that box had no NAS shares. R-120's row records the same deliberate choice (*"Floor untouched per publish-train rule 2"*). So the floor sitting at 0.214.0 was **policy being followed**, not neglect. **What it was functionally while it sat there:** not a live problem — every reporting box was at or above it — but a **safety net set two versions low**. The floor is what drags a box forward if it ever falls behind (restored from an old backup, reinstalled, long offline), and at 0.214.0 it would have pulled such a box only to two versions back, missing R-328's severity fix and R-335's follow-up. **THE MEASURED BLAST RADIUS — five reads, and the third and fifth are the findings.** **(1) Floor:** `0.216.0` @ 12:36:58Z, from the store. **(2) Per-customer overrides:** **zero** — all five `customer_configs` rows carry an empty `min_controller_version`, so nothing hides behind a lower override and the global applies to everyone. **(3) Every box's controller version:** `demo-felhom` **0.216.0**, `demo-hp` **0.216.0** — both AT the floor **now**; `drill-r50` **0.213.0** (status `blocked`, last report 2026-08-12, powered off/reverted) and `peti-felhom` **0.115.0** (host row DELETED) are **below** it but are **not reporting boxes**. **(4) Directives/holds:** no `managed floor HELD` line exists; the hub logged `[INFO] Global controller-version floor set to "0.216.0"`. **(5) THE FINDING — a controller DID auto-update after the raise.** `demo-felhom` had been on **0.214.0** since 2026-08-12 16:44 and the hub recorded `controller_updated — Controller frissítve: 0.214.0 → 0.216.0` at **12:37:07Z**, then `controller_started (0.216.0)` at 12:37:12Z — **nine seconds after the raise**, exactly the immediate action rule 2 documents. `demo-hp` was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move. **No error, warning or critical event followed** — the update completed and the controller came back up. **So the change was real, not inert: "every reporting box is at or above the floor" is true BECAUSE of the raise, not independently of it.** **Why it is safe by construction**, cited rather than asserted: `ResolveManagedFloor` (`hub/internal/store/store.go:2068`) sets `Held` and clears the floor entirely when `Floor > GoldenVersion` (the R-216 shape) — floor 0.216.0 **equals** golden 0.216.0, so that guard does not trip — and holds per-box when the box's agent is below the manifest's `MinAgent`, or unknown, or unparseable; both boxes report agent 0.129.0 against `MinAgent` 0.129.0, so the floor was served rather than held. That second guard remains armed for any box reporting with an old agent. **No ISO rebuild is required:** the golden is fetched at first boot from the hub's manifest, which is 0.216.0 — at the floor, not below it — and rule 5's `assert_golden_ge_floor` is a **build-time** gate for *future* builds (`scripts/iso/build-felhom-iso.sh:77`, called at **:267**) which **fails open with a warning** when its inputs are absent (`:78-82`: `if [[ -z "$golden" \| \| -z "$floor" ]]` → `log_warn "… UNENFORCED …"` → `return 0`), both confirmed in the script. **`peti-felhom` was NOT contacted and needs no contact** — from the PETI row: its host row was deleted 2026-07-15 and *"a report from a deleted host 401s and is not persisted"*, so it cannot receive a floor directive at all and the raise cannot reach it | **OPEN — NEW 2026-08-18.** Deliberately NOT closed: the task's closing condition was *all five reads clean, no directive served*, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later | — | **Confirm the 0.216.0 update on `demo-felhom` is healthy in normal operation** (it reported and restarted clean, but it has not yet run a full backup cycle on 0.216.0 at the time of writing), then close. Separately: `drill-r50` at 0.213.0 will be dragged to 0.216.0 by this floor if it is ever booted and reports — that is the floor doing its job, noted so it is not read as a surprise | CC | | **R-345** | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** Lines 21-22 of `docker-push`: `docker tag $(IMAGE):$(VERSION) $(IMAGE):latest` then `docker push $(IMAGE):latest`. `.claude/rules/hub.md:35` says *"Pin explicit versions, never `:latest`"* and `.claude/rules/manifests.md:15` repeats it. Verified present on `848368ec3`. Small, and the deployed manifests do pin a version, so nothing is currently broken by it — but a documented command that performs the prohibited action is a trap for whoever next reads the Makefile as the reference for how to publish, and a floating `:latest` on the registry is exactly the thing an emergency `kubectl set image` reaches for. Noticed while running the 2026-08-20 connections spike; unfiled until now. | **READY (XS) — NEW 2026-08-20** | — | Delete the two lines, or keep them behind an explicit opt-in target that says in a comment why it exists. Check whether a stale `:latest` tag already sits on `gitea.dooplex.hu/admin/felhom-hub` before deciding — an existing floating tag is the more dangerous half. | CC | @@ -611,7 +600,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC | | **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | | **R-367** | **The database dumps already written under the wrong name are stranded, and nothing will ever collect them.** R-355's fix sends `paperless-ngx`'s dump to the right place from now on; it does not move the ones already written. On `demo-hp` that is `/mnt/sys_drive/felhom-data/backups/primary/paperless/db-dumps/paperless-postgres.sql` (312 381 B, 2026-08-22 07:38, the last pre-fix cycle). **Nothing deletes them and that is by design, not by luck:** the F5 stale-primary prune (`backup.go:1248`) skips any directory whose name is not a deployed app, under the guard *"an undeployed app's last backup is still its restore point"* — verified still present after the fix. They are equally invisible to the off-site push, which resolves paths from the app's own unit. **They CAN be adopted, by hand:** move the file to `…/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql` and it becomes a readable restore point for that app. **It is deliberately not automatic.** The adopted dump would sit beside volume tars taken at a different time, i.e. an INCOHERENT pair — the exact shape R-43/R-44's coherence stamp exists to make visible — and a controller that silently relocates a customer's data on upgrade is a migration, not a fix. **Filed rather than done**, because whether a stale orphan is worth adopting at all is a judgement about one machine's history, not a rule. | **OPEN — LOW** | follows R-355 | Decide per box: adopt (and say the pair is skewed), or delete deliberately. Neither on an upgrade path. | Viktor rules, CC executes | - | **R-368** | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks//deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC | | **R-369** | **There are TWO registers, only one calls itself the source of truth, and work filed in the other is invisible to every standing rule that says "grep the register".** `OPEN-ITEMS.md` opens *"the single source of truth for open work"*; `ROADMAP.md` opens *"the prioritized decision log of planned/open work"*. Both hold open work. **Measured 2026-08-22: 72 `R-` ids appear in ROADMAP and not in OPEN-ITEMS; 29 of those are not marked shipped/closed/killed.** Most are feature ideas that arguably belong only in ROADMAP — but some are FINDINGS: R-30/R-31/R-32 (all marked P2-HIGH), R-35, R-40, R-76, R-79, R-25, R-49, R-10 and **R-107**. **The cost is measured, not hypothetical:** R-107 — *"No offsite action unpacks the named-volume tars Tier-3 captures on every run"* — was filed **2026-07-28, READY, severity M**, and re-stated in `07-backup-architecture.md:337,902`. It is absent from OPEN-ITEMS. On 2026-08-21 an overnight drill rediscovered it from scratch by planting files and watching them not come back, and it shipped as R-354 on 2026-08-22. **25 days.** **The dating is exact and makes it a rule violation rather than a gap in the rules:** OPEN-ITEMS.md was created and the template gained its *"single source of truth"* bullet on **2026-07-27** (`655b69f`); R-107 went into ROADMAP alone on **2026-07-28** (`070b0ce`) — the day after. **The risk is NOT double-minting** (ids are shared; OPEN-ITEMS' ceiling 368 exceeds ROADMAP's 331) — it is that a session greps one file, finds nothing, and redoes the work. | **OPEN — HIGH** | R-107 (ROADMAP-only), R-354 | Decide the division and make it mechanical: either one register, or a gate that fails when a ROADMAP row that is not shipped/killed has no OPEN-ITEMS counterpart. Triage the 29 first — most are ideas, a minority are findings. | **Viktor rules**, CC executes | | **R-371** | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | @@ -670,33 +658,26 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** | | **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. **MARKED LATENT 2026-09-01.** It is **harmless today** and the reason is precise: the credential CAN delete, so the removal really happens and the success it reports is accidentally true. **THE CONDITION THAT MAKES IT LIVE — the only one, so it is stated as a trigger and not as prose: the moment delete is withdrawn from the box.** That is exactly what R-95's remedy does, by either route (retention moved off-box, or an append-only transport via R-436). From that moment `resticStep`'s crash-lock self-heal is escalating with a call that cannot fail, in the one path that runs unattended against the customer's off-site history. **So this is a PRECONDITION on the R-95 build, not a follow-up to it** — settle it in the same change or the self-heal ships already broken. | **OPEN — LATENT; becomes live the moment delete is withdrawn (precondition on any R-95 build)** | | **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. **ANSWERED 2026-09-01 (DRILL, `audits/evidence-drill-r95-recovery-2026-09-01/`) — NO, AND THE PANEL READ IS NOT NEEDED.** The named-entry hypothesis was tested exhaustively and fails: **777,600 exact names** in the vendor-documented format `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes, **zero hits**, with a control proving the identical batch shape returns a path that does exist (6/6). **And there is a structural reason:** `df` reports `u629488-sub3` mounted on `/home` at **st_dev 0,82** while `/.zfs/snapshot` is **st_dev 0,276** — a different filesystem — and `/home/.zfs` does not exist. A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset owning that `.zfs`, not to the child at `/home`, so a correctly-named snapshot there **could not contain `felhom-repo`**. Three tools agree with controls in the same run (SFTP, the port-23 shell, `rsync --list-only`). The empty listing is not a display toggle hiding a reachable tree — from a sub-account there is no tree. **CONSEQUENCE: per-file recovery is not "operator-only", it is unreachable from the box entirely** → R-433. | **ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary** | -| **R-433** | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. | **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | +| **R-433** | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. **-- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628).** **Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box?** Hetzner: *"With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation"*, citing `docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots`; and separately *"A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken."* **So the Storage BOX documentation governs, not the Storage Share FAQ** - which is exactly the trap `provider-questions-2026-09-01.md` warned the reader about, and the answer came back on the right side of it. **File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it**, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). **HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on `u629488`. That is now the measurement this row needs, and it is cheap.** **Q2 - is `--append-only` enforced by Hetzner, or taken from what the client sends?** Hetzner: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, with `fluix.one/blog/hetzner-restic-append-only/`. **So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's `authorized_keys`** - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. **This is the answer R-95 has been blocked on** and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. **NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof.** What is owed is a real forced-command key on `u629488`, a restic `forget --prune` through it that is REFUSED, and a backup through it that still succeeds. | **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`**| | **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. **✅ CLOSED 2026-09-01, hub v0.111.1 — AND THE BLOCK ABOVE WAS WRONG, WHICH IS THE POINT WORTH KEEPING.** This row said the fix had to wait for R-433 to establish what IS true. **It did not, because the fix is a DELETION and not a REPLACEMENT.** The promise was withdrawn rather than swapped for a new one: *"The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box."* **That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite** — which is the whole reason it was not blocked. A replacement would have been. **Three tests in `hub/internal/monitor/offsite_r434_test.go`**, all driving the production path so they assert the sentence an operator RECEIVES: the withdrawal is present, the promise is absent in three shapes, `confirmed data loss` may appear only inside its negation, and the stored row must not drift from the delivered mail. **RED-PROOF: restoring the v0.111.0 sentence failed all three**, on every fragment, with the offending sentence printed. **One existing test was edited and it had caught this fix correctly** — `TestR431_FiresOnAMassDeletion` asserted `"NOT confirmed data loss"`; the fragment was REMOVED rather than updated so the wording keeps ONE home. | **CLOSED 2026-09-01 — hub v0.111.1; the "blocked on R-433" verdict above was mine and it was wrong** | | **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** | -| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. | **OPEN — ask the vendor before building anything** | +| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. **-- ANSWERED BY HETZNER 2026-09-22 (ticket #2026090103040671), and the lead is GOOD.** Asked directly whether `--append-only` is enforced by Hetzner or taken from what the client sends, support replied: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, pointing at `fluix.one/blog/hetzner-restic-append-only/`. **The mechanism is an SSH FORCED COMMAND pinned to the key in the Storage Box's `authorized_keys`** - so the flag is enforced on the server side, by a line WE write, and a client asking for a plain `rclone serve restic --stdio` on that key gets the forced one instead. **That is real prevention without a new machine, which is what this lead claimed and could not confirm.** **What it does NOT say, and must not be read as saying:** a key WITHOUT that prefix still gets a deleting server, so this protects a repository only if every key that can reach it carries the forced command - including any key an operator adds later. **NOTHING IS MEASURED YET.** Owed, and cheap: a forced-command key on `u629488`, a restic backup through it that SUCCEEDS, and a `forget --prune` through it that is REFUSED - with the refusal quoted. Until that runs this is a support answer, not a property of our backups. | **OPEN — ask the vendor before building anything** | | **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** | | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | -| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. - -| **R-626** | **[P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed.** FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. `navidrome` was removed through the product: the `remove_hdd_data:true` call was correctly refused `409` (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the `remove_hdd_data:false` call returned **200** with `volumes_removed: ['navidrome_navidrome_data']`. **Two seconds later a container carrying `com.docker.compose.project=navidrome` was CREATED** (`.Created = 19:12:20Z`), and Docker's `restart: unless-stopped` started it again at the next guest boot (`.StartedAt = 20:19:33Z`, the B5 power cut). Seven hours later: **`app.yaml` absent, `deployed=false`, one volume back, and the controller happily probing it — `Health probe navidrome: API GET :4533/ping → 200`.** **The customer-visible shape is the bad one:** *"I deleted that app and it came back"* — and it came back **blank**, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on `deployed`, which is exactly why the teardown found it and nothing else did. **WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container.** The controller was restarted several times later in the night and its log no longer reaches that moment — **the second time in one night that a restart destroyed the evidence of the thing that mattered** (see R-621). **Needs:** reproduce with a loop that removes an app and watches `docker events` for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. **And one instrument lesson worth keeping:** this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: `audits/update-night-2026-09-21/26-removed-app-came-back.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **WAITING-ON-OPERATOR — `09` §3b Q6; owner: CC once answered** | +| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. | **WAITING-ON-OPERATOR — `09` §3b Q6; owner: CC once answered** | | **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. | **WAITING-ON-OPERATOR — `09` §3b Q1–Q4 (own-edge half SHIPPED); owner: CC once answered** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. | **WAITING-ON-OPERATOR — `09` §3b Q7; owner: CC once answered** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | -| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. - -| **R-625** | **[P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one.** MEASURED 2026-09-21 (update night, leg B6). `glance` was HELD by a genuine unattended failed update. The drill catalog then published a **fixed newer version** — a real forward route, the thing a household would hope for. Afterwards the app page read **„Frissítés elérhető — ma" / „Update available — today"**, in both languages, with the Update button offered; pressing it answered **`409 reason='held'`** and the hold sentence. **The BEHAVIOUR is correct and is DESIGN, not a defect:** `09` §6.1 says the hold is `settings.RestoreHold` and that **a successful unit restore lifts an update hold (only that kind)** — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. **The DEFECT is that the page says otherwise.** R-524 settled precisely this shape for the Ahead case — *"(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled"* — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. **What it needs:** the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while `RestoreHold` stands. One verdict, read by both surfaces, exactly as R-524 did it. **AND A QUESTION FOR `09` §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER:** the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: `audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **READY — rank P3-LOW; owner: CC** | +| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. | **READY — rank P3-LOW; owner: CC** | | **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 **-- UPDATE NIGHT 2026-09-21:** bookstack's edge was walked again on 2026-09-21 and is again **half-proven**: the database half read back through `php artisan` with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). | **READY — rank P3-LOW; owner: CC** | | **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | | **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. | **READY — rank P2-MEDIUM; owner: CC** | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** | - | **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | | **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** | - - | **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** | | **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). **PARTLY SHIPPED in v0.242.0 (`d698ce3`), measured live on 9202 the same night:** the difference is computed and a fresh compose-created volume IS reported (`["opengist_opengist_data"]`), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (`docker volume create `, `restore.go:154`) carries no labels — compose still removes it and the response says `[]` (`audits/v0242-2026-09-14/19-R489-cause.txt`). **Remaining fix:** list by the `_` name prefix as well (union), or label the recreated volume as compose would. | **READY — rank P3-LOW; owner: CC (residual)** | | **R-492** | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** | @@ -803,11 +784,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | | **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** | | **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** | - | **R-615** | **[P3-LOW] Pointing a box at a different app catalog by `git.repo_url` alone is INERT — the box keeps fetching from the repository it first cloned.** FOUND 2026-09-21 by reading `sync.go` **before** running it, which is the only reason the update night's drill catalog worked at all. `Syncer.gitCloneOrPull` (`controller/internal/sync/sync.go:274-306`) clones **only when `/catalog-cache/.git` is absent**; on every later cycle it runs `git fetch --depth 1 origin ` + `git reset --hard origin/` **against the remote stored in the clone**, which `buildRepoURL` wrote at clone time. Changing `git.repo_url` in `controller.yaml` and restarting therefore changes **nothing**: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, `git -C /catalog-cache remote -v` still read `app-catalog-felhom.eu`; the box only followed the drill repo once the cache directory was removed. **Why it matters beyond a drill:** this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. **Nothing is wrong with the CACHING**, which is right; what is missing is that a changed `repo_url` must invalidate the clone. **Fix shape:** on start, compare `git.repo_url` with the clone's `origin` and re-clone when they differ (or `git remote set-url` + a full fetch); log which happened. A test that changes `repo_url` under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: `audits/update-night-2026-09-21/04-9202-config-pre.txt`, `05-9202-follows-drill.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-616** | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** | | **R-617** | **[P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through `repos/migrate` — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it".** FOUND 2026-09-21 creating the drill catalog. Both `~/.git-credentials` tokens carry `write:misc,write:notification,write:package,write:issue,write:repository`; `POST /api/v1/user/repos` requires `write:user` and answers **403**, and `POST /api/v1/admin/users//repos` requires `write:admin` and answers 403 too. `POST /api/v1/repos/migrate` with the same token answered **201** and created the private repository. **Why this is a row and not a note:** a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. **What it needs:** one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: `audits/update-night-2026-09-21/03-drill-repo.txt`. | **READY — rank P3-LOW; owner: CC (docs)** | -| **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | +| **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. | **READY — rank P1-HIGH; owner: CC (catalog)** | | **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-620** | **[P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so.** FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. `hub.enabled: false` there, and `Notifier.Publish` returns at `notify/notifier.go:269` — **before** any log line — as do `NotifyHealthChange` (:359) and four more entry points. Startup says it once (`[INFO] Notifier disabled (hub not configured)`) and then every later event, of every severity up to `critical`, vanishes without a word. **The measurable consequence tonight:** the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. **The consequence on a real box is smaller but not zero:** the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. **Fix shape:** one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: `audits/update-night-2026-09-21/11-notifier-disabled.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | @@ -816,7 +796,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-624** | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. | **READY — rank P3-LOW; owner: CC (catalog harness)** | | **R-625** | **[P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one.** MEASURED 2026-09-21 (update night, leg B6). `glance` was HELD by a genuine unattended failed update. The drill catalog then published a **fixed newer version** — a real forward route, the thing a household would hope for. Afterwards the app page read **„Frissítés elérhető — ma" / „Update available — today"**, in both languages, with the Update button offered; pressing it answered **`409 reason='held'`** and the hold sentence. **The BEHAVIOUR is correct and is DESIGN, not a defect:** `09` §6.1 says the hold is `settings.RestoreHold` and that **a successful unit restore lifts an update hold (only that kind)** — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. **The DEFECT is that the page says otherwise.** R-524 settled precisely this shape for the Ahead case — *"(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled"* — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. **What it needs:** the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while `RestoreHold` stands. One verdict, read by both surfaces, exactly as R-524 did it. **AND A QUESTION FOR `09` §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER:** the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: `audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-626** | **[P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed.** FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. `navidrome` was removed through the product: the `remove_hdd_data:true` call was correctly refused `409` (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the `remove_hdd_data:false` call returned **200** with `volumes_removed: ['navidrome_navidrome_data']`. **Two seconds later a container carrying `com.docker.compose.project=navidrome` was CREATED** (`.Created = 19:12:20Z`), and Docker's `restart: unless-stopped` started it again at the next guest boot (`.StartedAt = 20:19:33Z`, the B5 power cut). Seven hours later: **`app.yaml` absent, `deployed=false`, one volume back, and the controller happily probing it — `Health probe navidrome: API GET :4533/ping → 200`.** **The customer-visible shape is the bad one:** *"I deleted that app and it came back"* — and it came back **blank**, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on `deployed`, which is exactly why the teardown found it and nothing else did. **WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container.** The controller was restarted several times later in the night and its log no longer reaches that moment — **the second time in one night that a restart destroyed the evidence of the thing that mattered** (see R-621). **Needs:** reproduce with a loop that removes an app and watches `docker events` for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. **And one instrument lesson worth keeping:** this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: `audits/update-night-2026-09-21/26-removed-app-came-back.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | - +| **R-628** | **[P2-MEDIUM] An empty search of a mailbox I do not control was turned into a claim about what a THIRD PARTY had done, and it went into the register as fact.** FOUND 2026-09-22; the operator caught it within minutes by producing the thread. The `due_checks_gate` fired R-433 (*have Hetzner answered?*). Two searches of the felhom catch-all came back empty, and the emptiness was written into R-433 as **"Hetzner has not answered"** and **"there is no evidence the tickets were ever opened"**. Both were false: ticket **#2026090103040671** had been opened and answered. **THE PRECISE FAULT IS THE INFERENCE, NOT THE QUERY — and that distinction is the whole value of this row.** Re-run afterwards WITH a positive control: `from:monitoring@felhom.eu` returns **201 threads**, `in:anywhere … includeTrash` reaches SENT, TRASH and mail back to January — **the instrument works.** And the exact ticket number, the exact subject `Storage Box issue`, and `from:hetzner` each still return **nothing**. So the literal finding — *this correspondence is not in this mailbox* — was CORRECT. What was invented was the step from there to *Hetzner has not answered* and *the tickets were never opened*. **A mailbox I can read is not the only place a reply can be**, and the operator's own account is exactly where a support ticket he opened would land. An absent record in ONE place can never answer a question about what SOMEONE ELSE did. **This is R-96 rule 3 in a new surface, one step further out than R-607**: there the instrument reported a stale value as current; here a correct observation was promoted to a conclusion it could not carry. **It is sharper still because the same session, the night before, gave every fixture a negative control and every gate a red-proof — and then reached for a search with neither.** **THE RULE, in one line:** *an empty search may be reported as "absent from the place I looked", never as "it did not happen" — and only after a control query that MUST hit has been seen to hit.* **Done:** the rule is written into the Gmail-access memory, where the next session meets it before it searches rather than after. | **CLOSED 2026-09-22 — rule recorded; R-433 corrected with the real answers** | +| **R-627** | **[P2-MEDIUM] Nothing checked that the register is a well-formed table, so an append that ate two rows' state cells went unnoticed until a person read the file — and one row had been broken the same way for 45 days.** FOUND 2026-09-22. The 2026-09-21 update night appended measured results to eight rows with a regex that matched each row's trailing state cell; on **R-446** and **R-458** it consumed the cell and did not restore it, the cell reappeared as a stray FOURTH cell on a DUPLICATED copy of **R-626** and **R-625**, and a blank line was left between each pair. The register then reported **317 rows for 315 findings**, two rows carried no state at all, and two findings existed twice with contradictory state cells. **Nothing caught it:** `one_register_gate.py` compares this file against ROADMAP and `closed_register_gate.py` forbids an id in BOTH files — neither asks whether the file is a well-formed table, and neither notices an id duplicated WITHIN it. **THE RED-PROOF THEN FOUND AN OLDER INSTANCE NOBODY HAD SEEN: R-254 lost its state cell on 2026-08-08 (commit `59527d0`) and had rendered without a State column for 45 days.** **Closed the same day:** `scripts/register_shape_gate.py`, registered in `repo_gates.py` as gate 14 and reached by the pre-push hook, refusing a row that does not end with `|` (an eaten state cell), a duplicated id, or a blank line splitting the table; four decoys in `test_gate_decoys.py` — three convicting on the exact damage shapes and one asserting a healthy register still passes. **TWO THINGS THE RED-PROOF CORRECTED IN THE GATE ITSELF, kept because they are the finding's real content:** a first draft counted CELLS and convicted **125 innocent rows** — register cells carry literal `|` inside prose and shell snippets (`owner: CC | …`), so a row cannot be split on `|`, and a count that cannot be computed is not a check; and it skipped malformed rows before counting ids, reporting **5 duplicates where there were 2**. **Repaired:** both state cells restored from the stray cells that carried them, the two duplicate rows deleted, R-254's verdict sentence given its cell back, and **15 blank lines that split the register into 12 separate markdown tables** removed — every row's text byte-identical afterwards, proven by diff. **And the gate immediately earned itself:** the very next row inserted in this session (R-628) left a blank line behind and the gate refused it. | **CLOSED 2026-09-22 — gate 14, four decoys, register repaired 317→315** | +| **R-629** | **[P2-MEDIUM] The drill catalog sent the operator 47 CI-failure alarms in one night, and the drill method that created it did not mention CI at all.** FOUND 2026-09-22 while checking a different mailbox question — which is the only reason it was found. `admin/app-catalog-drill` was created on 2026-09-21 by `POST /api/v1/repos/migrate` from the live catalog, and a migrated repo inherits `has_actions: true`. Every drill push therefore ran the catalog's CI workflow, which failed immediately (the drill repo carries the workflow but the run has no meaningful gate context), and **each failure mailed `admin@felhom.eu`**: *"[felhom CI] gates FAILED in admin/app-catalog-drill"*. **47 runs, 47 alarms, all overnight, all unread and flagged IMPORTANT.** **WHY THIS MATTERS MORE THAN THE NOISE:** that mailbox is the operator's alarm channel, and R-168 made CI mail the thing that notices a bypassed gate. A night of throwaway failures from a repo nobody must ever act on trains the reader to skim exactly the sender that must never be skimmed — and it did it on the night the same mailbox was also carrying real `offsite_snapshots_dropped` and `offsite_proof_empty` alarms. **The drill method wrote down the fences it needed** — private repo, no customer box may follow it, reset at teardown — **and said nothing about CI, because nobody had run a drill repo through a CI-enabled Gitea before.** **FIXED 2026-09-22:** `has_actions` set to **false** on the drill repo (verified by re-reading the repo: `has_actions: False, private: True`), and `09` §6.5 now carries it as a step of creating a drill repo rather than as a thing to notice afterwards. **What is NOT done:** the 47 mails are still in the operator's inbox, unread — deleting another person's mail is not mine to do, and they are named here so they can be cleared in one search: `subject:"gates FAILED in admin/app-catalog-drill"`. | **CLOSED 2026-09-22 — actions disabled, method updated; the 47 mails are the operator's to clear** | | item | due (UTC) | what to measure | |---|---|---| -| R-433 | 2026-09-22 | Have Hetzner answered `runbooks/provider-questions-2026-09-01.md`? Record BOTH answers in R-433 and R-436, or re-date this row with the reason. The 2026-07-27 snapshot check sat unconfirmed for 36 days because it was never entered here — that is the scar this row exists to avoid repeating. | +| R-436 | 2026-10-06 | **The question is ANSWERED; what is owed is the MEASUREMENT.** Hetzner (ticket #2026090103040671, recorded in R-433) says `--append-only` is enforced by pinning `command="rclone serve restic --stdio --append-only path/to/repo"` to the key in the Storage Box's `authorized_keys`. Measure it on `u629488`: a forced-command key, a restic backup through it that SUCCEEDS, and a `forget --prune` through it that is REFUSED — quote the refusal. Record in R-436 and R-95. A support answer is not a property of our backups. | diff --git a/scripts/register_shape_gate.py b/scripts/register_shape_gate.py new file mode 100644 index 00000000..3612accb --- /dev/null +++ b/scripts/register_shape_gate.py @@ -0,0 +1,132 @@ +#!/usr/bin/env python3 +# -*- coding: utf-8 -*- +"""register_shape_gate.py — OPEN-ITEMS.md is a TABLE, and a table has a shape (R-627, 2026-09-22). + +WHAT IT CONVICTS ON (a FAIL, exit 1), four rules and nothing else: + + RULE 1 — every register row (`| **R-nnn** | …`) ends with `|`, so it HAS a state cell. + RULE 2 — reserved. **There is deliberately no cell-COUNT rule, and that is a measurement, not an + omission.** The first draft counted cells and convicted 125 innocent rows: register cells + carry literal `|` characters inside their prose and their shell snippets (`owner: CC | + …`, `--format '{{.Names}}|{{.Image}}'`), so splitting a row on `|` does not yield its + cells. A count that cannot be computed is not a check. RULE 1 catches the thing that + actually breaks — a row whose state cell was eaten — because such a row stops ending + with a pipe. + RULE 3 — no `R-` identifier appears on more than one row. + RULE 4 — no blank line sits between two register rows. A blank line ENDS a markdown table, so + the rows after it render as a separate table — and, worse, a script that walks "the + table" stops there. + +WHY IT EXISTS, and the cost that bought it. On 2026-09-21 the update night appended measured +results to eight existing rows with a regex that matched a row's trailing state cell. On two rows +(**R-446** and **R-458**) it consumed the state cell and did not put it back; the cell reappeared as +a stray FOURTH cell on a duplicated copy of a different row (**R-626** and **R-625**), and a blank +line was left between each pair. The register then reported **317 rows for 315 findings**, two rows +carried no state at all, and two findings existed twice with contradictory state cells. + +**Nothing caught it.** `one_register_gate.py` compares OPEN-ITEMS against ROADMAP; `closed_register_gate.py` +forbids an id in BOTH OPEN and CLOSED — neither asks whether the file is a well-formed table, and +neither notices an id duplicated WITHIN OPEN-ITEMS. It was found the next morning by a person +reading the file. **A register that silently mis-states how many findings exist is the one file in +this project that must not be able to do that**, because every standing rule says "grep the register +before minting" and every count in every report is read off it. + +WHAT THIS GATE CANNOT SEE — the residual holes, named rather than implied: + + 1. **A row whose state cell is WRONG but present.** Shape is not meaning. A row that says READY + when it is closed passes here and is `closed_register_gate.py`'s business. + 2. **A row whose text was truncated mid-sentence** but still ends with a valid state cell. Only a + human or a diff against the previous commit sees that. + 3. **The DUE-CHECKS block and the other tables in this file.** This gate reads register rows only + — lines that begin `| **R-`. The file's other tables are `due_checks_gate.py`'s business. + +USAGE + python3 scripts/register_shape_gate.py # the real register + python3 scripts/register_shape_gate.py # any file, for the red-proof +Exit: 0 clean · 1 convicted. +""" +import re +import sys +from pathlib import Path + +ROW = re.compile(r'^\| \*\*(R-\d+)\*\* \|') +REG = Path(__file__).resolve().parent.parent / "documentation" / "backlog" / "OPEN-ITEMS.md" + + +def cells(line): + """The cells of a markdown table row, as a reader sees them. + + A row is `| a | b | c |`. Stripping the outer pipes and splitting on `|` gives the cells; a row + that forgot its trailing pipe yields one fewer, which is exactly RULE 2's symptom, so RULE 1 and + RULE 2 are reported separately rather than collapsed into one confusing message. + """ + t = line.rstrip() + if not t.endswith("|"): + return None + return t.strip("|").split("|") + + +SEP = re.compile(r'^\|\s*:?-{2,}:?\s*(\|\s*:?-{2,}:?\s*)*\|\s*$') + + +def check(path): + lines = Path(path).read_text(encoding="utf-8").split("\n") + bad = [] + seen = {} + rows = 0 + + for i, line in enumerate(lines, start=1): + m = ROW.match(line) + if not m: + continue + rows += 1 + rid = m.group(1) + + # RULE 3 first, and ALWAYS — a row that also breaks another rule is still a row, and + # skipping it here would under-count the ids and report phantom duplicates. (Measured: the + # first draft skipped them and claimed 5 duplicates where there were 2.) + if rid in seen: + bad.append((i, rid, f"RULE 3 — duplicate: {rid} already has a row at line {seen[rid]}")) + else: + seen[rid] = i + + # RULE 1 — the row must end with `|`, i.e. it must still HAVE a state cell + if not line.rstrip().endswith("|"): + bad.append((i, rid, "RULE 1 — the row does not end with `|`, so it has no state cell. " + "This is exactly what an eaten state cell looks like, and what " + "R-446, R-458 (2026-09-21) and R-254 (2026-08-08, unnoticed for 45 " + "days) each looked like.")) + + # RULE 4 — look back over blanks to the previous non-blank line + j = i - 2 + blanks = 0 + while j >= 0 and lines[j].strip() == "": + blanks += 1 + j -= 1 + if blanks and j >= 0 and ROW.match(lines[j]): + bad.append((i, rid, f"RULE 4 — {blanks} blank line(s) between this row and " + f"{ROW.match(lines[j]).group(1)} at line {j+1}. A blank line ends " + f"the markdown table.")) + + return rows, len(seen), bad + + +def main(): + path = sys.argv[1] if len(sys.argv) > 1 else REG + rows, ids, bad = check(path) + print(f"register-shape: {path}") + print(f"register-shape: {rows} register row(s), {ids} distinct R- id(s)") + if rows != ids: + print(f"register-shape: {rows - ids} row(s) are DUPLICATES") + if not bad: + print("register-shape: OK — every row ends with a state cell, every id is unique, " + "and no blank line splits the table") + return 0 + print(f"register-shape: CONVICTED — {len(bad)} problem(s)") + for line_no, rid, why in bad: + print(f" line {line_no} {rid}: {why}") + return 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/repo_gates.py b/scripts/repo_gates.py index f856153d..059b3db2 100644 --- a/scripts/repo_gates.py +++ b/scripts/repo_gates.py @@ -22,6 +22,8 @@ Gates, in order (all must pass; **non-zero exit on any failure**): 11. one-register open work living outside OPEN-ITEMS.md (R-369) 12. closed-register a CLOSED row whose verdict still reads open, or an id in both (R-405) 13. observations a REPORT.md observation with no register row behind it (R-389) + 14. register-shape a register row whose state cell was eaten, a duplicated id, or a blank line + splitting the table — the register mis-stating how many findings exist (R-627) **THE `GATES` TABLE BELOW IS THE LIST; THIS IS A POINTER TO IT.** It drifted once already — it read eleven while thirteen were registered, from 2026-08-24 until 2026-09-01, so @@ -145,6 +147,11 @@ GATES = [ # R-389 — a live finding lived in a REPORT.md observations paragraph and nowhere else, and # REPORT.md is overwritten every session. Fast: stdlib file reads. ("observations", os.path.join(SCRIPTS, "observations_gate.py"), [ROOT], True, False), + # R-627 — on 2026-09-21 an append regex ate two rows' state cells, duplicated two others and + # split the table with blank lines; the register then reported 317 rows for 315 findings and + # nothing noticed until a person read the file. R-254 had been missing its state cell since + # 2026-08-08 — 45 days — for the same reason. Fast: one file read. + ("register-shape", os.path.join(SCRIPTS, "register_shape_gate.py"), [], True, False), # R-421 — every registered gate across all four repos ships with a decoy test, or is # named in the exemption list with its row. Registered LAST, after every runner was green: # a failing gate refuses every push, which is what instructions_gate learned the hard way. diff --git a/scripts/test_gate_decoys.py b/scripts/test_gate_decoys.py index 0fb3a3ca..3cfeae56 100644 --- a/scripts/test_gate_decoys.py +++ b/scripts/test_gate_decoys.py @@ -44,6 +44,9 @@ COVERS = { "reuse-refs": "a cited .go path that does not exist; the .md hole is asserted as R-422", "golden-currency": "R-410: an empty directory with a perfect name, checked by what it COUNTED", "closed-register": "a verdict cell reading open, and a row with no state cell at all", + "register-shape": ("R-627: a row whose state cell was EATEN (no trailing pipe), a DUPLICATED " + "id, and a blank line splitting the table — each the exact shape the " + "2026-09-21 append produced; plus the genuine article, which must pass"), "decoy-coverage": "a gate registered in a runner with no decoy and no exemption (its red-proof)", "guide-quote": ("R-596: SEVEN cases in scripts/test_guide_quote_gate.py, run from here so " "this suite stays the single entry point. The load-bearing one is " @@ -311,6 +314,70 @@ else: _n = _gq.stdout.strip().splitlines()[-1] if _gq.stdout.strip() else "?" print(" ok %-20s %s" % ("guide-quote", _n)) +# ── register-shape (R-627) ─────────────────────────────────────────────────────────────────────── +# +# THE DECOYS ARE THE REAL DAMAGE, not invented shapes. On 2026-09-21 an append regex ate two rows' +# state cells, left the cell behind as a stray extra cell on a DUPLICATE of another row, and put a +# blank line between each pair. The register then reported 317 rows for 315 findings. Each decoy +# below is one of those three shapes, written as a session would actually produce it. +# +# AND THE GENUINE ARTICLE IS ASSERTED TOO: a gate that rejects the damage and also rejects a healthy +# register is worse than the hole it replaces. +import importlib.util as _ilu +_spec = _ilu.spec_from_file_location("rsg", os.path.join(ROOT, "scripts", "register_shape_gate.py")) +_rsg = _ilu.module_from_spec(_spec) +_spec.loader.exec_module(_rsg) + +_HEALTHY = "\n".join([ + "| ID | What | State |", + "|---|---|---|", + "| **R-901** | **A finding.** Its text mentions `owner: CC | and a literal pipe` on purpose. | **READY — owner: CC** |", + "| **R-902** | **Another finding.** | **CLOSED 2026-09-01** |", + "", +]) + +def _rsg_case(name, text, must_convict, expect_rule=None): + global ran + ran += 1 + fp = os.path.join(ROOT, "scripts", ".decoy-register-%s.md" % name) + try: + with open(fp, "w", encoding="utf-8") as fh: + fh.write(text) + rows, ids, bad = _rsg.check(fp) + convicted = bool(bad) + if convicted != must_convict: + fails.append("register-shape [%s]: expected %s, got %s (%s)" + % (name, "CONVICT" if must_convict else "PASS", + "CONVICT" if convicted else "PASS", + "; ".join(w for _, _, w in bad)[:200] or "no findings")) + elif expect_rule and not any(expect_rule in w for _, _, w in bad): + fails.append("register-shape [%s]: convicted, but not on %s — on %s" + % (name, expect_rule, "; ".join(w.split(" — ")[0] for _, _, w in bad))) + else: + print(" ok %-20s %s" % ("register-shape", name)) + finally: + if os.path.exists(fp): + os.remove(fp) + +# the genuine article must PASS — including a row carrying a literal `|` in its prose, which is why +# there is no cell-count rule +_rsg_case("genuine", _HEALTHY, must_convict=False) + +# decoy 1 — the state cell eaten (R-446 / R-458 / R-254's shape) +_rsg_case("eaten-state-cell", + _HEALTHY.replace("| **CLOSED 2026-09-01** |", "").rstrip().rstrip("|").rstrip() + "\n", + must_convict=True, expect_rule="RULE 1") + +# decoy 2 — the same finding filed twice (R-625 / R-626's shape) +_rsg_case("duplicate-id", + _HEALTHY.rstrip() + "\n| **R-901** | **The same finding again.** | **READY** |\n", + must_convict=True, expect_rule="RULE 3") + +# decoy 3 — a blank line splitting the table +_rsg_case("blank-splits-table", + _HEALTHY.replace("| **R-902**", "\n| **R-902**"), + must_convict=True, expect_rule="RULE 4") + print() if fails: for f in fails: