From 7d0dffcf341137c130bb5e3ebddc0127b9341515 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 6 Oct 2026 13:09:04 +0200 Subject: [PATCH] ten answers done: R-444 seen working live (weekly trim + System page), closed; STATUS, CONTEXT, REPORT (150 -> 142; 2 opened, 10 closed) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 9 ++ REPORT.md | 84 ++++++++++++++++++- STATUS.md | 36 ++++---- .../ten-answers-2026-10-06/r444-live.txt | 6 ++ documentation/backlog/CLOSED-ITEMS.md | 8 ++ documentation/backlog/OPEN-ITEMS.md | 1 - 6 files changed, 123 insertions(+), 21 deletions(-) create mode 100644 documentation/audits/ten-answers-2026-10-06/r444-live.txt diff --git a/CONTEXT.md b/CONTEXT.md index 61b41695..32c3ac0c 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,15 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-06 (midday) — the operator's ten answers (`09` §3 139–148, „A" for all; CC's own pick differed on 140, 146, +> 147).** Register 150 → 142 (2 opened: R-890 vaultwarden ladder, R-891 a stale CLAUDE.md line; 10 closed). Releases: +> agent v0.149.0 (tag = `f277e61`, sha `6bcae9c2…`, bundle `e182c82d…` — new sudo rule `FELHOM_FSTRIM`; weekly +> `internal/fstrim`; `GET /host/crash-guard`; phantom WARN names `runbooks/pbs-phantom-cleanup.md`), hub v0.139.0 (System +> page „Last disk trim"; R-872 early-return log lines), controller v0.300.0 (R-645 version skip, R-856 crash-boot grace +> 15 min; MinAgent 0.131.0) + golden 0.300.0 (`fb2e9d42…`, pinned), catalog `1938921` (R-747, R-774, R-734, R-624). +> ep0: read-only listing, no phantom, nothing deleted. Bench LXC 9401 recreated on demo-hp (stopped). New full-run gate +> `iso-bootstrap` (image `felhom-iso-assistant:trixie` built on DooPlex). Report: `REPORT.md`. + > **2026-10-06 (morning) — the morning after (rulings `09` §3 137–138).** Register 164 → 150 (0 opened, 14 closed). > Installer 1.32.0 PUBLISHED (tag `installer-v1.32.0` = `32a1520833`, pins `dbede797`; public sha = tag). R-889: every > disk percent is `df`'s (`system.DFUsedPercent`, controller `bee2c2d`) — controller v0.299.0 (`ff4a99a`, MinAgent 0.131.0) diff --git a/REPORT.md b/REPORT.md index 4b9903e2..dea8f325 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,4 +1,82 @@ -# REPORT — the operator's ten answers (2026-10-06, in progress) +# REPORT — the operator's ten answers built (2026-10-06, rulings `09` §3 139–148) -This session builds the ten rulings of 2026-10-06 10:41 (`09` §3 decisions 139–148). The final report overwrites this file -at the end of the session. +| Part | Result | +|---|---| +| **A** — the box changes (1, 4, 5) | **done** — R-444 weekly trim (measured on demo-hp first, then built, delivered, seen working by itself), R-645 version skip, R-856 crash-boot grace; agent v0.149.0 + bundle, controller v0.300.0 + golden 0.300.0, hub v0.139.0, all on demo-hp, demo-felhom, Tester 1 | +| **B** — the backup leftovers (2) | **done** — runbook written; read-only listing of ep0: **no leftover exists**, nothing deleted, real counts unchanged; the box's warning names the runbook | +| **C** — the page sentences (6, 10) | **done** — mealie and Karakeep, hu + en, live in the catalog | +| **D** — the test tools (3, 7, 8, 9) | **done** — R-618 closed by ruling; R-734 marker list; R-624 bench seed proven on a recreated bench; R-502 full-run gate, first real run green with its decoy convicted | + +| Rows before | Rows after | Opened | Closed | +|---|---|---|---| +| **150** | **142** | **2** (R-890, R-891) | **10** | + +Closed: R-444, R-99, R-618, R-645, R-856, R-747, R-734, R-624, R-502, R-774. + +## Baselines and rulings + +Verified 11:07: felhom.eu `8f40b3ce26` (the operator's own „cleaned reports" commit on top of `fe998b8`), controller +`13bda270c3` (v0.299.0), agent `37e98f452b` (v0.148.0), catalog `d1a148408f`, register 150. The ten rulings were recorded +first (`09` §3 139–148, `8c65ff0c`), with the note that CC's own pick differed on 2, 8 and 9. + +**The operator's cleanup removed every `REPORT*.md`.** Two felhom.eu checks read `REPORT.md`: the decoy test now plants a +missing file and removes it again (`8c65ff0c`), and this file is the session's own report, as the repo rule says. + +## Part A + +- **R-444, measured first** (demo-hp, 09:14Z, `audits/ten-answers-2026-10-06/r444-measure.txt`): `pct fstrim 9201` rc 0 in + 24.4 s; thin pool 65.53 % → 33.40 %; 18/18 app probes 200, slowest 1.1 s. **Built** (agent `ee71abd`): weekly, due + Wednesday from 10:00 host-local, starts only 10:00–20:59, under the one-heavy-op gate (backup, restore-test, OS steps), + 3 tries a week, persisted, reported as `guest_disk_trim`; ONE sudo rule `/usr/sbin/pct ^fstrim [0-9]+$`. Hub + (`a411cde7`): System page „Last disk trim". **Live:** `sudo -l` on demo-hp and demo-felhom allows the trim and refuses + `--ignore-mountpoints`, `;x`, a second vmid and `pct destroy`. The job's first catch-up try ran 2 minutes after the agent + update and failed (the bundle had not landed — expected); the hourly retry at 12:56 local logged + `fstrim: guest 9201 trimmed 2.1 GiB in 2.4s` and the hub page read „10 min ago · 2.1 GiB" (`r444-live.txt`). +- **R-645** (controller `2d63714`): every night leg skips an app whose pin is not what it runs; one amber line on the + backups page. Red-proved on the hand-lift shape. **Residual, stated:** once the boot reconciler starts the app on the + new version, pin = running again — a design question, not built. +- **R-856** (controller `c393d85` + agent route `GET /host/crash-guard`): after a crash boot, app mails wait 15 min. + Live through the real route: demo-hp logged the normal 90 s because its last crash boot (2026-10-05) was not this + start; demo-felhom logged a clean boot. **The crash branch was not shown live** (no crash allowed); tests cover it. + +## Part B — ep0, read only + +Every snapshot of datastore `felhom-offsite`, both from the server's API and from the directories: 9 snapshots in 5 +namespaces, all ≥ 369,808,250 B, all verification `ok`, every directory with its manifest. **No phantom; nothing deleted.** +Runbook `runbooks/pbs-phantom-cleanup.md` (the list command, the two-part test, one `api delete` per proven phantom, the +before/after count control) + `runbooks/pbs-phantom-list.py`; the agent's WARN now names the runbook (`be398f9`). + +## Part C and D + +- **R-747, R-774** (catalog `ec72c9d`, live `1938921`): one first-step sentence each, hu + en, freeze updated. +- **R-734** (catalog `b0939cf`): the update test ignores listed marker files (immich's six `.immich`), each with a reason, + only when changed/added and ≤ 64 bytes; an unlisted file still counts (red-proved). +- **R-624** (catalog `aed80ee`): bench-only vaultwarden seed through its admin invite. The bench LXC 9401 no longer + existed; recreated by its recorded recipe (stopped again afterwards). Proven: admin sign-in, invite and registration + 200; the seed read back before and after; `.env` shredded, secret greps 0 with a working control; without the run flag + the admin route is not tried (`inconclusive`). The move used to drive it (1.36.0-alpine → 1.36.0) failed its health + check — the non-alpine image fails the alpine health check; that is the chosen target, not the seed. **New question: + R-890.** +- **R-502** (felhom.eu `9d39faab`): `iso-bootstrap`, full runs only, NOT CHECKED without docker or the image; the image + `felhom-iso-assistant:trixie` was built on DooPlex; **first real run: 73 checks green, the built-in decoy convicted.** +- **R-618:** closed by the ruling; nothing to build. + +## Said plainly + +- **CI run 1423 (felhom.eu, the hub-manifest commit) was red** on the golden gate: it read controller v0.300.0 a few minutes + before the golden 0.300.0 record was committed. Run 1425 on the next push was green. +- The agent update's first trim attempt failed by design (the sudo rule rides the bundle, which came 10 minutes later). + On a box that gets the agent but not the bundle, the trim warns weekly and its capability reads degraded. +- **R-891:** felhom.eu `CLAUDE.md` still says `--fast` runs every gate — untrue since `iso-bootstrap`. An instruction file; + left for the operator. + +## CI, last commit of every repo + +felhom.eu: this commit (checked by its commit after the push); controller `36088fd` → 1426 success; agent `cefdc73` → 1427 +success; catalog `65130c6` → 1428 success. + +## Teardown + +Drill VM: build guest destroyed, secrets shredded, off, `virgin`. Bench 9401: stopped (kept; its evidence copied off). +Scratch 9202: not used. ep0: read only. demo-hp / demo-felhom: the deliveries, one measured trim, `sudo -l` checks. +Hub: the deploy, the vouches, three floors per release. Scratch secrets shredded at the end. diff --git a/STATUS.md b/STATUS.md index 45665a69..f84bcc45 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,25 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-06 09:00: every box of ours healthy on hub 0.138.0, agent 0.148.0, controller 0.299.0; installer 1.32.0 -is public. The open-items list is at 150. Report: `REPORT-morning-after-2026-10-06.md`.** +**Updated 2026-10-06 13:15: every box of ours healthy on hub 0.139.0, agent 0.149.0, controller 0.300.0. The open-items +list is at 142. Report: `REPORT.md`.** + +## Midday (2026-10-06): your ten answers built; the list at 142 + +- **All ten are done** (your „A" on each). Released to demo-hp, demo-felhom and Tester 1: hub 0.139.0, agent 0.149.0 (and + its root files), controller 0.300.0 and a new install image. +- **The weekly disk trim works:** measured by hand first (the disk pool went from 65 % to 33 % full, the apps did not + notice), then the box's own weekly job trimmed by itself, and the System page shows it. +- **No broken backup leftovers exist on the backup server today**, so nothing was deleted. The steps are written down + for the day one appears. +- **Three of your answers differed from my picks (2, 8 and 9).** Your answer governs, and each is recorded. + +**Needs you (none urgent):** +1. **R-890** — a vaultwarden update can still not be written to the update list: the list needs a proof on the scratch box + too, and your ruling keeps the admin password on the test bench only. Pick: allow it on scratch box 9202 as well. If + nothing: vaultwarden updates stay manual. +2. **R-891** — one line in felhom.eu `CLAUDE.md` is now untrue (it says the quick test run covers every check; the new ISO + test runs only in full runs). Only you may edit that file. ## This morning (2026-10-06): your two answers done; the list at 150 @@ -13,21 +30,6 @@ is public. The open-items list is at 150. Report: `REPORT-morning-after-2026-10- - **The four waiting app fixes passed on the scratch box and are live** (visitor addresses for kimai, zipline, vikunja, nextcloud; nextcloud's health check). Each was tested against a control. -## Ten questions for you — answer „all as picked", or name the ones you want differently - -| # | Question | Option A | Option B | If you decide nothing | My pick | -|---|---|---|---|---|---| -| 1 | **R-444** — Should the boxes trim their customer guests' disks once a week, so deleted data stops filling the disk pool? | Yes: one new admin permission (`pct fstrim`), weekly, outside the night, measured once on demo-hp first | No: leave it; the pool keeps space nobody uses | Nothing changes; a full pool can one day stop every guest on that box | **A** | -| 2 | **R-99** — Should broken leftovers of aborted backups be deleted on the backup server? | Yes, by hand, by a runbook, when one is seen | No: they are detected, harmless, and do not affect what is kept | They stay; one small leftover per aborted upload | **B** | -| 3 | **R-618** — May an app update count Docker's own „healthy" as proof that the new version works? | No: keep our own check only | Yes: Docker's healthy can end a wait early | Nothing changes (A) | **A** | -| 4 | **R-645** — When you lift a held update by hand, how do we stop the night backup from saving the broken version over the good copy? | The night backup skips an app whose saved version is not what it runs | Lifting a hold by hand is refused; you use „Undo" instead | The good copy can be overwritten within seconds of a hand lift | **A** | -| 5 | **R-856** — After a crash restart, should app mails wait longer (about 15 minutes), so the household gets one message, not three? | Yes: a longer quiet time after a crash boot (two repositories change) | No: close it; one crash can send several true mails | Several mails after a crash, as now | **A** | -| 6 | **R-747** — Should mealie's app page tell the household that five wrong logins lock the account for 1–2 hours? | Yes: one sentence on the page | No | A locked household does not know why or for how long | **A** | -| 7 | **R-734** — Should the update test ignore small marker files an app rewrites at every start (immich)? | Yes: a per-app list of such files, each with a reason | No: keep marking it; immich updates then need a fresh full copy first | immich's night update is skipped where no fresh copy exists | **A** | -| 8 | **R-624** — May the update test hold an app's admin password to create test data in apps that refuse strangers (vaultwarden, zipline)? | Yes, on the test bench only | No: write down that these apps are tested by hand | No change; those two apps stay hand-tested | **B** | -| 9 | **R-502** — May a slow check that starts Docker containers run on DooPlex (the ISO's first-boot test)? | Yes, in full runs only (never on every push), with a clear „not checked" when Docker is missing | No: it stays a hand step of each ISO release | It stays a hand step; it can be forgotten | **B** | -| 10 | **R-774** — Should Karakeep's page say that its phone app sends crash reports to its makers? | Yes: one sentence on the page | No | The household is not told | **A** | - ## Tonight (2026-10-06): the list at 164 **What happened:** 36 rows closed, 1 opened. Released and delivered the normal way to demo-hp, demo-felhom and Tester 1: diff --git a/documentation/audits/ten-answers-2026-10-06/r444-live.txt b/documentation/audits/ten-answers-2026-10-06/r444-live.txt new file mode 100644 index 00000000..2c5dc862 --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/r444-live.txt @@ -0,0 +1,6 @@ +== R-444 live, 2026-10-06 +agent journal demo-hp: 11:56:18 local 'fstrim: guest 9201 trim FAILED after 0.0s … sudo: a password is required' (before the bundle; attempt 1 of 3) +bundle installed 12:06:31 local (BUNDLE DONE, capability probe 68/68) +agent journal demo-hp: 12:56:20 local 'fstrim: guest 9201 trimmed 2.1 GiB in 2.4s' bytes_trimmed=2238750720 mounts=2 +hub System page (a different channel), 11:07:01Z: demo-hp 'Last disk trim' = '10 min ago · 2.1 GiB' +thin pool after: data 33.44 %, vm-9201-disk-1 22.99 % diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 830ade8e..b79990e8 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,14 @@ --- +## 2026-10-06 (midday) — R-444 live + +The full text of every row below: `git show bfdea832:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** (P3) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 139): a weekly disk trim of each customer guest, daytime only | Measured by hand on demo-hp first (`audits/ten-answers-2026-10-06/r444-measure.txt`): `pct fstrim 9201` rc 0 in 24.4 s, thin pool 65.53 % → 33.40 %, 18/18 app probes 200. Built: agent `ee71abd` (v0.149.0, `internal/fstrim`; due Wednesday from 10:00, starts only 10:00–20:59, under the one-heavy-op gate, 3 tries a week, persisted, reported as `guest_disk_trim`) + ONE sudo rule `FELHOM_FSTRIM` `/usr/sbin/pct ^fstrim [0-9]+$` in the bundle + hub `a411cde7` (v0.139.0, System page „Last disk trim”). Live: `sudo -l` on demo-hp and demo-felhom allows `pct fstrim 9201` and refuses `--ignore-mountpoints`, `;x`, `9201 9202`, `pct destroy`; the job's first catch-up try failed before the bundle (expected), the hourly retry logged `fstrim: guest 9201 trimmed 2.1 GiB in 2.4s` and the hub page read „10 min ago · 2.1 GiB” (`r444-live.txt`). | + ## 2026-10-06 (midday) — the operator's ten answers built The full text of every row below: `git show 929e59e8:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index c60601b3..65f5893b 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -229,7 +229,6 @@ stopping line that lies. | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | | **R-242** | Box system & updates | P3 | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | — | — | operator | -| **R-444** | Box system & updates | P3 | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** A scheduled trim needs a new root grant (`pct fstrim `, widening what R-861 narrowed) and a new scheduled action on every customer guest's disks. Next: an operator yes on the grant and a cadence, then one measured run on demo-hp. **RULED 2026-10-06 10:41 (`09` §3 decision 139): A — weekly `pct fstrim` of the customer guest, outside the night; measured on demo-hp first.** Being built. | — | — | CC | | **R-468** | Box system & updates | P3 | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | — | — | CC | | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |