register: compress R-459 and R-467 to CLOSED-ITEMS (full text at ae59c31); REPORT sizes + CI
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -32,7 +32,7 @@ controller CHANGELOG header, and **the last four headers carry no such line**
|
||||
| repo | before | after (pushed to `main`) |
|
||||
|---|---|---|
|
||||
| app-catalog-felhom.eu | `b7ef0c4a09d6` | `eec1228` templates → `bd32830` gate → `3525e35` CHANGELOG/REPORT |
|
||||
| felhom.eu | `4b2e5608c227` | this push (two commits: documents + gate, then the register compression) |
|
||||
| felhom.eu | `4b2e5608c227` | `ae59c31` (documents + gate + evidence) → the compression commit on top |
|
||||
| felhom-controller | `155271672265` (v0.236.0) | **unchanged — no code.** Golden bake only. |
|
||||
|
||||
Highest R-id before: 467. Minted: **R-468** (waiver), **R-469** (engine-major rule expiry),
|
||||
@@ -203,6 +203,12 @@ the last per-release one. **No `--no-verify` anywhere in this session.**
|
||||
|
||||
No Go code changed in any repo; `go build/vet/test` not applicable.
|
||||
|
||||
`python3 scripts/unproven.py --summary`: 55 claims, walked 20 / partial 17 / built 14 / missing 4,
|
||||
**NOT WALKED 35 of 55 — unchanged** (nothing in this session walked a capability claim).
|
||||
|
||||
**CI, by `head_sha`:** app-catalog job **536** for `3525e35` → `completed` / `success`. The
|
||||
felhom.eu run for this push is quoted in the follow-up commit that records it.
|
||||
|
||||
## 7. Register — rows opened / closed, size
|
||||
|
||||
Closed: **R-459** (harness + live evidence), **R-467** (the bake). Narrowed: **R-242** (waiver half built;
|
||||
@@ -210,7 +216,8 @@ vouch half open). Opened: **R-468** (WATCHING — the waiver; renew ≤ 14 days
|
||||
first external install), **R-469** (BLOCKED on R-448 — remove the engine-major rule), **R-470**
|
||||
(READY — MinAgent header line), **R-471** (READY — observations decoy hole, pre-existing).
|
||||
Register size: **210 open rows / 169 closed before → 212 open / 171 closed after** (R-459 and R-467
|
||||
compressed to `CLOSED-ITEMS.md`; nothing deleted). OPEN-ITEMS.md bytes: see the compression commit.
|
||||
compressed to `CLOSED-ITEMS.md` in the second commit, citing `ae59c31` for the full text; nothing deleted).
|
||||
`OPEN-ITEMS.md`: 433 845 → 434988 bytes with the four new rows and two closures written out, → 429692 bytes after compression.
|
||||
|
||||
## 8. Evidence off the machine before every teardown
|
||||
|
||||
@@ -250,10 +257,10 @@ Helper files on demo-hp (`ctrl_pw`, `dh_user`, `dh_pat`, `upg.tgz`, `guest-setup
|
||||
section and not an appended second one. Reproduced on a clean worktree of `4b2e560`. **FILED: R-471.**
|
||||
2. **Four consecutive controller CHANGELOG headers carry no `MinAgent:` line** while the vouch runbook
|
||||
says to read it from the header. **FILED: R-470.**
|
||||
3. **`POST /api/stacks/<app>/restart` recreated only the changed service** (`bookstack-db` got a new
|
||||
ID; the `bookstack` app container kept its 4-hour uptime) — it is `compose up -d`, as
|
||||
`09-update-architecture.md` §1.3 records as a chosen behaviour. **NOT-A-FINDING: documented
|
||||
design; the message "restart completed" is accurate for a compose-level restart.**
|
||||
3. **The stack restart recreated only the changed service** — bookstack-db got a new ID, the bookstack
|
||||
app container kept its 4-hour uptime — because the restart is compose up -d, which the update
|
||||
architecture §1.3 records as a chosen behaviour.
|
||||
**NOT-A-FINDING: documented design, and "restart completed" is accurate for a compose-level restart.**
|
||||
4. **`mariadb:12.3` resolved to 12.3.3 in the harness and 12.3.2 on 9201** — the floating-pin class.
|
||||
**NOT-A-FINDING: already R-446.**
|
||||
5. **`target-selection.md` names `/mnt/nvme-1tb`, which does not exist on demo-hp.** **NOT-A-FINDING:
|
||||
|
||||
@@ -256,3 +256,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-408** | **`RestoreOffboxScratch` took no single-writer flag while a comment asserted every off-site operation did.** Shipped in controller **v0.232.0**. **Reasoning kept:** *the real deliverable is the WALK, not the acquire* — the sentence was false for months and nothing checked it, the ninth instance of this project's most-repeated class; *it is an AST pass and not `strings.Contains`, because a commented-out call still contains the string*; *adding a line to `offsiteExempt` is a deliberate act and belongs in the commit that adds it*; *the R-87 proof's exemption is kept HONEST by a second test that fails if that path ever gains `unlockStale`, routes through `resticStep`, or loses `--no-lock`*. | **CLOSED — SHIPPED** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-411** | **A background job deleted the lock of a live customer restore and logged it as a crash that did not happen.** Shipped in controller **v0.232.0**. Evidence: `audits/DRILL-soak-2026-08-31/phase1-lock-collision/`, `audits/R411-R414-2026-09-01/`. **Reasoning kept:** *`restic stats` TAKES a repository lock* — the fact nobody had, and the one that made the chain reachable; *`restic check` takes one too, `restic snapshots` and `restic list` do not*; *the fix was wider than the row — FOUR entry points were unflagged, three of them found by R-408's walk rather than by the report*; *the escalation in `resticStep` was NOT removed — real stale locks exist and it clears them; the defect was that a sibling could be live*. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.232.0, 2026-09-01) | full text: `git show 22e1c95:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-459** | **The skipped MariaDB conversion is STABLE but never self-resolving; converting costs 7 s and keeps the abort — RULED YES and shipped 2026-09-13: `MARIADB_AUTO_UPGRADE=1` on `bookstack-db`, `kimai-db`, `nextcloud-db`, `romm-db`** (catalog `eec1228`/`bd32830`/`3525e35`; no image moved, `catalog_since` untouched). Measured `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md`; proven `audits/r459-close-2026-09-13/` — harness E3/E3b `proven` with `engine_state_after` = `already upgraded to 12.3.3-MariaDB [exit=1]`, `skipped due to $MARIADB_AUTO_UPGRADE` 0 lines, C3 still `failed`; landed on demo-hp via the real 15-min cycle with both container IDs unchanged, one deliberate restart → `MariaDB upgrade not required`, `/login` 200. **Reasoning kept:** *ask the engine, not the log* — the entrypoint prints `MariaDB upgrade not required` on an unsupported downgrade too (R-464); `mariadb-upgrade --check-if-upgrade-is-needed` exit 0 = needed, 1 = not. *Not established, unchanged:* whether any MariaDB feature misbehaves on an UNCONVERTED datadir. *The precaution that keeps the setting inert until Slice 4:* the engine-major rule + gate, removal tracked as R-469. | **CLOSED 2026-09-13 — shipped in the catalog, PROVEN by harness and live** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-467** | **Controller v0.236.0 owed a golden — PAID 2026-09-13:** golden `0.236.0` baked (`GOLDEN_SHA256=58a3cc24…958bf`, 654 115 664 B), round-tripped from the DOWNLOADED bytes, `./etc/felhom-controller-image` says `felhom-controller:0.236.0`, hub dropdown agreed, three-field vouch re-read (`0.236.0` / `agent 0.130.0` / `min_agent 0.129.0`, not the R-216 shape), floor raised 0.232.0 → 0.236.0. Evidence `documentation/tests/golden-0.236.0-2026-09-13/`. **Reasoning kept:** *this bake carried FOUR unbaked releases and is the LAST per-release bake* — goldens are weekly and before any install from today (R-468); *the MinAgent line was missing from four headers* (R-470). | **CLOSED 2026-09-13 — baked, vouched, floor raised** | full text: `git show ae59c31:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -690,7 +690,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-456** | **[P3-LOW] A partly-dead stack is not a boot orphan, and that is written down nowhere.** MEASURED 2026-09-02 on demo-hp while validating v0.233.0: `docker rm -f bookstack` (leaving `bookstack-db` running) then a controller restart produced `Boot reconciliation: 1 boot-orphaned app(s) found: [bentopdf]` — **bookstack was NOT selected**, although the app container was gone and `desired_state: running` was recorded. Removing `bookstack-db` as well made the whole stack orphaned and the very next pass repaired it in 6.3 s. **So `bootrecon.isBootOrphan` requires the stack as a WHOLE to be down; one live member is enough to make it invisible to the reconciler.** **NOT called a defect, and the reason is part of the row:** `StateDegraded` IS in `IsDownState`, and the crash-loop/dead-app alarm path (`classifyRunStates`) does count a degraded stack as down — so the customer IS told; it is the automatic REPAIR that does not fire, and there may be a good reason (repairing half a stack while its DB is live is not obviously safe). **What is certain is that nobody has written the rule down**, so the next session re-derives it the same way this one did — by watching a reconciliation not happen, which is an absent observable and the weakest possible evidence. Either state the rule in `02-controller-module-map.md` with a test pinning it, or change it. Owner: **CC.** `tests/VALIDATION-update-slice12-2026-09-02.md` §2.2 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-459** | **[P3-LOW, was P2] MEASURED 2026-09-06 — the skipped MariaDB conversion is STABLE but never self-resolving, and fixing it costs 7 SECONDS and does NOT cost the ability to go back.** The consequence R-459 deliberately left unmeasured is now measured: `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md`. **(1) It does not degrade: 5 of 5 restarts of 12.3 on the 11.6 datadir, readback passed every time, `mariadb_upgrade_info` unchanged at `11.6.2-MariaDB`, and the entrypoint's line never escalated past `[Note]`.** **(2) It never heals either** — the engine answers `Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!` on every start, and will forever. **(3) THE TRADE THIS ROW WAS EXPECTED TO PRODUCE DOES NOT EXIST.** The fear was that converting properly would end the ability to abort. Measured: with `MARIADB_AUTO_UPGRADE=1` the conversion **succeeds** across the multi-major jump (so it is not a stepping problem either), takes **7 s**, takes its own system-database backup first (`system_mysql_backup_11.6.2-MariaDB.sql.zst`), and **putting 11.6 back afterwards still starts and serves the data**. **So the choice is no longer a trade; it is a cheap correction.** **WHAT REMAINS OPEN IS THE DECISION, NOT THE MEASUREMENT:** setting `MARIADB_AUTO_UPGRADE` changes how **four** apps upgrade — bookstack (the only one that has moved a major), kimai, nextcloud, romm — and R-459's original owner line reserves a fleet-wide env change for the operator. **Nothing was committed to any template**; the comparison arm used a scratch copy inside a throwaway guest. **NOT ESTABLISHED, and not smuggled in: whether any specific MariaDB FEATURE misbehaves on unconverted system tables.** This run exercised BookStack's normal read/write path only, over minutes, with one seeded record. **Rank dropped P2 → P3** because the failure mode is now bounded by measurement rather than unknown. Raised in `STATUS.md`. **✅ RULED YES 2026-09-13 and SHIPPED the same day** (catalog `eec1228`/`bd32830`/`3525e35`): `MARIADB_AUTO_UPGRADE=1` on `bookstack-db`, `kimai-db`, `nextcloud-db`, `romm-db`; `MARIADB_DISABLE_UPGRADE_BACKUP` unset; no image line moved so `catalog_since` did not. **Proven before it shipped** (throwaway LXC 9403 on demo-hp, destroyed; `audits/r459-close-2026-09-13/harness/`): C3 still `failed`; E3 and E3b `proven` with `engine_state_after` = `12.3.3-MariaDB | This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again. [exit=1]`; the entrypoint printed `Backing up system database to system_mysql_backup_11.6.2-MariaDB.sql.zst` → `Starting mariadb-upgrade` → `Finished mariadb-upgrade` (6 s) and `skipped due to $MARIADB_AUTO_UPGRADE` **0** times. **Watched as it landed** (`live-9201/`): the change travelled the real 15-minute cycle to demo-hp at 08:09:17 UTC — `[INFO] [sync] Updated bookstack/docker-compose.yml`, `stored definition refreshed with the delivered fix` — the live compose gained the setting, **both container IDs and `StartedAt` were identical to the pre-push baseline (nothing recreated)**, and the one deliberate `POST /api/stacks/bookstack/restart` logged `MariaDB upgrade not required`, the engine's own check said `already upgraded to 12.3.2-MariaDB [exit=1]`, and `/login` served 200. **The precaution that keeps the setting inert:** the engine-major rule + gate (R-469). **Still not established, unchanged:** whether any MariaDB feature misbehaves on an UNCONVERTED datadir — moot for boxes that convert from now on, live for any datadir converted before today (none: only bookstack ever moved a major, and its demo datadir was born on 12.3). | **CLOSED — 2026-09-13** |
|
||||
| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-461** | **[P3-LOW] `runbooks/target-selection.md` names a venue that does not exist and fences a fixture that is gone.** MEASURED 2026-09-06 on demo-hp while siting the R-449 guest. (a) The runbook says to put VM disks on a dir storage at **`/mnt/nvme-1tb`, at its root**. **There is no `/mnt/nvme-1tb`** — the 1 TB NVMe is mounted at **`/mnt/hdd_1`**, which is the enrolled user-data drive and the `felhom-backup` target, i.e. the same disk under a different path. The instruction's REASON is still exactly right (`local-lvm` is an over-subscribed thin pool backing the live guest 9201, and this run kept off it — `local-lvm` read **30.50 % before and after**), so the fence held; only its address is stale. (b) The runbook forbids destroying **`drill-r50` (VM 300)**, "the only drift fixture (R-93)". **`qm list` returns nothing on demo-hp** — there are no VMs at all. **The fence currently protects nothing, and R-93's premise that a drift fixture exists is false.** **Why it is a row and not a quiet edit:** a runbook that names a path nobody can find is one a session works around, and working around a safety instruction is how the instruction stops being followed. Both halves need checking against the box before the text is changed — (b) in particular may mean R-93 should be closed or reopened as "the drift fixture is gone", which is a different fact from "do not destroy it". Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §7 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 | **READY — rank P2-MEDIUM; owner: VIKTOR rules on scope, CC implements** |
|
||||
@@ -698,7 +697,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-465** | **[P3-LOW] `cfg.Paths.HDDPath` — the global that R-442 proved is set on NO box (demo-hp AND demo-felhom: 0 `hdd_path`, 0 `FELHOM_PATHS_*`) — still has SIX readers, each reading an always-empty value:** `report/builder.go:69`, `monitor/healthcheck.go:35`, `api/router.go` (system-info), `web/server.go:740`, `cmd/controller/main.go` (auto-discovery seed + metrics HDD path). Removal was silently inert for months on the very same read. Whether any of these is inert the same way — a report field that is always empty, a health check that never fires, a metric never collected — is a one-hour audit: for each reader, name the POSITIVE observable that must appear when it works and check it on the box. R-442 §5 said "do not delete it here"; this row is the audit it deferred. Owner: **CC.** `felhom-controller/REPORT.md` (v0.236.0, Observations 1) | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-466** | **[P3-LOW] Removing an app with „Mentési adatok törlése" ticked leaves its recovery unit's `compose/` + `manifest.json` on the drive.** The router passes only `backup.AppDBDumpPath(nsRoot, name)` to `RemoveStack`, so `backups/primary/<app>/db-dumps/` goes and the unit root keeps `compose/` (the app's `app.yaml` with the portable secret class at 0600) and `manifest.json`. MEASURED 2026-09-13 on demo-hp after the R-442 Scenario A removal — `audits/R442-2026-09-13/teardown-and-log.txt`, residue check: `backups/primary/nextcloud` still listed with the app gone; removed by hand at teardown. A customer who asked for the backups to go is left with the app's definition and a manifest. **Decide:** the button means the WHOLE unit (pass the unit root, under the same `backups/`-prefix guard) or stays db-dumps-only (then the modal must say so). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-467** | **[P3-LOW] Controller v0.236.0 owes a golden** (golden-notice gate, 2026-09-13): newest golden baked is **0.232.0**, so a machine installed now receives 0.232.0 and reaches 0.236.0 only by self-update. Bake + vouch per `runbooks/RUNBOOK-manual-build.md` §4.1 (the THREE-field vouch: `golden_version` + `agent_version` + `min_agent`); bake record to `documentation/tests/golden-<ver>-<date>/`. If the train deliberately skips this version, this row is the waiver the gate asks for. Owner: **CC.** **✅ PAID 2026-09-13 — golden `0.236.0` baked, published, round-trip verified, VOUCHED, floor RAISED 0.232.0 → 0.236.0.** Evidence `documentation/tests/golden-0.236.0-2026-09-13/`: bake `GOLDEN_SHA256=58a3cc24…958bf`, round trip 654 115 664 B same sha hashed from the DOWNLOADED bytes, `./etc/felhom-controller-image` out of the archive says `felhom-controller:0.236.0`, the hub's Day-0 dropdown read the same sha on its own code path; three-field vouch re-read from the page (`golden_version 0.236.0`, `agent_version 0.130.0`, `min_agent 0.129.0` — NOT the R-216 shape), R-120 banner absent; floor re-read `0.236.0`. **This bake carried FOUR unbaked releases (0.233.0…0.236.0)** and it is the LAST per-release bake — goldens are on a weekly cadence from today (R-468). Both demo guests already ran 0.236.0 by hand, so the floor moved nothing; the chain was last exercised 2026-09-01. `golden_currency_gate.py` went red → green on the same command. The MinAgent line was missing from four headers — R-470. | **CLOSED — 2026-09-13** |
|
||||
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver.yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. | **BLOCKED — on R-448; rank P3-LOW; owner: CC** |
|
||||
|
||||
Reference in New Issue
Block a user