docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
This commit is contained in:
@@ -261,3 +261,6 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-448** | **UPDATE ARC SLICE 4 — a guarded update: verified-backup precondition, abort path, truth at the moment of action.** Shipped controller **v0.237.0** (job) + **v0.238.0** (page) + **v0.238.1** (nightly legs skip an app mid-update, found live). Proven live on demo-hp 2026-09-13, scenarios A/B/E/F/H and the restore walk. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *the precondition is the existing verified backup, not a new copy* (ruling 2026-09-02) — `backup.Tier2UnitRestorePoint`, extracted from the backups page, not copied; *age a copy by its last SUCCESSFUL Tier-2 copy, never the manifest `created_at`* (measured: the manifest moves only on definition changes); *no automatic rollback — measured per-app, the route back is the restore*; *anything that writes a restore point skips an app that is held OR updating*. Open consequences: R-472, R-475, R-476. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-443** | **The Update button reported success over an app it had just broken.** Closed by slice 4 (v0.237.0): 202 `accepted, not completed`; the outcome exists only as `update_phase` after health. Pinned by `TestR443_UpdateIsNeverReportedCompleteSynchronously`. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a compose exit code is never a success signal* (spike §4: HTTP 200 over a crash loop). | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-439** | **The restore hold was not honoured by the update path.** Closed by slice 4 (v0.237.0): `update` joined the router's hold check; live Scenario H refused start/restart/update and the boot sweep. Evidence: `audits/slice4-2026-09-13/`. **Reasoning kept:** *a hold that only one path honours is not a hold* — the audit found the drive-return gate (restart + boot recreate) and the nightly volume dump ignoring any hold; all three fixed and red-proofed. | **CLOSED 2026-09-13** | full text: `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-470** | **Four controller CHANGELOG headers (v0.233.0–v0.236.0) carried no `MinAgent:` line, while the vouch and now the declared floor read it from the header.** Closed 2026-09-13 (`felhom-controller` `f946b0d`): the four headers backfilled with `**MinAgent: 0.129.0** (unchanged)` — v0.232.0's value, proven unchanged (no commit under `internal/agentapi` since 2026-09-01; highest `featureMinAgent` 0.129.0) — and `controller/scripts/minagent_header_gate.py` (fast, blocking) refuses a newest header without the line; a prose or code-span mention does not count (decoy; red-proof F). | **CLOSED 2026-09-13 — GATED** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-472** | **The golden cadence ruling and the hub's floor rule contradicted each other: a floor above the vouched golden delivered nothing.** Operator ruling 2026-09-13, hub **v0.112.0** (`f181efd`): a floor saved with the release's declared MinAgent is served above the golden under the same agent comparison; an undeclared one is still held and both forms refuse it (`floor_needs_min_agent`). Proven live: controller 0.239.0 reached demo-hp in 14 s and demo-felhom in 15 s from the save, hub `managed floor SERVED … from declared`. Evidence: `audits/rulings-r472-r475-2026-09-13/` 02, 03. **Reasoning kept:** *the manifest leads the floor inside the golden; above it, the release's own declared MinAgent does* (publish-train rule 1); a declaration binds to its exact floor. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-475** | **The update precondition was Tier-2-only, so an app with no second-drive copy could not be updated.** Operator ruling 2026-09-13, controller **v0.239.0** (`b93c154`): the first fresh copy in the order Tier 2, Tier 1, Tier 3 (bounded, unreachable = absent + WARN); `backup_max_age` applies to the chosen tier; nothing anywhere → back up first; refused only when no copy and no backup can be taken; the hold names the tier; a Tier-2 failure in the pre-backup is a WARN. Proven live on demo-hp: nothing anywhere → backed up first (04), Tier 1 alone (05), held naming „saját meghajtó” (07), restored from „helyi” (08). Red-proof M (age only on Tier 2) fails. **Reasoning kept:** *first FRESH copy, not first copy* — a stale mirror must not force a backup while the own unit is minutes old. Follow-ups: R-477..R-480. | **CLOSED 2026-09-13 — PROVEN-LIVE** | full text: `git show 2f5d3af:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -697,15 +697,16 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
|
||||
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
|
||||
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). | **READY — unblocked by R-448; rank P3-LOW; owner: CC (removal is a deliberate act)** |
|
||||
| **R-470** | **[P3-LOW] Four consecutive controller CHANGELOG headers (v0.233.0 … v0.236.0) carry NO `MinAgent:` line, and `RUNBOOK-manual-build.md` §4.1 step 5 tells the vouch to read it from the header.** Measured 2026-09-13 while baking golden 0.236.0: `sed -n 1,349p CHANGELOG.md | grep -c MinAgent` → **0**; the newest statement is v0.232.0's `**MinAgent: 0.129.0** (unchanged)`. The vouch used 0.129.0 on the ground that no later entry declares a change and `internal/agentapi/features.go` gates per feature, not per release — but a reader following the runbook literally finds nothing to read, which is the R-233 shape (a document pointing at a string that is not there). **Fix:** either every release header carries the line (the v0.232.0 convention), or the runbook says where MinAgent actually lives when a header omits it. One of the two, not both. | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-471** | **[P3-LOW] `observations_gate.py` reads the FIRST `Observations` section of `REPORT.md` and nothing after it, so an appended second section with an unmarked item passes — and the decoy that proves it has been reading "LIVE HOLE" at HEAD, unnoticed, because the decoy suite is run by hand.** MEASURED 2026-09-13 on a clean worktree of `4b2e560`: `python3 scripts/test_gate_decoys.py` → `FAIL: observations/R-419: decoy PASSED - LIVE HOLE (rc=0)`; the gate's own output shows it scanned `## 11. Observations — noticed, documented, NOT acted on` (3 items, all marked) and never reached the appended `## Observations` block the decoy planted. Whether the cause is "first heading wins" or a heading-shape filter is NOT established — only that the planted unmarked item was not seen. **Consequence:** a report with two observation sections gets the second one unchecked. **Two fixes, both owed:** scan every section whose heading contains `Observations`, and put `test_gate_decoys.py` where something runs it (it is the instrument for R-421 and nothing in `repo_gates.py` invokes it). | **READY — rank P3-LOW; owner: CC** |
|
||||
|
||||
| **R-472** | **[P2-MEDIUM] THE GOLDEN CADENCE RULING AND THE HUB'S FLOOR RULE CONTRADICT EACH OTHER: a floor raise delivers NOTHING until a golden carries the release.** The 2026-09-13 ruling (R-468) rests on *"every release still raises the floor, so both demo boxes keep getting each release in ~20 s; only the golden moves to a cadence"*. MEASURED THE SAME DAY, releasing controller v0.237.0: `POST /configuration/global-floor` → `0.237.0` (303, re-read), and the hub logged for BOTH boxes `managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0, so its agent requirement is unknown — vouch a golden carrying the floor's controller (publish-train rule 1) (controller floor withheld)`; the box logged `SetFloor: floor "0.236.0" → ""`; neither box moved in 8 minutes. Publish-train rule 1 (hub `ResolveManagedFloor`) was built so a controller is never pushed past the agent it needs, and it reads that requirement from the vouched golden. **So under a weekly golden, releases between bakes do not reach the fleet by floor at all** — they need a hand deploy. v0.237.0 and v0.238.0 were hand-deployed to both demo guests by the skill's documented route and the floor was put back to 0.236.0 (a floor held on every box is a standing dashboard flag). Evidence `audits/slice4-2026-09-13/live/00-floor-0.237.0.txt`. **This is the operator's call, not CC's:** (a) bake per release again (the treadmill R-468 ended), or (b) let the floor carry a release above the golden when its CHANGELOG header states an unchanged `MinAgent` — which needs R-470's header line to be reliable first. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-473** | **[P2-MEDIUM] The `glance` catalog template crash-loops on EVERY fresh install.** MEASURED 2026-09-13 on demo-hp (deployed as a throwaway for the slice-4 live test): `restarts=13 status=restarting`, the log repeating `parsing config: reading /app/config/glance.yml: open /app/config/glance.yml: no such file or directory`, the `glance_config` volume empty. The image does not create a default config and the template seeds none. The install page reports the deploy as started and the card then shows a restarting app. Not a version problem: the template's current tag v0.8.5 has the same config requirement (the catalog walk that pinned it never did a fresh install). **Fix shape:** seed a minimal `glance.yml` the way `gokapi`'s catalog entry seeds `config.json` (memory `gokapi-headless-setup`), then prove a fresh install lands healthy. Evidence `audits/slice4-2026-09-13/live/02-glance-abandoned.txt`. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-475** | **[P2-MEDIUM] The update precondition is Tier-2-ONLY, so an app with no Tier-2 copy cannot be updated at all — even when its primary unit or off-site copy could restore it.** Slice 4 (v0.237.0) gates the Update button on `backup.Tier2UnitRestorePoint`, exactly as specified. MEASURED on demo-hp 2026-09-13: `gokapi` and `nextcloud` have NO `cross_drive` record, so both are refused with the no-backup sentence. The classes this reaches: a box with no second target (`no_target`), an app whose customer switched Tier 2 off, and a freshly installed app before its first Tier-2 run. **It is a finding about the specification, not a defect in the build:** the primary recovery unit (same drive) and the off-site unit (Tier 3) are both restorable routes the predicate does not consider — the driveless-app class R-356 names. Options: accept Tier 2 as the only update route (and say so in the refusal), or widen the predicate to the primary unit with an explicit "same drive" caveat. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules on the route set, CC implements** |
|
||||
| **R-474** | **[P3-LOW] "Remove app" with *also delete backups* deletes only the `db-dumps` directory, and reports `volumes_removed: null` over a volume it DID remove.** MEASURED 2026-09-13 removing the glance throwaway on controller v0.237.0 (`remove_hdd_data:true, remove_backups:true`): response `{"removed":"glance","volumes_removed":null,"hdd_paths_removed":[],…}`; afterwards `docker volume ls` shows no glance volume, while `backups/primary/glance/{compose,manifest.json}` and the stack dir's `applied-compose.yml` remain. The router passes ONLY `backup.AppDBDumpPath(nsRoot, name)` (router.go, beside a comment that disk-tier backup "moved to the host agent" — stale since Tier 2 returned to the controller). So the recovery unit and any Tier-2 copy survive a removal that promised to delete backups, and the answer names no volume. Same class as R-442: the removal's answer does not describe what happened. (Removal also refuses a crash-looping app as "still running — stop it first", which is correct and was observed.) **REPRODUCED a second time the same day** removing the uptime-kuma throwaway: `volumes_removed: null`, and `backups/primary/uptime-kuma/{compose,manifest.json,volume-dumps}` plus `backups/secondary/uptime-kuma/recovery-unit` survived `remove_backups:true` (residue cleared by hand; `audits/slice4-2026-09-13/` `live/15-teardown.txt`). **REPRODUCED a third time 2026-09-13 (afternoon, v0.239.0)** removing the gokapi and actualbudget throwaways with `remove_backups:true`: both `backups/primary/<app>/` units survived (actualbudget's with its 74 KB volume tar), `volumes_removed: null` again, and both `app_backup` prefs remained (`10-teardown.txt`). It also fed R-478. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-476** | **[P3-LOW] The Mentések page names a Tier-2 copy's date from the unit MANIFEST, which moves only when the app's DEFINITION changes — so it can undersell a fresh copy by a day or more.** MEASURED on demo-hp 2026-09-13: bookstack's mirror held `bookstack-mariadb.sql` written 2026-09-13T00:30Z under a manifest dated 2026-09-12T02:15:29Z; the Tier-2 run succeeded at 01:30Z. A capture rewrites the manifest only when checksums, dump NAMES or controller version change, and nightly dumps keep their names. R-403 chose the package date so a PRESERVED package is never shown as fresh — correct — but for a normal run it is the older, flattering-in-reverse date. The update (slice 4) deliberately ages the copy by the last successful copy instead (`Tier2RestorePoint.ProvenCopyTime`), so the two can name different dates. Fix shape: record the unit's DATA time (newest dump mtime) in the manifest, or name `LastSuccess` when the leg was not preserved. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-477** | **[P3-LOW] The update's Tier-3 lookup pays its full 15 s bound on demo-hp, and the bound shows up as a WARN about a DIFFERENT app.** MEASURED 2026-09-13 (controller v0.239.0, Scenario K): the `checking` phase ran 15:25:45 → 15:26:00, and the only line in the window was `[WARN] [offbox] inventory: size of kimai's newest snapshot unknown: offbox stats b4350226: signal: killed`. Cause: `UpdateRestorePoints` calls `OffsiteInventoryList`, which runs one `snapshots` AND one `stats` per app; the update needs only the snapshot times, and the 15 s context killed a `stats`. **Consequence:** every update with no fresh Tier 2/Tier 1 copy waits up to 15 s before backing up, and the operator log blames kimai for an update of actualbudget. **Fix shape:** a snapshots-only lookup for the update path (no sizes), so the bound is a guard and not the normal duration. `audits/rulings-r472-r475-2026-09-13/04-K-no-unit-anywhere.txt` | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-478** | **[P3-LOW] A recovery unit LEFT BEHIND by a removed install counted as the fresh Tier-1 restore point of a reinstall.** MEASURED 2026-09-13 (v0.239.0, Scenario H): gokapi had been removed earlier; its unit (manifest 06:59:17Z) survived (R-474). A new gokapi was deployed at 15:30:46Z and updated at 15:31:10Z: `precondition met — Tier 1 (own recovery unit) copy from 2026-09-13T06:59:17Z (8h32m0s old)`. That unit belonged to the OLD install — a different `app.yaml` (subdomain, generated password). The capture sweep rewrote it only at 15:34:22Z. **Consequence:** in that window a failed update would name, and a restore would apply, the previous install's definition. **Fix shape:** R-474 deleting the unit on removal closes most of it; independently, a unit whose manifest predates the stack's current deploy should not count. `05-H-gokapi-tier1.txt`, `06-F-setup-unit-rewritten.txt` | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-479** | **[P2-MEDIUM] For an app whose data is a bind mount, the Tier-1 unit holds settings only — so the route back a Tier-1 hold names restores the definition and NOT the data.** MEASURED 2026-09-13 restoring gokapi from „helyi” after a held update: `a beállítások visszaálltak … FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem`. The restore message is honest; the HOLD sentence („Visszaállítható … saját meghajtó”) does not say it. The ruling (R-475) accepts any tier, in the order 2, 1, 3 — so for such an app a fresh Tier-1 unit is chosen ahead of an off-site snapshot that WOULD carry the data. **Consequence:** an update whose migration rewrote bind-mounted data has no data route back through the copy it named. **Decision-shaped:** either Tier 1 counts only for apps whose unit carries their data (DB dump / volume tar), or the hold sentence says "settings only" for that case. `07-F-hold-names-own-drive.txt`, `08-restore-from-helyi.txt` | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: VIKTOR rules, CC implements** |
|
||||
| **R-480** | **[P2-MEDIUM] After a successful restore, the app card still shows the failed update's sentence — which says the running app „leállítva marad”.** MEASURED 2026-09-13 on demo-hp (v0.239.0): the restore cleared the hold (`restore hold CLEARED for gokapi`), gokapi ran healthy, `hold_reason` was empty — but `update_phase=failed` and `update_error` stayed, and `GET /stacks` rendered the sentence inside gokapi's card (control: absent from actualbudget's card). It survived the app's removal too (`not_deployed … phase=failed err=…`). Present since v0.238.0's page; this morning's restore walk did not check the card text. **Fix shape:** a successful restore (and a removal) clears the stack's last update outcome. `09-card-after-restore.txt`, `10-teardown.txt` | **READY — rank P2-MEDIUM; owner: CC** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user