From 0e3d83103029c90919ae51534db2a446bb322e48 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 13 Sep 2026 17:54:27 +0200 Subject: [PATCH] REPORT: v0.239.0 any-tier update and MinAgent header gate, proven live; R-477..R-480 filed --- REPORT.md | 196 +++++++++++------------------------------------------- 1 file changed, 39 insertions(+), 157 deletions(-) diff --git a/REPORT.md b/REPORT.md index 11d1add..e7f5f5d 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,172 +1,54 @@ -# REPORT — Update arc slice 4: the Update button takes a backup first, and tells the truth (2026-09-13) +# REPORT — any backup tier lets an app update, and the release header must state its MinAgent (2026-09-13) *Overwritten each run. This records the most recent implementation only.* -> **Shipped as controller v0.237.0 (the guarded job) and v0.238.0 (the page), both live on both demo -> guests, every scenario A–H proven live on demo-hp — including the restore walk.** Four claims in the -> prompt turned out wrong or incomplete; they are first, in §1. +> **Shipped as controller v0.239.0 and delivered to both demo guests by the managed floor alone** (hub +> v0.112.0, declared MinAgent). No hand install. Scenarios G–M are tested; K, H, the failed update and +> the restore were walked live on demo-hp. Evidence: `felhom.eu/documentation/audits/rulings-r472-r475-2026-09-13/`. -## 1. Claims in the prompt that turned out wrong, named first +## 1. What changed -1. **"A release in this task raises the floor and does NOT need a bake."** Wrong, measured. The hub - HOLDS a floor above the vouched golden (publish-train rule 1): raising the floor to 0.237.0 made the - hub log `managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0 … - (controller floor withheld)`, the box logged `SetFloor: floor "0.236.0" → ""`, and neither box moved - in 8 minutes. Both releases were hand-deployed by the skill's route; the floor was put back to - 0.236.0. **This also falsifies the premise of this morning's golden-cadence ruling** — filed - **R-472**, an operator decision; the five documents that repeated the claim were corrected. -2. **"Is the restorable-unit predicate Tier-2-only?" — yes, and that is a finding.** The update - precondition is exactly what the backups page uses (`Tier2UnitRestorePoint`), so an app with no - Tier-2 copy cannot be updated at all — on demo-hp `gokapi` and `nextcloud` have no Tier-2 record. - The primary unit and the off-site unit are restorable routes it does not consider — **R-475**. -3. **"The proven copy date" is not one field.** The page names the unit MANIFEST's `created_at`, which - moves only when the app's definition changes — measured on demo-hp: bookstack's mirror held a - 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aging the copy by it would call a - fresh copy stale forever and a backup-first would not fix it. The update ages by the last SUCCESSFUL - copy instead; the page's date is **R-476**. -4. **"A person clears it by restoring, exactly as for R-379."** Not exactly: an R-379 hold is cleared - only by the operator CLI (`-clear-restore-hold`); nothing in the restore path cleared it. Slice 4 - makes a successful unit restore lift an **update** hold only, and leaves R-379's rule unchanged. +| Commit | What | +|---|---| +| `f946b0d` | `controller/scripts/minagent_header_gate.py` (fast, blocking): the newest `## vX.Y.Z` block must hold a line starting `**MinAgent: X.Y.Z**`. Decoy test; registered in GATES, FINGERPRINTS and COVERS. Backfill: v0.233.0–v0.236.0 now read `**MinAgent: 0.129.0** (unchanged)` — v0.232.0's value; zero commits under `internal/agentapi` since 2026-09-01, highest `featureMinAgent` 0.129.0 (R-470). | +| `b93c154` | v0.239.0 (R-475): `backup.Manager.UpdateRestorePoints` walks Tier 2 → 1 → 3 and returns the first copy the caller accepts; stacks' `freshRestorePoint` is the one age rule; nothing anywhere → back up first; refused only when no copy AND `CanBackUpApp` is false. `RunAppBackupNow` treats a Tier-2 failure as a WARN and marks the captured unit proven current. The hold stores `copy_tier` and names the tier; old holds keep their sentence. A successful off-site restore now clears an update hold. The backups page still calls `Tier2UnitRestorePoint`. | -Also: **the drive-return gate and the nightly volume dump started held apps** — "a hold that only one -path honours" was true of the R-379 hold too, before this slice. Both now honour it. +## 2. Tests and red-proofs -## 2. Baselines, commits, deployed versions +- Go gate green per commit: `go build ./... && go vet ./... && go test ./...` (28 packages ok). +- `controller_gates.py --fast` green; `test_controller_gates.py` and `test_gate_decoys.py` green. +- New tests: `internal/stacks/update_tiers_test.go`, `internal/backup/update_tiers_test.go`, + `cmd/controller/r475_wiring_test.go`, `scripts/test_minagent_header_gate.py`. -| repo | before | after | +| Red-proof | Mutation | Result | |---|---|---| -| felhom-controller | `155271672265` (v0.236.0) | `0d402f7` v0.237.0 → `129201a` v0.238.0 → `cbcca03` v0.238.1 | -| app-catalog-felhom.eu | `3525e35` | live-test commits and reverts, **no net change** (see its REPORT) | -| felhom.eu | `abe567e` | documents + register (this session's final push) | +| F | gate accepts a body mention | the decoy test fails; 4 tests ran | +| M | age checked only for Tier 2 | 3 M sub-cases fail; 5 RUN lines | +| L | refuse even with a copy | the L test fails | +| tail | the proven-current mark misses | fails on the mtime | +| off-site clear | call removed | the hold stays; test fails | -**Deployed:** `gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up … (healthy)` on guest 9201 of **both** -demo boxes, hand-deployed by the skill's route (the floor could not carry it — §1.1). Fleet floor left at -0.236.0, the vouched golden. +The first run of all four Go red-proofs was inert: the harness ran outside the module (`RUN lines=0`). +They were rerun from the module folder and each failed as intended. -## 3. The two config knobs +## 3. Live — demo-hp, endpoint-level -| key | default | meaning | -|---|---|---| -| `update.backup_max_age` | `24h` | a proven Tier-2 copy older than this is refreshed before the update | -| `update.health_timeout` | `5m` | how long the new version has to become healthy before the app is held | +| Step | Result | +|---|---| +| Delivery | floor 0.239.0 + MinAgent 0.129.0 saved 15:19:12Z; demo-hp on 0.239.0 at +14 s, demo-felhom at +15 s | +| K | throwaway actualbudget, Tier 2 off, unit removed: `found: none — backing up first`, Tier 2 skipped, Tier 1 chosen, done | +| H | gokapi: `precondition met — Tier 1 (own recovery unit)`, no backup, done | +| F-shape | never-healthy image: held; text ends „saját meghajtó, 2026-09-13 17:34."; start 409; record `copy_tier: 1` | +| Restore | `POST /backup/restore` from „helyi": hold CLEARED, gokapi v1.9.6 healthy | -Plus two fixed rules, stated as fixed: the disk floor is **2 GB** (image size unknown without a registry -query) and an app with no `.felhom.yml` health check must run **60 s** with no container restarting. +A first K attempt did not test K: the deploy had already written a Tier-1 unit, so Tier 1 was +chosen. It is kept as `04a-…`. The restore script's first read polled a 404 path and stopped early; the +second read is appended to `08-…`. -## 4. Live evidence — endpoint-level, on demo-hp, throwaway app uptime-kuma +## 4. Observations -All in `felhom.eu/documentation/audits/slice4-2026-09-13/live/`. Catalog syncs used `POST /api/sync`, the -dashboard's „Sablonok frissítése" button — the syncer's own entry point, not the 15-minute timer. - -**A — real upgrade 2.3.2 → 2.4.0** (`05-A-update.txt`): 202 `{"accepted":true,"completed":false}`; phases -`safety-dump → pulling → starting → verifying → done`. Verbatim: - -``` -update uptime-kuma: precondition met — proven copy from 2026-09-13T10:03:07Z (4m0s old, limit 24h0m0s) -update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op) -update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.0) -update uptime-kuma: healthy after 5s (the app's health check passed) -update uptime-kuma: DONE in 47s -``` - -Final API: `update_phase: done`, pinned and installed `louislam/uptime-kuma:2.4.0`; the card reads -„Naprakész". The safety-dump **path** is empty because uptime-kuma has no database container — the -no-op the design specifies; the path shape is `pre-restore---.sql` (R-361). - -**B — the copy is stale** (`06-B-setup.txt`, `07-B-update.txt`). **Stated test method:** the real knob -`update.backup_max_age: 2m` was appended to `controller.yaml` and the file restored byte-identical -afterwards (`sha256` prefix `042b71a3123d70da` before and after; no `update:` key remains). - -``` -update uptime-kuma: the proven copy is 7m0s old (limit 2m0s) — backing up first -update pre-backup for uptime-kuma: volume dump OK -update pre-backup for uptime-kuma: recovery unit captured (0 database dump(s)) -Tier 2 copied uptime-kuma → /mnt/felhom-drives/hdd_1/backups/secondary/uptime-kuma (19.9 KB, 0 leg(s), 0s) -update pre-backup for uptime-kuma: complete in 1.395s -update uptime-kuma: DONE in 8s -``` - -Tier-2 `last_success` moved to 10:09:51Z; the new volume tar is in both the primary unit and the mirror. - -**E — the tag does not exist** (`08-E-setup.txt`, `09-E-update.txt`): `update_error` is the Hungarian -sentence `Az új verzió letöltése nem sikerült, …`, no Docker stderr in it; `pin and definition PUT BACK -to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0)`. **The app untouched:** container -`cb541381d93e…`, `started=2026-09-13T10:09:50.950051403Z` — identical before and after; live and stored -definitions `2.4.0`; no pre-update copies left. - -**F — never healthy** (`11-F-setup.txt`, `12-F-update.txt`): catalog `alpine:3.20`. - -``` -update uptime-kuma FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run) -restore hold SET for uptime-kuma — the app stays stopped until it is cleared -``` - -Hold record `{"reason":"update_failed","copy_date":"2026-09-13T10:09:51Z","at":"2026-09-13T10:18:02Z"}`. -API and page: `A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, … Visszaállítható a(z) -2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon.` Card: `data-held="true"`, **0** lifecycle -buttons. App info, ASCII fragments with controls: `friss` ×3, `leáll` ×3, `Mentések` ×2, -`href="/backups/apps"` ×2. Container removed. - -**H — every path refuses the held app** (`13-H-refusals.txt`): `start`, `restart`, `update` → **409** -with the hold sentence each; after a controller restart: - -``` -[bootrecon] "uptime-kuma" is a boot orphan by intent but is HELD (held after a failed update (2026-09-13T10:18:02Z) — restore it from its backup to start it) — NOT starting it -[bootrecon] Boot reconciliation: nothing to start — 1 app(s) held (…): [uptime-kuma] -``` - -The drive-return gate cannot be exercised without unplugging a drive; it is covered by -`TestSlice4_DriveReturnGateSkipsAHeldApp` with its red-proof, not live. - -**The restore walk** (`14-restore-walk.txt`): `POST /backup/tier2/unit-restore` (the page's own form) → -302; restore-status `ok: true`, `A(z) uptime-kuma: 1 adatkötet visszaállítva — az alkalmazás -újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-13 12:09).` in 9.1 s. -Container `louislam/uptime-kuma:2.4.0 … (healthy)`, pinned 2.4.0, and -`uptime-kuma: restore completed — the update hold (set 2026-09-13T10:18:02Z) is CLEARED`. - -## 5. Tests, red-proofs, gate - -Full gate after each release: `go build ./...` rc 0, `go vet ./...` rc 0, `go test ./...` rc 0, **28 packages ok**. - -| red-proof | mutation | seen to fail | -|---|---|---| -| **1** (A) | health wait replaced by `true` | `the health wait was never reached; state: updating=false phase=done` | -| **2** (C) | preflight drops `!rp.Restorable` | `an app with no restorable copy must be REFUSED (no_backup), got ` | -| **3** (F) | the hold call skipped | `an app that did not come up must be HELD` | -| **4** (H, R-439) | `update` removed from the router's hold line | refused with the no-backup sentence instead of the hold's | -| nightly hold | isHeld skip removed from the volume dump | `the nightly volume dump touched a HELD app` | -| drive-return | appHeld check removed | `the drive-return gate tried to START a held app` | -| page ×2 | updating/held checks moved after `isOperational` | phase label missing, lifecycle buttons present | -| v0.238.1 | updating clause removed from isHeld | `a nightly leg touched an app MID-UPDATE` | - -**Three red-proofs first ran inertly and were fixed before they counted:** proof 1 did not compile, -proof 2's fixture was also unproven, so another check still refused, and proof 4 was masked by the -preflight's own hold check. Outputs: `audits/slice4-2026-09-13/redproofs/`. - -## 6. Rows - -Closed: **R-448, R-443, R-439**. Opened: **R-472** (floor vs golden, operator), **R-473** (glance -template), **R-474** (removal leaves backups, reports `volumes_removed: null` — reproduced twice), -**R-475** (Tier-2-only precondition, operator), **R-476** (page names the manifest date). **R-469** -unblocked, not lifted. Register 214 open / 174 closed. - -## 7. Teardown — three layers - -1. **Machine:** glance and uptime-kuma removed through the feature. The removal left backup directories - and applied-compose files (R-474), cleared by hand by named path; four test images removed by name, - no prune. The standing apps and bentopdf were never touched and are `Up … (healthy)`. -2. **Host:** nothing created on demo-hp's Proxmox layer — no guest, no storage. -3. **Hub:** `/hosts` lists exactly `demo-felhom-8363b5` and `demo-hp-bb76ea`; 0 customer configs; the - floor is back at `0.236.0`. Two app-deployed events from the throwaways reached the hub as ordinary events. - -## 8. Observations - -1. **The floor cannot carry a release past the vouched golden.** **FILED: R-472.** -2. **The glance template crash-loops on a fresh install.** **FILED: R-473.** -3. **App removal leaves the recovery unit and Tier-2 copy behind and names no volume.** **FILED: R-474.** -4. **The update precondition is Tier-2-only.** **FILED: R-475.** -5. **The page names a copy by its manifest date, which lags the data.** **FILED: R-476.** -6. **The controller CHANGELOG headers v0.233.0–v0.238.1 still need a MinAgent line** — these three - releases carry one. **NOT-A-FINDING: already R-470.** -7. **A live-test catalog tag is exposed to every new install while it stands.** - **NOT-A-FINDING: the alpine tag was reverted about two minutes after landing, neither demo box installed uptime-kuma in that window, and the practice is recorded in the catalog report.** +1. **The Tier-3 lookup pays its full 15 s bound on demo-hp and logs a misleading WARN about another app's snapshot size.** FILED: R-477 +2. **A leftover recovery unit from a REMOVED install counted as the fresh Tier-1 restore point of a reinstall.** FILED: R-478 +3. **For a bind-data app the Tier-1 unit holds settings only, so the route the hold names does not restore data.** FILED: R-479 +4. **After a successful restore the card keeps the failed update's sentence, which says the running app is stopped.** FILED: R-480 +5. **Removal with `remove_backups` again left both throwaways' recovery units and their backup prefs.** FILED: R-474