Files
felhom-controller/REPORT.md
T

11 KiB
Raw Blame History

REPORT — Update arc slice 4: the Update button takes a backup first, and tells the truth (2026-09-13)

Overwritten each run. This records the most recent implementation only.

Shipped as controller v0.237.0 (the guarded job) and v0.238.0 (the page), both live on both demo guests, every scenario A–H proven live on demo-hp — including the restore walk. Four claims in the prompt turned out wrong or incomplete; they are first, in §1.

1. Claims in the prompt that turned out wrong, named first

  1. "A release in this task raises the floor and does NOT need a bake." Wrong, measured. The hub HOLDS a floor above the vouched golden (publish-train rule 1): raising the floor to 0.237.0 made the hub log managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0 … (controller floor withheld), the box logged SetFloor: floor "0.236.0" → "", and neither box moved in 8 minutes. Both releases were hand-deployed by the skill's route; the floor was put back to 0.236.0. This also falsifies the premise of this morning's golden-cadence ruling — filed R-472, an operator decision; the five documents that repeated the claim were corrected.
  2. "Is the restorable-unit predicate Tier-2-only?" — yes, and that is a finding. The update precondition is exactly what the backups page uses (Tier2UnitRestorePoint), so an app with no Tier-2 copy cannot be updated at all — on demo-hp gokapi and nextcloud have no Tier-2 record. The primary unit and the off-site unit are restorable routes it does not consider — R-475.
  3. "The proven copy date" is not one field. The page names the unit MANIFEST's created_at, which moves only when the app's definition changes — measured on demo-hp: bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aging the copy by it would call a fresh copy stale forever and a backup-first would not fix it. The update ages by the last SUCCESSFUL copy instead; the page's date is R-476.
  4. "A person clears it by restoring, exactly as for R-379." Not exactly: an R-379 hold is cleared only by the operator CLI (-clear-restore-hold); nothing in the restore path cleared it. Slice 4 makes a successful unit restore lift an update hold only, and leaves R-379's rule unchanged.

Also: the drive-return gate and the nightly volume dump started held apps — "a hold that only one path honours" was true of the R-379 hold too, before this slice. Both now honour it.

2. Baselines, commits, deployed versions

repo before after
felhom-controller 155271672265 (v0.236.0) 0d402f7 v0.237.0 → 129201a v0.238.0 → cbcca03 v0.238.1
app-catalog-felhom.eu 3525e35 live-test commits and reverts, no net change (see its REPORT)
felhom.eu abe567e documents + register (this session's final push)

Deployed: gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up … (healthy) on guest 9201 of both demo boxes, hand-deployed by the skill's route (the floor could not carry it — §1.1). Fleet floor left at 0.236.0, the vouched golden.

3. The two config knobs

key default meaning
update.backup_max_age 24h a proven Tier-2 copy older than this is refreshed before the update
update.health_timeout 5m how long the new version has to become healthy before the app is held

Plus two fixed rules, stated as fixed: the disk floor is 2 GB (image size unknown without a registry query) and an app with no .felhom.yml health check must run 60 s with no container restarting.

4. Live evidence — endpoint-level, on demo-hp, throwaway app uptime-kuma

All in felhom.eu/documentation/audits/slice4-2026-09-13/live/. Catalog syncs used POST /api/sync, the dashboard's „Sablonok frissítése" button — the syncer's own entry point, not the 15-minute timer.

A — real upgrade 2.3.2 → 2.4.0 (05-A-update.txt): 202 {"accepted":true,"completed":false}; phases safety-dump → pulling → starting → verifying → done. Verbatim:

update uptime-kuma: precondition met — proven copy from 2026-09-13T10:03:07Z (4m0s old, limit 24h0m0s)
update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.0)
update uptime-kuma: healthy after 5s (the app's health check passed)
update uptime-kuma: DONE in 47s

Final API: update_phase: done, pinned and installed louislam/uptime-kuma:2.4.0; the card reads „Naprakész". The safety-dump path is empty because uptime-kuma has no database container — the no-op the design specifies; the path shape is pre-restore-<stamp>-<app>-<dbtype>.sql (R-361).

B — the copy is stale (06-B-setup.txt, 07-B-update.txt). Stated test method: the real knob update.backup_max_age: 2m was appended to controller.yaml and the file restored byte-identical afterwards (sha256 prefix 042b71a3123d70da before and after; no update: key remains).

update uptime-kuma: the proven copy is 7m0s old (limit 2m0s) — backing up first
update pre-backup for uptime-kuma: volume dump OK
update pre-backup for uptime-kuma: recovery unit captured (0 database dump(s))
Tier 2 copied uptime-kuma → /mnt/felhom-drives/hdd_1/backups/secondary/uptime-kuma (19.9 KB, 0 leg(s), 0s)
update pre-backup for uptime-kuma: complete in 1.395s
update uptime-kuma: DONE in 8s

Tier-2 last_success moved to 10:09:51Z; the new volume tar is in both the primary unit and the mirror.

E — the tag does not exist (08-E-setup.txt, 09-E-update.txt): update_error is the Hungarian sentence Az új verzió letöltése nem sikerült, …, no Docker stderr in it; pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0). The app untouched: container cb541381d93e…, started=2026-09-13T10:09:50.950051403Z — identical before and after; live and stored definitions 2.4.0; no pre-update copies left.

F — never healthy (11-F-setup.txt, 12-F-update.txt): catalog alpine:3.20.

update uptime-kuma FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
restore hold SET for uptime-kuma — the app stays stopped until it is cleared

Hold record {"reason":"update_failed","copy_date":"2026-09-13T10:09:51Z","at":"2026-09-13T10:18:02Z"}. API and page: A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, … Visszaállítható a(z) 2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon. Card: data-held="true", 0 lifecycle buttons. App info, ASCII fragments with controls: friss ×3, leáll ×3, Mentések ×2, href="/backups/apps" ×2. Container removed.

H — every path refuses the held app (13-H-refusals.txt): start, restart, update → 409 with the hold sentence each; after a controller restart:

[bootrecon] "uptime-kuma" is a boot orphan by intent but is HELD (held after a failed update (2026-09-13T10:18:02Z) — restore it from its backup to start it) — NOT starting it
[bootrecon] Boot reconciliation: nothing to start — 1 app(s) held (…): [uptime-kuma]

The drive-return gate cannot be exercised without unplugging a drive; it is covered by TestSlice4_DriveReturnGateSkipsAHeldApp with its red-proof, not live.

The restore walk (14-restore-walk.txt): POST /backup/tier2/unit-restore (the page's own form) → 302; restore-status ok: true, A(z) uptime-kuma: 1 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-13 12:09). in 9.1 s. Container louislam/uptime-kuma:2.4.0 … (healthy), pinned 2.4.0, and uptime-kuma: restore completed — the update hold (set 2026-09-13T10:18:02Z) is CLEARED.

5. Tests, red-proofs, gate

Full gate after each release: go build ./... rc 0, go vet ./... rc 0, go test ./... rc 0, 28 packages ok.

red-proof mutation seen to fail
1 (A) health wait replaced by true the health wait was never reached; state: updating=false phase=done
2 (C) preflight drops !rp.Restorable an app with no restorable copy must be REFUSED (no_backup), got <nil>
3 (F) the hold call skipped an app that did not come up must be HELD
4 (H, R-439) update removed from the router's hold line refused with the no-backup sentence instead of the hold's
nightly hold isHeld skip removed from the volume dump the nightly volume dump touched a HELD app
drive-return appHeld check removed the drive-return gate tried to START a held app
page ×2 updating/held checks moved after isOperational phase label missing, lifecycle buttons present
v0.238.1 updating clause removed from isHeld a nightly leg touched an app MID-UPDATE

Three red-proofs first ran inertly and were fixed before they counted: proof 1 did not compile, proof 2's fixture was also unproven, so another check still refused, and proof 4 was masked by the preflight's own hold check. Outputs: audits/slice4-2026-09-13/redproofs/.

6. Rows

Closed: R-448, R-443, R-439. Opened: R-472 (floor vs golden, operator), R-473 (glance template), R-474 (removal leaves backups, reports volumes_removed: null — reproduced twice), R-475 (Tier-2-only precondition, operator), R-476 (page names the manifest date). R-469 unblocked, not lifted. Register 214 open / 174 closed.

7. Teardown — three layers

  1. Machine: glance and uptime-kuma removed through the feature. The removal left backup directories and applied-compose files (R-474), cleared by hand by named path; four test images removed by name, no prune. The standing apps and bentopdf were never touched and are Up … (healthy).
  2. Host: nothing created on demo-hp's Proxmox layer — no guest, no storage.
  3. Hub: /hosts lists exactly demo-felhom-8363b5 and demo-hp-bb76ea; 0 customer configs; the floor is back at 0.236.0. Two app-deployed events from the throwaways reached the hub as ordinary events.

8. Observations

  1. The floor cannot carry a release past the vouched golden. FILED: R-472.
  2. The glance template crash-loops on a fresh install. FILED: R-473.
  3. App removal leaves the recovery unit and Tier-2 copy behind and names no volume. FILED: R-474.
  4. The update precondition is Tier-2-only. FILED: R-475.
  5. The page names a copy by its manifest date, which lags the data. FILED: R-476.
  6. The controller CHANGELOG headers v0.233.0–v0.238.1 still need a MinAgent line — these three releases carry one. NOT-A-FINDING: already R-470.
  7. A live-test catalog tag is exposed to every new install while it stands. NOT-A-FINDING: the alpine tag was reverted about two minutes after landing, neither demo box installed uptime-kuma in that window, and the practice is recorded in the catalog report.