Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
11 KiB
REPORT — Update arc slice 4: the Update button takes a backup first, and tells the truth (2026-09-13)
Overwritten each run. This records the most recent implementation only.
Shipped as controller v0.237.0 (the guarded job) and v0.238.0 (the page), both live on both demo guests, every scenario A–H proven live on demo-hp — including the restore walk. Four claims in the prompt turned out wrong or incomplete; they are first, in §1.
1. Claims in the prompt that turned out wrong, named first
- "A release in this task raises the floor and does NOT need a bake." Wrong, measured. The hub
HOLDS a floor above the vouched golden (publish-train rule 1): raising the floor to 0.237.0 made the
hub log
managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0 … (controller floor withheld), the box loggedSetFloor: floor "0.236.0" → "", and neither box moved in 8 minutes. Both releases were hand-deployed by the skill's route; the floor was put back to 0.236.0. This also falsifies the premise of this morning's golden-cadence ruling — filed R-472, an operator decision; the five documents that repeated the claim were corrected. - "Is the restorable-unit predicate Tier-2-only?" — yes, and that is a finding. The update
precondition is exactly what the backups page uses (
Tier2UnitRestorePoint), so an app with no Tier-2 copy cannot be updated at all — on demo-hpgokapiandnextcloudhave no Tier-2 record. The primary unit and the off-site unit are restorable routes it does not consider — R-475. - "The proven copy date" is not one field. The page names the unit MANIFEST's
created_at, which moves only when the app's definition changes — measured on demo-hp: bookstack's mirror held a 2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aging the copy by it would call a fresh copy stale forever and a backup-first would not fix it. The update ages by the last SUCCESSFUL copy instead; the page's date is R-476. - "A person clears it by restoring, exactly as for R-379." Not exactly: an R-379 hold is cleared
only by the operator CLI (
-clear-restore-hold); nothing in the restore path cleared it. Slice 4 makes a successful unit restore lift an update hold only, and leaves R-379's rule unchanged.
Also: the drive-return gate and the nightly volume dump started held apps — "a hold that only one path honours" was true of the R-379 hold too, before this slice. Both now honour it.
2. Baselines, commits, deployed versions
| repo | before | after |
|---|---|---|
| felhom-controller | 155271672265 (v0.236.0) |
0d402f7 v0.237.0 → 129201a v0.238.0 → cbcca03 v0.238.1 |
| app-catalog-felhom.eu | 3525e35 |
live-test commits and reverts, no net change (see its REPORT) |
| felhom.eu | abe567e |
documents + register (this session's final push) |
Deployed: gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up … (healthy) on guest 9201 of both
demo boxes, hand-deployed by the skill's route (the floor could not carry it — §1.1). Fleet floor left at
0.236.0, the vouched golden.
3. The two config knobs
| key | default | meaning |
|---|---|---|
update.backup_max_age |
24h |
a proven Tier-2 copy older than this is refreshed before the update |
update.health_timeout |
5m |
how long the new version has to become healthy before the app is held |
Plus two fixed rules, stated as fixed: the disk floor is 2 GB (image size unknown without a registry
query) and an app with no .felhom.yml health check must run 60 s with no container restarting.
4. Live evidence — endpoint-level, on demo-hp, throwaway app uptime-kuma
All in felhom.eu/documentation/audits/slice4-2026-09-13/live/. Catalog syncs used POST /api/sync, the
dashboard's „Sablonok frissítése" button — the syncer's own entry point, not the 15-minute timer.
A — real upgrade 2.3.2 → 2.4.0 (05-A-update.txt): 202 {"accepted":true,"completed":false}; phases
safety-dump → pulling → starting → verifying → done. Verbatim:
update uptime-kuma: precondition met — proven copy from 2026-09-13T10:03:07Z (4m0s old, limit 24h0m0s)
update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.0)
update uptime-kuma: healthy after 5s (the app's health check passed)
update uptime-kuma: DONE in 47s
Final API: update_phase: done, pinned and installed louislam/uptime-kuma:2.4.0; the card reads
„Naprakész". The safety-dump path is empty because uptime-kuma has no database container — the
no-op the design specifies; the path shape is pre-restore-<stamp>-<app>-<dbtype>.sql (R-361).
B — the copy is stale (06-B-setup.txt, 07-B-update.txt). Stated test method: the real knob
update.backup_max_age: 2m was appended to controller.yaml and the file restored byte-identical
afterwards (sha256 prefix 042b71a3123d70da before and after; no update: key remains).
update uptime-kuma: the proven copy is 7m0s old (limit 2m0s) — backing up first
update pre-backup for uptime-kuma: volume dump OK
update pre-backup for uptime-kuma: recovery unit captured (0 database dump(s))
Tier 2 copied uptime-kuma → /mnt/felhom-drives/hdd_1/backups/secondary/uptime-kuma (19.9 KB, 0 leg(s), 0s)
update pre-backup for uptime-kuma: complete in 1.395s
update uptime-kuma: DONE in 8s
Tier-2 last_success moved to 10:09:51Z; the new volume tar is in both the primary unit and the mirror.
E — the tag does not exist (08-E-setup.txt, 09-E-update.txt): update_error is the Hungarian
sentence Az új verzió letöltése nem sikerült, …, no Docker stderr in it; pin and definition PUT BACK to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0). The app untouched: container
cb541381d93e…, started=2026-09-13T10:09:50.950051403Z — identical before and after; live and stored
definitions 2.4.0; no pre-update copies left.
F — never healthy (11-F-setup.txt, 12-F-update.txt): catalog alpine:3.20.
update uptime-kuma FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
restore hold SET for uptime-kuma — the app stays stopped until it is cleared
Hold record {"reason":"update_failed","copy_date":"2026-09-13T10:09:51Z","at":"2026-09-13T10:18:02Z"}.
API and page: A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, … Visszaállítható a(z) 2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon. Card: data-held="true", 0 lifecycle
buttons. App info, ASCII fragments with controls: friss ×3, leáll ×3, Mentések ×2,
href="/backups/apps" ×2. Container removed.
H — every path refuses the held app (13-H-refusals.txt): start, restart, update → 409
with the hold sentence each; after a controller restart:
[bootrecon] "uptime-kuma" is a boot orphan by intent but is HELD (held after a failed update (2026-09-13T10:18:02Z) — restore it from its backup to start it) — NOT starting it
[bootrecon] Boot reconciliation: nothing to start — 1 app(s) held (…): [uptime-kuma]
The drive-return gate cannot be exercised without unplugging a drive; it is covered by
TestSlice4_DriveReturnGateSkipsAHeldApp with its red-proof, not live.
The restore walk (14-restore-walk.txt): POST /backup/tier2/unit-restore (the page's own form) →
302; restore-status ok: true, A(z) uptime-kuma: 1 adatkötet visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-13 12:09). in 9.1 s.
Container louislam/uptime-kuma:2.4.0 … (healthy), pinned 2.4.0, and
uptime-kuma: restore completed — the update hold (set 2026-09-13T10:18:02Z) is CLEARED.
5. Tests, red-proofs, gate
Full gate after each release: go build ./... rc 0, go vet ./... rc 0, go test ./... rc 0, 28 packages ok.
| red-proof | mutation | seen to fail |
|---|---|---|
| 1 (A) | health wait replaced by true |
the health wait was never reached; state: updating=false phase=done |
| 2 (C) | preflight drops !rp.Restorable |
an app with no restorable copy must be REFUSED (no_backup), got <nil> |
| 3 (F) | the hold call skipped | an app that did not come up must be HELD |
| 4 (H, R-439) | update removed from the router's hold line |
refused with the no-backup sentence instead of the hold's |
| nightly hold | isHeld skip removed from the volume dump | the nightly volume dump touched a HELD app |
| drive-return | appHeld check removed | the drive-return gate tried to START a held app |
| page ×2 | updating/held checks moved after isOperational |
phase label missing, lifecycle buttons present |
| v0.238.1 | updating clause removed from isHeld | a nightly leg touched an app MID-UPDATE |
Three red-proofs first ran inertly and were fixed before they counted: proof 1 did not compile,
proof 2's fixture was also unproven, so another check still refused, and proof 4 was masked by the
preflight's own hold check. Outputs: audits/slice4-2026-09-13/redproofs/.
6. Rows
Closed: R-448, R-443, R-439. Opened: R-472 (floor vs golden, operator), R-473 (glance
template), R-474 (removal leaves backups, reports volumes_removed: null — reproduced twice),
R-475 (Tier-2-only precondition, operator), R-476 (page names the manifest date). R-469
unblocked, not lifted. Register 214 open / 174 closed.
7. Teardown — three layers
- Machine: glance and uptime-kuma removed through the feature. The removal left backup directories
and applied-compose files (R-474), cleared by hand by named path; four test images removed by name,
no prune. The standing apps and bentopdf were never touched and are
Up … (healthy). - Host: nothing created on demo-hp's Proxmox layer — no guest, no storage.
- Hub:
/hostslists exactlydemo-felhom-8363b5anddemo-hp-bb76ea; 0 customer configs; the floor is back at0.236.0. Two app-deployed events from the throwaways reached the hub as ordinary events.
8. Observations
- The floor cannot carry a release past the vouched golden. FILED: R-472.
- The glance template crash-loops on a fresh install. FILED: R-473.
- App removal leaves the recovery unit and Tier-2 copy behind and names no volume. FILED: R-474.
- The update precondition is Tier-2-only. FILED: R-475.
- The page names a copy by its manifest date, which lags the data. FILED: R-476.
- The controller CHANGELOG headers v0.233.0–v0.238.1 still need a MinAgent line — these three releases carry one. NOT-A-FINDING: already R-470.
- A live-test catalog tag is exposed to every new install while it stands. NOT-A-FINDING: the alpine tag was reverted about two minutes after landing, neither demo box installed uptime-kuma in that window, and the practice is recorded in the catalog report.