Files
felhom-controller/REPORT.md
T

173 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — Update arc slice 4: the Update button takes a backup first, and tells the truth (2026-09-13)
*Overwritten each run. This records the most recent implementation only.*
> **Shipped as controller v0.237.0 (the guarded job) and v0.238.0 (the page), both live on both demo
> guests, every scenario A–H proven live on demo-hp — including the restore walk.** Four claims in the
> prompt turned out wrong or incomplete; they are first, in §1.
## 1. Claims in the prompt that turned out wrong, named first
1. **"A release in this task raises the floor and does NOT need a bake."** Wrong, measured. The hub
HOLDS a floor above the vouched golden (publish-train rule 1): raising the floor to 0.237.0 made the
hub log `managed floor HELD for demo-hp: held: floor 0.237.0 is ABOVE the vouched golden 0.236.0 …
(controller floor withheld)`, the box logged `SetFloor: floor "0.236.0" → ""`, and neither box moved
in 8 minutes. Both releases were hand-deployed by the skill's route; the floor was put back to
0.236.0. **This also falsifies the premise of this morning's golden-cadence ruling** — filed
**R-472**, an operator decision; the five documents that repeated the claim were corrected.
2. **"Is the restorable-unit predicate Tier-2-only?" — yes, and that is a finding.** The update
precondition is exactly what the backups page uses (`Tier2UnitRestorePoint`), so an app with no
Tier-2 copy cannot be updated at all — on demo-hp `gokapi` and `nextcloud` have no Tier-2 record.
The primary unit and the off-site unit are restorable routes it does not consider — **R-475**.
3. **"The proven copy date" is not one field.** The page names the unit MANIFEST's `created_at`, which
moves only when the app's definition changes — measured on demo-hp: bookstack's mirror held a
2026-09-13T00:30Z dump under a manifest dated 2026-09-12T02:15:29Z. Aging the copy by it would call a
fresh copy stale forever and a backup-first would not fix it. The update ages by the last SUCCESSFUL
copy instead; the page's date is **R-476**.
4. **"A person clears it by restoring, exactly as for R-379."** Not exactly: an R-379 hold is cleared
only by the operator CLI (`-clear-restore-hold`); nothing in the restore path cleared it. Slice 4
makes a successful unit restore lift an **update** hold only, and leaves R-379's rule unchanged.
Also: **the drive-return gate and the nightly volume dump started held apps** — "a hold that only one
path honours" was true of the R-379 hold too, before this slice. Both now honour it.
## 2. Baselines, commits, deployed versions
| repo | before | after |
|---|---|---|
| felhom-controller | `155271672265` (v0.236.0) | `0d402f7` v0.237.0 → `129201a` v0.238.0 → `cbcca03` v0.238.1 |
| app-catalog-felhom.eu | `3525e35` | live-test commits and reverts, **no net change** (see its REPORT) |
| felhom.eu | `abe567e` | documents + register (this session's final push) |
**Deployed:** `gitea.dooplex.hu/admin/felhom-controller:0.238.1 Up … (healthy)` on guest 9201 of **both**
demo boxes, hand-deployed by the skill's route (the floor could not carry it — §1.1). Fleet floor left at
0.236.0, the vouched golden.
## 3. The two config knobs
| key | default | meaning |
|---|---|---|
| `update.backup_max_age` | `24h` | a proven Tier-2 copy older than this is refreshed before the update |
| `update.health_timeout` | `5m` | how long the new version has to become healthy before the app is held |
Plus two fixed rules, stated as fixed: the disk floor is **2 GB** (image size unknown without a registry
query) and an app with no `.felhom.yml` health check must run **60 s** with no container restarting.
## 4. Live evidence — endpoint-level, on demo-hp, throwaway app uptime-kuma
All in `felhom.eu/documentation/audits/slice4-2026-09-13/live/`. Catalog syncs used `POST /api/sync`, the
dashboard's „Sablonok frissítése" button — the syncer's own entry point, not the 15-minute timer.
**A — real upgrade 2.3.2 → 2.4.0** (`05-A-update.txt`): 202 `{"accepted":true,"completed":false}`; phases
`safety-dump → pulling → starting → verifying → done`. Verbatim:
```
update uptime-kuma: precondition met — proven copy from 2026-09-13T10:03:07Z (4m0s old, limit 24h0m0s)
update safety dump for uptime-kuma: the app has no database — nothing to copy (no-op)
update uptime-kuma: pin advanced to the catalog's current definition (uptime-kuma=louislam/uptime-kuma:2.4.0)
update uptime-kuma: healthy after 5s (the app's health check passed)
update uptime-kuma: DONE in 47s
```
Final API: `update_phase: done`, pinned and installed `louislam/uptime-kuma:2.4.0`; the card reads
„Naprakész". The safety-dump **path** is empty because uptime-kuma has no database container — the
no-op the design specifies; the path shape is `pre-restore-<stamp>-<app>-<dbtype>.sql` (R-361).
**B — the copy is stale** (`06-B-setup.txt`, `07-B-update.txt`). **Stated test method:** the real knob
`update.backup_max_age: 2m` was appended to `controller.yaml` and the file restored byte-identical
afterwards (`sha256` prefix `042b71a3123d70da` before and after; no `update:` key remains).
```
update uptime-kuma: the proven copy is 7m0s old (limit 2m0s) — backing up first
update pre-backup for uptime-kuma: volume dump OK
update pre-backup for uptime-kuma: recovery unit captured (0 database dump(s))
Tier 2 copied uptime-kuma → /mnt/felhom-drives/hdd_1/backups/secondary/uptime-kuma (19.9 KB, 0 leg(s), 0s)
update pre-backup for uptime-kuma: complete in 1.395s
update uptime-kuma: DONE in 8s
```
Tier-2 `last_success` moved to 10:09:51Z; the new volume tar is in both the primary unit and the mirror.
**E — the tag does not exist** (`08-E-setup.txt`, `09-E-update.txt`): `update_error` is the Hungarian
sentence `Az új verzió letöltése nem sikerült, …`, no Docker stderr in it; `pin and definition PUT BACK
to the pre-update version (uptime-kuma=louislam/uptime-kuma:2.4.0)`. **The app untouched:** container
`cb541381d93e…`, `started=2026-09-13T10:09:50.950051403Z` — identical before and after; live and stored
definitions `2.4.0`; no pre-update copies left.
**F — never healthy** (`11-F-setup.txt`, `12-F-update.txt`): catalog `alpine:3.20`.
```
update uptime-kuma FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state restarting) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
restore hold SET for uptime-kuma — the app stays stopped until it is cleared
```
Hold record `{"reason":"update_failed","copy_date":"2026-09-13T10:09:51Z","at":"2026-09-13T10:18:02Z"}`.
API and page: `A(z) uptime-kuma frissítése 2026-09-13 12:18-kor nem sikerült, … Visszaállítható a(z)
2026-09-13 12:09-i biztonsági mentésből a Mentések oldalon.` Card: `data-held="true"`, **0** lifecycle
buttons. App info, ASCII fragments with controls: `friss` ×3, `leáll` ×3, `Mentések` ×2,
`href="/backups/apps"` ×2. Container removed.
**H — every path refuses the held app** (`13-H-refusals.txt`): `start`, `restart`, `update` → **409**
with the hold sentence each; after a controller restart:
```
[bootrecon] "uptime-kuma" is a boot orphan by intent but is HELD (held after a failed update (2026-09-13T10:18:02Z) — restore it from its backup to start it) — NOT starting it
[bootrecon] Boot reconciliation: nothing to start — 1 app(s) held (…): [uptime-kuma]
```
The drive-return gate cannot be exercised without unplugging a drive; it is covered by
`TestSlice4_DriveReturnGateSkipsAHeldApp` with its red-proof, not live.
**The restore walk** (`14-restore-walk.txt`): `POST /backup/tier2/unit-restore` (the page's own form) →
302; restore-status `ok: true`, `A(z) uptime-kuma: 1 adatkötet visszaállítva — az alkalmazás
újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-09-13 12:09).` in 9.1 s.
Container `louislam/uptime-kuma:2.4.0 … (healthy)`, pinned 2.4.0, and
`uptime-kuma: restore completed — the update hold (set 2026-09-13T10:18:02Z) is CLEARED`.
## 5. Tests, red-proofs, gate
Full gate after each release: `go build ./...` rc 0, `go vet ./...` rc 0, `go test ./...` rc 0, **28 packages ok**.
| red-proof | mutation | seen to fail |
|---|---|---|
| **1** (A) | health wait replaced by `true` | `the health wait was never reached; state: updating=false phase=done` |
| **2** (C) | preflight drops `!rp.Restorable` | `an app with no restorable copy must be REFUSED (no_backup), got <nil>` |
| **3** (F) | the hold call skipped | `an app that did not come up must be HELD` |
| **4** (H, R-439) | `update` removed from the router's hold line | refused with the no-backup sentence instead of the hold's |
| nightly hold | isHeld skip removed from the volume dump | `the nightly volume dump touched a HELD app` |
| drive-return | appHeld check removed | `the drive-return gate tried to START a held app` |
| page ×2 | updating/held checks moved after `isOperational` | phase label missing, lifecycle buttons present |
| v0.238.1 | updating clause removed from isHeld | `a nightly leg touched an app MID-UPDATE` |
**Three red-proofs first ran inertly and were fixed before they counted:** proof 1 did not compile,
proof 2's fixture was also unproven, so another check still refused, and proof 4 was masked by the
preflight's own hold check. Outputs: `audits/slice4-2026-09-13/redproofs/`.
## 6. Rows
Closed: **R-448, R-443, R-439**. Opened: **R-472** (floor vs golden, operator), **R-473** (glance
template), **R-474** (removal leaves backups, reports `volumes_removed: null` — reproduced twice),
**R-475** (Tier-2-only precondition, operator), **R-476** (page names the manifest date). **R-469**
unblocked, not lifted. Register 214 open / 174 closed.
## 7. Teardown — three layers
1. **Machine:** glance and uptime-kuma removed through the feature. The removal left backup directories
and applied-compose files (R-474), cleared by hand by named path; four test images removed by name,
no prune. The standing apps and bentopdf were never touched and are `Up … (healthy)`.
2. **Host:** nothing created on demo-hp's Proxmox layer — no guest, no storage.
3. **Hub:** `/hosts` lists exactly `demo-felhom-8363b5` and `demo-hp-bb76ea`; 0 customer configs; the
floor is back at `0.236.0`. Two app-deployed events from the throwaways reached the hub as ordinary events.
## 8. Observations
1. **The floor cannot carry a release past the vouched golden.** **FILED: R-472.**
2. **The glance template crash-loops on a fresh install.** **FILED: R-473.**
3. **App removal leaves the recovery unit and Tier-2 copy behind and names no volume.** **FILED: R-474.**
4. **The update precondition is Tier-2-only.** **FILED: R-475.**
5. **The page names a copy by its manifest date, which lags the data.** **FILED: R-476.**
6. **The controller CHANGELOG headers v0.233.0–v0.238.1 still need a MinAgent line** — these three
releases carry one. **NOT-A-FINDING: already R-470.**
7. **A live-test catalog tag is exposed to every new install while it stands.**
**NOT-A-FINDING: the alpine tag was reverted about two minutes after landing, neither demo box installed uptime-kuma in that window, and the practice is recorded in the catalog report.**