- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two live-only defects, §6.4 part 1 SHIPPED. - Capability map: a failed update is undone by the box - PROVEN-LIVE. - Live evidence on 9202: three apps undone by the product with seeds before the backup, after it and seconds before the press read back; cut-off copy held honestly; power cut during the undo resumed; manual press after undo. - Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643 ruled; R-646 opened. STATUS asks the floor question. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
6.0 KiB
REPORT — the undo: bake-off, build (controller v0.263.0 → v0.263.2), live proof
2026-09-23 (afternoon). Repos touched: felhom-controller (v0.263.0, v0.263.1, v0.263.2),
felhom.eu (docs, register, capability map, STATUS, CONTEXT, evidence), admin/app-catalog-drill
(drill commits, reset to live main at the end). app-catalog-felhom.eu, felhom-agent, hub: untouched.
Architecture read first and named: documentation/architecture/09-update-architecture.md (§3 decisions
11–20, §4, §6.1, §6.1a, §6.4). Baselines verified live: controller b9deec19077b, agent
d9864a94bf62, felhom.eu 4c92beab8fdb, catalog cfcfe5278428 — all as the brief said.
1. Not done, or changed
| item | state |
|---|---|
| rulings 19–20 | recorded (09 §3), commit 5a349d9 |
| Part 1 — bake-off | done in ~40 min of the 2 h cap; folder copy chosen |
| Part 2 — build | done — but in THREE releases, not one. v0.263.0 failed its first two live proofs honestly (HELD, data put back). Two defects only the live box could show; each fixed + red-proofed + released: v0.263.1 (the undo's probe never ran on an app the current probe held unhealthy), v0.263.2 (the "old" .felhom.yml was already the new one — it flows in on every catalog sync; now recorded at pin time). 0.263.0/0.263.1 ran only on 9202 and were removed from it. |
| the household mail + operator event of decision 15 | not built — 09 §6.4 part 2 (R-606). The page line and the hold sentence are built. |
| the floor | not raised — the operator's question (STATUS) |
| immich/nextcloud rate test | nextcloud deployed and was used (185 MB MariaDB volume, 300 files through WebDAV); immich not tried |
| cut-off copy on vikunja in the bake-off | could not be cut (2.9 MB finishes before a kill lands); proven on docmost and romm instead; in the LIVE proof the marker was removed from a vikunja copy mid-update |
Claims in the brief that turned out wrong or unmeasured:
- "All three apps keep their data only in named volumes" — true from compose AND on disk for these three; but romm also has two BIND folders (roms, resources), empty throughout — never copied, by rule.
- "A volume copy with the containers stopped is consistent for both engines" — measured true: docmost (PostgreSQL) and romm (MariaDB) came back with ledgers equal, three times each.
- "Immich or Nextcloud deploys on 9202" — Nextcloud did; Immich was not tried.
- Not in the brief, and the most important: "the old
.felhom.yml" does not exist at update time. It is replaced on every catalog sync (09§5.4). The build now records it when a version is pinned. R-646 for apps pinned before that.
2. The bake-off (audits/undo-bakeoff-2026-09-23/README.md)
Both methods passed every row on docmost, romm, vikunja (seeds A and B back, ledgers equal, cut-off
detected before anything moves, ≤ 5.3 s extra downtime). Folder copy chosen (decision 19): an app
with no database server gets no dump, so dump-and-load would need the folder copy anyway. Rate: 185 MB
in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s (warm cache); 5 GB ≈ 12 s here, 25–50 s on a cold or spinning
disk. Found on the way: the undo must never read the recovery unit (R-645), and killing docker run
does not stop the copy container.
3. Red-proofs (each seen failing, then restored)
| # | mutation | test that failed |
|---|---|---|
| 1 | the undo call removed (v0.262.1's shape) | TestUndo_FailedUpdateIsPutBackWithItsData — held=true phase="failed" |
| 2 | copy validation removed | TestUndo_CutOffCopyIsRefusedBeforeAnythingMoves — state half |
| 3 | old-probe rule removed | TestUndo_UsesTheOldProbe — not_started |
| 3b | the old probe read from the stack dir (v0.263.1's shape) | TestUndo_UsesTheOldProbe |
| 4 | the undoing recovery arm removed |
TestUndo_PowerCutDuringTheUndoResumesIt — resumed=[] |
| 5a | last_update_undone not recorded |
TestUndo_FailedUpdateIsPutBackWithItsData — got <nil> |
| 5b | the page line removed from the handler | TestUndo_PageSaysTheBoxPutTheAppBack (hu and en) |
| 6 | the hold prefix ignored | TestUndo_HoldSentenceSaysTheUndoWasTriedAndTheDataState |
| 7 | R-642: "completed" back | TestR642_StartIsNeverReportedCompleted |
| 8 | a successful update keeps the undone note | TestUndo_ASuccessfulUpdateEndsTheUndoneNote |
| 9 | the undo probe switched off on an unhealthy app (v0.263.0's shape) |
TestUndo_OldProbeRunsOnAnAppTheCurrentProbeMarkedUnhealthy — not healthy within 1s (last: state unhealthy), the live message verbatim |
| 10 | the pin advance records no probe | TestUndo_PinningRecordsThatVersionsProbe |
go build ./... && go vet ./... && go test ./... rc=0 before each of the three commits; controller
gates rc=0 (the docker -v gate needed four named-volume mounts allowlisted with their why).
4. Live proof (audits/undo-live-2026-09-23/README.md) — endpoint-level, both languages
Three apps undone by the product with seeds A, B, C back and ledgers equal (30–52 s of undo); page
line hu/en; cut-off copy → HOLD untouched (prefix in the box's language, hu and en quoted); power cut
during undoing → resumed and undone; a person's press after an undo → done, note cleared; removal
deleted kept copies; R-642 Start answer.
5. Rows
Closed (4, moved to CLOSED-ITEMS.md): R-637, R-639, R-641, R-642. Narrowed: R-638, R-640 (to
the restore paths). Ruled: R-643 (decision 20). Noted: R-645. Opened: R-645 (earlier this
session, filed with the bake-off) and R-646 (apps pinned before v0.263.2). Open rows 337 → 335.
6. Teardown — three layers
Machine (9202): the three apps and every copy removed through the product; test images removed by
name; controller.yaml restored and read back; catalog cache on live cfcfe52; 9202 stays on
controller 0.263.2 (self-update off, fleet floor untouched). Host: nothing provisioned; 9202 stopped
and started once for the power cut. Hub: untouched. Drill repo: reset to live main.