The undo, built and proven live: controller v0.263.2 (09 decision 15)
gates / gates (push) Successful in 25s

- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
  live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
  the backup, after it and seconds before the press read back; cut-off copy
  held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
  ruled; R-646 opened. STATUS asks the floor question.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 12:25:13 +02:00
parent 5a349d9884
commit 05ea21e918
36 changed files with 2727 additions and 124 deletions
+62 -100
View File
@@ -1,119 +1,81 @@
# REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan
# REPORT — the undo: bake-off, build (controller v0.263.0 → v0.263.2), live proof
2026-09-23. Repos touched: **felhom.eu** (docs, register, STATUS, CONTEXT, evidence),
**app-catalog-felhom.eu** (`scripts/` only — the memory watch), **admin/app-catalog-drill** (drill
commits, reset to live `main` at the end). **felhom-controller, felhom-agent, hub: read only.**
Architecture read first and named: `documentation/architecture/09-update-architecture.md` (all of it),
`07-backup-architecture.md` §6.
Baselines verified live before starting: controller `b9deec19077b`, agent `d9864a94bf62`, felhom.eu
`267dcad01bcf`, catalog `02844ae0a579` — all equal to the brief.
2026-09-23 (afternoon). Repos touched: **felhom-controller** (v0.263.0, v0.263.1, v0.263.2),
**felhom.eu** (docs, register, capability map, STATUS, CONTEXT, evidence), **admin/app-catalog-drill**
(drill commits, reset to live `main` at the end). **app-catalog-felhom.eu, felhom-agent, hub: untouched.**
Architecture read first and named: `documentation/architecture/09-update-architecture.md` (§3 decisions
11–20, §4, §6.1, §6.1a, §6.4). Baselines verified live: controller `b9deec19077b`, agent
`d9864a94bf62`, felhom.eu `4c92beab8fdb`, catalog `cfcfe5278428` — all as the brief said.
---
## 1. Not done, or changed from the brief — first
## 1. Not done, or changed
| item | state |
|---|---|
| Part 0 — rulings into `09` | **done**, commit `805ad1e` (documents only) |
| Part 1 — the undo, four cases + the wrong case | **done, with three changes.** (a) **The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder:** romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) **Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around.** The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0). |
| Part 2 — the ladder | **done** |
| Part 3 — the memory watch + red-proof | **done**; harness v2. `C3`, the harness's standing negative control, was **NOT run**: its template's `container_name: privatebin` collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass. |
| Part 4 — the build plan | **done**, `09` §6.4 — with **one open point for the operator** the brief did not expect (§6) |
| rulings 19–20 | **recorded** (`09` §3), commit `5a349d9` |
| Part 1 — bake-off | **done in ~40 min** of the 2 h cap; folder copy chosen |
| Part 2 — build | **done — but in THREE releases, not one.** v0.263.0 failed its first two live proofs honestly (HELD, data put back). Two defects only the live box could show; each fixed + red-proofed + released: **v0.263.1** (the undo's probe never ran on an app the current probe held `unhealthy`), **v0.263.2** (the "old" `.felhom.yml` was already the new one — it flows in on every catalog sync; now recorded at pin time). 0.263.0/0.263.1 ran only on 9202 and were removed from it. |
| the household mail + operator event of decision 15 | **not built** — `09` §6.4 part 2 (R-606). The page line and the hold sentence are built. |
| the floor | **not raised** — the operator's question (STATUS) |
| immich/nextcloud rate test | **nextcloud deployed and was used** (185 MB MariaDB volume, 300 files through WebDAV); immich not tried |
| cut-off copy on vikunja in the bake-off | **could not be cut** (2.9 MB finishes before a kill lands); proven on docmost and romm instead; in the LIVE proof the marker was removed from a vikunja copy mid-update |
**Claims in the brief that turned out wrong:**
1. *"A safety dump exists for every app class"* — **false.** An app with no database server gets none
(`update safety dump for vikunja: the app has no database — nothing to copy (no-op)`), measured.
2. *"The box keeps a git clone of the catalog"* with history — **false.** Depth 1 on both demo
guests, `rev-list --count HEAD` = 1; `sync.go:283`/`:300` clone and fetch `--depth 1`.
3. *"`stacks.update_window` is unread"* — **true, and the grep was widened** from `config.go` +
`setup/handlers.go` to the whole controller repo: the only other hits are
`configs/controller.yaml.example` and the i18n base file. No Go code reads it.
4. The register held **329** row lines by `grep -c '^| \*\*R-'`, not 326; the highest id was R-636 as stated.
**Claims in the brief that turned out wrong or unmeasured:**
1. *"All three apps keep their data only in named volumes"* — true from compose AND on disk for these
three; but romm also has two BIND folders (roms, resources), empty throughout — never copied, by rule.
2. *"A volume copy with the containers stopped is consistent for both engines"* — **measured true**:
docmost (PostgreSQL) and romm (MariaDB) came back with ledgers equal, three times each.
3. *"Immich or Nextcloud deploys on 9202"* — Nextcloud did; Immich was not tried.
4. **Not in the brief, and the most important:** *"the old `.felhom.yml`"* does not exist at update
time. It is replaced on every catalog sync (`09` §5.4). The build now records it when a version is
pinned. R-646 for apps pinned before that.
## 2. Part 0 — the rulings
## 2. The bake-off (`audits/undo-bakeoff-2026-09-23/README.md`)
`09` §3 gains decisions **11–18** in the existing shape; decision 3's second half and §6.1's abort
paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept,
headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices
table updated. Register: R-450, R-451, R-446, R-463 cite the decisions.
Both methods passed every row on docmost, romm, vikunja (seeds A and B back, ledgers equal, cut-off
detected before anything moves, ≤ 5.3 s extra downtime). **Folder copy chosen (decision 19):** an app
with no database server gets no dump, so dump-and-load would need the folder copy anyway. Rate: 185 MB
in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s (warm cache); 5 GB ≈ 12 s here, 25–50 s on a cold or spinning
disk. Found on the way: the undo must never read the recovery unit (R-645), and killing `docker run`
does not stop the copy container.
## 3. Part 1 — the undo, by hand
## 3. Red-proofs (each seen failing, then restored)
Full evidence and tables: `documentation/audits/update-rulings-2026-09-23/README.md`.
| # | mutation | test that failed |
|---|---|---|
| 1 | the undo call removed (v0.262.1's shape) | `TestUndo_FailedUpdateIsPutBackWithItsData` — `held=true phase="failed"` |
| 2 | copy validation removed | `TestUndo_CutOffCopyIsRefusedBeforeAnythingMoves` — state `half` |
| 3 | old-probe rule removed | `TestUndo_UsesTheOldProbe` — `not_started` |
| 3b | the old probe read from the stack dir (v0.263.1's shape) | `TestUndo_UsesTheOldProbe` |
| 4 | the `undoing` recovery arm removed | `TestUndo_PowerCutDuringTheUndoResumesIt` — `resumed=[]` |
| 5a | `last_update_undone` not recorded | `TestUndo_FailedUpdateIsPutBackWithItsData` — `got <nil>` |
| 5b | the page line removed from the handler | `TestUndo_PageSaysTheBoxPutTheAppBack` (hu and en) |
| 6 | the hold prefix ignored | `TestUndo_HoldSentenceSaysTheUndoWasTriedAndTheDataState` |
| 7 | R-642: "completed" back | `TestR642_StartIsNeverReportedCompleted` |
| 8 | a successful update keeps the undone note | `TestUndo_ASuccessfulUpdateEndsTheUndoneNote` |
| 9 | the undo probe switched off on an `unhealthy` app (v0.263.0's shape) | `TestUndo_OldProbeRunsOnAnAppTheCurrentProbeMarkedUnhealthy` — `not healthy within 1s (last: state unhealthy)`, the live message verbatim |
| 10 | the pin advance records no probe | `TestUndo_PinningRecordsThatVersionsProbe` |
| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (volume, no DB server) |
|---|---|---|---|
| held after | 95.3 s | 102.6 s | 93.3 s |
| old version on migrated data | **refuses** | **refuses** | starts |
| safety dump | 135 816 B, DB only | 62 943 B, DB only | **none** |
| product loader (`ImportDump` semantics) | **FAILS**, rc 3, foreign keys | rc 0, **12 tables left behind** | — |
| fixed load (empty schema + copy, one transaction) | rc 0, 1.38 s | (product loader sufficed) | — |
| undo → healthy on the OLD probe | ≈ 16 s | ≈ 38 s | ≈ 1 s |
| data before / after the backup | yes / **yes** | yes / **yes** | yes / **yes** (attachment too) |
`go build ./... && go vet ./... && go test ./...` rc=0 before each of the three commits; controller
gates rc=0 (the `docker -v` gate needed four named-volume mounts allowlisted with their why).
**Does a product path load a safety dump back?** Yes — `rollbackSafetyDump`, but only the off-site
restore calls it; the update never reads its own dump, and `failAndHold` deletes the pre-update
definition copies.
## 4. Live proof (`audits/undo-live-2026-09-23/README.md`) — endpoint-level, both languages
**The wrong case:** PostgreSQL + product loader → rc 3, nothing changed (honest hold). **PostgreSQL +
the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints,
0 indexes) that still shows 42 tables and 48 ledger rows** — no hold, a dishonest success. MariaDB +
half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker
the truncated ones lack; `ValidateDump` does not check it.
Three apps undone by the product with seeds A, B, C back and ledgers equal (30–52 s of undo); page
line hu/en; cut-off copy → HOLD `untouched` (prefix in the box's language, hu and en quoted); power cut
during `undoing` → resumed and undone; a person's press after an undo → `done`, note cleared; removal
deleted kept copies; R-642 Start answer.
## 4. Part 2 — the ladder
## 5. Rows
One press on vikunja two steps behind: **2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran.** The box cannot see
2.4.0 (depth-1 clone). Recommended format: `update_ladder:` in `.felhom.yml`, intermediate steps with
their own definition — **not** git history, because romm's image-moving commit is the definition that
OOM-looped on demo-hp. Full comparison in the audit.
**Closed (4, moved to `CLOSED-ITEMS.md`):** R-637, R-639, R-641, R-642. **Narrowed:** R-638, R-640 (to
the restore paths). **Ruled:** R-643 (decision 20). **Noted:** R-645. **Opened:** R-645 (earlier this
session, filed with the bake-off) and R-646 (apps pinned before v0.263.2). **Open rows 337 → 335.**
## 5. Part 3 — the memory watch
## 6. Teardown — three layers
`upgrade-test.py` v2: after a successful readback, `--soak` seconds (default 600) of light load
(4 callers), sampling every 15 s the kernel's `oom_kill` counter read host-side from the container's
cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → `failed`;
peak > 80 % → mark `memory_tight`. New `Romm` fixture; edges `M1` (current template) and `M1old`
(the template as promoted, `15f9ebf`).
- **M1old (red-proof): `failed` — first OOM kill at +76 s**, peak 512 MiB = 100 % of the limit,
restarts 0 (the container kept running — the shape that hid it on demo-hp), abort `refuses`.
**Ten minutes is ample for this failure under load.**
- **M1 (positive control):** **proven** + mark `memory_tight` — 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open.
## 6. Part 4 — the build plan
`09` §6.4: ten parts, **≈ 22 evenings**, recommended order undo → sentences in the household's language
+ notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 →
R-625 → PostgreSQL conversion; fleet view deferred by ruling. **One open point needs the operator
(R-643):** as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate
(W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside
its own four-hour window.
## 7. Rows
**Opened (8):** R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the
restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted
on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have
no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg,
operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here).
**Updated:** R-446, R-450, R-451, R-462, R-463. **Closed:** none. **329 → 337.**
## 8. Teardown — three layers
| layer | state |
|---|---|
| **machine — guest 9202** | `controller.yaml` restored from the saved copy and read back identical (live catalog, no `update:` block); catalog cache re-cloned from the live repo (`02844ae`); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; `/opt/upg` and every temp file in `/root` removed; the eight test images removed **by name** — `vikunja:2.4.0` was **absent**, i.e. never pulled, which is the jump seen from a second side; **no `prune`**. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). `82-teardown-guest.txt` |
| **host — demo-hp** | nothing provisioned; only transient `/tmp` files, removed |
| **hub** | nothing touched — no hub call was made |
| **drill repo** | reset to live `main` `02844ae0a579`; `has_actions: false`; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. `81-teardown-drill-repo.txt` |
**One instrumentation slip, recorded:** a background watcher and the first teardown both wrote the same
temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file
held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run
is the one recorded.
**Fences:** DooPlex, Peti's box, ep0, the demo guests' apps, `drill-r50`, `tester-1` and the hub were
not touched. The live catalog's `main` was `02844ae0a579` before and after.
**`unproven.py --summary`:** walked 20 / partial 17 / built 14 / missing 4 — **not walked 35 of 55, unchanged.**
Machine (9202): the three apps and every copy removed through the product; test images removed by
name; `controller.yaml` restored and read back; catalog cache on live `cfcfe52`; **9202 stays on
controller 0.263.2** (self-update off, fleet floor untouched). Host: nothing provisioned; 9202 stopped
and started once for the power cut. Hub: untouched. Drill repo: reset to live `main`.