Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
gates / gates (push) Successful in 26s

- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
  (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
  back with data written before AND after the backup. The product's loader
  cannot do it: over a migrated PG database it fails on the new tables'
  foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
  rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
  Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.

No product code. Live catalog untouched; 9202 back on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 08:54:41 +02:00
parent 805ad1e962
commit 4c92beab8f
67 changed files with 5977 additions and 162 deletions
+94 -136
View File
@@ -1,161 +1,119 @@
# REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0
# REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan
2026-09-21 (evening). Repos touched: **felhom.eu** (floor, docs, register, evidence),
**felhom-controller v0.261.0**, **app-catalog-felhom.eu** (two drill pairs, both reverted).
Architecture read first and named: `documentation/architecture/09-update-architecture.md` §3, §3b,
§4, §6.1, §6.2, §6.4, §8.
2026-09-23. Repos touched: **felhom.eu** (docs, register, STATUS, CONTEXT, evidence),
**app-catalog-felhom.eu** (`scripts/` only — the memory watch), **admin/app-catalog-drill** (drill
commits, reset to live `main` at the end). **felhom-controller, felhom-agent, hub: read only.**
Architecture read first and named: `documentation/architecture/09-update-architecture.md` (all of it),
`07-backup-architecture.md` §6.
Baselines verified live before starting: controller `b9deec19077b`, agent `d9864a94bf62`, felhom.eu
`267dcad01bcf`, catalog `02844ae0a579` — all equal to the brief.
---
## 1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session
## 1. Not done, or changed from the brief — first
| item | state |
|---|---|
| **Part 0** floor to 0.260.0 | **done** |
| **Part 1** the three cuts | **done, but NOT as specified.** All three landed in `verifying`, never in `starting` — see §2. The brief allowed this explicitly and asked that it be said. |
| **Part 2** the lock + `reason` on the wire | **done**, five red-proofs, proven live |
| **Part 3** the unattended night | **done for the success night and the no-retry proof. The unattended HOLD was NOT produced** — see §4. |
| **Part 4** docs and rows | **done** |
| the caller script's ≤150-line budget | **167 lines.** Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. |
| Scenario C run by the measuring agent | **run by the coordinator instead.** The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. |
| a second drill bump/revert pair | **used.** The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (`ae08a037fd68`) added one real edge and one deliberately failing edge for Part 3. |
| Part 0 — rulings into `09` | **done**, commit `805ad1e` (documents only) |
| Part 1 — the undo, four cases + the wrong case | **done, with three changes.** (a) **The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder:** romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) **Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around.** The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0). |
| Part 2 — the ladder | **done** |
| Part 3 — the memory watch + red-proof | **done**; harness v2. `C3`, the harness's standing negative control, was **NOT run**: its template's `container_name: privatebin` collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass. |
| Part 4 — the build plan | **done**, `09` §6.4 — with **one open point for the operator** the brief did not expect (§6) |
**Two instrumentation failures of my own, recorded because they cost evidence:**
1. The first unattended run's stdout was piped through `tail`, which buffers, and the run was later
killed — **the caller's own log for the Scenario F press was lost.** The outcome survived on the
box; the log did not. The second run wrote straight to a file.
2. A background security review flagged the deliberately-broken `vikunja → alpine:3.20` catalog edge
as a supply-chain change. **It was right to.** Accepted deliberately — no customer or demo box runs
vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class
control — but the window is now closed by the revert, and it is named here rather than left in a
tool notification.
**Claims in the brief that turned out wrong:**
1. *"A safety dump exists for every app class"* — **false.** An app with no database server gets none
(`update safety dump for vikunja: the app has no database — nothing to copy (no-op)`), measured.
2. *"The box keeps a git clone of the catalog"* with history — **false.** Depth 1 on both demo
guests, `rev-list --count HEAD` = 1; `sync.go:283`/`:300` clone and fetch `--depth 1`.
3. *"`stacks.update_window` is unread"* — **true, and the grep was widened** from `config.go` +
`setup/handlers.go` to the whole controller repo: the only other hits are
`configs/controller.yaml.example` and the i18n base file. No Go code reads it.
4. The register held **329** row lines by `grep -c '^| \*\*R-'`, not 326; the highest id was R-636 as stated.
---
## 2. Part 0 — the rulings
## 2. Claims in the brief that turned out wrong
`09` §3 gains decisions **11–18** in the existing shape; decision 3's second half and §6.1's abort
paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept,
headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices
table updated. Register: R-450, R-451, R-446, R-463 cite the decisions.
**§2.4's open question — does `backupMgr.IsRunning()` cover the update's `backing-up` phase?**
**YES.** `RunAppBackupNow` calls `acquireRunning` (`internal/backup/update_guard.go:333`), so that one
phase was already protected. The gap was `checking`, `safety-dump`, `pinning`, `pulling`, `starting`
and `verifying`. **The live lock probe landed in `safety-dump`**, i.e. squarely in the previously
unprotected window rather than in the one that was already covered. Everything else §2.4 asserted
held at source.
## 3. Part 1 — the undo, by hand
**§2.3 — the resume path was READ, not measured. It is now measured**, three times, and it does what
it said.
Full evidence and tables: `documentation/audits/update-rulings-2026-09-23/README.md`.
**The stopped-guest byte path in my own brief to the measuring agent was wrong.**
`/var/lib/lxc/9202/rootfs/var/lib/felhom/...` is an empty mountpoint while 9202 is stopped, because
`mp0` is a separate raw volume; `pct mount 9202` does attach it. I had corrected one trap and
introduced a second. The positive control caught it.
**"Four qualifying apps exist" — held.** vikunja, uptime-kuma, wishlist, glance: single-container, no
database sidecar, none on either demo box, all four target tags verified to exist upstream first.
---
## 3. Part 0 — the floor
Raised to **0.260.0** with MinAgent **0.131.0** declared (above the vouched golden 0.258.0, so the
declaration carries it — §3 decision 7). Hub log:
```
[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0")
[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
```
Blast radius, read from the hub before saving: **3 boxes below** — `drill-r50` (0.213.0, BLOCKED),
`peti-felhom` (0.115.0, DOWN), `tester-1` (0.245.0, DOWN). None was reachable, so none moved; they
take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor.
---
## 4. Part 1 — the three cuts (R-610, CLOSED)
Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first.
| | A — vikunja | B — uptime-kuma | C — wishlist |
| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (volume, no DB server) |
|---|---|---|---|
| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 |
| cut | `pct stop` | `pct stop` | **controller container only** |
| phase at decision | `starting` | `verifying` | `starting` |
| cut latency | 3 759 ms | 3 016 ms | **1 675 ms** |
| phase it died in | `verifying` | `verifying` | `verifying` |
| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s |
| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s |
| four observables | **agree** | **agree** | **agree** |
| seeded data | **read back intact** | **read back intact** | not re-read (gap) |
| held after | 95.3 s | 102.6 s | 93.3 s |
| old version on migrated data | **refuses** | **refuses** | starts |
| safety dump | 135 816 B, DB only | 62 943 B, DB only | **none** |
| product loader (`ImportDump` semantics) | **FAILS**, rc 3, foreign keys | rc 0, **12 tables left behind** | — |
| fixed load (empty schema + copy, one transaction) | rc 0, 1.38 s | (product loader sufficed) | — |
| undo → healthy on the OLD probe | ≈ 16 s | ≈ 38 s | ≈ 1 s |
| data before / after the backup | yes / **yes** | yes / **yes** | yes / **yes** (attachment too) |
**The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.**
vikunja's own log: `Ran all migrations successfully` / `Vikunja version v2.6.0` at **12:28:26.881
UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema
migration had already been applied to the customer's database when the power went. Recovery resumed
**forward**, so old-binary-on-migrated-database never happened — **but this branch is one step from
it**, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it.
**Does a product path load a safety dump back?** Yes — `rollbackSafetyDump`, but only the off-site
restore calls it; the update never reads its own dump, and `failAndHold` deletes the pre-update
definition copies.
**Instrument limit:** `starting` lasts well under a second on this box. Three attempts, two cut
mechanisms, all landed in `verifying`. No phase was faked. `RecoverUpdates` handles `starting` and
`verifying` in **one branch**, so all three exercise the arm under test. A cut inside `starting`
itself needs an in-process fault injector.
**The wrong case:** PostgreSQL + product loader → rc 3, nothing changed (honest hold). **PostgreSQL +
the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints,
0 indexes) that still shows 42 tables and 48 ledger rows** — no hold, a dishonest success. MariaDB +
half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker
the truncated ones lack; `ValidateDump` does not check it.
**Scenario C also showed the other apps were undisturbed** by the controller restart — uptime-kuma,
vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating
restarted.
## 4. Part 2 — the ladder
---
One press on vikunja two steps behind: **2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran.** The box cannot see
2.4.0 (depth-1 clone). Recommended format: `update_ladder:` in `.felhom.yml`, intermediate steps with
their own definition — **not** git history, because romm's image-moving commit is the definition that
OOM-looped on demo-hp. Full comparison in the audit.
## 5. Part 2 — v0.261.0, proven live with a control
## 5. Part 3 — the memory watch
See `felhom-controller/REPORT.md` for the code. The live proof is the part worth repeating:
`upgrade-test.py` v2: after a successful readback, `--soak` seconds (default 600) of light load
(4 callers), sampling every 15 s the kernel's `oom_kill` counter read host-side from the container's
cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → `failed`;
peak > 80 % → mark `memory_tight`. New `Romm` fixture; edges `M1` (current template) and `M1old`
(the template as promoted, `15f9ebf`).
| probe | result |
- **M1old (red-proof): `failed` — first OOM kill at +76 s**, peak 512 MiB = 100 % of the limit,
restarts 0 (the container kept running — the shape that hid it on demo-hp), abort `refuses`.
**Ten minutes is ample for this failure under load.**
- **M1 (positive control):** **proven** + mark `memory_tight` — 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open.
## 6. Part 4 — the build plan
`09` §6.4: ten parts, **≈ 22 evenings**, recommended order undo → sentences in the household's language
+ notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 →
R-625 → PostgreSQL conversion; fleet view deferred by ruling. **One open point needs the operator
(R-643):** as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate
(W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside
its own four-hour window.
## 7. Rows
**Opened (8):** R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the
restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted
on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have
no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg,
operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here).
**Updated:** R-446, R-450, R-451, R-462, R-463. **Closed:** none. **329 → 337.**
## 8. Teardown — three layers
| layer | state |
|---|---|
| manual self-update, **no** app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the **agent** refusal |
| manual self-update, app update **in flight** (`safety-dump`) | „Egy alkalmazás frissítése éppen folyamatban van…" — **our** refusal |
| **machine — guest 9202** | `controller.yaml` restored from the saved copy and read back identical (live catalog, no `update:` block); catalog cache re-cloned from the live repo (`02844ae`); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; `/opt/upg` and every temp file in `/root` removed; the eight test images removed **by name** — `vikunja:2.4.0` was **absent**, i.e. never pulled, which is the jump seen from a second side; **no `prune`**. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). `82-teardown-guest.txt` |
| **host — demo-hp** | nothing provisioned; only transient `/tmp` files, removed |
| **hub** | nothing touched — no hub call was made |
| **drill repo** | reset to live `main` `02844ae0a579`; `has_actions: false`; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. `81-teardown-drill-repo.txt` |
**The sentence changed.** Guest 9202 has no host agent, so `TriggerUpdate` refuses either way — which
makes it the perfect negative control, because the new check sits *before* the agent check. Two
sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and
**restored to `false`** from a copy taken first; verified by re-reading the file.
**One instrumentation slip, recorded:** a background watcher and the first teardown both wrote the same
temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file
held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run
is the one recorded.
**The reverse direction — an app update refused while the controller swaps — is NOT staged live.** It
is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this
guest does not have. Stated as a gap.
**Fences:** DooPlex, Peti's box, ep0, the demo guests' apps, `drill-r50`, `tester-1` and the hub were
not touched. The live catalog's `main` was `02844ae0a579` before and after.
---
## 6. Part 3 — the unattended night (R-611, CLOSED)
**Success:** `uptime-kuma` 2.5.0 → 2.5.1 applied with nobody pressing anything.
**No-retry:** after the revert left all four apps *ahead* of the catalog, the caller pressed each
**exactly once**, was refused `downgrade` (terminal), and pressed nothing across two further passes.
`never_again=['glance','uptime-kuma','vikunja','wishlist']`, `outcomes={}`. That is R-524 and R-609
working together, unattended.
**The unattended HOLD was never produced, and the reason matters:** the only failing edge available
(`vikunja → alpine:3.20`) was **correctly refused by the within-a-major rule before it was ever
attempted**. The rule that makes automatic updates safe is the same rule that refuses the obvious way
to break one. Measuring it needs an image that passes the version test and still fails health.
**§3b Q4 therefore still rests on the ATTENDED hold from slice 4.**
---
## 7. Rows, and the catalog
**Closed:** R-608, R-609, R-610, R-611. **Opened:** R-612 (P1 — wishlist unusable on a fresh install
and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 —
stale update phase survives a redeploy). **Corrected:** R-520's closing pointer now names R-610.
Register 302 → **307**.
**Catalog:** DRILL `573e41f5`, DRILL 2 `ae08a037`, **REVERT `f5f6a152`**. Every `image:` line in
`templates/` is byte-identical to the pre-drill `ff9717d3` — `git diff` over those paths is **0
lines**. `catalog_since` reads 2026-09-21 on the four rather than the older dates, because the gate
requires an image move to carry the day's date in either direction and a revert is a move.
**Teardown, three layers.** *Machine:* guest 9202 left running on v0.261.0 with the four throwaway
apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for
the remaining Q4 work and removing them would cost the next session the seeding. *Host:* demo-hp
untouched apart from 9202; **guest 9201 never touched.** *Hub:* nothing provisioned, nothing
enrolled; 9202 reports to no hub by design. `/tmp/.ctlpw` shredded.
**`unproven.py --summary`:** walked 20 / partial 17 / built 14 / missing 4 — **not walked 35 of 55, unchanged.**