Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came back with data written before AND after the backup. The product's loader cannot do it: over a migrated PG database it fails on the new tables' foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads rc 0 into an empty database. No-DB apps have no last-second copy. - Part 2: one press jumps A -> C; the box's catalog clone is depth 1. Ladder format recommended: update_ladder in .felhom.yml, not git history. - Part 3: memory watch red-proof results (harness change in the catalog repo). - Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643). - Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT. No product code. Live catalog untouched; 9202 back on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,161 +1,119 @@
|
||||
# REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0
|
||||
# REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan
|
||||
|
||||
2026-09-21 (evening). Repos touched: **felhom.eu** (floor, docs, register, evidence),
|
||||
**felhom-controller v0.261.0**, **app-catalog-felhom.eu** (two drill pairs, both reverted).
|
||||
Architecture read first and named: `documentation/architecture/09-update-architecture.md` §3, §3b,
|
||||
§4, §6.1, §6.2, §6.4, §8.
|
||||
2026-09-23. Repos touched: **felhom.eu** (docs, register, STATUS, CONTEXT, evidence),
|
||||
**app-catalog-felhom.eu** (`scripts/` only — the memory watch), **admin/app-catalog-drill** (drill
|
||||
commits, reset to live `main` at the end). **felhom-controller, felhom-agent, hub: read only.**
|
||||
Architecture read first and named: `documentation/architecture/09-update-architecture.md` (all of it),
|
||||
`07-backup-architecture.md` §6.
|
||||
Baselines verified live before starting: controller `b9deec19077b`, agent `d9864a94bf62`, felhom.eu
|
||||
`267dcad01bcf`, catalog `02844ae0a579` — all equal to the brief.
|
||||
|
||||
---
|
||||
|
||||
## 1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session
|
||||
## 1. Not done, or changed from the brief — first
|
||||
|
||||
| item | state |
|
||||
|---|---|
|
||||
| **Part 0** floor to 0.260.0 | **done** |
|
||||
| **Part 1** the three cuts | **done, but NOT as specified.** All three landed in `verifying`, never in `starting` — see §2. The brief allowed this explicitly and asked that it be said. |
|
||||
| **Part 2** the lock + `reason` on the wire | **done**, five red-proofs, proven live |
|
||||
| **Part 3** the unattended night | **done for the success night and the no-retry proof. The unattended HOLD was NOT produced** — see §4. |
|
||||
| **Part 4** docs and rows | **done** |
|
||||
| the caller script's ≤150-line budget | **167 lines.** Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. |
|
||||
| Scenario C run by the measuring agent | **run by the coordinator instead.** The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. |
|
||||
| a second drill bump/revert pair | **used.** The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (`ae08a037fd68`) added one real edge and one deliberately failing edge for Part 3. |
|
||||
| Part 0 — rulings into `09` | **done**, commit `805ad1e` (documents only) |
|
||||
| Part 1 — the undo, four cases + the wrong case | **done, with three changes.** (a) **The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder:** romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) **Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around.** The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0). |
|
||||
| Part 2 — the ladder | **done** |
|
||||
| Part 3 — the memory watch + red-proof | **done**; harness v2. `C3`, the harness's standing negative control, was **NOT run**: its template's `container_name: privatebin` collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass. |
|
||||
| Part 4 — the build plan | **done**, `09` §6.4 — with **one open point for the operator** the brief did not expect (§6) |
|
||||
|
||||
**Two instrumentation failures of my own, recorded because they cost evidence:**
|
||||
1. The first unattended run's stdout was piped through `tail`, which buffers, and the run was later
|
||||
killed — **the caller's own log for the Scenario F press was lost.** The outcome survived on the
|
||||
box; the log did not. The second run wrote straight to a file.
|
||||
2. A background security review flagged the deliberately-broken `vikunja → alpine:3.20` catalog edge
|
||||
as a supply-chain change. **It was right to.** Accepted deliberately — no customer or demo box runs
|
||||
vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class
|
||||
control — but the window is now closed by the revert, and it is named here rather than left in a
|
||||
tool notification.
|
||||
**Claims in the brief that turned out wrong:**
|
||||
1. *"A safety dump exists for every app class"* — **false.** An app with no database server gets none
|
||||
(`update safety dump for vikunja: the app has no database — nothing to copy (no-op)`), measured.
|
||||
2. *"The box keeps a git clone of the catalog"* with history — **false.** Depth 1 on both demo
|
||||
guests, `rev-list --count HEAD` = 1; `sync.go:283`/`:300` clone and fetch `--depth 1`.
|
||||
3. *"`stacks.update_window` is unread"* — **true, and the grep was widened** from `config.go` +
|
||||
`setup/handlers.go` to the whole controller repo: the only other hits are
|
||||
`configs/controller.yaml.example` and the i18n base file. No Go code reads it.
|
||||
4. The register held **329** row lines by `grep -c '^| \*\*R-'`, not 326; the highest id was R-636 as stated.
|
||||
|
||||
---
|
||||
## 2. Part 0 — the rulings
|
||||
|
||||
## 2. Claims in the brief that turned out wrong
|
||||
`09` §3 gains decisions **11–18** in the existing shape; decision 3's second half and §6.1's abort
|
||||
paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept,
|
||||
headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices
|
||||
table updated. Register: R-450, R-451, R-446, R-463 cite the decisions.
|
||||
|
||||
**§2.4's open question — does `backupMgr.IsRunning()` cover the update's `backing-up` phase?**
|
||||
**YES.** `RunAppBackupNow` calls `acquireRunning` (`internal/backup/update_guard.go:333`), so that one
|
||||
phase was already protected. The gap was `checking`, `safety-dump`, `pinning`, `pulling`, `starting`
|
||||
and `verifying`. **The live lock probe landed in `safety-dump`**, i.e. squarely in the previously
|
||||
unprotected window rather than in the one that was already covered. Everything else §2.4 asserted
|
||||
held at source.
|
||||
## 3. Part 1 — the undo, by hand
|
||||
|
||||
**§2.3 — the resume path was READ, not measured. It is now measured**, three times, and it does what
|
||||
it said.
|
||||
Full evidence and tables: `documentation/audits/update-rulings-2026-09-23/README.md`.
|
||||
|
||||
**The stopped-guest byte path in my own brief to the measuring agent was wrong.**
|
||||
`/var/lib/lxc/9202/rootfs/var/lib/felhom/...` is an empty mountpoint while 9202 is stopped, because
|
||||
`mp0` is a separate raw volume; `pct mount 9202` does attach it. I had corrected one trap and
|
||||
introduced a second. The positive control caught it.
|
||||
|
||||
**"Four qualifying apps exist" — held.** vikunja, uptime-kuma, wishlist, glance: single-container, no
|
||||
database sidecar, none on either demo box, all four target tags verified to exist upstream first.
|
||||
|
||||
---
|
||||
|
||||
## 3. Part 0 — the floor
|
||||
|
||||
Raised to **0.260.0** with MinAgent **0.131.0** declared (above the vouched golden 0.258.0, so the
|
||||
declaration carries it — §3 decision 7). Hub log:
|
||||
|
||||
```
|
||||
[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0")
|
||||
[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
```
|
||||
|
||||
Blast radius, read from the hub before saving: **3 boxes below** — `drill-r50` (0.213.0, BLOCKED),
|
||||
`peti-felhom` (0.115.0, DOWN), `tester-1` (0.245.0, DOWN). None was reachable, so none moved; they
|
||||
take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor.
|
||||
|
||||
---
|
||||
|
||||
## 4. Part 1 — the three cuts (R-610, CLOSED)
|
||||
|
||||
Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first.
|
||||
|
||||
| | A — vikunja | B — uptime-kuma | C — wishlist |
|
||||
| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (volume, no DB server) |
|
||||
|---|---|---|---|
|
||||
| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 |
|
||||
| cut | `pct stop` | `pct stop` | **controller container only** |
|
||||
| phase at decision | `starting` | `verifying` | `starting` |
|
||||
| cut latency | 3 759 ms | 3 016 ms | **1 675 ms** |
|
||||
| phase it died in | `verifying` | `verifying` | `verifying` |
|
||||
| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s |
|
||||
| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s |
|
||||
| four observables | **agree** | **agree** | **agree** |
|
||||
| seeded data | **read back intact** | **read back intact** | not re-read (gap) |
|
||||
| held after | 95.3 s | 102.6 s | 93.3 s |
|
||||
| old version on migrated data | **refuses** | **refuses** | starts |
|
||||
| safety dump | 135 816 B, DB only | 62 943 B, DB only | **none** |
|
||||
| product loader (`ImportDump` semantics) | **FAILS**, rc 3, foreign keys | rc 0, **12 tables left behind** | — |
|
||||
| fixed load (empty schema + copy, one transaction) | rc 0, 1.38 s | (product loader sufficed) | — |
|
||||
| undo → healthy on the OLD probe | ≈ 16 s | ≈ 38 s | ≈ 1 s |
|
||||
| data before / after the backup | yes / **yes** | yes / **yes** | yes / **yes** (attachment too) |
|
||||
|
||||
**The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.**
|
||||
vikunja's own log: `Ran all migrations successfully` / `Vikunja version v2.6.0` at **12:28:26.881
|
||||
UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema
|
||||
migration had already been applied to the customer's database when the power went. Recovery resumed
|
||||
**forward**, so old-binary-on-migrated-database never happened — **but this branch is one step from
|
||||
it**, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it.
|
||||
**Does a product path load a safety dump back?** Yes — `rollbackSafetyDump`, but only the off-site
|
||||
restore calls it; the update never reads its own dump, and `failAndHold` deletes the pre-update
|
||||
definition copies.
|
||||
|
||||
**Instrument limit:** `starting` lasts well under a second on this box. Three attempts, two cut
|
||||
mechanisms, all landed in `verifying`. No phase was faked. `RecoverUpdates` handles `starting` and
|
||||
`verifying` in **one branch**, so all three exercise the arm under test. A cut inside `starting`
|
||||
itself needs an in-process fault injector.
|
||||
**The wrong case:** PostgreSQL + product loader → rc 3, nothing changed (honest hold). **PostgreSQL +
|
||||
the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints,
|
||||
0 indexes) that still shows 42 tables and 48 ledger rows** — no hold, a dishonest success. MariaDB +
|
||||
half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker
|
||||
the truncated ones lack; `ValidateDump` does not check it.
|
||||
|
||||
**Scenario C also showed the other apps were undisturbed** by the controller restart — uptime-kuma,
|
||||
vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating
|
||||
restarted.
|
||||
## 4. Part 2 — the ladder
|
||||
|
||||
---
|
||||
One press on vikunja two steps behind: **2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran.** The box cannot see
|
||||
2.4.0 (depth-1 clone). Recommended format: `update_ladder:` in `.felhom.yml`, intermediate steps with
|
||||
their own definition — **not** git history, because romm's image-moving commit is the definition that
|
||||
OOM-looped on demo-hp. Full comparison in the audit.
|
||||
|
||||
## 5. Part 2 — v0.261.0, proven live with a control
|
||||
## 5. Part 3 — the memory watch
|
||||
|
||||
See `felhom-controller/REPORT.md` for the code. The live proof is the part worth repeating:
|
||||
`upgrade-test.py` v2: after a successful readback, `--soak` seconds (default 600) of light load
|
||||
(4 callers), sampling every 15 s the kernel's `oom_kill` counter read host-side from the container's
|
||||
cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → `failed`;
|
||||
peak > 80 % → mark `memory_tight`. New `Romm` fixture; edges `M1` (current template) and `M1old`
|
||||
(the template as promoted, `15f9ebf`).
|
||||
|
||||
| probe | result |
|
||||
- **M1old (red-proof): `failed` — first OOM kill at +76 s**, peak 512 MiB = 100 % of the limit,
|
||||
restarts 0 (the container kept running — the shape that hid it on demo-hp), abort `refuses`.
|
||||
**Ten minutes is ample for this failure under load.**
|
||||
- **M1 (positive control):** **proven** + mark `memory_tight` — 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open.
|
||||
|
||||
## 6. Part 4 — the build plan
|
||||
|
||||
`09` §6.4: ten parts, **≈ 22 evenings**, recommended order undo → sentences in the household's language
|
||||
+ notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 →
|
||||
R-625 → PostgreSQL conversion; fleet view deferred by ruling. **One open point needs the operator
|
||||
(R-643):** as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate
|
||||
(W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside
|
||||
its own four-hour window.
|
||||
|
||||
## 7. Rows
|
||||
|
||||
**Opened (8):** R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the
|
||||
restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted
|
||||
on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have
|
||||
no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg,
|
||||
operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here).
|
||||
**Updated:** R-446, R-450, R-451, R-462, R-463. **Closed:** none. **329 → 337.**
|
||||
|
||||
## 8. Teardown — three layers
|
||||
|
||||
| layer | state |
|
||||
|---|---|
|
||||
| manual self-update, **no** app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the **agent** refusal |
|
||||
| manual self-update, app update **in flight** (`safety-dump`) | „Egy alkalmazás frissítése éppen folyamatban van…" — **our** refusal |
|
||||
| **machine — guest 9202** | `controller.yaml` restored from the saved copy and read back identical (live catalog, no `update:` block); catalog cache re-cloned from the live repo (`02844ae`); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; `/opt/upg` and every temp file in `/root` removed; the eight test images removed **by name** — `vikunja:2.4.0` was **absent**, i.e. never pulled, which is the jump seen from a second side; **no `prune`**. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). `82-teardown-guest.txt` |
|
||||
| **host — demo-hp** | nothing provisioned; only transient `/tmp` files, removed |
|
||||
| **hub** | nothing touched — no hub call was made |
|
||||
| **drill repo** | reset to live `main` `02844ae0a579`; `has_actions: false`; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. `81-teardown-drill-repo.txt` |
|
||||
|
||||
**The sentence changed.** Guest 9202 has no host agent, so `TriggerUpdate` refuses either way — which
|
||||
makes it the perfect negative control, because the new check sits *before* the agent check. Two
|
||||
sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and
|
||||
**restored to `false`** from a copy taken first; verified by re-reading the file.
|
||||
**One instrumentation slip, recorded:** a background watcher and the first teardown both wrote the same
|
||||
temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file
|
||||
held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run
|
||||
is the one recorded.
|
||||
|
||||
**The reverse direction — an app update refused while the controller swaps — is NOT staged live.** It
|
||||
is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this
|
||||
guest does not have. Stated as a gap.
|
||||
**Fences:** DooPlex, Peti's box, ep0, the demo guests' apps, `drill-r50`, `tester-1` and the hub were
|
||||
not touched. The live catalog's `main` was `02844ae0a579` before and after.
|
||||
|
||||
---
|
||||
|
||||
## 6. Part 3 — the unattended night (R-611, CLOSED)
|
||||
|
||||
**Success:** `uptime-kuma` 2.5.0 → 2.5.1 applied with nobody pressing anything.
|
||||
|
||||
**No-retry:** after the revert left all four apps *ahead* of the catalog, the caller pressed each
|
||||
**exactly once**, was refused `downgrade` (terminal), and pressed nothing across two further passes.
|
||||
`never_again=['glance','uptime-kuma','vikunja','wishlist']`, `outcomes={}`. That is R-524 and R-609
|
||||
working together, unattended.
|
||||
|
||||
**The unattended HOLD was never produced, and the reason matters:** the only failing edge available
|
||||
(`vikunja → alpine:3.20`) was **correctly refused by the within-a-major rule before it was ever
|
||||
attempted**. The rule that makes automatic updates safe is the same rule that refuses the obvious way
|
||||
to break one. Measuring it needs an image that passes the version test and still fails health.
|
||||
**§3b Q4 therefore still rests on the ATTENDED hold from slice 4.**
|
||||
|
||||
---
|
||||
|
||||
## 7. Rows, and the catalog
|
||||
|
||||
**Closed:** R-608, R-609, R-610, R-611. **Opened:** R-612 (P1 — wishlist unusable on a fresh install
|
||||
and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 —
|
||||
stale update phase survives a redeploy). **Corrected:** R-520's closing pointer now names R-610.
|
||||
Register 302 → **307**.
|
||||
|
||||
**Catalog:** DRILL `573e41f5`, DRILL 2 `ae08a037`, **REVERT `f5f6a152`**. Every `image:` line in
|
||||
`templates/` is byte-identical to the pre-drill `ff9717d3` — `git diff` over those paths is **0
|
||||
lines**. `catalog_since` reads 2026-09-21 on the four rather than the older dates, because the gate
|
||||
requires an image move to carry the day's date in either direction and a revert is a move.
|
||||
|
||||
**Teardown, three layers.** *Machine:* guest 9202 left running on v0.261.0 with the four throwaway
|
||||
apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for
|
||||
the remaining Q4 work and removing them would cost the next session the seeding. *Host:* demo-hp
|
||||
untouched apart from 9202; **guest 9201 never touched.** *Hub:* nothing provisioned, nothing
|
||||
enrolled; 9202 reports to no hub by design. `/tmp/.ctlpw` shredded.
|
||||
**`unproven.py --summary`:** walked 20 / partial 17 / built 14 / missing 4 — **not walked 35 of 55, unchanged.**
|
||||
|
||||
Reference in New Issue
Block a user