c85262111c
gates / gates (push) Successful in 23s
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all down or blocked; both demo boxes SERVED. Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s after the cut decision and the seeded data read back intact — so the branch that is one step from old-binary-on-migrated-database is now evidence, not argument. Instrument limit stated: `starting` lasts under a second; all three landed in `verifying`, which RecoverUpdates handles in the same branch. Part 3 (R-611) — the night the previous session skipped without saying so. An app updated with nobody pressing anything; a terminally-refused app was pressed exactly once and never again over three passes. The unattended HOLD was NOT produced: the within-a-major rule correctly refused the broken edge before it was attempted, so Q4 still rests on the attended hold from slice 4. Said plainly rather than implied. Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard), R-614 (stale update phase survives a redeploy). R-520's pointer corrected. Catalog: two drill pairs, both reverted; every image line byte-identical to ff9717d3. The alpine:3.20 negative control a security review flagged is cleared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
162 lines
9.5 KiB
Markdown
162 lines
9.5 KiB
Markdown
# REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0
|
|
|
|
2026-09-21 (evening). Repos touched: **felhom.eu** (floor, docs, register, evidence),
|
|
**felhom-controller v0.261.0**, **app-catalog-felhom.eu** (two drill pairs, both reverted).
|
|
Architecture read first and named: `documentation/architecture/09-update-architecture.md` §3, §3b,
|
|
§4, §6.1, §6.2, §6.4, §8.
|
|
|
|
---
|
|
|
|
## 1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session
|
|
|
|
| item | state |
|
|
|---|---|
|
|
| **Part 0** floor to 0.260.0 | **done** |
|
|
| **Part 1** the three cuts | **done, but NOT as specified.** All three landed in `verifying`, never in `starting` — see §2. The brief allowed this explicitly and asked that it be said. |
|
|
| **Part 2** the lock + `reason` on the wire | **done**, five red-proofs, proven live |
|
|
| **Part 3** the unattended night | **done for the success night and the no-retry proof. The unattended HOLD was NOT produced** — see §4. |
|
|
| **Part 4** docs and rows | **done** |
|
|
| the caller script's ≤150-line budget | **167 lines.** Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. |
|
|
| Scenario C run by the measuring agent | **run by the coordinator instead.** The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. |
|
|
| a second drill bump/revert pair | **used.** The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (`ae08a037fd68`) added one real edge and one deliberately failing edge for Part 3. |
|
|
|
|
**Two instrumentation failures of my own, recorded because they cost evidence:**
|
|
1. The first unattended run's stdout was piped through `tail`, which buffers, and the run was later
|
|
killed — **the caller's own log for the Scenario F press was lost.** The outcome survived on the
|
|
box; the log did not. The second run wrote straight to a file.
|
|
2. A background security review flagged the deliberately-broken `vikunja → alpine:3.20` catalog edge
|
|
as a supply-chain change. **It was right to.** Accepted deliberately — no customer or demo box runs
|
|
vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class
|
|
control — but the window is now closed by the revert, and it is named here rather than left in a
|
|
tool notification.
|
|
|
|
---
|
|
|
|
## 2. Claims in the brief that turned out wrong
|
|
|
|
**§2.4's open question — does `backupMgr.IsRunning()` cover the update's `backing-up` phase?**
|
|
**YES.** `RunAppBackupNow` calls `acquireRunning` (`internal/backup/update_guard.go:333`), so that one
|
|
phase was already protected. The gap was `checking`, `safety-dump`, `pinning`, `pulling`, `starting`
|
|
and `verifying`. **The live lock probe landed in `safety-dump`**, i.e. squarely in the previously
|
|
unprotected window rather than in the one that was already covered. Everything else §2.4 asserted
|
|
held at source.
|
|
|
|
**§2.3 — the resume path was READ, not measured. It is now measured**, three times, and it does what
|
|
it said.
|
|
|
|
**The stopped-guest byte path in my own brief to the measuring agent was wrong.**
|
|
`/var/lib/lxc/9202/rootfs/var/lib/felhom/...` is an empty mountpoint while 9202 is stopped, because
|
|
`mp0` is a separate raw volume; `pct mount 9202` does attach it. I had corrected one trap and
|
|
introduced a second. The positive control caught it.
|
|
|
|
**"Four qualifying apps exist" — held.** vikunja, uptime-kuma, wishlist, glance: single-container, no
|
|
database sidecar, none on either demo box, all four target tags verified to exist upstream first.
|
|
|
|
---
|
|
|
|
## 3. Part 0 — the floor
|
|
|
|
Raised to **0.260.0** with MinAgent **0.131.0** declared (above the vouched golden 0.258.0, so the
|
|
declaration carries it — §3 decision 7). Hub log:
|
|
|
|
```
|
|
[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0")
|
|
[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
|
[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
|
```
|
|
|
|
Blast radius, read from the hub before saving: **3 boxes below** — `drill-r50` (0.213.0, BLOCKED),
|
|
`peti-felhom` (0.115.0, DOWN), `tester-1` (0.245.0, DOWN). None was reachable, so none moved; they
|
|
take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor.
|
|
|
|
---
|
|
|
|
## 4. Part 1 — the three cuts (R-610, CLOSED)
|
|
|
|
Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first.
|
|
|
|
| | A — vikunja | B — uptime-kuma | C — wishlist |
|
|
|---|---|---|---|
|
|
| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 |
|
|
| cut | `pct stop` | `pct stop` | **controller container only** |
|
|
| phase at decision | `starting` | `verifying` | `starting` |
|
|
| cut latency | 3 759 ms | 3 016 ms | **1 675 ms** |
|
|
| phase it died in | `verifying` | `verifying` | `verifying` |
|
|
| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s |
|
|
| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s |
|
|
| four observables | **agree** | **agree** | **agree** |
|
|
| seeded data | **read back intact** | **read back intact** | not re-read (gap) |
|
|
|
|
**The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.**
|
|
vikunja's own log: `Ran all migrations successfully` / `Vikunja version v2.6.0` at **12:28:26.881
|
|
UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema
|
|
migration had already been applied to the customer's database when the power went. Recovery resumed
|
|
**forward**, so old-binary-on-migrated-database never happened — **but this branch is one step from
|
|
it**, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it.
|
|
|
|
**Instrument limit:** `starting` lasts well under a second on this box. Three attempts, two cut
|
|
mechanisms, all landed in `verifying`. No phase was faked. `RecoverUpdates` handles `starting` and
|
|
`verifying` in **one branch**, so all three exercise the arm under test. A cut inside `starting`
|
|
itself needs an in-process fault injector.
|
|
|
|
**Scenario C also showed the other apps were undisturbed** by the controller restart — uptime-kuma,
|
|
vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating
|
|
restarted.
|
|
|
|
---
|
|
|
|
## 5. Part 2 — v0.261.0, proven live with a control
|
|
|
|
See `felhom-controller/REPORT.md` for the code. The live proof is the part worth repeating:
|
|
|
|
| probe | result |
|
|
|---|---|
|
|
| manual self-update, **no** app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the **agent** refusal |
|
|
| manual self-update, app update **in flight** (`safety-dump`) | „Egy alkalmazás frissítése éppen folyamatban van…" — **our** refusal |
|
|
|
|
**The sentence changed.** Guest 9202 has no host agent, so `TriggerUpdate` refuses either way — which
|
|
makes it the perfect negative control, because the new check sits *before* the agent check. Two
|
|
sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and
|
|
**restored to `false`** from a copy taken first; verified by re-reading the file.
|
|
|
|
**The reverse direction — an app update refused while the controller swaps — is NOT staged live.** It
|
|
is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this
|
|
guest does not have. Stated as a gap.
|
|
|
|
---
|
|
|
|
## 6. Part 3 — the unattended night (R-611, CLOSED)
|
|
|
|
**Success:** `uptime-kuma` 2.5.0 → 2.5.1 applied with nobody pressing anything.
|
|
|
|
**No-retry:** after the revert left all four apps *ahead* of the catalog, the caller pressed each
|
|
**exactly once**, was refused `downgrade` (terminal), and pressed nothing across two further passes.
|
|
`never_again=['glance','uptime-kuma','vikunja','wishlist']`, `outcomes={}`. That is R-524 and R-609
|
|
working together, unattended.
|
|
|
|
**The unattended HOLD was never produced, and the reason matters:** the only failing edge available
|
|
(`vikunja → alpine:3.20`) was **correctly refused by the within-a-major rule before it was ever
|
|
attempted**. The rule that makes automatic updates safe is the same rule that refuses the obvious way
|
|
to break one. Measuring it needs an image that passes the version test and still fails health.
|
|
**§3b Q4 therefore still rests on the ATTENDED hold from slice 4.**
|
|
|
|
---
|
|
|
|
## 7. Rows, and the catalog
|
|
|
|
**Closed:** R-608, R-609, R-610, R-611. **Opened:** R-612 (P1 — wishlist unusable on a fresh install
|
|
and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 —
|
|
stale update phase survives a redeploy). **Corrected:** R-520's closing pointer now names R-610.
|
|
Register 302 → **307**.
|
|
|
|
**Catalog:** DRILL `573e41f5`, DRILL 2 `ae08a037`, **REVERT `f5f6a152`**. Every `image:` line in
|
|
`templates/` is byte-identical to the pre-drill `ff9717d3` — `git diff` over those paths is **0
|
|
lines**. `catalog_since` reads 2026-09-21 on the four rather than the older dates, because the gate
|
|
requires an image move to carry the day's date in either direction and a revert is a move.
|
|
|
|
**Teardown, three layers.** *Machine:* guest 9202 left running on v0.261.0 with the four throwaway
|
|
apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for
|
|
the remaining Q4 work and removing them would cost the next session the seeding. *Host:* demo-hp
|
|
untouched apart from 9202; **guest 9201 never touched.** *Hub:* nothing provisioned, nothing
|
|
enrolled; 9202 reports to no hub by design. `/tmp/.ctlpw` shredded.
|