Update arc: the undo and ladder spiked, the build plan for the 2026-09-23 rulings
gates / gates (push) Successful in 26s

- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
  (PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
  back with data written before AND after the backup. The product's loader
  cannot do it: over a migrated PG database it fails on the new tables'
  foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
  rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
  Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.

No product code. Live catalog untouched; 9202 back on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 08:54:41 +02:00
parent 805ad1e962
commit 4c92beab8f
67 changed files with 5977 additions and 162 deletions
@@ -777,6 +777,27 @@ the harness has proven it, is slice 6's.
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
### 6.1a The undo (decision 15) — SPIKED BY HAND 2026-09-23, not built
Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made
to fail a deliberately wrong probe, each held by today's product, each then undone by hand.
| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (SQLite in a volume) |
|---|---|---|---|
| old version on the migrated data, nothing loaded | **refuses** (migration ledger) | **refuses** (alembic revision) | starts and serves |
| the safety dump | DB only, holds the post-backup write | DB only, holds it | **none — no-op** |
| undo, load + start → healthy | **≈ 16 s** | **≈ 38 s** | ≈ 1 s |
| data written before AND after the backup read back | yes / yes | yes / yes | yes / yes |
**The undo works — and not with the loader the product has.** `ImportDump` over a migrated
PostgreSQL database FAILS (the new version's foreign keys block the dump's own drops); over MariaDB it
succeeds and leaves the new version's tables behind. The load that worked empties the schema and loads
the copy in one transaction. **And a truncated PostgreSQL copy loads with exit 0 into an EMPTY
database** — the copy's completion marker must be checked first, which `ValidateDump` does not do. A
product path that loads a safety dump back exists (`rollbackSafetyDump`), but only the off-site
restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices
them.
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
@@ -816,10 +837,19 @@ evidence; Slice 6 puts the same shape in the catalog.
"seed_read_before": true, "seed_read_after": true, "healthy_after": true,
"migration_observed": "verbatim log line, or null",
"abort": "starts-and-serves | refuses | starts-data-gone | not-attempted",
"memory": {"soak_s": 600, "containers": {"<name>": {"limit": 0, "peak": 0, "peak_pct": 0.0,
"oom_kills": 0, "restarts": 0}}, "first_kill": null},
"marks": ["memory_tight"],
"abort_detail": "the refusal quoted verbatim, or null",
"duration_s": 0, "measured_at": "RFC3339", "evidence": "relative path"}
```
**Harness version 2 (2026-09-23, R-635) adds `memory` and `marks`.** After a successful readback the
harness runs the new version for `--soak` seconds (default 600) under light load and reads the
kernel's own `oom_kill` counter host-side. A kill or a restart turns `proven` into `failed`; a peak
above 80 % of the compose limit adds `memory_tight`. The two fields are the test record's memory half
(decision 13).
**`inconclusive` is a first-class verdict and must never be collapsed into `failed`.** "We could not
measure it" and "it does not work" are different facts, and only one of them is about the app.
**`migration_observed` is a quoted line, never an inference from timing** — the value of both the
@@ -900,7 +930,16 @@ spend) or the leg gets almost no time. That is a build choice inside decision 11
**How far — one step at a time (decision 14).** A box two steps behind applies step A→B, then B→C,
each the full guarded update, each with its own health check and undo. A failed step stops the
ladder for that app. The ladder's format is §6.4 / `audits/update-rulings-2026-09-23/`.
ladder for that app. **Measured 2026-09-23:** today one press jumps A → C and B never runs, and the
box cannot see B at all — its catalog clone is `--depth 1` (`sync.go:283`, `:300`; one commit
visible on both demo guests).
**The ladder's format — recommended, not ruled** (`audits/update-rulings-2026-09-23/README.md` Part 2):
an `update_ladder:` list in `.felhom.yml`, one entry per step — `from`/`to` refs per service, the
digest per ref, the test record, the marks — and, for every step but the last, the step's OWN
complete definition in `templates/<app>/steps/<to>.yml`. **Not the git history:** romm's image-moving
commit `15f9ebf` is the definition that OOM-looped on demo-hp; the step that works is its images with
the later `f4eb94f` template, and no commit holds that pair.
**When it fails — undo, then hold only if the undo fails (decision 15).** The household is told on
the app page and by **one** mail, in the box's language (R-606 is a precondition — an automatic
@@ -949,7 +988,47 @@ Rank stays P3-LOW at two enrolled boxes. It rises with the fleet, and §2 of the
that looks like today: the only way to answer *"is the fleet current?"* was to read both boxes' files
by hand.
### 6.4 The update night — a drill brief outline, costed from R-462's real numbers
### 6.4 The build order for the 2026-09-23 rulings (PLAN — each part returns to the operator for go/no-go)
Costed in **CC-evenings** (one evening ≈ one unattended session: build, red-proofs, live proof on 9202,
release). Written from the two spikes and the memory watch of 2026-09-23
(`audits/update-rulings-2026-09-23/`), not from source reading alone. **Risk to customer data** is
what the part can do to a household's data if it is wrong, not how likely that is.
| # | part | rulings / rows | cost | depends on | risk to customer data |
|---|---|---|---|---|---|
| **1** | **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none |
| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
| **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
| **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
| **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
| **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **8** | **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |
**Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total
for 1–10: ≈ 22 evenings.** The automatic caller (7) is deliberately late: it is the only part that
acts with nobody watching, and every part before it is what makes that safe. Parts 8 and 9 are small
and independent and can fill any short evening.
**The one open point the build cannot settle alone — part 7, inside decision 11.** The ruled chain is
*off-site copy → updates → full-system backup*. Today the off-site leg starts at **W+105m** and the
full-system backup's gate opens at **W+2h** (`quiesce.go` `gateOpenOffsetMin = 120`, span to W+6h). So
the update leg has **at most 15 minutes**, and none on a night the off-site copy runs long — while one
step takes ~1 min when it works and ~2–6 min when it fails and is undone.
| option | cost |
|---|---|
| **the full-system backup waits for the update leg, inside its own window; the leg stops starting new steps at W+5h** | the full-system backup starts later on update nights, still inside its four-hour window, with an hour kept; one more interlock between two nightly jobs |
| the leg stops at W+2h as the chain stands | ≤ 15 min a night — about ten steps on a good night, none on a slow one; a box far behind takes weeks to climb |
**Recommendation: the first.** It keeps the ruling's order and its promise that the full-system backup
is never skipped for an update; only the start time inside its existing window moves.
### 6.4.1 (record) The update night — the drill brief that preceded the rulings, costed and re-costed
**The ruling is decision 6: all 53 apps, through the nightly rotation.** This is an ORDER inside that
ruling, not a scope change. The database apps go first because they are the ones where a wrong answer