night 2026-10-04: the guest undo runbook (proven 48 s), R-866, A3 in progress
gates / gates (push) Successful in 30s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 22:37:25 +02:00
parent a1a32bbd74
commit 76ff0654ea
4 changed files with 73 additions and 1 deletions
@@ -0,0 +1,6 @@
BLOCK 20:26:09
host: 000blocked
guest: 000blocked
PASS 20:26:10
selftest=os-update: desired state: hub: transport error: Get "https://hub.felhom.eu/api/v1/hosts/tester-1-d70be4/desired-state": dial tcp 192.168.0.192:443: connect: invalid argument
PASS-END 20:26:11
+1
View File
@@ -410,6 +410,7 @@ stopping line that lies.
| **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC |
| **R-586** | Process & tooling | P3 | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** FOUND 2026-09-18 starting slice 4 (R-559): running `scripts/iso/test/bootstrap-modes.sh` unchanged at `183727db9c44` reported `FAIL: R-496: banner painted to the console seam` and `FAIL: R-496: banner names the Tulajdonosi jelmondat`. **Cause:** the script paints each banner with `> "$CONSOLE_DEV"`. On a real console that is a character device and truncation is a no-op, so every paint appears; pointed at a plain file, as the harness did, each paint TRUNCATES. Commit `c033b3b` (ISO 1.28.0, R-535, 2026-09-16) added `print_bound_banner`, which paints immediately after the pairing banner in the same invocation — from that commit the pairing banner was wiped before the check read it. `c033b3b` did not touch the harness. **Why it survived: the harness is in NO gate and NO CI run** — not in `repo_gates.py`, not in `.gitea/`; it runs only when a person runs it, and between 09-16 and 09-18 nobody did. FIXED in the same session: the harness points `FELHOM_CONSOLE_DEV` at a FIFO with a background reader, which restores device semantics (opening a FIFO with `>` truncates nothing) and lets a test see EVERY paint — which slice 4's golden checks then needed anyway. Production code untouched. **What is still open: the harness remains outside every gate.** It needs a container, so it cannot join `repo_gates.py --fast`, which is what both the pre-push hook and CI run — meaning a non-fast entry would still never execute. **Fix shape:** either give CI a container-capable job that runs it, or make the ISO release gate's G16 the place it is required (done for G16 this session — so it now runs at least once per ISO release, which is better than never but later than a push). | **READY - rank P2-MED; owner: CC** **Re-ranked 2026-10-03: P2->P3: test harness gap; now required at each ISO release gate.** | — | — | CC |
| **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC |
| **R-866** | Process & tooling | P4 | **The debug OS pass (`felhom-agent --selftest=os-update`) cannot run while the hub is unreachable: it fetches the plan from the hub and stops, where the daemon's own leg uses its saved plan.** 2026-10-04 night drill A3 (Tester 1 box, hub blocked): `selftest=os-update: desired state: hub: transport error … connect: invalid argument`. So "an OS pass with the hub away" can be exercised only by the daemon's leg, which the 20-hour gap holds back on a box that just ran. Fix direction: the selftest falls back to the last saved block (and says so), or a documented way to trigger the daemon's leg. `audits/night-2026-10-04/accidents/a3-*.txt` | **READY — owner: CC** | — | — | CC |
| **R-93** | Process & tooling | P4 | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture".** | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
| **R-129** | Process & tooling | P4 | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC |
| **R-161** | Process & tooling | P4 | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **OPEN** — **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator |
@@ -0,0 +1,65 @@
# Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A)
> **When:** the hub mailed `os_update_health_failed` for the **guest** layer, or the box's apps broke right after a
> guest OS pass, and installing the previous versions one by one is not enough.
> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host.
> **Owner design:** `architecture/11-os-updates.md` §5.6 (guest row), `09` §3 decision 81 (option A, by hand; nothing
> new built). **Proved** on demo-felhom 2026-10-04 night: **48 s down**, every app healthy, the package versions back
> to the "before" ones — evidence `audits/night-2026-10-04/undo/`.
**What it costs the household:** every app is down for under a minute (measured 48 s), and **everything written to
the guest after the backup is lost** — app data in volumes and databases since that backup. Data on the household's
drive (`/mnt/felhom-drives`, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it
soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the
bad pass. After that, prefer putting single packages back (`os-updates-host-undo.md`, the same `snapshot.debian.org`
method works inside the guest with `pct exec`).
There is no product route: the agent has no executor for a restore over a live guest (`restore_overwrite` is
classified, not built) and `RUNBOOK-manual-guest-restore.md` is for a REPLACED host. This page is the route.
## 1. Pick the archive
The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is
normally the newest one on the box's primary backup tier:
```bash
pvesm list <primary-tier-storage> | grep -- "-9201-" | tail -3 # e.g. felhom-backup, local
journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)" # the order, with times
```
Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box):
```bash
tar -xOf /var/lib/vz/dump/<archive>.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)'
```
## 2. Restore in place (this is the downtime)
```bash
date -u +%T # down from
pct shutdown 9201 --timeout 60
pct restore 9201 <storage>:backup/<archive>.tar.zst --force --storage local-lvm
pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)' # mp8 + mp9 binds, onboot 1, the hook, no lock
pct start 9201
```
- `--force` replaces the running guest's volumes; the new volumes get NEW names (`vm-9201-disk-2/3` instead of
`disk-0/1`, measured) — nothing on the box pinned the old names.
- While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured:
`guest is stopped but LOCKED — leaving it to the stale-lock path`).
- `--storage` must be the storage the guest's volumes were on (`local-lvm` on the appliances; check `pct config` first).
## 3. Read back (all must hold)
```bash
pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}" # every app "(healthy)"; measured 15 s after start
pct exec 9201 -- dpkg-query -W <the packages the bad pass changed> # the BEFORE versions
```
The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing
else (measured: no mail).
## 4. Before the next night
The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System
page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.
@@ -78,4 +78,4 @@ installs the new version again. Write the hold into the register row that tracks
- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built).
- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's
own repository keeps old versions; that undo is not written yet.
- The customer guest: its undo is last night's whole-guest backup (decision 81).
- The customer guest: its undo is last night's whole-guest backup (decision 81) — `os-updates-guest-undo.md` (proven 2026-10-04 night, 48 s down).