night 2026-10-04: the guest undo runbook (proven 48 s), R-866, A3 in progress
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,6 @@
|
||||
BLOCK 20:26:09
|
||||
host: 000blocked
|
||||
guest: 000blocked
|
||||
PASS 20:26:10
|
||||
selftest=os-update: desired state: hub: transport error: Get "https://hub.felhom.eu/api/v1/hosts/tester-1-d70be4/desired-state": dial tcp 192.168.0.192:443: connect: invalid argument
|
||||
PASS-END 20:26:11
|
||||
@@ -410,6 +410,7 @@ stopping line that lies.
|
||||
| **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC |
|
||||
| **R-586** | Process & tooling | P3 | **[P2-MED] The ISO bootstrap harness captured the console to a FILE, so each banner erased the one before it — two checks were RED for two days and nobody saw.** FOUND 2026-09-18 starting slice 4 (R-559): running `scripts/iso/test/bootstrap-modes.sh` unchanged at `183727db9c44` reported `FAIL: R-496: banner painted to the console seam` and `FAIL: R-496: banner names the Tulajdonosi jelmondat`. **Cause:** the script paints each banner with `> "$CONSOLE_DEV"`. On a real console that is a character device and truncation is a no-op, so every paint appears; pointed at a plain file, as the harness did, each paint TRUNCATES. Commit `c033b3b` (ISO 1.28.0, R-535, 2026-09-16) added `print_bound_banner`, which paints immediately after the pairing banner in the same invocation — from that commit the pairing banner was wiped before the check read it. `c033b3b` did not touch the harness. **Why it survived: the harness is in NO gate and NO CI run** — not in `repo_gates.py`, not in `.gitea/`; it runs only when a person runs it, and between 09-16 and 09-18 nobody did. FIXED in the same session: the harness points `FELHOM_CONSOLE_DEV` at a FIFO with a background reader, which restores device semantics (opening a FIFO with `>` truncates nothing) and lets a test see EVERY paint — which slice 4's golden checks then needed anyway. Production code untouched. **What is still open: the harness remains outside every gate.** It needs a container, so it cannot join `repo_gates.py --fast`, which is what both the pre-push hook and CI run — meaning a non-fast entry would still never execute. **Fix shape:** either give CI a container-capable job that runs it, or make the ISO release gate's G16 the place it is required (done for G16 this session — so it now runs at least once per ISO release, which is better than never but later than a push). | **READY - rank P2-MED; owner: CC** **Re-ranked 2026-10-03: P2->P3: test harness gap; now required at each ISO release gate.** | — | — | CC |
|
||||
| **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC |
|
||||
| **R-866** | Process & tooling | P4 | **The debug OS pass (`felhom-agent --selftest=os-update`) cannot run while the hub is unreachable: it fetches the plan from the hub and stops, where the daemon's own leg uses its saved plan.** 2026-10-04 night drill A3 (Tester 1 box, hub blocked): `selftest=os-update: desired state: hub: transport error … connect: invalid argument`. So "an OS pass with the hub away" can be exercised only by the daemon's leg, which the 20-hour gap holds back on a box that just ran. Fix direction: the selftest falls back to the last saved block (and says so), or a documented way to trigger the daemon's leg. `audits/night-2026-10-04/accidents/a3-*.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-93** | Process & tooling | P4 | `drill-r50` is both a blocked customer and the only drift fixture **FACT 2026-09-13 (R-461): the fixture is GONE — `qm list` is empty on both demo boxes, so neither option is available and the row's premise no longer holds; the operator decides whether that closes it or reopens it as "build a drift fixture".** | READY (XS) | — | Retire it for a synthetic fixture, or unblock + silence per-customer | CC |
|
||||
| **R-129** | Process & tooling | P4 | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC |
|
||||
| **R-161** | Process & tooling | P4 | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **OPEN** — **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator |
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
# Undo a failed OS update on the customer guest — restore the whole-guest backup in place (decision 81, R-842 option A)
|
||||
|
||||
> **When:** the hub mailed `os_update_health_failed` for the **guest** layer, or the box's apps broke right after a
|
||||
> guest OS pass, and installing the previous versions one by one is not enough.
|
||||
> **Who:** the operator, or CC with the operator's word, as root on the box's Proxmox host.
|
||||
> **Owner design:** `architecture/11-os-updates.md` §5.6 (guest row), `09` §3 decision 81 (option A, by hand; nothing
|
||||
> new built). **Proved** on demo-felhom 2026-10-04 night: **48 s down**, every app healthy, the package versions back
|
||||
> to the "before" ones — evidence `audits/night-2026-10-04/undo/`.
|
||||
|
||||
**What it costs the household:** every app is down for under a minute (measured 48 s), and **everything written to
|
||||
the guest after the backup is lost** — app data in volumes and databases since that backup. Data on the household's
|
||||
drive (`/mnt/felhom-drives`, a bind mount) is NOT in the whole-guest backup and is NOT touched by the restore. So use it
|
||||
soon after the bad pass — the OS leg runs right after the night backup, so last night's backup is minutes older than the
|
||||
bad pass. After that, prefer putting single packages back (`os-updates-host-undo.md`, the same `snapshot.debian.org`
|
||||
method works inside the guest with `pct exec`).
|
||||
|
||||
There is no product route: the agent has no executor for a restore over a live guest (`restore_overwrite` is
|
||||
classified, not built) and `RUNBOOK-manual-guest-restore.md` is for a REPLACED host. This page is the route.
|
||||
|
||||
## 1. Pick the archive
|
||||
|
||||
The newest whole-guest backup taken BEFORE the bad pass. The pass runs right after the primary backup, so it is
|
||||
normally the newest one on the box's primary backup tier:
|
||||
|
||||
```bash
|
||||
pvesm list <primary-tier-storage> | grep -- "-9201-" | tail -3 # e.g. felhom-backup, local
|
||||
journalctl -u felhom-agent --since -1d | grep -E "backup: completed|osupdate: (START|DONE)" # the order, with times
|
||||
```
|
||||
|
||||
Check the archive's config BEFORE restoring (the binds must be this host's — they are, on the same box):
|
||||
|
||||
```bash
|
||||
tar -xOf /var/lib/vz/dump/<archive>.tar.zst ./etc/vzdump/pct.conf | grep -E '^(rootfs|mp[0-9]|onboot|hookscript)'
|
||||
```
|
||||
|
||||
## 2. Restore in place (this is the downtime)
|
||||
|
||||
```bash
|
||||
date -u +%T # down from
|
||||
pct shutdown 9201 --timeout 60
|
||||
pct restore 9201 <storage>:backup/<archive>.tar.zst --force --storage local-lvm
|
||||
pct config 9201 | grep -E '^(rootfs|mp[0-9]|onboot|hookscript|lock)' # mp8 + mp9 binds, onboot 1, the hook, no lock
|
||||
pct start 9201
|
||||
```
|
||||
|
||||
- `--force` replaces the running guest's volumes; the new volumes get NEW names (`vm-9201-disk-2/3` instead of
|
||||
`disk-0/1`, measured) — nothing on the box pinned the old names.
|
||||
- While the guest is locked by the restore, the agent's guest-power watchdog leaves it alone (measured:
|
||||
`guest is stopped but LOCKED — leaving it to the stale-lock path`).
|
||||
- `--storage` must be the storage the guest's volumes were on (`local-lvm` on the appliances; check `pct config` first).
|
||||
|
||||
## 3. Read back (all must hold)
|
||||
|
||||
```bash
|
||||
pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}" # every app "(healthy)"; measured 15 s after start
|
||||
pct exec 9201 -- dpkg-query -W <the packages the bad pass changed> # the BEFORE versions
|
||||
```
|
||||
|
||||
The hub: the box's next host report (≤ 15 min) arrives; the household's timeline shows "Controller elindult" — nothing
|
||||
else (measured: no mail).
|
||||
|
||||
## 4. Before the next night
|
||||
|
||||
The next night's OS leg would install the same bad versions again. Switch the box's OS updates OFF on the hub's System
|
||||
page (or move it out of ring 0) until a fixed release exists, and say so in the box's event log.
|
||||
@@ -78,4 +78,4 @@ installs the new version again. Write the hold into the register row that tracks
|
||||
- A kernel: the fast lane never installs one (R14). The kernel lane is `11` §5.6 (not built).
|
||||
- A Proxmox-origin package (pve-*, qemu-server, …): never installed by the fast lane (R12/origin rule). Proxmox's
|
||||
own repository keeps old versions; that undo is not written yet.
|
||||
- The customer guest: its undo is last night's whole-guest backup (decision 81).
|
||||
- The customer guest: its undo is last night's whole-guest backup (decision 81) — `os-updates-guest-undo.md` (proven 2026-10-04 night, 48 s down).
|
||||
|
||||
Reference in New Issue
Block a user