From 87cb9233903c9533ad0165064c80774c860e08d0 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 17:17:35 +0200 Subject: [PATCH] ruling 3 recorded: the restart brake is blind to a slow crash loop (R-539) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit It sits beside the 3-in-15-minutes budget in the host-agent design, because that sentence and its measured blind spot belong together: four kills 20 minutes apart were all restarted, none accumulated, and the only trace was an info event that mails nobody. The budget itself is unchanged. Also recorded: why Part C.2's re-issue button is correctly hidden for a customer with no host, the venue's storage reconciliation (nvme-scratch IS /mnt/hdd_1), the built ISO landing on demo-hp byte-identical, and my own scp/-P mistake that cost a golden bake — including why its token-leak check reported a false hit on an empty needle. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- documentation/architecture/03-host-agent.md | 11 +++++++++++ .../phaseC-reissue.txt | 11 +++++++++++ .../phaseE-golden-bake.txt | 17 +++++++++++++++++ .../phaseE-venue.txt | 13 +++++++++++++ 4 files changed, 52 insertions(+) create mode 100644 documentation/audits/evidence-backup-promise-2026-09-16/phaseE-venue.txt diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 87be928a..80df4fd2 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -110,6 +110,17 @@ by verb**: neither lose nor invent one. New goldens run the controller with `--restart always`, which covers only a Docker daemon restart. +**[RULING 2026-09-16, operator] The brake catches a FAST crash loop and is blind to a SLOW one — a +second, slower counter is owed (R-539).** MEASURED the same day on the drill box: the controller was +killed four times, each kill 20 minutes after the last. Every one was restarted (dashboard back in +61 s / 41 s / 61 s), and **none of them accumulated**, because the budget window is 15 minutes — so +the 30-minute pause was never armed and the only trace was an `info` `controller_restarted_by_agent` +event, which mails nobody. A box whose controller dies every 20 minutes is therefore restarted +for ever, quietly. **The 3-in-15-minutes budget is unchanged**; what is owed is a second counter over +a long window (N restarts in 24 h) raising a `warning`. Not built in the 2026-09-16 task, which was +told to measure the budget rather than change it. Evidence: +`audits/evidence-drill-0243-2026-09-16/phase2-f9dprime.txt`. + Signed payloads carry a **nonce + expiry** (anti-replay: a captured "restore" job cannot be re-injected later) and a target binding (host + guest id) so a signature can't be retargeted. Notification-on-destructive-op is an **audit signal, never the guard** — a compromised hub diff --git a/documentation/audits/evidence-backup-promise-2026-09-16/phaseC-reissue.txt b/documentation/audits/evidence-backup-promise-2026-09-16/phaseC-reissue.txt index 5254981e..6fda8a46 100644 --- a/documentation/audits/evidence-backup-promise-2026-09-16/phaseC-reissue.txt +++ b/documentation/audits/evidence-backup-promise-2026-09-16/phaseC-reissue.txt @@ -1 +1,12 @@ ## 2026-09-16T15:12:32Z Part C.2 — the PBS re-issue on tester-1's REAL state, now that the grant exists +## WHY THE RE-ISSUE BUTTON IS NOT ON THE PAGE — read from source, not guessed: +## `config_form_body.html:182` renders the „Re-issue PBS credentials" button inside +## `{{if .PBSDR.Provisioned}}`. tester-1 currently has NO host (the drill box was torn down on +## 2026-09-16), so there is no provisioned DR descriptor and the control is correctly hidden. +## The hub said the same thing in its own log at 17:05:09: +## „pbsdr: DR tier ON for tester-1, no host enrolled yet — the descriptor applies once the +## cascade is ready (host -> WG peer -> apply)". +## CONSEQUENCE: Part C.2 cannot be walked against tester-1 today, and that is the product behaving +## correctly rather than a blocked step. The grant is proven end-to-end by the FRESH BOX in Part E: +## its WG registration must provision `felhom-pbs` with no hand — which is exactly the path that +## failed on 2026-09-16 with „missing Datastore.Modify". R-534 and R-511 stay open until that run. diff --git a/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-golden-bake.txt b/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-golden-bake.txt index bf942e7f..c7a20fcb 100644 --- a/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-golden-bake.txt +++ b/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-golden-bake.txt @@ -12,3 +12,20 @@ token-leak control — planted copy must be 1: 1; committed log must be 0: 0 qemu exited; disk reverted to virgin ## bake done 2026-09-16T15:13:40Z +## 2026-09-16T15:14:36Z golden 0.244.0 — bake in the drill VM (§4.0/§4.1 of RUNBOOK-manual-build) + reverted to virgin (this also proves no qemu holds the qcow2) + cold boot started +## MY OWN MISTAKE, recorded because it cost a bake (first attempt, 15:13Z): +## 1) `scp` takes `-P `; I passed the ssh flag set, whose `-p 2222` means preserve-times and +## made scp treat „2222" as a source filename. Neither the script nor the token landed in the VM, +## and the bake died in 16 s with „/root/build-golden.sh: No such file or directory". +## 2) The token-leak check then reported „1" and it was NOT a leak: with the token file ABSENT, +## `grep -c -F "$(cat ...)"` becomes `grep -c -F ""`, which matches every line. An instrument +## that reports a hit when its needle is empty is not a measurement — the same class this repo +## keeps re-learning. The re-run guards both: it aborts unless BOTH files land non-empty. + ssh up: pve-manager/9.2.2/b9984c6d90a4bd80 (running kernel: 7.0.2-6-pve) + template: debian-13-standard_13.6-1_amd64.tar.zst + template downloaded + script + token landed in the VM (both non-empty) + bake launched as a transient unit at 2026-09-16T15:15:37Z + token leak check on the unit (must be 0, and the token is non-empty so the grep cannot match everything): 0 diff --git a/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-venue.txt b/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-venue.txt new file mode 100644 index 00000000..e62a23a0 --- /dev/null +++ b/documentation/audits/evidence-backup-promise-2026-09-16/phaseE-venue.txt @@ -0,0 +1,13 @@ +## 2026-09-16 the venue, reconciled rather than assumed +## The task's fence says „VM disks on /mnt/hdd_1 at its root, never local-lvm"; demo-hp's storage list +## shows no storage NAMED hdd_1. They are the same place: +## dir: nvme-scratch path /mnt/hdd_1 content images,rootdir is_mountpoint yes +## /dev/nvme0n1 938G 14G used 877G free mounted on /mnt/hdd_1 +## So the fresh box's disk goes on storage `nvme-scratch`, which IS /mnt/hdd_1 at its root. Recorded +## because the names differ and a session that trusts the name alone would either refuse to proceed +## or reach for local-lvm. +## The built image is on the venue, byte-identical: +## local sha256 a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635 +## remote sha256 a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635 +## /var/lib/vz/template/iso/felhom-installer-1.28.0-pve9.2-1.iso (1 705 322 496 B) +## The published 1.27.1 image is still there and still the published one — nothing was overwritten.