ruling 3 recorded: the restart brake is blind to a slow crash loop (R-539)
gates / gates (push) Successful in 24s

It sits beside the 3-in-15-minutes budget in the host-agent design, because that
sentence and its measured blind spot belong together: four kills 20 minutes apart
were all restarted, none accumulated, and the only trace was an info event that
mails nobody. The budget itself is unchanged.

Also recorded: why Part C.2's re-issue button is correctly hidden for a customer
with no host, the venue's storage reconciliation (nvme-scratch IS /mnt/hdd_1), the
built ISO landing on demo-hp byte-identical, and my own scp/-P mistake that cost a
golden bake — including why its token-leak check reported a false hit on an empty
needle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 17:17:35 +02:00
parent f2250ca31f
commit 87cb923390
4 changed files with 52 additions and 0 deletions
@@ -110,6 +110,17 @@ by verb**:
neither lose nor invent one. New goldens run the controller with `--restart always`, which covers
only a Docker daemon restart.
**[RULING 2026-09-16, operator] The brake catches a FAST crash loop and is blind to a SLOW one — a
second, slower counter is owed (R-539).** MEASURED the same day on the drill box: the controller was
killed four times, each kill 20 minutes after the last. Every one was restarted (dashboard back in
61 s / 41 s / 61 s), and **none of them accumulated**, because the budget window is 15 minutes — so
the 30-minute pause was never armed and the only trace was an `info` `controller_restarted_by_agent`
event, which mails nobody. A box whose controller dies every 20 minutes is therefore restarted
for ever, quietly. **The 3-in-15-minutes budget is unchanged**; what is owed is a second counter over
a long window (N restarts in 24 h) raising a `warning`. Not built in the 2026-09-16 task, which was
told to measure the budget rather than change it. Evidence:
`audits/evidence-drill-0243-2026-09-16/phase2-f9dprime.txt`.
Signed payloads carry a **nonce + expiry** (anti-replay: a captured "restore" job cannot be
re-injected later) and a target binding (host + guest id) so a signature can't be retargeted.
Notification-on-destructive-op is an **audit signal, never the guard** — a compromised hub
@@ -1 +1,12 @@
## 2026-09-16T15:12:32Z Part C.2 — the PBS re-issue on tester-1's REAL state, now that the grant exists
## WHY THE RE-ISSUE BUTTON IS NOT ON THE PAGE — read from source, not guessed:
## `config_form_body.html:182` renders the „Re-issue PBS credentials" button inside
## `{{if .PBSDR.Provisioned}}`. tester-1 currently has NO host (the drill box was torn down on
## 2026-09-16), so there is no provisioned DR descriptor and the control is correctly hidden.
## The hub said the same thing in its own log at 17:05:09:
## „pbsdr: DR tier ON for tester-1, no host enrolled yet — the descriptor applies once the
## cascade is ready (host -> WG peer -> apply)".
## CONSEQUENCE: Part C.2 cannot be walked against tester-1 today, and that is the product behaving
## correctly rather than a blocked step. The grant is proven end-to-end by the FRESH BOX in Part E:
## its WG registration must provision `felhom-pbs` with no hand — which is exactly the path that
## failed on 2026-09-16 with „missing Datastore.Modify". R-534 and R-511 stay open until that run.
@@ -12,3 +12,20 @@
token-leak control — planted copy must be 1: 1; committed log must be 0: 0
qemu exited; disk reverted to virgin
## bake done 2026-09-16T15:13:40Z
## 2026-09-16T15:14:36Z golden 0.244.0 — bake in the drill VM (§4.0/§4.1 of RUNBOOK-manual-build)
reverted to virgin (this also proves no qemu holds the qcow2)
cold boot started
## MY OWN MISTAKE, recorded because it cost a bake (first attempt, 15:13Z):
## 1) `scp` takes `-P <port>`; I passed the ssh flag set, whose `-p 2222` means preserve-times and
## made scp treat „2222" as a source filename. Neither the script nor the token landed in the VM,
## and the bake died in 16 s with „/root/build-golden.sh: No such file or directory".
## 2) The token-leak check then reported „1" and it was NOT a leak: with the token file ABSENT,
## `grep -c -F "$(cat ...)"` becomes `grep -c -F ""`, which matches every line. An instrument
## that reports a hit when its needle is empty is not a measurement — the same class this repo
## keeps re-learning. The re-run guards both: it aborts unless BOTH files land non-empty.
ssh up: pve-manager/9.2.2/b9984c6d90a4bd80 (running kernel: 7.0.2-6-pve)
template: debian-13-standard_13.6-1_amd64.tar.zst
template downloaded
script + token landed in the VM (both non-empty)
bake launched as a transient unit at 2026-09-16T15:15:37Z
token leak check on the unit (must be 0, and the token is non-empty so the grep cannot match everything): 0
@@ -0,0 +1,13 @@
## 2026-09-16 the venue, reconciled rather than assumed
## The task's fence says „VM disks on /mnt/hdd_1 at its root, never local-lvm"; demo-hp's storage list
## shows no storage NAMED hdd_1. They are the same place:
## dir: nvme-scratch path /mnt/hdd_1 content images,rootdir is_mountpoint yes
## /dev/nvme0n1 938G 14G used 877G free mounted on /mnt/hdd_1
## So the fresh box's disk goes on storage `nvme-scratch`, which IS /mnt/hdd_1 at its root. Recorded
## because the names differ and a session that trusts the name alone would either refuse to proceed
## or reach for local-lvm.
## The built image is on the venue, byte-identical:
## local sha256 a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635
## remote sha256 a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635
## /var/lib/vz/template/iso/felhom-installer-1.28.0-pve9.2-1.iso (1 705 322 496 B)
## The published 1.27.1 image is still there and still the published one — nothing was overwritten.