From f96f93081d7864ceb1ac4bbac12d5488b2ae5fea Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 16 Sep 2026 13:18:50 +0200 Subject: [PATCH] drill 0.243.0: Phase 3 morning-after, off-site not walked (R-534), R-528 and R-511 updated The off-site restore onto 9202 cannot be walked: this box never had an off-site tier, because the re-issue fails on the endpoint token's missing Datastore.Modify grant. Read-only listing of ep0 shows ns/tester-1/ct empty both before and after the drill, with ns/demo-hp/ct as the positive control. Nothing on ep0 was written, removed or pruned. R-528 re-measured on a second, different box: all three OOM signals silent again. R-511 records that its shipped fix is sound and inert until the grant is given. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../audits/DRILL-prove-fixes-0243-2026-09-16.md | 16 ++++++++++++++++ .../phase2-f10.txt | 3 +++ .../phase3-morning-after.txt | 16 ++++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 4 ++-- 4 files changed, 37 insertions(+), 2 deletions(-) diff --git a/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md b/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md index 35b2beac..20ff7045 100644 --- a/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md +++ b/documentation/audits/DRILL-prove-fixes-0243-2026-09-16.md @@ -88,3 +88,19 @@ tier 2 / tier 3, both unset. The design is not the defect. The defects are that „DB + Konfig + Adatok" and prints the drive size beside it (**R-537**), and that a restore reports success while leaving the app listing files it cannot open — and wipes the app's own trash, which still held every byte (**R-538**). + +## Phase 3 — the morning after + +**Every app healthy, labels true.** Front doors at 11:15Z: PrivateBin 200, Vaultwarden 200, Paperless +302 (its login redirect), Nextcloud `status.php` 200, file manager 200 at `files.enkicsifelhom.hu`. +All seven stacks read `running`. Controller **0.243.0** on the page, agent **0.131.0** on the hub, the +customer row reads Version 0.243.0 / Floor v0.242.0 — a floor is a minimum, so that is correct, and the +box runs the golden this drill baked. + +**Off-site restore onto 9202 — NOT WALKED, and not for time.** There is nothing to restore from. The +off-site tier was never provisioned on this box: the re-issue fails on the endpoint token's missing +`Datastore.Modify` grant (**R-534**), so no descriptor and no upload path were ever created. Read-only +listing of ep0 confirms it: `ns/tester-1/ct` is **empty** both before (09:12Z) and after (11:14Z) the +drill, while the neighbouring `ns/demo-hp/ct` lists group 9201 — the positive control proving the +listing method shows groups when they exist. **Nothing on ep0 was written, removed or pruned by this +session.** diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt index 46eb0ef8..64aaa5fb 100644 --- a/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase2-f10.txt @@ -83,3 +83,6 @@ ## therefore BY DESIGN. What is NOT by design is the page calling tier 1 "DB + Konfig + Adatok" and ## printing "Adatlemez 65.1 MB" next to it, and a restore reporting plain success while leaving the ## app inconsistent. Those two are the rows. + trash re-check at 2026-09-16T11:17:05Z, after the app restarted twice since the restore: + PROPFIND trash http=207 + entries listed: 1 (1 = only the trash root itself, so the trash is EMPTY to the customer) diff --git a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt index 923aab46..f1037eba 100644 --- a/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt +++ b/documentation/audits/evidence-drill-0243-2026-09-16/phase3-morning-after.txt @@ -17,3 +17,19 @@ traefik running deployed=False subdomain= vaultwarden running deployed=True subdomain=vault controller version on the page: 0.243.0 + CORRECTION to the line above: fajlok.enkicsifelhom.hu was MY guess and is not the address. + file manager re-checked at its real address: -> + hub host page, version facts: + Agent -> r-1-652049 ONLINE Host ID tester-1-652049 Customer Tester 1 Agent Version 0.131.0 PBS wrapper matches vouched 104db0a4401f… Enrolled 1h ago Last Report 9 min ago Desired Generation + agent -> lhom-pbs critical ok backup tier felhom-pbs readable by the agent (archive listing, restore-test candidacy) pve:store-grant:local ok backup tier local readable by the agent (archiv + 0.131 -> INE Host ID tester-1-652049 Customer Tester 1 Agent Version 0.131.0 PBS wrapper matches vouched 104db0a4401f… Enrolled 1h ago Last Report 9 min ago Desired Generation 1 Vitals CPU + file manager, CORRECT route (it has no subdomain of its own; it is served by the dashboard): + https://felhom.enkicsifelhom.hu/files/ (logged in) -> 404 + https://felhom.enkicsifelhom.hu/files/ (no session) -> 302 <- negative control, must not be 200 + version labels, measured: + controller on the page: 0.243.0 ; hub customer row: Version 0.243.0, Floor v0.242.0 (a floor is a minimum) + agent on the hub host page: 0.131.0 ; PBS wrapper matches vouched 104db0a4401f... + file manager, THIRD and correct attempt - the dashboard's own link says files.enkicsifelhom.hu: + https://files.enkicsifelhom.hu/ -> 200 + (my two earlier lines, fajlok.enkicsifelhom.hu 404 and /files/ 404, were MY wrong guesses, + not product faults. The address was read from the dashboard page, not invented.) diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 7ec060e1..f35ef5c8 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -710,7 +710,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-507** | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** | | **R-508** | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** | | **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). **SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL:** triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as `selfbind_link_sent` and shown on the Setup tab. Unit-proven with a red-proof (`TestSelfBind_EmailSetOnWaitingCustomerSendsLink`). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 `deprovision` + Cloudflare), which is fenced without the operator's word. **Closes on:** one real mail from either trigger. | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** | -| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes. | **READY — rank P2-MEDIUM; owner: CC (hub)** | +| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes **CONSEQUENCE MEASURED 2026-09-16:** the code fix shipped for this row is sound and currently INERT - the adopt path runs and the ENDPOINT refuses it for a missing `Datastore.Modify` grant (**R-534**). On the drill box the result is that a fresh install has NO off-site tier at all, and read-only listing of ep0 shows `ns/tester-1/ct` empty before and after the whole drill. This row stays open until R-534's grant is given and a re-issue is seen to succeed. | **READY — rank P2-MEDIUM; owner: CC (hub)** | | **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** | | **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | @@ -726,7 +726,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** | | **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** | | **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** | -| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend. | **READY — rank P2-MEDIUM; owner: CC** | +| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** | | **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | | **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** | | **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |