From 8c7f882d1ce0e1bf49fab3e7a24ba23286279476 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 14 Sep 2026 16:40:59 +0200 Subject: [PATCH] drill 0242: teardown complete in three layers; R-501 filed Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed, ~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next full-list push. Evidence pulled before the destroy. R-501: the documented CI-check recipe reads only the last jobs page, which is not in id order. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-drill-fresh-install-0242.md | 31 +++++++++++++++++-- STATUS.md | 2 +- .../journal.md | 14 +++++++++ .../teardown-layer3-hub.txt | 29 +++++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 1 + 5 files changed, 73 insertions(+), 4 deletions(-) create mode 100644 documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/teardown-layer3-hub.txt diff --git a/REPORT-drill-fresh-install-0242.md b/REPORT-drill-fresh-install-0242.md index 26a5a676..5763df08 100644 --- a/REPORT-drill-fresh-install-0242.md +++ b/REPORT-drill-fresh-install-0242.md @@ -62,9 +62,10 @@ entry. Full list with harness slips: findings doc §3. | R-499 | P2 | „already in the PBS backup, nothing to do" on a box with no PBS | | R-498 | P3 | 52 of 53 app pages say literal `wiki.DOMAIN` | | R-500 | P3 | dashboard backup time in UTC, backup pages in local time | +| R-501 | P3 | the documented CI-check recipe reads only the last jobs page and can miss a run (hygiene, found at push) | **R-469 was not touched.** R-214 reproduced (recorded, row unchanged). -**Register: 200 → 208 table rows. Opened 8, closed 0.** +**Register: 200 → 209 table rows. Opened 9 (R-493…R-501), closed 0.** ## 5. Teardown — three layers @@ -72,11 +73,28 @@ entry. Full list with harness slips: findings doc §3. |---|---|---| | **1. machine** | `qm list`: VM 330 running, 8.6 G in `/mnt/hdd_1/images/330` | `qm destroy 330 --purge` → `qm list` empty; `images/330` gone; `images/9202` untouched | | **2. host** | `nvme-scratch` used 19 059 372 KiB · `local` used 24 768 200 KiB · ISO present · `/root/drill0242` 40 files | `nvme-scratch` **10 130 532 KiB** (≈8.5 GiB returned) · `local` **23 013 832 KiB** (the 1.7 GiB ISO) · ISO 0 · scratch dir shredded and removed. `nvme-scratch` storage itself stays — it hosts 9202. `vmbr9` pre-existed | -| **3. hub** | customer `drill0242`, host `drill0242-3f4b42` ONLINE, `deletable:false`, `wg_peer_bound:true`, `recovery_present:true` | **PENDING — see §5b** | +| **3. hub** | customer `drill0242`, host `drill0242-3f4b42` ONLINE, `deletable:false`, `wg_peer_bound:true`, `recovery_present:true` | **customer and host DELETED** — customer page 404, host page 404, both lists 0 (control customer listed 1); ep0 peer `10.77.0.5` **gone** at the next peer push (control peer present) — §5b | ### 5b. The hub record -*(Filled in when the host has aged to stale and the delete cascade has run.)* +**Disposition: DELETED.** Not retained as a fixture, not blocked. + +| UTC | observable | +|---|---| +| 14:34:45 | hub: `Host staleness: drill0242-3f4b42 ok → stale (host_stale)` · `Operator email sent for drill0242/host_stale` — a **true** alarm, caused by the VM destroy | +| 14:34:53 | `/hosts/drill0242-3f4b42/delete-impact` → `"deletable":true,"status":"stale","wg_peer_bound":true,"recovery_present":true` | +| 14:35:15 | `POST /configs/drill0242/delete` `ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=drill0242 expect_hosts=1` → **303 `/configs?flash=deleted`** | +| 14:35:16–17 | `customer DELETE cascade started … (journal #17, 1 host(s))` · `host drill0242-3f4b42 deleted (escrow DEMOTED to retained custody)` · `tenantsync: deprovision ok for drill0242 (ns=drill0242, existed=false)` · `reset drill0242: PBS tenancy deprovisioned` · `[claim] reset to unclaimed` · `residue purged (reports=9 app_telemetry=16 … notif_prefs=1 selfbind_tokens=1 appliance_registrations=1)` · **`customer DELETE cascade COMPLETE for drill0242 (journal #17) — full teardown`** | +| 14:35:2x | `/customers/drill0242` **404** · `/hosts/drill0242-3f4b42` **404** · `drill0242` on `/configs` **0** (control `enkisfelhom` **1**) · on `/hosts` **0** | +| 14:35:15 / 14:37:21 | ep0 `wg show all allowed-ips`: `10.77.0.5/32` present **1**, **1** (read-only; control `10.77.0.3/32` present) | +| 14:39:30 | hub: `wgsync: pushed 4 peers to 167.233.158.164:22` (was 5) | +| 14:40:00 | ep0: `10.77.0.5/32` **0**, config files naming it **0**; control `10.77.0.3/32` **1** | + +**Observation, not filed:** the delete does not trigger an immediate peer push, so the ep0 peer +outlived the customer by 4 m 14 s, until the periodic full-list push (`wgsync/reconciler.go:19-23`, +declarative by design). **The retained escrow custody the host delete mentions is empty here** — +`escrow_present:false`; no ceremony ran — and the customer purge is the step that removes it anyway. +`tenantsync … existed=false` confirms no ep0 PBS namespace was ever created (DR tier off). **Append-only and staying, by design:** the hub's event stream (`controller_started`, `app_removed`, `claim_lockout`, the bind and enrol lines) and the operator e-mail for the lockout. @@ -97,6 +115,13 @@ event, by design. Unchanged: 55 claims, NOT WALKED 35 of 55. It reads its own claim list, not the new map row. +## 7b. CI for my push + +`38848ff` (felhom.eu `main`): **job id 581 `gates` — completed, conclusion `success`**, 14:18:23Z +(run id 582, run_number 335 on `actions/tasks`). Found by scanning every page of `actions/jobs` and +matching `head_sha`: the documented last-page recipe did not list it for 8 minutes because the list is +not in id order — **R-501**. Register now **200 → 209** table rows (opened 9, closed 0). + ## 8. Observations - The deploy page's poll has no `degraded` branch; 16 s of stale step text on BookStack. diff --git a/STATUS.md b/STATUS.md index 6a0cc449..02a8c428 100644 --- a/STATUS.md +++ b/STATUS.md @@ -23,7 +23,7 @@ literally; and the dashboard showing the backup two hours off from the backup pa dashboard once in a way a volunteer could not — that is the single intervention. I fixed nothing; this run only records. -**Rows.** Opened 8, closed 0. Register table rows: 200 before, 208 after. +**Rows.** Opened 9 (eight from the walk, one about how I check the build server), closed 0. Register table rows: 200 before, 209 after. **Needs you.** (1) **Decide how a new customer reaches their dashboard**: the hub creates the web address automatically, or the box offers a home-network address that works with no setup. **If you do diff --git a/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md b/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md index b06f6a57..aff94be4 100644 --- a/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md +++ b/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/journal.md @@ -480,3 +480,17 @@ the lock is honest about its length and lifts on time. Two properties recorded, household out for 15 minutes, by design against guessing; and the lockout e-mail went to the **operator only** (`Operator email sent for drill0242/claim_lockout`), while the screen tells the customer. + +## Teardown — three layers (`teardown-before.txt`, `teardown-layer1.txt`, `teardown-layer2.txt`, `teardown-layer3-hub.txt`) + +Evidence was copied off the box **before** the destroy (`box-logs-final/`, 14:16, secret sweep 0). + +| layer | UTC | result | +|---|---|---| +| **1. machine** | 14:16:23 | `qm destroy 330 --purge` → `qm list` empty; `/mnt/hdd_1/images/330` gone; `images/9202` untouched | +| **2. host** | 14:16:27 | ISO removed (0 left); `/root/drill0242` shredded, removed; `nvme-scratch` used **19 059 372 → 10 130 532 KiB**; `local` **24 768 200 → 23 013 832 KiB**; `local-lvm` 44.17 % before and after; 9201 and 9202 running; `vmbr9` pre-existed | +| **3. hub** | 14:35:15 | host stale at 14:34:45 (true `host_stale` operator mail); **customer DELETE cascade COMPLETE** (journal #17); customer and host pages 404; lists 0 (control 1) | +| 3b. ep0 peer | 14:40:00 | `10.77.0.5/32` removed at the 14:39:30 periodic push (5 → 4 peers); 4 m 14 s after the delete; control peer present | + +**Disposition of `drill0242`: DELETED.** The hub's event stream and the three operator mails +(`claim_lockout`, `host_stale`, and the bind/enrol log lines) are append-only and stay, by design. diff --git a/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/teardown-layer3-hub.txt b/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/teardown-layer3-hub.txt new file mode 100644 index 00000000..70f0a179 --- /dev/null +++ b/documentation/audits/evidence-drill-fresh-install-0242-2026-09-14/teardown-layer3-hub.txt @@ -0,0 +1,29 @@ +=== LAYER 3 — the hub 2026-09-14T14:35:15Z +ep0 WG peer for 10.77.0.5 BEFORE: 1 +delete POST: HTTP/1.1 303 See Other Location: /configs?flash=deleted + +customer page: 404 +host page: 404 +configs list mentions drill0242: 0 (control, a live customer 'enkisfelhom': 1) +hosts list mentions drill0242: 0 +ep0 WG peer for 10.77.0.5 AFTER: 1 (control, demo-hp 10.77.0.3: 1) +--- hub log +2026/09/14 16:34:45 [INFO] Host staleness: drill0242-3f4b42 ok → stale (host_stale) +2026/09/14 16:34:46 [INFO] Operator email sent for drill0242/host_stale +2026/09/14 16:35:16 [INFO] customer DELETE cascade started for drill0242 (journal #17, 1 host(s)) +2026/09/14 16:35:16 [INFO] delete drill0242: host drill0242-3f4b42 deleted (escrow DEMOTED to retained custody) +2026/09/14 16:35:17 [INFO] tenantsync: deprovision ok for drill0242 (ns=drill0242, existed=false) +2026/09/14 16:35:17 [INFO] reset drill0242: PBS tenancy deprovisioned +2026/09/14 16:35:17 [INFO] [claim] reset to unclaimed for drill0242 (customer RESET) — next onboarding mints a fresh code +2026/09/14 16:35:17 [INFO] delete drill0242: residue purged (reports=9 app_telemetry=16 app_log_tails=0 log_tail_requests=0 notif_prefs=1 selfbind_tokens=1 appliance_registrations=1) +2026/09/14 16:35:17 [INFO] customer DELETE cascade COMPLETE for drill0242 (journal #17) — full teardown +=== ep0 re-check 14:37:21 +peer 10.77.0.5: 1 control 10.77.0.3: 1 +peer config files naming 10.77.0.5: 1 +--- hub log wg lines +2026/09/14 16:29:30 [INFO] wgsync: pushed 5 peers to 167.233.158.164:22 +2026/09/14 16:34:30 [INFO] wgsync: pushed 5 peers to 167.233.158.164:22 +=== ep0 re-check after the next wgsync push 14:40:00 +peer 10.77.0.5: 0 control 10.77.0.3: 1 +config files naming 10.77.0.5: 0 +2026/09/14 16:39:30 [INFO] wgsync: pushed 4 peers to 167.233.158.164:22 diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index bc547a81..cdbbcb75 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -704,6 +704,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-498** | **[P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal `wiki.DOMAIN` — the placeholder is never filled in.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a **wiki.DOMAIN** címet a böngészőben", PrivateBin's „Nyisd meg a **paste.DOMAIN** címet". `grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml` → **52 of 53** templates carry it in `first_steps`; the controller renders the string as text (`internal/stacks/metadata.go`). A stranger reading their first instruction meets a word that is not an address. **Fix shape:** substitute the stack's real `SUBDOMAIN.DOMAIN` at render time (one place in the controller), with a render test per template that fails on a literal `DOMAIN`. | **READY — rank P3-LOW; owner: CC** | | **R-499** | **[P2-MEDIUM] Every app without a data drive is told its data is „already in the full system backup (PBS)" and there is „nothing to do" — on a box with no PBS, whose only whole-box copy sits on the same disk.** MEASURED 2026-09-14 on a fresh 0.242.0 box with the DR tier off (drill step 7): `GET /stacks/bookstack/backup` renders „Ennek az alkalmazásnak az adatai a belső rendszerlemezen vannak, amelyek **már szerepelnek a teljes rendszermentésben (PBS)** … ehhez az alkalmazáshoz nincs külön teendő." The sentence sits under `{{if not .IsHDDApp}}` in `controller/internal/web/templates/tier2_config.html:20-26` and consults nothing about where the whole-guest backup goes. On the same box `/backups` says, correctly, „Helyi tároló (local)" and „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem." **Two pages of one product contradict each other, and the reassuring one is the false one.** **Fix shape:** branch the sentence on the box's actual whole-guest target (PBS vs local, same-disk flag the overview already computes); a render test per branch. | **READY — rank P2-MEDIUM; owner: CC** | | **R-500** | **[P3-LOW] The dashboard shows the last backup in UTC while every backup page shows it in local time — two different clock times for one backup.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 12): the same backup (`last_run 2026-09-14T13:41:37Z`) reads „Utolsó mentés: **2026-09-14 13:41**" on `/dashboard` and „Utolsó adatbázis mentés: **2026-09-14 15:41** (most)" on `/backups/apps`. Cause, from source: `controller/internal/web/templates/dashboard.html:154` renders `{{.BackupStatus.LastRun.Format "2006-01-02 15:04"}}` with no conversion to the box's zone, while the backup pages go through the zone-aware helpers in `funcmap.go`. A household comparing the two screens sees a backup two hours apart from itself. **Fix shape:** format through the same zone-aware helper; a render test that pins a non-UTC zone and asserts the local hour. | **READY — rank P3-LOW; owner: CC** | +| **R-501** | **[P3-LOW] The documented "confirm your CI run" recipe reads only the LAST page of the jobs list, and that list is not in id order — so it can report a run as missing that exists and passed.** MEASURED 2026-09-14 for felhom.eu commit `38848ff`: `actions/jobs?limit=1` → `total_count 335`; the recipe's page `T/50+1` = page 7 held ids 473…570 and **no match for 8 minutes**; a scan of all seven pages found the job on **page 6** (ids 380…581): `job id=581 name=gates status=completed conclusion=success completed_at 2026-09-14T14:18:23Z`. Pages are not sorted (page 1 ids 1…273, page 2 84…208). `actions/tasks` listed the same run first (`id 582, run_number 335, success`). **Fix shape:** the recipe in `felhom.eu/CLAUDE.md` (end-of-session checklist) scans every page and matches `head_sha`; say so in the same sentence that warns about the id offset (R-417). | **READY — rank P3-LOW; owner: CC** |