Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed, ~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next full-list push. Evidence pulled before the destroy. R-501: the documented CI-check recipe reads only the last jobs page, which is not in id order. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -62,9 +62,10 @@ entry. Full list with harness slips: findings doc §3.
|
|||||||
| R-499 | P2 | „already in the PBS backup, nothing to do" on a box with no PBS |
|
| R-499 | P2 | „already in the PBS backup, nothing to do" on a box with no PBS |
|
||||||
| R-498 | P3 | 52 of 53 app pages say literal `wiki.DOMAIN` |
|
| R-498 | P3 | 52 of 53 app pages say literal `wiki.DOMAIN` |
|
||||||
| R-500 | P3 | dashboard backup time in UTC, backup pages in local time |
|
| R-500 | P3 | dashboard backup time in UTC, backup pages in local time |
|
||||||
|
| R-501 | P3 | the documented CI-check recipe reads only the last jobs page and can miss a run (hygiene, found at push) |
|
||||||
|
|
||||||
**R-469 was not touched.** R-214 reproduced (recorded, row unchanged).
|
**R-469 was not touched.** R-214 reproduced (recorded, row unchanged).
|
||||||
**Register: 200 → 208 table rows. Opened 8, closed 0.**
|
**Register: 200 → 209 table rows. Opened 9 (R-493…R-501), closed 0.**
|
||||||
|
|
||||||
## 5. Teardown — three layers
|
## 5. Teardown — three layers
|
||||||
|
|
||||||
@@ -72,11 +73,28 @@ entry. Full list with harness slips: findings doc §3.
|
|||||||
|---|---|---|
|
|---|---|---|
|
||||||
| **1. machine** | `qm list`: VM 330 running, 8.6 G in `/mnt/hdd_1/images/330` | `qm destroy 330 --purge` → `qm list` empty; `images/330` gone; `images/9202` untouched |
|
| **1. machine** | `qm list`: VM 330 running, 8.6 G in `/mnt/hdd_1/images/330` | `qm destroy 330 --purge` → `qm list` empty; `images/330` gone; `images/9202` untouched |
|
||||||
| **2. host** | `nvme-scratch` used 19 059 372 KiB · `local` used 24 768 200 KiB · ISO present · `/root/drill0242` 40 files | `nvme-scratch` **10 130 532 KiB** (≈8.5 GiB returned) · `local` **23 013 832 KiB** (the 1.7 GiB ISO) · ISO 0 · scratch dir shredded and removed. `nvme-scratch` storage itself stays — it hosts 9202. `vmbr9` pre-existed |
|
| **2. host** | `nvme-scratch` used 19 059 372 KiB · `local` used 24 768 200 KiB · ISO present · `/root/drill0242` 40 files | `nvme-scratch` **10 130 532 KiB** (≈8.5 GiB returned) · `local` **23 013 832 KiB** (the 1.7 GiB ISO) · ISO 0 · scratch dir shredded and removed. `nvme-scratch` storage itself stays — it hosts 9202. `vmbr9` pre-existed |
|
||||||
| **3. hub** | customer `drill0242`, host `drill0242-3f4b42` ONLINE, `deletable:false`, `wg_peer_bound:true`, `recovery_present:true` | **PENDING — see §5b** |
|
| **3. hub** | customer `drill0242`, host `drill0242-3f4b42` ONLINE, `deletable:false`, `wg_peer_bound:true`, `recovery_present:true` | **customer and host DELETED** — customer page 404, host page 404, both lists 0 (control customer listed 1); ep0 peer `10.77.0.5` **gone** at the next peer push (control peer present) — §5b |
|
||||||
|
|
||||||
### 5b. The hub record
|
### 5b. The hub record
|
||||||
|
|
||||||
*(Filled in when the host has aged to stale and the delete cascade has run.)*
|
**Disposition: DELETED.** Not retained as a fixture, not blocked.
|
||||||
|
|
||||||
|
| UTC | observable |
|
||||||
|
|---|---|
|
||||||
|
| 14:34:45 | hub: `Host staleness: drill0242-3f4b42 ok → stale (host_stale)` · `Operator email sent for drill0242/host_stale` — a **true** alarm, caused by the VM destroy |
|
||||||
|
| 14:34:53 | `/hosts/drill0242-3f4b42/delete-impact` → `"deletable":true,"status":"stale","wg_peer_bound":true,"recovery_present":true` |
|
||||||
|
| 14:35:15 | `POST /configs/drill0242/delete` `ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=drill0242 expect_hosts=1` → **303 `/configs?flash=deleted`** |
|
||||||
|
| 14:35:16–17 | `customer DELETE cascade started … (journal #17, 1 host(s))` · `host drill0242-3f4b42 deleted (escrow DEMOTED to retained custody)` · `tenantsync: deprovision ok for drill0242 (ns=drill0242, existed=false)` · `reset drill0242: PBS tenancy deprovisioned` · `[claim] reset to unclaimed` · `residue purged (reports=9 app_telemetry=16 … notif_prefs=1 selfbind_tokens=1 appliance_registrations=1)` · **`customer DELETE cascade COMPLETE for drill0242 (journal #17) — full teardown`** |
|
||||||
|
| 14:35:2x | `/customers/drill0242` **404** · `/hosts/drill0242-3f4b42` **404** · `drill0242` on `/configs` **0** (control `enkisfelhom` **1**) · on `/hosts` **0** |
|
||||||
|
| 14:35:15 / 14:37:21 | ep0 `wg show all allowed-ips`: `10.77.0.5/32` present **1**, **1** (read-only; control `10.77.0.3/32` present) |
|
||||||
|
| 14:39:30 | hub: `wgsync: pushed 4 peers to 167.233.158.164:22` (was 5) |
|
||||||
|
| 14:40:00 | ep0: `10.77.0.5/32` **0**, config files naming it **0**; control `10.77.0.3/32` **1** |
|
||||||
|
|
||||||
|
**Observation, not filed:** the delete does not trigger an immediate peer push, so the ep0 peer
|
||||||
|
outlived the customer by 4 m 14 s, until the periodic full-list push (`wgsync/reconciler.go:19-23`,
|
||||||
|
declarative by design). **The retained escrow custody the host delete mentions is empty here** —
|
||||||
|
`escrow_present:false`; no ceremony ran — and the customer purge is the step that removes it anyway.
|
||||||
|
`tenantsync … existed=false` confirms no ep0 PBS namespace was ever created (DR tier off).
|
||||||
|
|
||||||
**Append-only and staying, by design:** the hub's event stream (`controller_started`, `app_removed`,
|
**Append-only and staying, by design:** the hub's event stream (`controller_started`, `app_removed`,
|
||||||
`claim_lockout`, the bind and enrol lines) and the operator e-mail for the lockout.
|
`claim_lockout`, the bind and enrol lines) and the operator e-mail for the lockout.
|
||||||
@@ -97,6 +115,13 @@ event, by design.
|
|||||||
|
|
||||||
Unchanged: 55 claims, NOT WALKED 35 of 55. It reads its own claim list, not the new map row.
|
Unchanged: 55 claims, NOT WALKED 35 of 55. It reads its own claim list, not the new map row.
|
||||||
|
|
||||||
|
## 7b. CI for my push
|
||||||
|
|
||||||
|
`38848ff` (felhom.eu `main`): **job id 581 `gates` — completed, conclusion `success`**, 14:18:23Z
|
||||||
|
(run id 582, run_number 335 on `actions/tasks`). Found by scanning every page of `actions/jobs` and
|
||||||
|
matching `head_sha`: the documented last-page recipe did not list it for 8 minutes because the list is
|
||||||
|
not in id order — **R-501**. Register now **200 → 209** table rows (opened 9, closed 0).
|
||||||
|
|
||||||
## 8. Observations
|
## 8. Observations
|
||||||
|
|
||||||
- The deploy page's poll has no `degraded` branch; 16 s of stale step text on BookStack.
|
- The deploy page's poll has no `degraded` branch; 16 s of stale step text on BookStack.
|
||||||
|
|||||||
@@ -23,7 +23,7 @@ literally; and the dashboard showing the backup two hours off from the backup pa
|
|||||||
dashboard once in a way a volunteer could not — that is the single intervention. I fixed nothing; this
|
dashboard once in a way a volunteer could not — that is the single intervention. I fixed nothing; this
|
||||||
run only records.
|
run only records.
|
||||||
|
|
||||||
**Rows.** Opened 8, closed 0. Register table rows: 200 before, 208 after.
|
**Rows.** Opened 9 (eight from the walk, one about how I check the build server), closed 0. Register table rows: 200 before, 209 after.
|
||||||
|
|
||||||
**Needs you.** (1) **Decide how a new customer reaches their dashboard**: the hub creates the web
|
**Needs you.** (1) **Decide how a new customer reaches their dashboard**: the hub creates the web
|
||||||
address automatically, or the box offers a home-network address that works with no setup. **If you do
|
address automatically, or the box offers a home-network address that works with no setup. **If you do
|
||||||
|
|||||||
@@ -480,3 +480,17 @@ the lock is honest about its length and lifts on time. Two properties recorded,
|
|||||||
household out for 15 minutes, by design against guessing; and the lockout e-mail went to the
|
household out for 15 minutes, by design against guessing; and the lockout e-mail went to the
|
||||||
**operator only** (`Operator email sent for drill0242/claim_lockout`), while the screen tells the
|
**operator only** (`Operator email sent for drill0242/claim_lockout`), while the screen tells the
|
||||||
customer.
|
customer.
|
||||||
|
|
||||||
|
## Teardown — three layers (`teardown-before.txt`, `teardown-layer1.txt`, `teardown-layer2.txt`, `teardown-layer3-hub.txt`)
|
||||||
|
|
||||||
|
Evidence was copied off the box **before** the destroy (`box-logs-final/`, 14:16, secret sweep 0).
|
||||||
|
|
||||||
|
| layer | UTC | result |
|
||||||
|
|---|---|---|
|
||||||
|
| **1. machine** | 14:16:23 | `qm destroy 330 --purge` → `qm list` empty; `/mnt/hdd_1/images/330` gone; `images/9202` untouched |
|
||||||
|
| **2. host** | 14:16:27 | ISO removed (0 left); `/root/drill0242` shredded, removed; `nvme-scratch` used **19 059 372 → 10 130 532 KiB**; `local` **24 768 200 → 23 013 832 KiB**; `local-lvm` 44.17 % before and after; 9201 and 9202 running; `vmbr9` pre-existed |
|
||||||
|
| **3. hub** | 14:35:15 | host stale at 14:34:45 (true `host_stale` operator mail); **customer DELETE cascade COMPLETE** (journal #17); customer and host pages 404; lists 0 (control 1) |
|
||||||
|
| 3b. ep0 peer | 14:40:00 | `10.77.0.5/32` removed at the 14:39:30 periodic push (5 → 4 peers); 4 m 14 s after the delete; control peer present |
|
||||||
|
|
||||||
|
**Disposition of `drill0242`: DELETED.** The hub's event stream and the three operator mails
|
||||||
|
(`claim_lockout`, `host_stale`, and the bind/enrol log lines) are append-only and stay, by design.
|
||||||
|
|||||||
+29
@@ -0,0 +1,29 @@
|
|||||||
|
=== LAYER 3 — the hub 2026-09-14T14:35:15Z
|
||||||
|
ep0 WG peer for 10.77.0.5 BEFORE: 1
|
||||||
|
delete POST: HTTP/1.1 303 See Other Location: /configs?flash=deleted
|
||||||
|
|
||||||
|
customer page: 404
|
||||||
|
host page: 404
|
||||||
|
configs list mentions drill0242: 0 (control, a live customer 'enkisfelhom': 1)
|
||||||
|
hosts list mentions drill0242: 0
|
||||||
|
ep0 WG peer for 10.77.0.5 AFTER: 1 (control, demo-hp 10.77.0.3: 1)
|
||||||
|
--- hub log
|
||||||
|
2026/09/14 16:34:45 [INFO] Host staleness: drill0242-3f4b42 ok → stale (host_stale)
|
||||||
|
2026/09/14 16:34:46 [INFO] Operator email sent for drill0242/host_stale
|
||||||
|
2026/09/14 16:35:16 [INFO] customer DELETE cascade started for drill0242 (journal #17, 1 host(s))
|
||||||
|
2026/09/14 16:35:16 [INFO] delete drill0242: host drill0242-3f4b42 deleted (escrow DEMOTED to retained custody)
|
||||||
|
2026/09/14 16:35:17 [INFO] tenantsync: deprovision ok for drill0242 (ns=drill0242, existed=false)
|
||||||
|
2026/09/14 16:35:17 [INFO] reset drill0242: PBS tenancy deprovisioned
|
||||||
|
2026/09/14 16:35:17 [INFO] [claim] reset to unclaimed for drill0242 (customer RESET) — next onboarding mints a fresh code
|
||||||
|
2026/09/14 16:35:17 [INFO] delete drill0242: residue purged (reports=9 app_telemetry=16 app_log_tails=0 log_tail_requests=0 notif_prefs=1 selfbind_tokens=1 appliance_registrations=1)
|
||||||
|
2026/09/14 16:35:17 [INFO] customer DELETE cascade COMPLETE for drill0242 (journal #17) — full teardown
|
||||||
|
=== ep0 re-check 14:37:21
|
||||||
|
peer 10.77.0.5: 1 control 10.77.0.3: 1
|
||||||
|
peer config files naming 10.77.0.5: 1
|
||||||
|
--- hub log wg lines
|
||||||
|
2026/09/14 16:29:30 [INFO] wgsync: pushed 5 peers to 167.233.158.164:22
|
||||||
|
2026/09/14 16:34:30 [INFO] wgsync: pushed 5 peers to 167.233.158.164:22
|
||||||
|
=== ep0 re-check after the next wgsync push 14:40:00
|
||||||
|
peer 10.77.0.5: 0 control 10.77.0.3: 1
|
||||||
|
config files naming 10.77.0.5: 0
|
||||||
|
2026/09/14 16:39:30 [INFO] wgsync: pushed 4 peers to 167.233.158.164:22
|
||||||
@@ -704,6 +704,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
|||||||
| **R-498** | **[P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal `wiki.DOMAIN` — the placeholder is never filled in.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a **wiki.DOMAIN** címet a böngészőben", PrivateBin's „Nyisd meg a **paste.DOMAIN** címet". `grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml` → **52 of 53** templates carry it in `first_steps`; the controller renders the string as text (`internal/stacks/metadata.go`). A stranger reading their first instruction meets a word that is not an address. **Fix shape:** substitute the stack's real `SUBDOMAIN.DOMAIN` at render time (one place in the controller), with a render test per template that fails on a literal `DOMAIN`. | **READY — rank P3-LOW; owner: CC** |
|
| **R-498** | **[P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal `wiki.DOMAIN` — the placeholder is never filled in.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a **wiki.DOMAIN** címet a böngészőben", PrivateBin's „Nyisd meg a **paste.DOMAIN** címet". `grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml` → **52 of 53** templates carry it in `first_steps`; the controller renders the string as text (`internal/stacks/metadata.go`). A stranger reading their first instruction meets a word that is not an address. **Fix shape:** substitute the stack's real `SUBDOMAIN.DOMAIN` at render time (one place in the controller), with a render test per template that fails on a literal `DOMAIN`. | **READY — rank P3-LOW; owner: CC** |
|
||||||
| **R-499** | **[P2-MEDIUM] Every app without a data drive is told its data is „already in the full system backup (PBS)" and there is „nothing to do" — on a box with no PBS, whose only whole-box copy sits on the same disk.** MEASURED 2026-09-14 on a fresh 0.242.0 box with the DR tier off (drill step 7): `GET /stacks/bookstack/backup` renders „Ennek az alkalmazásnak az adatai a belső rendszerlemezen vannak, amelyek **már szerepelnek a teljes rendszermentésben (PBS)** … ehhez az alkalmazáshoz nincs külön teendő." The sentence sits under `{{if not .IsHDDApp}}` in `controller/internal/web/templates/tier2_config.html:20-26` and consults nothing about where the whole-guest backup goes. On the same box `/backups` says, correctly, „Helyi tároló (local)" and „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem." **Two pages of one product contradict each other, and the reassuring one is the false one.** **Fix shape:** branch the sentence on the box's actual whole-guest target (PBS vs local, same-disk flag the overview already computes); a render test per branch. | **READY — rank P2-MEDIUM; owner: CC** |
|
| **R-499** | **[P2-MEDIUM] Every app without a data drive is told its data is „already in the full system backup (PBS)" and there is „nothing to do" — on a box with no PBS, whose only whole-box copy sits on the same disk.** MEASURED 2026-09-14 on a fresh 0.242.0 box with the DR tier off (drill step 7): `GET /stacks/bookstack/backup` renders „Ennek az alkalmazásnak az adatai a belső rendszerlemezen vannak, amelyek **már szerepelnek a teljes rendszermentésben (PBS)** … ehhez az alkalmazáshoz nincs külön teendő." The sentence sits under `{{if not .IsHDDApp}}` in `controller/internal/web/templates/tier2_config.html:20-26` and consults nothing about where the whole-guest backup goes. On the same box `/backups` says, correctly, „Helyi tároló (local)" and „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem." **Two pages of one product contradict each other, and the reassuring one is the false one.** **Fix shape:** branch the sentence on the box's actual whole-guest target (PBS vs local, same-disk flag the overview already computes); a render test per branch. | **READY — rank P2-MEDIUM; owner: CC** |
|
||||||
| **R-500** | **[P3-LOW] The dashboard shows the last backup in UTC while every backup page shows it in local time — two different clock times for one backup.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 12): the same backup (`last_run 2026-09-14T13:41:37Z`) reads „Utolsó mentés: **2026-09-14 13:41**" on `/dashboard` and „Utolsó adatbázis mentés: **2026-09-14 15:41** (most)" on `/backups/apps`. Cause, from source: `controller/internal/web/templates/dashboard.html:154` renders `{{.BackupStatus.LastRun.Format "2006-01-02 15:04"}}` with no conversion to the box's zone, while the backup pages go through the zone-aware helpers in `funcmap.go`. A household comparing the two screens sees a backup two hours apart from itself. **Fix shape:** format through the same zone-aware helper; a render test that pins a non-UTC zone and asserts the local hour. | **READY — rank P3-LOW; owner: CC** |
|
| **R-500** | **[P3-LOW] The dashboard shows the last backup in UTC while every backup page shows it in local time — two different clock times for one backup.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 12): the same backup (`last_run 2026-09-14T13:41:37Z`) reads „Utolsó mentés: **2026-09-14 13:41**" on `/dashboard` and „Utolsó adatbázis mentés: **2026-09-14 15:41** (most)" on `/backups/apps`. Cause, from source: `controller/internal/web/templates/dashboard.html:154` renders `{{.BackupStatus.LastRun.Format "2006-01-02 15:04"}}` with no conversion to the box's zone, while the backup pages go through the zone-aware helpers in `funcmap.go`. A household comparing the two screens sees a backup two hours apart from itself. **Fix shape:** format through the same zone-aware helper; a render test that pins a non-UTC zone and asserts the local hour. | **READY — rank P3-LOW; owner: CC** |
|
||||||
|
| **R-501** | **[P3-LOW] The documented "confirm your CI run" recipe reads only the LAST page of the jobs list, and that list is not in id order — so it can report a run as missing that exists and passed.** MEASURED 2026-09-14 for felhom.eu commit `38848ff`: `actions/jobs?limit=1` → `total_count 335`; the recipe's page `T/50+1` = page 7 held ids 473…570 and **no match for 8 minutes**; a scan of all seven pages found the job on **page 6** (ids 380…581): `job id=581 name=gates status=completed conclusion=success completed_at 2026-09-14T14:18:23Z`. Pages are not sorted (page 1 ids 1…273, page 2 84…208). `actions/tasks` listed the same run first (`id 582, run_number 335, success`). **Fix shape:** the recipe in `felhom.eu/CLAUDE.md` (end-of-session checklist) scans every page and matches `head_sha`; say so in the same sentence that warns about the id offset (R-417). | **READY — rank P3-LOW; owner: CC** |
|
||||||
|
|
||||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||||
|
|||||||
Reference in New Issue
Block a user