walk5 venue TORN DOWN — census 168 rows -> 67, and R-244 grew by 30 as predicted
gates / gates (push) Failing after 13s
gates / gates (push) Failing after 13s
Operator-confirmed. Stopped under a name guard (demo-hp carries its own 9201), aged past the hub's stale_threshold read from the DEPLOYED ConfigMap (30m), and polled delete-impact until deletable:true — treating an empty response as retry, never as success. Cascade + qm destroy --purge, guarded a second time. Every layer verified absent against a positive control that must survive and does: VM 300 drill-r50 and demo-hp's own guest 9201 still there; ep0 namespaces demo-felhom + demo-hp still there; wg peers .2 .3 .4 .250 still on the live wg0; hub rows for demo-felhom, demo-hp, peti-felhom untouched. 16.64 GiB returned against 17 G measured. RECORDED FOR THE NEXT TEARDOWN: the WG peer is removed on a ~5-minute SCHEDULE, not by the cascade. Immediately after the delete the hub row was gone while 10.77.0.5 was still on ep0's live wg0; wgsync had last run 37 seconds before the cascade, and the next push (4 peers) removed it, verified on the live interface at 16:57:07Z. The previous ledger checked this after it had already converged, so it read as instantaneous — a teardown that checks too soon would file a false finding. R-244 grew by 30 rows (app_log_issues), PREDICTED in the pre-run enumeration rather than discovered afterwards. Running total across torn-down venues ~101. Nothing here claims a clean teardown. Storage Box layer evidenced from the hub's own deprovision log: the HETZNER_API token in ~/.config/credentials cannot see box 611421 (subaccounts -> 404, storage_boxes -> 200 with 0 entries) — it is scoped to another project.
This commit is contained in:
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,109 @@
|
||||
# Teardown ledger — `walk5`, the fifth walk's venue (2026-08-08)
|
||||
|
||||
**Done.** Supersedes `teardown-owed.md` in the same directory, which was the pre-agreed enumeration.
|
||||
Operator-confirmed before anything was deleted. Predecessor: `teardown-finalwalk-2026-08-07.md`.
|
||||
|
||||
**Matched on IDENTITY, never on size** — customer id → host id → sub-account → PBS namespace → wg peer.
|
||||
|
||||
---
|
||||
|
||||
## The cascade
|
||||
|
||||
VM 325 was **stopped at 16:31:30Z** under a `qm config` name guard (`demo-hp` also carries its own
|
||||
guest 9201) and aged past the hub's `stale_threshold` — **`30m`, read from the deployed `hub-config`
|
||||
ConfigMap, not assumed**. `GET /hosts/walk5-4bada5/delete-impact` was polled until it answered
|
||||
`deletable:true` (**16:51:56Z**, `status: stale`), **treating an empty response as a retry and never
|
||||
as success**.
|
||||
|
||||
`POST /configs/walk5/delete` with all six gates (`ack_hosts`, `ack_reset`, `ack_purge`,
|
||||
`confirm_id=walk5` typed, `expect_hosts=1`) → **303 `/configs?flash=deleted`**.
|
||||
|
||||
**The hub's leg-by-leg log is the evidence, not the flash:**
|
||||
|
||||
```
|
||||
18:52:12 customer DELETE cascade started for walk5 (journal #15, 1 host(s))
|
||||
18:52:12 delete walk5: host walk5-4bada5 deleted (escrow DEMOTED to retained custody)
|
||||
18:52:19 [offsite] deprovisioned shared sub-account 285351 for walk5 (repo data destroyed)
|
||||
18:52:20 tenantsync: deprovision ok for walk5 (ns=walk5, existed=true)
|
||||
18:52:20 [claim] reset to unclaimed for walk5 (customer RESET)
|
||||
18:52:20 delete walk5: residue purged (reports=30 app_telemetry=48 app_log_tails=0
|
||||
log_tail_requests=0 notif_prefs=1 selfbind_tokens=1 appliance_registrations=1)
|
||||
18:52:20 customer DELETE cascade COMPLETE for walk5 (journal #15) — full teardown
|
||||
```
|
||||
|
||||
`existed=true` on the PBS leg is the positive observable that it deleted something rather than
|
||||
no-op'ing on an absent namespace. Then, guarded a second time on `qm config 325` reading
|
||||
`name: walk5-appliance`: **`qm destroy 325 --purge` at 16:52:41Z**.
|
||||
|
||||
---
|
||||
|
||||
## After — every layer, each with a positive control that must survive, and does
|
||||
|
||||
| Layer | Item | After | Positive control |
|
||||
|---|---|---|---|
|
||||
| machine | VM 325 + 4 disks | **absent**; `/mnt/nvme-1tb/images/` **empty** | VM **300 `drill-r50`** still listed |
|
||||
| host | — | — | `demo-hp`'s **own guest 9201 still running** |
|
||||
| hub | `hosts`, `customer_configs`, `host_escrow`, `wg_peers`, `guests`, `reports`, … | **0 rows** | `demo-felhom` hosts=1 · `demo-hp` hosts=1 · `peti-felhom` config=1 |
|
||||
| off-site | Storage Box sub-account **285351** (`u629488-sub4`) | **deprovisioned, repo data destroyed** (hub log) | — *(see the access note below)* |
|
||||
| off-site | `ep0` PBS namespace `walk5` | **gone** | namespaces `demo-felhom`, `demo-hp` both present |
|
||||
| network | WireGuard peer **10.77.0.5** | **gone from the LIVE `wg show`** | peers `.2 .3 .4 .250` all still on `wg0` |
|
||||
|
||||
**Space returned, measured.** `c11-scratch` used **23 715 036 → 6 262 060 KiB = 17 452 976 KiB
|
||||
(16.64 GiB)** against the 17 G measured on disk before deletion. Available 909 638 052 → 927 091 028.
|
||||
|
||||
### ⚠ The WG peer is removed on a SCHEDULE, not by the cascade — check the live interface, not the DB
|
||||
|
||||
Immediately after the cascade the hub's `wg_peers` row was gone **while `10.77.0.5` was still on
|
||||
`ep0`'s live `wg0`**. That is **not** residue: `wgsync` pushes every ~5 minutes and had last run at
|
||||
**16:51:35Z, 37 seconds before the cascade**. The next push —
|
||||
`18:56:35 wgsync: pushed 4 peers` — removed it, **verified on the live interface at 16:57:07Z**.
|
||||
|
||||
**Recorded because the previous ledger checked this layer after it had already converged and so read
|
||||
as instantaneous.** A teardown that checks `wg show` inside ~5 minutes of the cascade will see a peer
|
||||
that is on its way out, and filing that as a finding would be wrong. **The DB row is not the
|
||||
evidence; the live interface is — and it needs a moment.**
|
||||
|
||||
---
|
||||
|
||||
## Residue — counted, not claimed away (R-244)
|
||||
|
||||
**Full-schema census** (every table, every column, `like '%walk5%'`, no truncation):
|
||||
**168 rows before → 67 rows after.** 101 removed.
|
||||
|
||||
| table | rows | verdict |
|
||||
|---|---|---|
|
||||
| `events` | 21 | **by design** — *"provenance/events are NEVER wiped; audit outlives every tier"* |
|
||||
| `notification_log` | 14 | **by design** |
|
||||
| `host_deletions` | 1 | **by design** — the deletion tombstone |
|
||||
| `customer_resets` | 1 | **by design** — the journal row (#15) |
|
||||
| **`app_log_issues`** | **30** | **THE GAP — R-244, exactly as predicted before the run** |
|
||||
|
||||
**R-244 grew by 30 rows, as this ledger's own pre-run enumeration said it would.** The prediction is
|
||||
the point: the previous ledger's *"0 occurrences"* came from a narrower query and did not survive a
|
||||
full census. **Nothing here claims a clean teardown.**
|
||||
|
||||
---
|
||||
|
||||
## Access note, per the standing rule that a "no access" claim must name its attempts
|
||||
|
||||
The Storage Box sub-account could **not** be verified directly from DooPlex. Tried, with
|
||||
`HETZNER_API` from `~/.config/credentials`: `GET /v1/storage_boxes/611421/subaccounts` → **404
|
||||
`not_found`**; `GET /v1/storage_boxes` → **200 with `total_entries: 0`**. **That token is scoped to a
|
||||
different project than box 611421** — the hub holds its own. The layer is therefore evidenced by the
|
||||
hub's own `[offsite] deprovisioned shared sub-account 285351 … (repo data destroyed)`, matched on the
|
||||
id recorded at provisioning. **Recorded so the next teardown does not re-derive it.**
|
||||
|
||||
---
|
||||
|
||||
## Not deleted, deliberately
|
||||
|
||||
The **escrow was DEMOTED to retained custody**, not destroyed (the cascade's own wording). That is
|
||||
R-198's retention behaving as designed. The customer record itself is gone; what survives is the
|
||||
sealed material under the operator's retained-custody path, plus the audit rows above.
|
||||
|
||||
## One thing that died with the box
|
||||
|
||||
During v0.207.0's live validation a **test retrieval passphrase was planted** on this appliance, and
|
||||
its settings were mutated (app removed, a drive deregistered, `storage_paths` emptied). All of it was
|
||||
on the machine destroyed above. **The `walk5` retrieval passphrase was also compromised earlier** — it
|
||||
reached a session transcript, which is R-249 — and it is now gone with the customer record.
|
||||
Reference in New Issue
Block a user