Files
felhom.eu/documentation/tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md
T
admin c1dec41328
gates / gates (push) Failing after 13s
walk5 venue TORN DOWN — census 168 rows -> 67, and R-244 grew by 30 as predicted
Operator-confirmed. Stopped under a name guard (demo-hp carries its own 9201),
aged past the hub's stale_threshold read from the DEPLOYED ConfigMap (30m), and
polled delete-impact until deletable:true — treating an empty response as retry,
never as success. Cascade + qm destroy --purge, guarded a second time.

Every layer verified absent against a positive control that must survive and does:
VM 300 drill-r50 and demo-hp's own guest 9201 still there; ep0 namespaces
demo-felhom + demo-hp still there; wg peers .2 .3 .4 .250 still on the live wg0;
hub rows for demo-felhom, demo-hp, peti-felhom untouched. 16.64 GiB returned
against 17 G measured.

RECORDED FOR THE NEXT TEARDOWN: the WG peer is removed on a ~5-minute SCHEDULE,
not by the cascade. Immediately after the delete the hub row was gone while
10.77.0.5 was still on ep0's live wg0; wgsync had last run 37 seconds before the
cascade, and the next push (4 peers) removed it, verified on the live interface at
16:57:07Z. The previous ledger checked this after it had already converged, so it
read as instantaneous — a teardown that checks too soon would file a false finding.

R-244 grew by 30 rows (app_log_issues), PREDICTED in the pre-run enumeration
rather than discovered afterwards. Running total across torn-down venues ~101.
Nothing here claims a clean teardown.

Storage Box layer evidenced from the hub's own deprovision log: the HETZNER_API
token in ~/.config/credentials cannot see box 611421 (subaccounts -> 404,
storage_boxes -> 200 with 0 entries) — it is scoped to another project.
2026-08-07 18:58:55 +02:00

6.0 KiB

Teardown ledger — walk5, the fifth walk's venue (2026-08-08)

Done. Supersedes teardown-owed.md in the same directory, which was the pre-agreed enumeration. Operator-confirmed before anything was deleted. Predecessor: teardown-finalwalk-2026-08-07.md.

Matched on IDENTITY, never on size — customer id → host id → sub-account → PBS namespace → wg peer.


The cascade

VM 325 was stopped at 16:31:30Z under a qm config name guard (demo-hp also carries its own guest 9201) and aged past the hub's stale_threshold30m, read from the deployed hub-config ConfigMap, not assumed. GET /hosts/walk5-4bada5/delete-impact was polled until it answered deletable:true (16:51:56Z, status: stale), treating an empty response as a retry and never as success.

POST /configs/walk5/delete with all six gates (ack_hosts, ack_reset, ack_purge, confirm_id=walk5 typed, expect_hosts=1) → 303 /configs?flash=deleted.

The hub's leg-by-leg log is the evidence, not the flash:

18:52:12  customer DELETE cascade started for walk5 (journal #15, 1 host(s))
18:52:12  delete walk5: host walk5-4bada5 deleted (escrow DEMOTED to retained custody)
18:52:19  [offsite] deprovisioned shared sub-account 285351 for walk5 (repo data destroyed)
18:52:20  tenantsync: deprovision ok for walk5 (ns=walk5, existed=true)
18:52:20  [claim] reset to unclaimed for walk5 (customer RESET)
18:52:20  delete walk5: residue purged (reports=30 app_telemetry=48 app_log_tails=0
          log_tail_requests=0 notif_prefs=1 selfbind_tokens=1 appliance_registrations=1)
18:52:20  customer DELETE cascade COMPLETE for walk5 (journal #15) — full teardown

existed=true on the PBS leg is the positive observable that it deleted something rather than no-op'ing on an absent namespace. Then, guarded a second time on qm config 325 reading name: walk5-appliance: qm destroy 325 --purge at 16:52:41Z.


After — every layer, each with a positive control that must survive, and does

Layer Item After Positive control
machine VM 325 + 4 disks absent; /mnt/nvme-1tb/images/ empty VM 300 drill-r50 still listed
host demo-hp's own guest 9201 still running
hub hosts, customer_configs, host_escrow, wg_peers, guests, reports, … 0 rows demo-felhom hosts=1 · demo-hp hosts=1 · peti-felhom config=1
off-site Storage Box sub-account 285351 (u629488-sub4) deprovisioned, repo data destroyed (hub log) (see the access note below)
off-site ep0 PBS namespace walk5 gone namespaces demo-felhom, demo-hp both present
network WireGuard peer 10.77.0.5 gone from the LIVE wg show peers .2 .3 .4 .250 all still on wg0

Space returned, measured. c11-scratch used 23 715 036 → 6 262 060 KiB = 17 452 976 KiB (16.64 GiB) against the 17 G measured on disk before deletion. Available 909 638 052 → 927 091 028.

⚠ The WG peer is removed on a SCHEDULE, not by the cascade — check the live interface, not the DB

Immediately after the cascade the hub's wg_peers row was gone while 10.77.0.5 was still on ep0's live wg0. That is not residue: wgsync pushes every ~5 minutes and had last run at 16:51:35Z, 37 seconds before the cascade. The next push — 18:56:35 wgsync: pushed 4 peers — removed it, verified on the live interface at 16:57:07Z.

Recorded because the previous ledger checked this layer after it had already converged and so read as instantaneous. A teardown that checks wg show inside ~5 minutes of the cascade will see a peer that is on its way out, and filing that as a finding would be wrong. The DB row is not the evidence; the live interface is — and it needs a moment.


Residue — counted, not claimed away (R-244)

Full-schema census (every table, every column, like '%walk5%', no truncation): 168 rows before → 67 rows after. 101 removed.

table rows verdict
events 21 by design"provenance/events are NEVER wiped; audit outlives every tier"
notification_log 14 by design
host_deletions 1 by design — the deletion tombstone
customer_resets 1 by design — the journal row (#15)
app_log_issues 30 THE GAP — R-244, exactly as predicted before the run

R-244 grew by 30 rows, as this ledger's own pre-run enumeration said it would. The prediction is the point: the previous ledger's "0 occurrences" came from a narrower query and did not survive a full census. Nothing here claims a clean teardown.


Access note, per the standing rule that a "no access" claim must name its attempts

The Storage Box sub-account could not be verified directly from DooPlex. Tried, with HETZNER_API from ~/.config/credentials: GET /v1/storage_boxes/611421/subaccounts404 not_found; GET /v1/storage_boxes200 with total_entries: 0. That token is scoped to a different project than box 611421 — the hub holds its own. The layer is therefore evidenced by the hub's own [offsite] deprovisioned shared sub-account 285351 … (repo data destroyed), matched on the id recorded at provisioning. Recorded so the next teardown does not re-derive it.


Not deleted, deliberately

The escrow was DEMOTED to retained custody, not destroyed (the cascade's own wording). That is R-198's retention behaving as designed. The customer record itself is gone; what survives is the sealed material under the operator's retained-custody path, plus the audit rows above.

One thing that died with the box

During v0.207.0's live validation a test retrieval passphrase was planted on this appliance, and its settings were mutated (app removed, a drive deregistered, storage_paths emptied). All of it was on the machine destroyed above. The walk5 retrieval passphrase was also compromised earlier — it reached a session transcript, which is R-249 — and it is now gone with the customer record.