hub v0.137.0 source + burn-down round 2 in felhom.eu: R-277 R-581 R-600 R-544 R-855 R-134 R-92 R-292 R-599 R-725 R-728 (hub), R-819 R-857 R-555 R-364 R-587 (gates/tools), R-571 R-129 R-124-runbook (docs); 28 rows closed incl. catalog + agent v0.147.0 rows, R-350 merged into R-132, R-888 opened, R-887 mechanism (249 -> 222)
gates / gates (push) Successful in 2m3s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 19:43:37 +02:00
parent e0d884565f
commit 557629d2bf
65 changed files with 2680 additions and 122 deletions
+8 -2
View File
@@ -108,8 +108,7 @@ felhom-pve, and **moved off DooPlex**. This page exists because that ruling sat
`disconnected` forever). **That warning is about one storage, not the box.** `/mnt/hdd_1` is also
the `felhom-backup` target and the enrolled user-data drive, so remove scratch storages when done.
- **Forbidden:** do not destroy or unblock **`drill-r50` (VM 300)** — the only drift fixture (R-93). **Measured 2026-09-13 (R-461): `qm list` is EMPTY on both demo-hp and demo-felhom — the VM does not exist anywhere, so the fence currently protects nothing and R-93's premise is gone. The fence stays as written for the day someone rebuilds it; do not read its presence here as evidence the fixture exists.**
(Access: the docs say no baked SSH key and G1 break-glass, but a key authenticated on 2026-07-31 —
**R-129**, unresolved.)
(Access: DooPlex's own key works — `ssh demo-hp`; G1 break-glass is the fallback. `operations/nodes.md`, R-129.)
### `demo-felhom` — N100 · **Tier 0**
@@ -200,6 +199,13 @@ Create and destroy freely **on a Tier 0 host**. Two exceptions: **`drill-r50`**
scratch; and scratch **customers** outlive their VMs in the hub — delete those too, or they accumulate
(`sess-c` and `sess-d` were both left behind before anyone noticed).
**The hub-side teardown waits for the stale threshold — plan for it (R-599).** After the VM is destroyed
the hub still reads its host as ONLINE until the last report is older than `alerting.stale_threshold`
(the deployed value is in `manifests/hub.yaml`, not the code default). Until then both the host delete
and `POST /configs/<id>/delete` answer **409** — by design, a live agent would get 401s for ever. The
409 body names the last report's age and the time deletion opens; wait for that time, then delete. A
409 inside that window is not a failed delete.
#### A fixture may prove a mechanism. Only a fresh box may prove a path.
Rebuilding from scratch every time is waste; reusing a box is legitimate — but not for every claim.