hub v0.137.0 source + burn-down round 2 in felhom.eu: R-277 R-581 R-600 R-544 R-855 R-134 R-92 R-292 R-599 R-725 R-728 (hub), R-819 R-857 R-555 R-364 R-587 (gates/tools), R-571 R-129 R-124-runbook (docs); 28 rows closed incl. catalog + agent v0.147.0 rows, R-350 merged into R-132, R-888 opened, R-887 mechanism (249 -> 222)
gates / gates (push) Successful in 2m3s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 19:43:37 +02:00
parent e0d884565f
commit 557629d2bf
65 changed files with 2680 additions and 122 deletions
@@ -224,6 +224,12 @@ bake's own acceptance check. The real guard was never weak: the script's own
never deleted by a bake (the publish step's pre-delete targets only its own version), so rolling
back is a form submission, not a rebuild.
**A re-bake under the SAME version: re-vouch at once (R-857).** The publish pre-deletes and re-uploads
that version's package, so the hub keeps vouching the FIRST bake's sha, which the registry no longer
holds — a fresh install in that window fails its sha check. Re-select the golden and Save right after
the publish. Put the evidence in `golden-<VER>-<DATE>-rebake/`; `golden_currency_gate.py` reads the
newest bake of a version (a later date, then the suffixed directory of the same date).
A golden is only needed when a publish train wants fresh installs current — demo deploys never need it.
**Since 2026-09-13 the cadence is §4.2: weekly, and before any drill or fresh install.**
The full 0.188.0 run, with the observables: `documentation/audits/tester-gate-golden-0.188.0-2026-07-31.md`.
@@ -43,6 +43,10 @@ recovered with the household's recovery code — the same as restoring from ep0)
`/etc/pve/priv/storage/<id>.enc`, then append the storage entry to `/etc/pve/storage.cfg`:
`pbs: <id>` / `datastore ep0-copy` / `server <DooPlex>` / `content backup` / `fingerprint <DooPlex PBS cert>` /
`namespace <customer>` / `username root@pam!<name>`.
**The namespace comes from the hub's Recipe** (`pbs.namespace`, read WITH `pbs.namespace_state`): `resolved` and a
name → `namespace <name>`; `resolved` and an **empty** value → the ROOT namespace: write **no** `namespace` line (and
pass no `--ns` to `proxmox-backup-client`). Agents before v0.147.0 wrote the word `root` there — no namespace has that
name, so treat a recorded `root` as empty (R-124). `unknown` → read the namespace off the copy's `ns/` folder above.
**Do not use `pvesm add pbs` without `--password`:** it validates with the password from its command line, fails 401,
and on failure DELETES the `.pw`/`.enc` files you placed (measured). Passing `--password` puts the token on argv.
3. `pvesm list <id>` → the household's snapshots (measured: 2 s). Then restore with the safe script (R-834):
+8 -2
View File
@@ -108,8 +108,7 @@ felhom-pve, and **moved off DooPlex**. This page exists because that ruling sat
`disconnected` forever). **That warning is about one storage, not the box.** `/mnt/hdd_1` is also
the `felhom-backup` target and the enrolled user-data drive, so remove scratch storages when done.
- **Forbidden:** do not destroy or unblock **`drill-r50` (VM 300)** — the only drift fixture (R-93). **Measured 2026-09-13 (R-461): `qm list` is EMPTY on both demo-hp and demo-felhom — the VM does not exist anywhere, so the fence currently protects nothing and R-93's premise is gone. The fence stays as written for the day someone rebuilds it; do not read its presence here as evidence the fixture exists.**
(Access: the docs say no baked SSH key and G1 break-glass, but a key authenticated on 2026-07-31 —
**R-129**, unresolved.)
(Access: DooPlex's own key works — `ssh demo-hp`; G1 break-glass is the fallback. `operations/nodes.md`, R-129.)
### `demo-felhom` — N100 · **Tier 0**
@@ -200,6 +199,13 @@ Create and destroy freely **on a Tier 0 host**. Two exceptions: **`drill-r50`**
scratch; and scratch **customers** outlive their VMs in the hub — delete those too, or they accumulate
(`sess-c` and `sess-d` were both left behind before anyone noticed).
**The hub-side teardown waits for the stale threshold — plan for it (R-599).** After the VM is destroyed
the hub still reads its host as ONLINE until the last report is older than `alerting.stale_threshold`
(the deployed value is in `manifests/hub.yaml`, not the code default). Until then both the host delete
and `POST /configs/<id>/delete` answer **409** — by design, a live agent would get 401s for ever. The
409 body names the last report's age and the time deletion opens; wait for that time, then delete. A
409 inside that window is not a failed delete.
#### A fixture may prove a mechanism. Only a fresh box may prove a path.
Rebuilding from scratch every time is waste; reusing a box is legitimate — but not for every claim.