docs: supervisor (03), node_* ruling (08, CONTEXT), per-tier page + tier skip (07), self-bind triggers + PBS-DR lifecycle (05), settings after install (02), park + No TLS Verify runbooks, volunteer prerequisites
gates / gates (push) Successful in 18s

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-15 10:42:53 +02:00
parent d8cd4d4412
commit 4c4e3b3a3f
9 changed files with 109 additions and 1 deletions
@@ -94,6 +94,21 @@ by verb**:
- **An operator signature is always required** to destroy/overwrite any resource holding the only/primary copy of customer data — live-guest destroy, storage detach/wipe, restore-overwrite, decommission — *regardless of whether it arrives as a job or as a desired-state delta*. A compromised hub cannot forge them because the signing key is **not held by the hub** (it lives with the operator / a separate signing path; the hub only queues opaque signed blobs).
- **Data-bearing-ness is agent-internal evidence, never a caller's claim (slice 8C).** For a customer-driven storage op (`POST /disks/format`, §6) the agent **inspects the actual device** (filesystem signature / partition table / partitions / mount, conservative — ambiguous → data-bearing) to decide the class. A blank device → benign self-serve `mkfs`; a data-bearing device → `ClassStorageWipe` → this gate → `pending_signature`. The **destructive completion of a data-bearing wipe is slice 10** (the operator-signed path); 8C refuses it. This mirrors the provenance rule above: just as the scratch tag is agent-internal (never hub-sourced), data-bearing-ness is agent-observed (never controller-asserted) — a compromised controller cannot relabel a data-bearing drive "blank" to walk the gate.
- **Healing a crashed controller is non-destructive by construction:** it is reconstructable from its image + the guest's persistent volume, so "redeploy" = restart the LXC / `docker compose up -d` **inside the existing guest** — never a guest destroy. (v0.33 precedent: `watchdog.go` restarts stopped stacks, it never destroys the guest.)
- **The controller supervisor is that sentence made real (R-523, agent v0.131.0).** Measured
2026-09-15: after `docker kill`, Docker restarts neither an `unless-stopped` nor an `always`
container, and the golden's `felhom-controller-bootstrap.service` is a oneshot that watches nothing.
So every 30 s the agent checks, for each felhom-pool guest it provisioned
(`/var/lib/felhom-agent/guests/<vmid>/bootstrap` exists) that is running, whether
`felhom-controller` is running; on the **second** consecutive "no" it runs
`systemctl restart felhom-controller-bootstrap.service` in the guest — the swap's own restart, over
the same two sudoers grants. **Guards:** not during a controller swap; not when parked
(`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on the host); not on a stopped,
locked or vzdump-busy guest; not on an unknown docker answer; and no thrash — 3 restarts in 15 minutes
stop restarts for 30 minutes. **Events:** the agent has no event channel; its host report carries a
`controller_supervisor` stanza, and the hub mints `controller_restarted_by_agent` (info) and
`controller_crashloop` (error), both operator-only, keyed on timestamps so an agent restart can
neither lose nor invent one. New goldens run the controller with `--restart always`, which covers
only a Docker daemon restart.
Signed payloads carry a **nonce + expiry** (anti-replay: a captured "restore" job cannot be
re-injected later) and a target binding (host + guest id) so a signature can't be retargeted.