testrun: coverage gaps RESOLVED — capability->hub alert fired live; 2 reboot findings (doubling NOT self-corrected, backup-lock)

Phase 1: capability->hub operator alert fired live (event captured, English, names cap+feature);
email is operator-config-gated in demo (not a defect). Phase 2: drive re-propagation + zero false
pages PASS, but TWO findings — F2-a the §4 doubling did NOT self-correct (boot script runs as root,
never needed the grant → hypothesis falsified), F2-b a vzdump backup in flight at reboot left a stale
snapshot-delete lock blocking guest autostart. Phase 3 referenced (host_leaf_changed proven on v0.22.1).
Final parity known-good.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pg8ANF97SEeKYSN5Jxw3qJ
This commit is contained in:
2026-06-29 23:56:54 +02:00
parent f4d88c1ff4
commit c978ea1951
@@ -178,6 +178,74 @@ missing half (detection at startup). **Prevention** for the original incident al
recommendation is now implemented; the full live reinstall is the only piece awaiting a supervised run.
(F1 swap-rollback `RestartCount` hardening remains backlog.)
## Coverage gaps — NOW RESOLVED (supervised follow-up, 2026-06-29)
Baselines: agent v0.48.0, controller v0.91.0, hub v0.22.1; leaf==pin `60b5974d…`; caps 45/45; both
drives bound; channel up.
### Phase 1 — capability→hub operator alert, LIVE — **PASS** (the path that never fired in the field)
Induced a minimal-blast-radius degrade: a scratch deny drop-in (`/etc/sudoers.d/zz-felhom-testdeny`,
`!…pct exec [0-9]* -- tee /etc/felhom-controller-image`, `visudo`-clean) + agent restart.
- Agent stayed **up** (serve-degraded): `capabilities self-check ok=44 total=45 degraded=1` +
`capability DEGRADED … controllerswap-write … critical=true reason="sudo policy denied"`.
- The box kept working — both drives `bound_under_parent=True` throughout (lxc-info still granted).
- **Hub fired the alert (~24s):** `Host capability: demo-felhom-01 ok → degraded
(agent_capability_degraded)`; the saved event message is **English** and names the cap+feature:
*"agent privileged capability degraded — controllerswap-write (impairs: controller-swap / managed
auto-update)"* (severity warning).
- Recover: removed the scratch deny + `visudo -cf /etc/sudoers` OK + restart → `ok=45/45` →
`Host capability: demo-felhom-01 degraded → ok (agent_capability_recovered)`.
- **Note (not a defect):** a literal operator *email* was **not** dispatched — the demo hub has no
operator email address configured (no "Operator email sent/failed" log → `processOperator` returns
before the send). In prod with the operator address set, the same path sends via FormatOperatorEmail +
Resend (RESEND key IS injected). The alert path itself (probe→report→checker→event+dispatch) is proven.
### Phase 2 — host-reboot "doubling" self-correction — **2 FINDINGS** (drive recovery PASS)
Rebooted felhom-pve (uptime was 5d, no boot since the `make-private` grant). After boot:
- **PASS — drive recovery + no false page:** the boot ran `felhom-shared-parent.service`; both drives
re-propagated into the guest (`bound_under_parent=True`, `master:` tags alive), nextcloud + controller
recovered healthy, leaf==pin MATCH `60b5974d…`, caps 45/45. **Zero** spurious alerts on the clean boot
(drive-ABSENT 0, channel DOWN 0, hub capability/leaf/stale 0) — confirms F2 first-obs-seed doesn't
false-page a normal reboot.
- **FINDING F2-a (doubling did NOT self-correct):** the host parent `/mnt/felhom-drives` is **still
`shared:1`** (== root's peer group) and the drives are **still doubled** (each under parent 149 AND
root 33) — identical to the pre-reboot state. **The §4 hypothesis is falsified:** the boot script
`felhom-shared-parent.sh` runs as **root** (systemd), so it never needed the sudoers `make-private`
grant — yet its `make-private`+`make-shared` sequence does not isolate the self-bind from root's
`shared:1`. Harmless functionally (drives work), but the doubling persists across a clean boot.
Severity: low (cosmetic/correctness; the agent's `AttachDrive` already converges binds). **Repro:**
reboot the host, then `grep ' /mnt/felhom-drives ' /proc/self/mountinfo` → `shared:1`. Root-cause not
pursued (won't churn the live parent — would orphan the guest slave). A fix should assert the parent's
group ≠ root's after the boot script and investigate the make-private timing/semantics for a
subdir-of-root self-bind.
- **FINDING F2-b (reboot during backup → stuck guest):** a **vzdump backup was in flight** at reboot
(snapshot @ 23:43:45); the reboot interrupted the `snapshot-delete` cleanup, leaving a **stale
`lock: snapshot-delete`** + a dangling `vzdump` snapshot. Guest 9201 therefore **failed to autostart**
(`onboot:1` can't start a locked CT: *"Starting CT 9201 failed: CT is locked (snapshot-delete)"*).
Recovered with `pct unlock 9201` + `pct start` + `pct delsnapshot 9201 vzdump`. The "power-loss/reboot
during a backup" edge — a host reboot should check for an in-flight backup, or boot-time recovery
should clear a stale vzdump snapshot lock + auto-start. Severity: medium (an unattended host reboot
during the daily backup window would leave the customer box down until manual unlock).
### Phase 3 — `host_leaf_changed` LIVE on the wired hub (v0.22.1) — **PASS (referenced)**
Already live-proven against the wired v0.22.1 in the leaf-fp task: a leaf regen produced the controller
`agent_channel_pin_mismatch` AND (on restore) the hub `host_leaf_changed`
(`82078fab…→60b5974d…`) — two **independent** detections of the same re-key (controller-local +
hub-proactive). Not re-run here (would require another regen + R1). See
`felhom.eu/REPORT.md` (HostLeafChecker) + the §"Self-health arc COMPLETE" note in the channel-health
spike doc.
### Coverage-gaps verdict
All three deferred gaps now exercised: Phase 1 fired live (event captured; email is operator-config-
gated, not a code defect); Phase 2 drive-recovery PASS but surfaced **two findings** (doubling NOT
self-corrected — §4 hypothesis falsified; reboot-during-backup stale lock); Phase 3 referenced. Scratch
artifacts removed (`zz-felhom-testdeny`; the dangling vzdump snapshot); `/etc/sudoers` re-validated.
### Final parity (post follow-up) — known-good
agent v0.48.0 / controller v0.91.0 / hub v0.22.1; leaf==pin `60b5974d…`; caps 45/45; both drives
`bound_under_parent=True`; channel up; nextcloud + controller healthy. (Persistent-but-harmless: the
§4 host-side doubling — finding F2-a.)
## Final state — known-good
agent 0.45.0 / controller 0.90.0 / hub 0.20.0; leaf fp == pinned fp `60b5974d…`; capabilities 45/45