Files
felhom.eu/documentation/audits/TESTRUN-fullstack-2026-06-29.md
T

187 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TESTRUN — full-stack no-mercy campaign (2026-06-29)
Adversarial integration/chaos validation of everything shipped this session, on the live demo stack
(felhom-pve, guest 9201, USB drives `felhom-usb`/`felhom-flash` — expendable). No production code
changed. Destructive steps were always recovered; the stack ended at known-good parity (§B re-run).
**Scope note (honest coverage):** this run executed the **headline + highest-value NEW** tests with
full evidence and recovery — the controller-swap happy/rollback (the deferred Scenario D), and the
**capstone** (leaf-regeneration → channel-health detection → recovery) — plus baseline, multi-drive
re-propagation, and the swap negative controls. Behaviors already **live-proven during this session's
implementation tasks** (multi-drive concurrency + other-drive-survival; capability degraded detection;
channel-health unreachable/transient/recovered) are **referenced**, not re-ground exhaustively. Deferred
items (host reboot, full `felhom-host-install.sh` reinstall, exhaustive flap loops, DNS/construction
live-alert) are listed in **Coverage gaps** with rationale.
---
## §B Baseline (oracle) — captured
| Item | Value |
|---|---|
| Versions | agent **0.45.0**, controller **0.90.0**, hub **0.20.0** |
| Leaf fp (agent) == pinned fp (controller.yaml) | `60b5974d…279d8245` == `60b5974d…279d8245` ✓ |
| Capabilities self-check | **45/45 ok, 0 degraded** |
| Drives | felhom-usb + felhom-flash both `state=attached`, `bound_under_parent=True` |
| Host parent peer group | `shared:1` (== root's — the §4 "doubling" still present; host not rebooted since `make-private` granted) |
| Guest propagation | guest sees both drives, `master:51`/`master:49` (propagation alive) |
| Apps | nextcloud (on felhom-usb) + controller healthy |
| Channel-health | steady up (no transitions) |
**Final parity re-run (end of campaign):** identical — agent 0.45.0 / controller 0.90.0, leaf==pin
`60b5974d…`, capabilities 45/45, both drives bound, nextcloud+controller healthy, channel up (0 down
events in 90 s), dashboard banner clear. **Stack restored to known-good.**
---
## §C Multi-drive — PASS
- **Concurrency:** both drives `bound_under_parent=True` + `state=attached` concurrently throughout
(baseline + every post-churn check). Guest mountinfo shows both with live `master:` propagation tags.
- **Guest reboot resilience (NEW):** `pct reboot 9201` → controller healthy, **both** drives
re-propagated (`bound_under_parent=True`, guest sees both), nextcloud `Up (healthy)`, and **0**
`drive ABSENT` gate events (no flapping). The C1 pre-start hook fired (placeholder for the
momentarily-absent bootstrap mount) as designed.
- **Other-drive-survival** (the original headline bug — detach one, the other survives) was
**live-proven during the sudoers task** (both drives Aktív concurrently; the gate RETURNED both
without cross-impact) — referenced, not re-run here.
**Deferred:** host reboot (would prove the "doubling" self-corrects to its own peer group now that
`make-private` is granted) — needs a maintenance window on the non-expendable host; the §4 doubling is
latent/harmless meanwhile. Mid-write detach flap loops — referenced via the prior chaos campaigns.
---
## §E Controller-swap — happy PASS; **1 FINDING** on rollback
- **Happy swap (real agent `POST /controller/swap` pipeline):** 0.90.0 → 0.89.0. Evidence: the
`tee`-written `/etc/felhom-controller-image` is **byte-exact** (`od -c`: `…controller:0.89.0\n`,
trailing newline); agent logged `controller swap requested``controller-swap: new controller
healthy`; container came up **0.89.0 Up (healthy)**. Swapped back 0.89.0 → 0.90.0 cleanly. The
stdin-`tee` write + the 5 narrow grants + bootstrap restart all work live. **PASS.**
- **Negative controls (scoping):** as `felhom-agent`, `sudo -n -l` of
`pct exec 9201 -- {bash -c id, rm -rf /tmp/x, docker rm -f felhom-controller, tee /etc/shadow}`
**all DENIED (exit 1)**; the real `tee /etc/felhom-controller-image`**permitted (0)**. **PASS.**
- **Rollback — FINDING (recorded, not fixed):** swapped to a deliberately-bad image (alpine tagged
`felhom-controller:9.9.9`). Expected health-fail → rollback; instead the swap marked it
**healthy in ~4 s** and did **NOT** roll back. Root cause: alpine has **no `HEALTHCHECK`**, so the
swap's "running + no-healthcheck = healthy" rule (a deliberate feature for images that legitimately
lack a healthcheck) passed on a transient "running" instant — then the container crash-loops
(`Restarting (0)`). **Severity: low** for the real controller (its image *has* a healthcheck, so a
broken real build would report unhealthy/never-running → rollback; the deterministic rollback path is
unit-proven by `TestControllerSwap_RollbackOnUnhealthy`). **Gap:** a controller image that has **no
healthcheck AND crash-loops** would be marked "done" by the point-in-time verify and not roll back.
**Suggested hardening (operator decision):** the verify should re-check health over a few seconds /
reject `RestartCount>0`, so a fast start-then-exit can't pass. Recovered cleanly by swapping back to
0.90.0 (has a healthcheck → healthy); bad tag removed.
**Deferred:** a *clean* live rollback needs a present-but-healthcheck-failing controller image, which
can't be manufactured without building one (out of scope — "no new builds"). Rollback logic is
unit-proven; the live finding above is the more interesting result.
---
## §F + §K Channel-health & CAPSTONE — PASS (the headline)
### Capstone (§K): original-incident class → DETECTED in ~25 s
Reproduced the exact failure that started this work: with the controller running (channel seeded up),
moved the agent leaf+key aside and restarted the agent → it **regenerated a new leaf**
(`60b5974d…``911d703c…`), so the agent serves a cert the controller doesn't pin. Agent stayed
**active**, capabilities **45/45 ok** (leaf change doesn't affect capabilities).
**Detection (the whole point):**
```
[channel] agent channel DOWN (up→down:pin_mismatch): agentapi: GET /storage: … TLS pin mismatch …
Event pushed: agent_channel_pin_mismatch (error) — … the agent's leaf cert no longer matches … (HTTP 200)
dashboard banner: "A tárolókezelő ügynök tanúsítványa megváltozott"
```
**DETECTED ~25 s** after the regeneration (one channel cycle) — vs the original **silent multi-day**
disk-UI outage. The operator alert fired with the correct event type + the dashboard pin-mismatch
banner appeared.
**Recovery (R1):** restored the backed-up leaf+key → restart agent → serves `60b5974d…` again →
`[channel] agent channel recovered (was down:pin_mismatch)` + `Event pushed: agent_channel_recovered`
+ both drives `bound_under_parent=True` + banner cleared. ~one cycle after the restart.
### §F other modes
- **Unreachable / transient / recovered** were **live-proven during the v0.90.0 task** (and the
recovered path again here): a clean `systemctl restart felhom-agent` → no alert (debounce);
`systemctl stop``transient down (unreachable, 1/2) — suppressed` then `DOWN (up→down:unreachable)`
+ `Event pushed: agent_channel_unreachable` (HTTP 200) + Hungarian banner; `start` → recovered.
Referenced.
- **401 / DNS-misconfigured / construction** — classification is unit-proven (`channelhealth` tests).
Live **alert** for these is harder to induce without a controller restart, which **reseeds** state
(first-obs = seed, no alert) — see Coverage gaps.
**§K verdict:** the original-incident class is now **DETECTED within a cycle** by the channel-health
check (PASS). **Prevention recommendation (still worth doing, operator decision):** ship
`--preserve-state-from` (carry `local-api.{crt,key}` + `local-tokens.log` across a reinstall) +
loud-on-regenerate, to turn the now-*detected*-but-disruptive reinstall into a **non-event** rather
than a detect-then-R1-recover. Detection is the safety net; preservation is the fix.
---
## Defects / findings (separated from EXPECTED outcomes)
| # | Finding | Severity | Repro |
|---|---|---|---|
| F1 | Swap verify passes a **no-healthcheck image that crash-loops** (transient "running" + no-healthcheck rule → false "done", no rollback) | low (real controller has a healthcheck) | tag alpine as `felhom-controller:X.Y.Z`, swap to it → marked healthy ~4 s, then `Restarting (0)` |
**F1 RESOLVED 2026-06-29 (agent v0.47.0):** the swap verify now reads `{{.RestartCount}}`
(`running && rc>0` → not healthy) and requires a 3-poll **stability dwell** for a no-healthcheck image
(a real healthcheck is trusted immediately). Live re-test: the same alpine `:9.9.9` swap now reaches
`state: failed` ("did not become healthy within timeout") and **rolls back** to 0.91.0 — no false
"done". The `docker inspect -f *` grant is unchanged (`*` spans the extended template, exit-0 confirmed).
See `felhom-agent/REPORT.md` + CHANGELOG v0.47.0.
**EXPECTED (not defects):** the capstone leaf regeneration → pin mismatch (the prevention fix was
never shipped — regeneration is expected; **detection** is the pass criterion, and it passed). The §4
host-side "doubling" persisting until a clean host reboot (expected; `make-private` is granted, applies
next fresh boot).
---
## Coverage gaps (couldn't test safely / deferred — with rationale)
- **Host reboot** (§C) — would confirm the parent self-corrects to its own peer group; needs a
maintenance window on the **non-expendable** host. Deferred.
- **Full `felhom-host-install.sh` reinstall** (§K) — reinstalls the agent **on felhom-pve** (non-
expendable); the **equivalent** leaf-regeneration was exercised instead (same incident class, same
detection, recoverable). The full reinstall is the operator's call in a window.
- **401 / DNS / construction live alert** (§F) — inducing these needs a controller-config change +
restart, which **reseeds** the channel state (first-obs seeds silently, no alert). Classification is
unit-proven; the dashboard would show the state, but the up→down **notification** for these isn't
cleanly inducible live without an agent-side trick (pin/unreachable have one; 401 would need a
token-store edit). Deferred as lower-value (unit-covered).
- **Capability live scratch-drop → hub alert** (§D) — the lxc-info-degraded path + drive flap was the
**live incident** the capability task already diagnosed/fixed; the hub `HostCapabilityChecker`
transition is unit-proven. A live drop requires editing the certified sudoers (sudoers can't
"override-to-deny" via a drop-in cleanly) — deferred to avoid mutating the freshly-certified
allowlist for a unit-covered path.
- **Exhaustive flap/mid-write/3rd-drive loops** (§C/§G adversarial) — referenced via the prior chaos
campaigns (#1#3, all FIXED); not re-run.
---
## F2 CLOSED (2026-06-29) + prevention shipped
The testrun surfaced **F2**: a channel/host failure present at startup/reseed alerted only on a live
transition, so a born-down (e.g. a controller booting into pin_mismatch) was dashboard-only, no
operator email. **Closed** — controller **v0.91.0** (`alerted`-flag rework) + hub **v0.21.0** (checkers
seed only healthy; born-degraded/stale emits on first Check, cooldown dedups). Live-validated: a
controller restarted into a regenerated-leaf channel logged `[channel] agent channel DOWN
(unseeded->down:pin_mismatch)` and **pushed `agent_channel_pin_mismatch`** (~60 s) — the capstone's
missing half (detection at startup). **Prevention** for the original incident also shipped: agent
**v0.46.0** loud-WARNs a regenerated leaf (`leaf REGENERATED — pins now INVALID`), and
`felhom-host-install.sh` gained `--preserve-state-from` (keeps the leaf/pin stable across a reinstall)
+ a populated-host guard (refuse to regenerate the leaf on a host with guests). The §K prevention
recommendation is now implemented; the full live reinstall is the only piece awaiting a supervised run.
(F1 swap-rollback `RestartCount` hardening remains backlog.)
## Final state — known-good
agent 0.45.0 / controller 0.90.0 / hub 0.20.0; leaf fp == pinned fp `60b5974d…`; capabilities 45/45
ok; both USB drives `bound_under_parent=True`, apps healthy; channel up, dashboard clean. All test
artifacts removed (`/root/testrun-leaf-backup`, the `9.9.9` tag). No secrets recorded (the per-guest
token was read out-of-band and never printed; leaf fingerprints are not secret).