Files
felhom-agent/REPORT.md
T

16 KiB
Raw Blame History

REPORT — TASK-D Part 3: the guest-network watchdog (R-54) · felhom-agent v0.91.2 → v0.92.1

Date: 2026-07-21 · Trunk, pushed to main. Baseline: 08b55a1 (clean, == origin/main). Deployed, running on felhom-pve, and STOP-2 RAN — the incident was replayed and the watchdog prevented it (§6b). Every claim below was observed.


1. What this closes

INCIDENT-guest-dhclient-killed-2026-07-20.md §5, "OPEN RISK": the guest's DHCP client is unsupervised. ifupdown starts it once at boot and nothing restarts it. When it was killed on 2026-07-20 the guest kept working for another ~80 minutes on its unexpired lease; only at expiry did the address and default route vanish, taking the Cloudflare tunnel, hub reports, catalog sync and the controller→agent channel with them — a 1h15m outage in which every observable signal said healthy for the first 80 minutes.

So the design consequence is the whole feature: liveness of the DHCP client is itself a probe. The watchdog flags a DHCP guest unhealthy on pgrep -x dhclient alone, while the address and route are still perfectly present. Waiting for the IP to disappear is waiting out exactly that silent window — and the red-proof reproduces it (§5).

Host tier is not a preference: a guest with no default route cannot repair its own default route.


2. Shipped

internal/guestnet (probe.go / watchdog.go / report.go), built on the wg-tunnel + storage watchdog loop shape, started with go wd.Watch(ctx) like selfheal.

  • Four fixed-shape pct exec probes, all constant argv + the numeric vmid: address, default route, /etc/network/interfaces mode, pgrep -x dhclient. No shell anywhere; no guest-supplied data is ever interpolated into a command.
  • Heal = the incident's restored invocation, verbatim, logged at INFO before it runs: pct exec <vmid> -- dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0 A test pins that argv element by element.
  • Dampers (this runs a privileged command inside a customer's container, so it is built to under-act): two CONSECUTIVE bad probes before any heal, ≥10 min between heals per guest, ≤3 heals/hour, and observe-only while the guest — or the agent itself — has been up under 3 minutes.
  • Refuses to act on: a static guest (dhclient must never fight a static config; a static guest missing its address is reported loudly and left to R-50), an unknown interface mode, a guest it cannot probe, and an ownership-unproven guest list. The guest source is the pool-verified ListLXC ∩ felhom-pool (audit A1) — never a bare ListLXC, which under a broad token would run dhclient inside a co-tenant's container.
  • A failed PROBE is never a dead client. pgrep exits 1 with EMPTY stderr on no-match; anything on stderr means the probe itself failed → unknown. Without that rule a missing pgrep would heal forever.
  • Healthy cycles log a Debug line. v0.91.2's lesson, one day old: if the quiet path is silent, "no alarms" and "never probed" are the same evidence.
  • Not in the errc fan-out — a watchdog over customer guests must never be able to terminate the agent. A test asserts that, because joining the fan-out would also make the shutdown drain bound off by one.
  • Report block: GuestNetStatus on HostReport (guest_net, omitempty), additive and stored opaquely hub-side like pbs_dr / wireguard. No hub code was touched.

Two deliberate deviations from TASK-D, both stated up front:

  1. GuestNetStatus, not WireGuestNet. In this repo Wire* is the DOWN direction (WireDesiredState / WirePBSDR — what the hub sends the agent); UP-direction report stanzas are *Status. WireGuestNet on HostReport would have been the only report block named against the convention.
  2. A sudoers change was required — see §4. The brief said none was needed.

Config guest_net is this repo's first default-ON gate. Every other gate defaults to false because those features reach outward (an offsite endpoint, an OOB tunnel) and enrolling a box by an update would be wrong. This one looks only inward at guests the agent already owns, and the failure it prevents exists on every box today. A watchdog that must be remembered per box is a watchdog that is missing on the box that needed it. Opt-out is explicit: "guest_net": {"disable": true}.


3. Phase-0 probes

P3 (watchdog ground truth) — DONE 2026-07-21, live from guest 9201. These exact bytes are the parser fixtures, including the literal backslash ip -o emits and the trailing space on the route:

ip -4 -o addr show dev eth0 → 2: eth0    inet 192.168.0.104/24 brd 192.168.0.255 scope global dynamic eth0\       valid_lft 4916sec preferred_lft 4916sec
ip route show default       → default via 192.168.0.1 dev eth0
pgrep -x dhclient           → 235839                       (rc=0; rc=1 + EMPTY stderr when absent)
ps -o args= -C dhclient     → dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0
/etc/network/interfaces     → iface eth0 inet dhcp
pct exec <bad vmid>         → rc=2, stderr "Configuration file 'nodes/demo-felhom/lxc/9999.conf' does not exist"

The live ps line confirms the running client's argv is byte-identical to the incident's restored invocation — i.e. the heal reproduces the guest's own boot-time command, not an approximation. A docker-bridge-only route table is also pinned as a negative (172.17.0.0/16 dev docker0 … must never read as a default route — those were the exact leftovers in the incident).

P4 (config surface + report pattern + loop precedent) — DONE by reading the tree; the wg loop's interval/damping/logging shape and the pbs_dr stanza's collector-seam pattern are what this copies.


4. The live finding: three of four probes had no sudoers grant

The agent runs non-root; Privileged.Mode=sudo fails closed with no prompt. The first sweep after deploying v0.92.0 logged:

level=WARN msg="guestnet: guest network not actionable — reporting only" vmid=9201
  state=unknown mode=unknown has_ip=true has_route=false
  detail="dhclient liveness probe failed: sudo: a password is required"

The watchdog behaved exactly as designed — it reported unknown and healed nothing rather than acting blind — but it was blind. The existing allowlist granted only lanresolver's address read (pct exec [0-9]* -- ip -4 -o addr show dev eth0), which is why has_ip=true while route, mode and liveness all failed.

Fixed in v0.92.1: a FELHOM_GUESTNET alias with four fixed vectors (route, interfaces, pgrep, heal). Every argument after the numeric vmid is a literal, so nothing the guest or the hub says can widen the grant. The address read is not duplicated — it stays FELHOM_DNSMASQ's; one command, one grant.

Plus four guestnet-* capability rows, so a host that has not taken the new sudoers file is VISIBLE as degraded instead of silently watchdog-less. Deliberately non-critical: a missing grant must not page an operator for every box on rollout day (the R-50b amber-fleet lesson).

v0.92.0 is superseded, not overwritten — do not vouch it. It was published before this was found, so its binary lacks the capability rows and its release lacks the sudoers file. A published version stays immutable (the v0.91.0 → v0.91.1 precedent).


5. Tests and red-proofs

Green gate: go build ./... && go vet ./... && go test ./... — all packages ok except the known flake TestGenerateRecoveryCode_EntropyAndFormat (internal/escrow), which fails when the wordlist yields a hyphenated word (drop-down → 11 tokens instead of 10). Confirmed pre-existing: internal/escrow has not been touched since v0.88.0 and this task changes nothing there; observed 3/8 runs, consistent with the documented ~1/5.

New: internal/guestnet/watchdog_test.go (16 cases), internal/hub/collect_guestnet_test.go (2), cmd/felhom-agent/guestnet_wiring_test.go (2).

# Red-proof Mutation Result
E detect on process liveness, not address presence classify's DHCP arm reverted to IP-presence-only FAIL ×5. The decisive one: classify = "healthy", want "unhealthy" for the July-20 fixture, detail="address, default route and dhclient all present", and heal ran 0 times. That is the 80-minute silent window, reproduced exactly. Restored, green.
W the watchdog must be wired SetGuestNetReporter and go gnWatchdog.Watch(ctx) both commented out FAIL with both reasons named — "the guest_net stanza would never reach the hub (the exact v0.91.0 inert-seam defect)" and "it would be constructed, reported on, and never probe anything". Restored, green.

Every damper is asserted as an exact count, and the load-bearing assertions are the negatives — a static guest, an unprobeable guest, a boot-race guest (young guest AND young agent), a failed probe tool and an ownership-unproven guest list must each record zero heal calls. The ceilings are driven by an injected clock over a scripted 10 hours of permanent failure: ≤30 heals total, and never a second heal inside the 10-minute cool-off.

Seam discipline (§9 rule 6) — three production-path tests: the guest_net stanza is asserted through the real Collect (and asserted ABSENT from the wire when no reporter is wired, so "not wired" and "found nothing" can never look identical); and the main.go wiring is an AST walk for the construction, the reporter call and the started goroutine. The AST form is deliberate — a strings.Contains version of the twin test in felhom-controller passed its own red-proof, because a commented-out call still contains the string.


6. Live validation (method: journald + capability self-check on felhom-pve)

Step Evidence
Publish 0.92.0 AGENT_SHA256=b1302790d412d22e969936ff52e3ee33e3edc111b8426364a21cdb1c5127ca6a, round-trip GET verified — superseded, do not vouch
Publish 0.92.1 AGENT_SHA256=7424bc1c3c533eff9157e15a18d4635c624931f5a479a48126de77a94e6a3d4d, round-trip GET verified
Deploy visudo -c parsed OK → sudoers installed 0440 root:root (backup /root/felhom-agent.sudoers.bak-preR54) + binary installed (backup felhom-agent.bak-0.91.2-preR54) → felhom-agent --version = 0.92.1, service active
Capability self-check ok=68 total=68 degraded=0 inactive=0 (was 64/64 before the four guestnet-* rows) — the sudoers grant is proven from the agent's own side, not assumed
Watchdog start INFO guestnet: watchdog starting interval=1m0s min_heal_interval=10m0s max_heals_per_hour=3 settle=3m0s
Healthy cycle level=DEBUG msg="guestnet: guest network healthy" vmid=9201 mode=dhcp has_route=true dhclient_alive=true (12:34:15 CEST)
Default-ON proven /etc/felhom-agent/agent.json has no guest_net key at all — the watchdog runs on defaults, which is the whole point of the inverted gate

Note for the operator: log_level on felhom-pve was temporarily raised to debug to capture that Debug line (backup at /root/agent.json.bak-debuglevel). It is still debug, deliberately, so STOP-2's heal chain is visible in journald. Revert it to info after STOP-2.


6b. STOP-2 — the incident replay (operator-present, 2026-07-21)

The 2026-07-20 kill, repeated deliberately. The /proc/<pid>/cgroup check the incident produced was applied before the kill — the script refuses unless the pid's cgroup is guest 9201's, which is the rule that would have prevented the original outage:

GATE OK — pid 336708 belongs to guest 9201 (0::/lxc/9201/ns/.lxc)
KILL at 2026-07-21T10:43:18Z / 12:43:18 CEST
after kill: no dhclient running; address 192.168.0.104/24 STILL PRESENT (valid_lft 4998s); default route STILL PRESENT

That second line is the whole point: the box looked perfectly healthy, with ~83 minutes of lease left before any symptom would appear.

Time (CEST) Event
12:43:15 DEBUG guest network healthy … dhclient_alive=true — last good cycle
12:43:18 kill -9
12:44:15 detected in 57 s, on process liveness aloneunhealthy (first bad probe — not acting yet) bad_probes=1 required=2, detail "dhclient is not running — the lease will not be renewed (the 2026-07-20 failure mode; address still present, renewal already dead)"
12:45:15 second consecutive bad probe → healing, then running heal command … cmd="pct exec 9201 -- dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0"
12:45:18 guest network healed ip=192.168.0.104 has_route=true dhclient_alive=true heals_last_hour=1

Healed 120 seconds after the kill — roughly 80 minutes before the outage would have begun. The strongest evidence is therefore what did not happen, verified from inside the guest:

Check Result
cloudflared Up 29 hours — the tunnel never dropped, never reconnected
DNS gitea.dooplex.hu OK
hub 302 in 0.16 s
https://felhom.demo-felhom.eu/ 302 in 0.25 s
new dhclient pid 2043719, cgroup=0::/lxc/9201/ns/.lxc, argv byte-identical to the original

On 2026-07-20 every one of those was dead for 1h15m. This time the outage was prevented, not detected.

The negative leg — 30 healthy minutes, run in full. From the heal at 12:45:18 to 13:16:15: 30 healthy Debug cycles (exactly one per minute — the cadence is precise), 0 heals, 0 WARN, 0 ERROR. So the damper is not a mute and the quiet path is not silence: "no alarms" stayed continuously distinguishable from "not probing", which is the v0.91.2 lesson holding in production.

An unplanned damper proof, against a REAL transient. STOP-1's guest reboot at 12:53 caught the guest mid-boot:

12:53:16 INFO  guest network unhealthy (first bad probe — not acting yet) detail="no IPv4 address on eth0" bad_probes=1
12:54:15 DEBUG guest network healthy …
12:54:15 INFO  guest network recovered previous_state=unhealthy

One bad probe, no action, then recovery. The two-consecutive-probes rule and the boot-race guard did exactly what they exist for — a booting guest was not injected with a dhclient — and this was a genuine transient, not a scripted one.

6c. What could NOT be verified here

The guest_net stanza on the wire. --selftest=hub builds its own one-shot collector and never wires the guestnet reporter, so it would print a misleading absence — using it as evidence would be worse than having none. (That divergence between the selftest collector and the daemon's is itself worth a small fix; the same comment already exists in the code for the pbs reporter.) The stanza is proven through the production Collect path by test, and SetGuestNetReporter is proven called by the AST wiring test, but seeing it in a real hub report needs the operator UI.

7. Deliverables

  • c0966d7 — v0.92.0: the watchdog, the report block, config, tests.
  • 0e8fd81 — the sudoers grant + capability rows (the live finding).
  • 98adb72 — v0.92.1: supersede + version bump.
  • Published: felhom-agent 0.92.1 / 7424bc1c3c533eff… (0.92.0 superseded).
  • Live on felhom-pve: binary 0.92.1 + the new sudoers file.

8. Operator actions outstanding

  1. STOP-2DONE 2026-07-21, passed (§6b).
  2. STOP-3 — Day-0 manifest Agent → 0.92.1 / 7424bc1c3c533eff9157e15a18d4635c624931f5a479a48126de77a94e6a3d4d (wrapper sha unchanged — configs/felhom-pbs-apply was not touched). MinAgent is the operator's call; 0.92.1 is safe once felhom-pve has converged, which it has.
  3. Revert log_level to infoDONE (13:17 CEST, after the quiet window closed; agent restarted clean, caps 68/68 ok, degraded=0, watchdog back up). The temporary raise is recorded here only so the journald volume change is explainable; /root/agent.json.bak-debuglevel remains as the pre-change copy.