Files
felhom-controller/REPORT.md
T
admin 92d670e8a6 controller v0.120.0: dead-app alerting (fix-3) + ring revision (fix-6) — CLOSES CAMPAIGN-3 (docs+CHANGELOG+REPORT+CONTEXT)
fix-3: deadapp-check job -> self-clearing WARN banner + one app_start_failed hub
event per running->down transition. fix-6: ring 1000->5000, periodic spam->TRACE
(ring-dropped), atomic SSD spill/load across restart. Live-validated on 9201+hub 0.48.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CDMFpFx84pfviCTVuGGhf
2026-07-12 10:31:12 +02:00

4.4 KiB

REPORT — v0.120.0: dead-app alerting (fix-3) + debug-ring revision (fix-6) — CLOSES CAMPAIGN-3

Date: 2026-07-12 · Version: controller v0.120.0 (from v0.119.0) · MinAgent: 0.81.0 (UNCHANGED) · Pairs with: hub v0.48.0 (accepts app_start_failed) · Deployed: guest 9201 (0.120.0 healthy). Source: felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md.

What shipped

  • fix-3 (MED) — a dead deployed app is loud. A deadapp-check job (every 30 s, after a 90 s boot grace) scans the deployed apps: stopped/exited ones (stacks.IsDownState) raise a self-clearing WARN dashboard banner (grouped above 3) and fire an app_start_failed hub event ONCE per running→down transition (Notifier.NotifyAppStartFailures; the hub owns cooldown — no controller timer).
  • fix-6 (MED) — the post-incident window survives. Ring cap 1000→5000 (display cap raised to match); periodic scheduler/refresh success lines demoted to a [TRACE] level the ring drops at write-time (failures never TRACE); atomic JSON-lines spill to <DataDir>/debug-ring.log (SSD only) every 30 s + on shutdown, loaded back on boot. Corruption-safe.
  • hub v0.48.0: app_start_failed added to allowedEventTypes + customerMessages (else the event 400s at ingest — the known allowlist gotcha).

Tests + red-proofs (all green)

  • fix-3: dead-app banner present + self-clears (companion: skip SetDeadAppAlerts → silent → fail); grouped above threshold; IsDownState table (stopped/exited down; starting/unhealthy/restarting/ deploying/paused/unknown NOT); notifier one-event-per-transition (companion: drop tracking → fires each cycle → fail); first-seen-down (dead-at-boot) fires; healthy never fires.
  • fix-6: TRACE dropped from ring while a WARN failure is kept (companion: demote too broadly → failure lost → fail); spill→load round-trip; load keeps newest N; corrupt/truncated spill loads valid + never panics; missing file no-op.

Live validation (demo 9201 + hub)

  • fix-3: docker stop seerr → within a cycle the dashboard banner "Telepített alkalmazás nem fut: Jellyseerr (stopped)" appeared AND the hub logged exactly ONE app_start_failed event (Event from demo-felhom: app_start_failed); across 3 down-cycles still ONE event (anti-spam); docker start seerr → the banner self-cleared (0 banners). This makes the campaign's 4-hour silent death impossible.
  • fix-6: the ring held 0 periodic-spam lines (TRACE demotion); a controller restart PRESERVED the pre-restart window — oldest entry unchanged across systemctl restart (total 390→464, oldest 08:11:28 both sides), with a 63 KB spill on the persistent SSD data volume (survives container recreation, never a NAS path).

CAMPAIGN-3 finding ledger — RESOLVED

Finding Fix Ships in
F12 (CRIT) boot ordering cycle; F11/F10/F9 (HIGH) automount re-arm; F2/F1 residue agent boot/recovery plane + appliance self-heal agent v0.85.0
F7 (HIGH) in-place volume-dump truncation; F6/F5 (LOW) single-copy + stale dirs atomic dumps + tier-2 for volume-only + stale sweep controller v0.118.0
F8 (MED) storage-health contradiction; F4 (LOW) mapped_uid share-row classifier fusion + range check controller v0.119.0
fix-3 (MED) silent dead app; fix-6 (MED) ring window loss dead-app alerting + ring cap/spill/spam controller v0.120.0 (+ hub v0.48.0)

CAMPAIGN-3 is closed.

Standing follow-ups (explicit)

  • Agent-ring persistence — the agent's own in-memory ring has the same restart-wipe gap; deferred (this task fixed the CONTROLLER ring only, to avoid an agent train).
  • F13 (HIGH, from Task A) — an active nfs4 under the mp8 bind can fail PVE's rbind with rc255; deferred (needs a pre-start idle-unmount design or an idmapped nfs mount).
  • Backup-locality option B — retarget NAS tier-1 to a local drive (operator chose A/keep locality).
  • The agentless-on-proxmox2 cluster gap on Peti's box (roadmap).
  • The publish train — agent 0.85 + controller 0.118/0.119/0.120 + hub 0.48 + MinAgent 0.81 + the journal-group one-liner + temp-creds deletion delivers this whole wave to Peti (on 0.113 / agent 0.81 today; reaches it at his next train).

Box state at wrap

controller 0.120.0 healthy on 9201; hub 0.48.0 live (Synced/Healthy); 8 apps healthy; seerr recovered; debug-ring spill on the SSD data volume; no test residue.