controller v0.120.0: dead-app alerting (fix-3) + ring revision (fix-6) — CLOSES CAMPAIGN-3 (docs+CHANGELOG+REPORT+CONTEXT)
fix-3: deadapp-check job -> self-clearing WARN banner + one app_start_failed hub event per running->down transition. fix-6: ring 1000->5000, periodic spam->TRACE (ring-dropped), atomic SSD spill/load across restart. Live-validated on 9201+hub 0.48. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CDMFpFx84pfviCTVuGGhf
This commit is contained in:
@@ -1,59 +1,68 @@
|
||||
# REPORT — v0.119.0: storage-health coherence (F8) + mapped_uid validation (F4)
|
||||
# REPORT — v0.120.0: dead-app alerting (fix-3) + debug-ring revision (fix-6) — CLOSES CAMPAIGN-3
|
||||
|
||||
**Date:** 2026-07-12 · **Version:** controller v0.119.0 (from v0.118.0) · **MinAgent:** 0.81.0 (UNCHANGED)
|
||||
· **Deployed:** guest 9201 (`0.119.0` healthy) · **Source:** `felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md`.
|
||||
|
||||
## §3 design fork — decision
|
||||
|
||||
Took the recommended **option B (controller-only)**: the share row reuses the shipped v0.117.0
|
||||
consuming-namespace classifier (`system.ClassifyPathFS`) — the exact ground truth the stacks-page stub
|
||||
badge already reads. No agent change, no new probing surface, and by construction the row and the
|
||||
stacks badge can never disagree (single source). Option A (agent-side export-level probe) was not built.
|
||||
**Date:** 2026-07-12 · **Version:** controller v0.120.0 (from v0.119.0) · **MinAgent:** 0.81.0 (UNCHANGED)
|
||||
· **Pairs with:** hub v0.48.0 (accepts `app_start_failed`) · **Deployed:** guest 9201 (`0.120.0` healthy).
|
||||
**Source:** `felhom.eu/documentation/audits/CAMPAIGN-3-2026-07-11.md`.
|
||||
|
||||
## What shipped
|
||||
|
||||
- **F8 (MED) — one classification, two surfaces.** `networkStorageItems` now fuses the agent's health
|
||||
with the namespace classification via `fuseNetHealth`: a new `stub` state overrides a benign idle/ok
|
||||
when the consuming namespace sees local disk at `Where`; a whole-server `unreachable` still wins over
|
||||
stub; autofs-healthy / network / `unknown` leave the agent health intact (no manufactured fault, no
|
||||
force-mount). Row badge for `stub` = "Hibás — az alkalmazások nem a NAS-t látják". The share row and
|
||||
the stacks/dashboard badge now derive from ONE classifier.
|
||||
- **F4 (LOW) — mapped_uid/gid range check at the door.** `handleNetStorageAdd` validates the container
|
||||
uid/gid (1..65533) after the `<=0` default, before the job — out of range → friendly Hungarian 400,
|
||||
nothing installed. Catches the campaign's `101000` (a host-side mapped value) that used to leak a raw
|
||||
`agent_error`.
|
||||
- **fix-3 (MED) — a dead deployed app is loud.** A `deadapp-check` job (every 30 s, after a 90 s boot
|
||||
grace) scans the deployed apps: `stopped`/`exited` ones (`stacks.IsDownState`) raise a self-clearing
|
||||
WARN dashboard banner (grouped above 3) and fire an `app_start_failed` hub event ONCE per running→down
|
||||
transition (`Notifier.NotifyAppStartFailures`; the hub owns cooldown — no controller timer).
|
||||
- **fix-6 (MED) — the post-incident window survives.** Ring cap 1000→5000 (display cap raised to
|
||||
match); periodic scheduler/refresh success lines demoted to a `[TRACE]` level the ring drops at
|
||||
write-time (failures never TRACE); atomic JSON-lines spill to `<DataDir>/debug-ring.log` (SSD only)
|
||||
every 30 s + on shutdown, loaded back on boot. Corruption-safe.
|
||||
- **hub v0.48.0:** `app_start_failed` added to `allowedEventTypes` + `customerMessages` (else the event
|
||||
400s at ingest — the known allowlist gotcha).
|
||||
|
||||
## Tests + red-proofs (all green)
|
||||
|
||||
- F8 fusion table: idle+stub→stub (the contradiction resolved), ok+stub→stub, **idle+autofs→idle**
|
||||
(the over-eager autofs=stub mutant fails here), unreachable+stub→unreachable (server wins),
|
||||
idle+unknown→idle (no manufactured fault). End-to-end `networkStorageItems` stub fusion (companion:
|
||||
drop the fuse call → row shows raw agent health → fail).
|
||||
- F4: uid 101000 → 400 + friendly message, agent never reached (companion: drop the check → reaches the
|
||||
agent → fail); 65534 → 400; 1000 / 65533 / 0-defaults pass the range check.
|
||||
- fix-3: dead-app banner present + self-clears (companion: skip SetDeadAppAlerts → silent → fail);
|
||||
grouped above threshold; `IsDownState` table (stopped/exited down; starting/unhealthy/restarting/
|
||||
deploying/paused/unknown NOT); notifier one-event-per-transition (companion: drop tracking → fires
|
||||
each cycle → fail); first-seen-down (dead-at-boot) fires; healthy never fires.
|
||||
- fix-6: TRACE dropped from ring while a WARN failure is kept (companion: demote too broadly → failure
|
||||
lost → fail); spill→load round-trip; load keeps newest N; corrupt/truncated spill loads valid + never
|
||||
panics; missing file no-op.
|
||||
|
||||
## Live validation (demo 9201, sim-NAS rails — exportfs only)
|
||||
## Live validation (demo 9201 + hub)
|
||||
|
||||
- **F8 the contradiction, killed:** baseline healthy → row `ok`, no stub badge. `exportfs -u` while idle
|
||||
+ drop the mount → the SHARE ROW showed `health=stub` ("Hibás — az alkalmazások nem a NAS-t látják")
|
||||
AND the stacks page showed the stub badge (4) — the two surfaces AGREE (previously: row "Készenlét" +
|
||||
stacks stub = contradiction). `reachable:true` throughout (the server-level dial is still green — the
|
||||
exact F8 blindness, now correctly overridden). Re-export → row cleared back to `ok`/"Elérhető"
|
||||
(healthy idle NOT downgraded — the regression).
|
||||
- **F4:** `mapped_uid:101000` → 400 + the friendly message, registry unchanged (no `c5uid`), no host
|
||||
unit/dir residue; `mapped_uid:1000` → 200, passed the range check (then failed later at the
|
||||
unreachable probe as designed, rolled back clean).
|
||||
- **fix-3:** `docker stop seerr` → within a cycle the dashboard banner "Telepített alkalmazás nem fut:
|
||||
Jellyseerr (stopped)" appeared AND the hub logged exactly ONE `app_start_failed` event
|
||||
(`Event from demo-felhom: app_start_failed`); across 3 down-cycles still ONE event (anti-spam);
|
||||
`docker start seerr` → the banner self-cleared (0 banners). This makes the campaign's 4-hour silent
|
||||
death impossible.
|
||||
- **fix-6:** the ring held 0 periodic-spam lines (TRACE demotion); a controller restart PRESERVED the
|
||||
pre-restart window — oldest entry unchanged across `systemctl restart` (total 390→464, oldest
|
||||
08:11:28 both sides), with a 63 KB spill on the persistent SSD data volume (survives container
|
||||
recreation, never a NAS path).
|
||||
|
||||
## NOT live-validated / standing items
|
||||
## CAMPAIGN-3 finding ledger — RESOLVED
|
||||
|
||||
- `unreachable`-wins live (a genuine server-down IP on a registered share) — unit-tested only; the F8
|
||||
live proof used the export-level cut (the actual finding).
|
||||
- Task D remains queued: fix-3 boot-time app-start-failure alerting + the ring wrap/count revision
|
||||
(fix-6 6.5-min horizon under load).
|
||||
- Peti's box (controller 0.113 / agentless-on-proxmox2) reaches 0.119 (+0.85/0.118) at his next train.
|
||||
- No publish/floor movement; agent untouched (MinAgent 0.81.0).
|
||||
| Finding | Fix | Ships in |
|
||||
|---|---|---|
|
||||
| F12 (CRIT) boot ordering cycle; F11/F10/F9 (HIGH) automount re-arm; F2/F1 residue | agent boot/recovery plane + appliance self-heal | **agent v0.85.0** |
|
||||
| F7 (HIGH) in-place volume-dump truncation; F6/F5 (LOW) single-copy + stale dirs | atomic dumps + tier-2 for volume-only + stale sweep | **controller v0.118.0** |
|
||||
| F8 (MED) storage-health contradiction; F4 (LOW) mapped_uid | share-row classifier fusion + range check | **controller v0.119.0** |
|
||||
| fix-3 (MED) silent dead app; fix-6 (MED) ring window loss | dead-app alerting + ring cap/spill/spam | **controller v0.120.0** (+ hub v0.48.0) |
|
||||
|
||||
CAMPAIGN-3 is closed.
|
||||
|
||||
## Standing follow-ups (explicit)
|
||||
|
||||
- **Agent-ring persistence** — the agent's own in-memory ring has the same restart-wipe gap; deferred
|
||||
(this task fixed the CONTROLLER ring only, to avoid an agent train).
|
||||
- **F13 (HIGH, from Task A)** — an active nfs4 under the mp8 bind can fail PVE's rbind with rc255;
|
||||
deferred (needs a pre-start idle-unmount design or an idmapped nfs mount).
|
||||
- **Backup-locality option B** — retarget NAS tier-1 to a local drive (operator chose A/keep locality).
|
||||
- **The agentless-on-proxmox2 cluster gap** on Peti's box (roadmap).
|
||||
- **The publish train** — agent 0.85 + controller 0.118/0.119/0.120 + hub 0.48 + MinAgent 0.81 + the
|
||||
journal-group one-liner + temp-creds deletion delivers this whole wave to Peti (on 0.113 / agent 0.81
|
||||
today; reaches it at his next train).
|
||||
|
||||
## Box state at wrap
|
||||
|
||||
controller 0.119.0 healthy on 9201; nas-media healthy + `ok`; registry = nas-media only (no test
|
||||
residue); all NAS apps healthy; NFS re-exported.
|
||||
controller 0.120.0 healthy on 9201; hub 0.48.0 live (Synced/Healthy); 8 apps healthy; seerr recovered;
|
||||
debug-ring spill on the SSD data volume; no test residue.
|
||||
|
||||
Reference in New Issue
Block a user