Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HxLA1mZurFq9kt8hneFeCs
5.8 KiB
felhom.eu — task reports
Overwrite this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in hub/CHANGELOG.md.
REPORT — hub v0.23.0: host root-disk pressure monitoring + alert
Date: 2026-06-30
Task: TASK.md — host root-disk pressure monitoring + alert (hub checker). Implementation class, hub-only.
Baseline (confirmed)
| value | |
|---|---|
Hub main before |
v0.22.1 (commit f133355) |
| Hub after | v0.23.0 |
| Commits | 897997c (code) · 77c61ff (manifest 0.22.1→0.23.0) |
| Live image | gitea.dooplex.hu/admin/felhom-hub:0.23.0 (ArgoCD felhom Synced/Healthy) |
What was built
A hub monitor checker that, on the existing 60s sweep, reads each host's latest root-fs disk_percent
(reported by the agent, already stored) and emits an operator alert when it crosses a warning (default
90%) or critical (default 95%) threshold — closing the silent-failure gap behind the felhom-pve
incident (vzdump piling under /var/lib/vz/dump until the host root filled and PVE/logging/agent writes
could start failing, with no alert).
Files changed (commit 897997c)
internal/monitor/host_disk.go(NEW) —HostDiskChecker. Sibling ofHostCapabilityChecker/HostLeafChecker; rank-based bands (ok→warning→critical); born/persistent (already-breached hosts left UNSEEDED → firstCheckemits — the F2 lesson); de-escalation/recovery re-arm silently; thresholds hub-config overridable with a defaults/sanity guard (normalizeDiskThresholds).internal/store/store.go—GetHostDiskUsage()+HostDiskRow(latest report per host;disk_percentcolumn + total/used bytes fromreport_json; no schema migration).internal/notify/templates.go— Hungarian customer templates forhost_disk_warning/_critical.internal/api/handler.go— both types added toallowedEventTypes.cmd/hub/main.go—alerting.host_disk_warn_percent/host_disk_crit_percentconfig; registerhostDiskCheckeron the 60s tick.
Design decision worth flagging (deviation from the task)
The task said "critical → severity critical". The live dispatcher (notify/dispatcher.go) only routes
severity == "warning" || "error" — a "critical" severity is silently dropped. So the critical band
maps to severity error (and the operator email's 🔴), which is what actually alerts. The event type
is still host_disk_critical. Asserted by TestHostDiskChecker_Severity.
Distinct from the guest disk alert (§8/§7-D)
Event types are host_disk_warning/host_disk_critical (the Proxmox HOST root fs) — never the controller's
GUEST disk_warning/disk_critical (its own cgroup view); the two never dedup or mask each other.
Tests — go build ./... && go vet ./... && go test ./... GREEN
TestHostDiskChecker_Bands— seed ok (no event), 80→nothing, ok→warning, steady (no re-emit), warning→critical escalation, recovery clears, re-arm (a fresh breach alerts again), type ishost_disk_*.TestHostDiskChecker_Severity— warning band →warning, critical band →error.TestHostDiskChecker_BornPersistent— a host already at 97% at init → firstCheckemitshost_disk_critical. Companion red-proof: a seed-all (transition-only) model — emulated by pre-seeding the breached host's state tocritical— stays silent on the born-breach; the real (unseeded) design emits. That gap is the bug this design fixes.TestHostDiskChecker_ThresholdDefaults— unset/invalid/misordered config → 90/95 (no silence/inversion).
Deploy + live validation (real operator alert)
Deployed via the repo's GitOps path (image build+push on 180; manifests/hub.yaml 0.22.1→0.23.0; ArgoCD
hard-refresh → OutOfSync → sync → Synced/Healthy). Used the manifest path, not bare kubectl set image, because the repo convention is that a bare set-image is reverted on the next sync.
Method: the cheapest real path from §13 — felhom-pve's host root was already at ~96–100% (the actual hazard), so no artificial fill and no threshold change were needed. Startup + first-tick logs:
[INFO] felhom-hub 0.23.0 starting
[INFO] Host disk checker initialized: warn=90% crit=95%, 0 ok seeded, 1 already-breached left unseeded (first Check emits)
[INFO] Host disk: demo-felhom-01 root 96% unknown→critical (host_disk_critical)
[INFO] Operator email sent for demo-felhom/host_disk_critical
The born-persistent alert fired on cycle 1 and a real operator email was sent via Resend (the
send-success log line — same proof standard as the prior capability alert). Pod Running 1/1, no errors.
Hazard relieved (operational)
The host root was genuinely full (85G of vzdump backups — ~18 copies of demo guest 9201 at ~3.9G each, on the root fs). I pruned the old backups, keeping the 2 most recent (regenerable demo-guest backups): root 100% → 24% (freed ~69G). The recovery/re-arm then reflects on the agent's next 15-min report (disk now below warn); the recovery logic itself is unit-tested.
Observations / follow-ups (noted, not built)
- Per-storage
StorageTargetsworst-fill alerting is the easy follow-on (a dedicated dump/backup storage filling). Host rootdisk_percentalready covered the observed vzdump-on-root case, so root is the v1 headline. - The retention fix — a provisioning-side
prune-backups(vzdump--maxfiles/prune-backups) default so a box can't refill its own root — is a separate small operational follow-up (agent/golden side). The manual prune above was a one-off relief, not the durable fix.
No secrets in any committed file. Operator email address is the operator's own (not a secret); the Resend
key stays injected from Secret/resend-api.