14 KiB
REPORT — v0.191.0 / .1 / .2: warn before the wall comes down (R-167 · R-158 · R-174)
Overwritten per the standing rule; the prior contents (v0.190.0 / R-157 A · R-170 · R-171, same
day) have their durable record in CHANGELOG.md.
Session: 2026-08-02 · Repos: felhom-controller v0.190.0 → v0.191.2, felhom.eu hub
v0.88.0 → v0.89.0 + docs. felhom-agent untouched (Part 3 read its golden-build script and
changed nothing in it).
1. Baselines — one drifted
| Repo | §1 said | Found on arrival | Match? |
|---|---|---|---|
felhom-controller |
95eb5c2c1af4, v0.190.0 |
95eb5c2c1af4, v0.190.0 |
yes |
felhom.eu |
d5774d318941, hub v0.87.0 |
8ef92a3f, hub v0.88.0 |
NO |
hub v0.88.0 had already shipped (R-172, the WAL fix) between the task being written and this session. Target corrected to hub v0.89.0. Highest register ID in use was R-173, not R-171.
2. One decision taken with the operator before any code was written
Reconnaissance found that disk_warning / disk_critical were a complete customer pipeline with no
producer — in the hub's allowedEventTypes, carrying Hungarian customerMessages, in
settings.DefaultEnabledEvents, with a UI checkbox (event_disk_alerts) — and grep across all four
repos found zero emitters. The only thing reacting to guest disk pressure was healthcheck.go:338,
folding it into a generic health_degraded at 90% for registered storage paths only.
The task said to mint a new customer-facing type. Doing so would have put a near-duplicate beside an
inert pair — and §5's own reason for rejecting backup_failed (it is customer-enabled and carries
Hungarian copy) is the argument for reusing this one, because D-c routes the fill warning to the
customer. Operator chose: wire the existing pair. So only ONE new hub type was minted (the
operator one), and the sixth built-but-never-wired instance was closed rather than joined by a seventh.
3. Files changed and commits
6 modified, 4 added. controller/cmd/controller/main.go,
internal/backup/{appstop_marker,backup,recovery_unit}.go, internal/notify/notifier.go,
cmd/controller/appstop_wiring_test.go, CHANGELOG.md, REUSE.md · new:
internal/fillwatch/{fillwatch.go,fillwatch_test.go},
internal/backup/{appstop_drivegate_test.go,recovery_unit_notify_test.go}
Commit on main |
Subject |
|---|---|
cf48214 |
v0.191.0 — warn before the wall comes down (R-167, R-158, R-174) |
5adae4d |
v0.191.1 — the fill check also runs at startup (R-167) |
9a3c485 |
v0.191.2 — a quiet fill check now says so (R-167) |
4. Tests
1157 → 1184 (+27). Full suite green in both repos (go build ./... && go vet ./... && go test ./...); all eight controller gates and all five felhom.eu gates OK. Hub: 574 → 579.
Red-proofs — each mutated, observed failing, restored
| # | What was mutated | Result |
|---|---|---|
| A1 | the errors.Is(err, ErrStartRefused) branch removed from Recover |
4 tests FAIL — refusals collapse into Failed and alarm |
| A2 | SetStarter(gatedAppStopStarter{…}) → SetStarter(stackMgr) (the v0.189.0 wiring) |
TestMainWiresGatedAppStopStarter FAILS |
| C | the m.unitNotify(...) call removed from the capture loop |
2 tests FAIL — the seam-never-wired shape |
| E | both edge guards removed | 3 tests FAIL — warns twice; the dead zone breaks |
| G | recovery_unit_capture_failed removed from operatorOnlyEvents |
2 tests FAIL, one reading "a customer was emailed the OPERATOR-ONLY …" |
| H | clear thresholds edited into equality with warn | TestThresholdsKeepTheirHysteresisGap FAILS |
| J | fillWatcher.SetNotify( commented out (string still in the file) |
TestMainWiresTheFillWatcher FAILS — proves AST, not substring |
One red-proof was a FALSE GREEN on the first attempt, and it is reported rather than quietly
redone. Red-proof E's first mutation removed only if next == prev { continue }; a second guard
(next <= prev) still caught the repeat, so every test passed and the "proof" proved nothing. Redone
removing both guards. This is exactly the trap CLAUDE.md records — a mutation or a -run filter
that yields a green and reads like a completed proof.
One pre-existing test was TIGHTENED, not loosened. TestMainReportsTheInterruptedOperation
asserted the guard was if appStopRecovery != nil. R-174 makes that insufficient — Recover now
returns non-nil for a refusal-only recovery — so the test now requires Alarming() in the condition.
5. Thresholds and cadence, with their justification
| value | why | |
|---|---|---|
| warn | used ≥ 85% OR free < 5 GiB | a percentage alone lies at both ends of this fleet's real size range: 85% of a 20 G mp1 leaves 3 G — less than one DB-backed app's unit (~2× its data, §7.5) — while 85% of a 4 TB drive leaves 600 G |
| critical | used ≥ 95% OR free < 2 GiB | below 2 GiB a volume tar of almost any real app fails: "the next backup will not complete", not "it is getting tight" |
| clear | used ≤ 75% AND free ≥ 7 GiB | both must hold; the gap to warn is the hysteresis dead zone, so a filesystem on the line does not flap. Pinned with a real margin, not merely an inequality |
| cadence | daily 03:30 + once 90 s after startup | a fill is slow-moving, so a shorter interval buys no earlier warning; 03:30 precedes the nightly app-data legs, so a customer about to lose a backup hears about it with a night's margin. The startup run exists because neither Daily nor Every fires on registration — §7 |
The live proof vindicated the two-term design: the critical crossing fired on the free-byte term (1.7 GB) while the disk was only 91% used. A percentage-only rule would have missed it entirely.
6. The exact Hungarian customer copy, for review as copy
disk_warning (severity warning) — as rendered live on 9201:
A(z) „Rendszer- és mentési terület” tároló 85% foglalt — 4,7 GB szabad hely maradt. Kérjük, szabadíts fel helyet, mielőtt megtelik: törölj felesleges fájlokat, vagy csatlakoztass új meghajtót. Ha megtelik, a biztonsági mentések meghiúsulnak.
disk_critical (severity critical) — as rendered live on 9201:
A(z) „Rendszer- és mentési terület” tároló kritikusan megtelt: 1,7 GB szabad hely maradt (91% foglalt). A biztonsági mentések és az alkalmazások írásai bármikor meghiúsulhatnak. Kérjük, mielőbb szabadíts fel helyet: törölj felesleges fájlokat, vagy csatlakoztass új meghajtót.
Decimal comma throughout (4,7 GB); the storage is named by its label, never its path; both
end in an action. No emoji (asserted by a test; emoji_gate passes).
7. Two defects found by the validation itself, both fixed and shipped
- v0.191.1 — the check was reachable only on its daily schedule. Neither
sched.Dailynorsched.Everyfires on registration; both wait for their first tick. A box that BOOTS already over the line would have stayed silent for up to 24 h — the R-100 shape, and the same gap the hub's own checkers avoid by leaving already-breached keys unseeded at init. Now also runs once 90 s after startup; safe because the check is edge-triggered against persisted state, so an already-warned filesystem stays silent. - v0.191.2 — a quiet run was unreadable as evidence. After the customer had been warned, a
restart produced zero
fillwatchlines — equally consistent with "ran and chose silence" and "never ran". For an edge-triggered check the quiet run is the healthy steady state, so that ambiguity is permanent, not rare.Checknow logs a per-run summary (checked N filesystem(s), M unreadable/skipped, K notification(s); bands: …), counting unreadable separately so a silently unreadable drive cannot read as "all fine".
8. Live evidence — three flows, guest 9201 on felhom-pve
Method: endpoint/log-level on real hardware, exercising the production paths — browser automation
is not available on DooPlex. The hub's notification_log was read from a WAL-aware copy of
/data/hub.db (all three files — R-172's lesson); the copies were deleted afterwards.
Flow 1 — the guard refuses, keeps the marker, does not alarm. Both hdd_1 mounts held unmounted
against the agent's ~60 s heal, per the documented method:
[WARN] [appstop] refusing to restart "calibre-web" after an interrupted operation:
drive /mnt/felhom-drives/hdd_1 is not a live mountpoint
[WARN] [appstop] crash recovery: NOT restarting calibre-web — start refused by a deliberate
holder: … the marker is KEPT and the holder owns the restart
[WARN] [appstop] crash recovery: 1 app(s) were deliberately NOT restarted (drive absent) —
KEEPING the marker; this is the gate working, not a fault: [calibre-web]
[WARN] [appstop] …HELD… — not alarming: … restarted=[] held_by_drive=[calibre-web]
Marker retained byte-identical; docker ps -a --filter name=calibre-web → 0; no
backup_failed event. Scenario B, on a later boot with the drive live: restarted calibre-web after the interrupted an app-data backup (volume dump) and the marker was cleared.
Flow 2 — the operator event arrives; the customer is refused. Two real capture failures:
Event pushed: recovery_unit_capture_failed (error) — Recovery unit capture FAILED for "calibre-web"
… /mnt/felhom-drives/hdd_1: 24.5/93.9 GB used (26%), 64.6 GB free. Error: … permission denied
Hub notification_log — the positive observable:
customer | recovery_unit_capture_failed | skipped | operator_only
operator | recovery_unit_capture_failed | sent |
Flow 3 — the customer is warned, once, in Hungarian. /mnt/sys_drive filled with fallocate:
| run | state | fired | log |
|---|---|---|---|
| 1 | 90% / 4.7 GB | disk_warning |
ok → warning … notifying the customer |
| 2 | unchanged | nothing | (edge trigger held) |
| 3 | 91% / 1.7 GB | disk_critical |
warning → critical … |
| 4 | file removed | nothing | critical → ok … cleared silently, re-armed; state emptied |
Hub: customer | disk_warning | sent and customer | disk_critical | sent, Hungarian rendered.
Exactly two events across four runs. Run 2's silence is meaningful only because run 3 proves the
startup check executes every boot — needing that inference is what motivated v0.191.2.
9. Deployed versions
$ ssh felhom-pve "pct exec 9201 -- docker ps --filter name=felhom-controller …"
gitea.dooplex.hu/admin/felhom-controller:0.191.2 | Up (healthy)
Hub: ArgoCD app felhom Synced / Healthy after a deliberate hard-refresh + sync (never
kubectl set image); deploy/hub rolled out; image gitea.dooplex.hu/admin/felhom-hub:0.89.0;
startup log clean.
10. Teardown
R167-FILLTEST.bin deleted (/mnt/sys_drive back to 2.0 G used / 45 G free); the app-stop marker
cleared by a successful recovery; /root/hold-unmount.sh + /tmp/hold.log removed from felhom-pve;
the secret-bearing local hub.db* copies deleted. All 15 containers on 9201 healthy.
One self-inflicted incident, disclosed in full. Restoring the drive by hand I ran
mount --bind /mnt/hdd_1 /mnt/felhom-drives/hdd_1, omitting the felhom-data segment the agent
uses (felhom-agent/internal/localapi/intermediary.go:15 —
mount --bind /mnt/<name>/felhom-data /mnt/felhom-drives/<name>). The guest then saw the drive root
instead of the namespace root, which an unprivileged guest cannot write to — producing real
permission denied capture failures. They served as flow 2's evidence, but they were caused by me,
not found by me, and the report says so. Separately, a manual docker compose up -d from a stack
dir started containers without the controller-injected env (there is no .env on disk — the
controller injects at exec time), which recreated immich-server with a blank DB password and left it
in a 28P01 auth-failure restart loop. Both were repaired through production paths: the bind
corrected, then the stacks brought back by the controller's own StartStack and boot reconciler
([bootrecon] … 1 app(s) recovered in 1 attempt(s): [immich]). Lesson, learned twice in one
session: on a box the agent manages, restore state by letting the agent and controller do it.
11. Register rows
Opened: R-174 (closed same session), R-175, R-176, R-177 — each ID established free by
grep -ro "R-17n\b" documentation/ *.md → 0 hits. Closed: R-158 (by R-167 — no second row was
filed for the same wire), R-167, R-174. Updated and still open: R-165 (gains the spike's M1-M5
and the operator question).
12. CI and --no-verify
--no-verify was NOT used. Every push ran .githooks/pre-push (gates --fast) and it passed.
CI checked by PULL, matching head_sha to each commit — CI emails only on failure, so a green
that was never looked at is an assumption. All three commits green:
| Commit | Task id | Run # | Conclusion |
|---|---|---|---|
cf48214 (v0.191.0) |
31 | 11 | success |
5adae4d (v0.191.1) |
34 | 12 | success |
9a3c485 (v0.191.2) |
35 | 13 | success |
The felhom.eu side (hub v0.89.0, the manifest bump and the docs) is in that repo's REPORT.md §6 —
also all green.
13. Observations — noticed, documented, NOT acted on
- R-177 — no operator-triggerable "run this scheduler job now" path. It cost a controller restart per observation here, and costs the same on a support call.
- R-175 —
07-backup-architecture.md§7.5 states one box's size bound as if it were the fleet's. - R-176 — two R-165 prerequisites are unmeasured (a pre-merge-archive restore-test; the in-place migration rehearsal).
- The drive gate's return branch failed to restart both apps (
[WARN] [gate] restart calibre-web: … exit code 1) at a moment when a manual start succeeded seconds later. Not filed as a row: it happened during my incorrect bind, so the likeliest cause is mine and the evidence is contaminated. Worth a look if it recurs on a clean box.