Files
felhom.eu/REPORT-hub-blindness.md
T
admin 104ef34f57
gates / gates (push) Successful in 15s
gates: the today-override announces itself; malformed no longer swallowed
Part 6 of the hub-blindness task, separable and done rather than dropped.

Both due_checks_gate.py and instructions_gate.py read FELHOM_GATE_TODAY so
their suites can control "today", and neither said so. A shell that still has
it exported -- exactly what a session doing gate-test work leaves behind --
made both gates evaluate against a fabricated date and pass in SILENCE. That
is this project's own named failure class: an instrument that can quietly
return the wrong answer is not a measurement. The seam is legitimate and
stays; the silence was the defect.

Both now print a loud line naming the variable, its value, and that the real
date is being ignored, before any verdict.

And instructions_gate.py no longer swallows a MALFORMED override: it used to
fall through to the real date without a word while due_checks_gate.py already
exited 2 on the same input -- one variable, two gates, disagreeing about what
a mistake means. Both exit 2 now.

Tests extended in both suites (42 and 73 assertions, green). Red-proof: the
announcement was deleted from due_checks_gate.py and its two assertions were
seen failing, then reverted.

Also adds REPORT-hub-blindness.md (topic-suffixed; the shared REPORT.md is
left alone per the parallel-session rule).
2026-08-18 19:32:35 +02:00

12 KiB
Raw Blame History

REPORT — the hub says something when it loses sight of the off-site endpoints (2026-08-18)

Shipped: hub v0.106.0, deployed and verified. Both box checkers now carry a second, independent reachability signal with paired all-clears. The fill logic is untouched. Part 6 was done, not dropped.

NOT proven live — see §7. No real or constructed outage has exercised the emit path.


1. Confirmed baselines

item value
felhom.eu main @ start 78a244bf0f… — matches the prompt's anchor
hub version in → out v0.105.0 → v0.106.0 (read from hub/CHANGELOG.md head)
scripts/ version in → out due_checks_gate.py v1.0.0 → v1.0.1
highest R in use at start R-343, so R-339 / R-340 free as specified

Had the three target files moved? No. Verified by hash before editing:

61a16462756099d2fa60dd0c50aeac8c  internal/monitor/pbsdr_box.go
d12b3e665d729a0291d2397895f23d1f  internal/monitor/offsite_box.go
702d17487ffb4ac17a9d18050bb546b7  internal/notify/dispatcher.go

All landmarks in §5 of the prompt resolved as described; nothing was stale.

2. Files created / modified

Created: hub/internal/monitor/box_reachability_test.go, hub/internal/notify/dispatcher_box_reachability_test.go, REPORT-hub-blindness.md.

Modified: hub/internal/monitor/pbsdr_box.go, hub/internal/monitor/offsite_box.go, hub/internal/notify/dispatcher.go, hub/cmd/hub/main.go, hub/CHANGELOG.md, hub/internal/monitor/{pbsdr_box_test.go,offsite_box_test.go} (new constructor arg), manifests/hub.yaml, CONTEXT.md, STATUS.md, documentation/backlog/OPEN-ITEMS.md, documentation/architecture/00-capability-map.md, scripts/due_checks_gate.py, scripts/instructions_gate.py, scripts/test_due_checks_gate.py, scripts/test_instructions_gate.py, scripts/CHANGELOG.md.

3. Commits pushed to main

commit contents
ab2262c91c1e976ae3b983e9df6b45dabfcb9d23 the code, tests, and register/doc edits
c03f629d43ed3d175e2c8561486cadc9a432640f manifests/hub.yaml 0.105.0 → 0.106.0 — the change that actually deploys
(Part 6 commit — see §10) the today-override announcement + scripts/CHANGELOG.md

4. Tests and the three red-proofs

All named tests pass. Groups A–F in internal/monitor/box_reachability_test.go, Group G in internal/notify/dispatcher_box_reachability_test.go:

test result
TestPBSDRBox_Unreachable_SustainedOutage (A) PASS
TestPBSDRBox_Unreachable_BlipBelowThreshold (B) PASS
TestPBSDRBox_Unreachable_Recovery (C) PASS
TestPBSDRBox_UsageUnsupported_IsNotBlindness (D) PASS
TestPBSDRBox_BornBlind_StillReports (E) PASS
TestOffsiteBox_Unreachable_AndRecovery (F) PASS
TestPBSDRBox_ZeroCapacitySuccess_ClearsBlindness (§8 truth table) PASS
TestBoxRecovery_ReachesTheOperatorDespiteInfoSeverity (G) PASS
TestBoxUnreachable_ReachesTheOperatorOnItsOwnSeverity (G) PASS
TestBoxRecovery_PairedWithTheCorrectDownType (G) PASS

None of the three red-proofs passed on the first attempt — each turned its test red, and each did so for the reason under test, which I checked in the message rather than in the count.

Red-proof 1 — threshold 3 → 1. Group B seen failing:

box_reachability_test.go:132: two failed windows emitted [pbsdr_box_unreachable pbsdr_box_unreachable], want silence below the threshold

The message names the premature events, not an incidental error. Reverted (defaultBoxUnreachableWindows = 3 restored).

Red-proof 2 — remove the ErrUsageUnsupported counter guard (deleted its early return so the branch falls through). Group D seen failing:

box_reachability_test.go:213: ErrUsageUnsupported emitted [pbsdr_box_unreachable × 8] — an expected pre-update condition must never alert

The message names the unexpected event type, as the prompt required — not merely a count. Reverted.

Red-proof 3 — remove pbsdr_box_recovered from recoveredPairedDownTypes. Group G seen failing, and the first failure is the end-to-end mail assertion, which is what proves the test exercises the wiring rather than the map:

dispatcher_box_reachability_test.go:41: pbsdr_box_recovered: operator mails = 0, want 1 — the all-clear must reach the operator; 0 means the recoveredPairedDownTypes entry is missing and "info" was dropped by the severity gate

Reverted. grep -rn MUTATED internal/ returns nothing.

5. Test count

go test ./... — 21 packages, all green (18 with tests, 3 with none). internal/monitor gained 7 tests; internal/notify gained 3. go build ./... and go vet ./... clean.

Repo gates: 10/10 OK, rc=0.

6. Deployed version and the wiring evidence

ArgoCD app "felhom":  Synced  Healthy   rev=c03f629d43ed3d175e2c8561486cadc9a432640f
pod:                  hub-654bbc8fbc-9wld9   1/1 Running
running image:        gitea.dooplex.hu/admin/felhom-hub:0.106.0

The required post-deploy check — both constructor log lines carrying the threshold:

19:29:07 [INFO] Offsite pool-box checker initialized: box=611714 fill warn=80% crit=90%,
                oversub warn=2.00x, unreachable after 3 consecutive failed reads, refresh 15m0s
19:29:07 [INFO] PBS-DR box checker initialized: fill warn=80% crit=90%,
                unreachable after 3 consecutive failed reads, refresh 15m0s

Two lines, both carrying the threshold — the parameter reached both checkers. Their absence would have meant the config was inert however green the tests were. Note this also exercised the absent-key path: box_unreachable_windows is deliberately not in any deployed config, so both checkers fell back to the documented default of 3, which is what the log shows.

7. NOT yet live-validated — explicitly

No real or constructed endpoint outage has exercised the emit path end to end. Everything in §4 is an injected fake with a scripted error and an injected clock. What is proven: the checkers emit the right events with the right details, and the dispatcher routes both new *_recovered types to a real operator mail. What is not proven: that a genuine ep0 or Hetzner failure produces those errors in the shape the checkers expect.

A real outage cannot be manufactured without making ep0 or the Hetzner API unreachable, and ep0 is Tier 2 protected — that was not done. The constructed-outage option, named but not performed: point the tenantsync client at a blackholed address on a scratch hub instance and let three windows elapse.

8. Teardown

This run provisioned nothing. No VM, no container beyond the hub's own rolling deployment, no drill target, no scratch guest. All three layers N/A. ep0, both demo boxes and the drill VM were untouched, as were the agent's credential-consume and self-heal paths.

9. Register rows

R-339 opened and marked SHIPPED (hub v0.106.0), with PROVEN-LIVE explicitly still owed and an instruction not to close it on the unit tests.

R-340 opened, READY (M) — the honest boundary: the hub's ep0 read is the usage op, which rides the local API daemon, and the 2026-08-18 incident explicitly cleared that daemon while the HTTPS proxy on 8007 was wedged. R-339's check would have shown green for all 9 h 37 m of the outage that motivated it. Overlap with R-336's remaining half is noted so whichever runs second reuses the first's evidence rather than re-measuring a protected machine.

R-336's next-step cell corrected. The replacement text, verbatim:

CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist. This cell used to read "PVE storage status is the prime suspect, and its interval is tunable". The first half is right and the second half is false. pvestatd stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (pvesm set <id> --disable 1) around the backup window, and that is substantially more than a tuning knob: it collides with felhom-agent/internal/pbsdr/manager.go's health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. Doc-only correction — no agent code was changed. The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3.

10. Part 6 — DONE, not dropped

Both gates now announce the FELHOM_GATE_TODAY override loudly before any verdict, and instructions_gate.py no longer swallows a malformed one — it used to fall through to the real date in silence while due_checks_gate.py already exited 2 on the same input, so one variable had two gates disagreeing about what a mistake means. Both exit 2 now.

Tests extended in both suites (42 and 73 assertions, all green). Red-proof: the announcement was deleted from due_checks_gate.py and its two assertions were seen failing — P6: valid override is announced and P6: the announcement says the real date is being ignored — then reverted.

11. Gate and CI status

python3 scripts/repo_gates.py → rc=0, all ten gates OK, including due-checks.

The due-checks gate did NOT refuse this push. R-341's first check comes due 2026-08-19 UTC and this work ran on 2026-08-18 (17:18–19:30 UTC), so the gate reported "2 dated check(s) pending, none due yet" throughout. No row was cleared, no date edited, no --no-verify used. Every push today went through the armed hook.

CI run IDs and conclusions: see §3's commits — reported in the closing summary.

12. unproven.py --summary

where felhom stands — 55 claims, verified_on 2026-08-09
  walked   23
  partial  14   (6 cite evidence, 8 prose only)
  built    14   (0 cite evidence, 14 prose only)
  missing  4   (0 cite evidence, 4 prose only)
  NOT WALKED: 32 of 55

No number moved. Correct: this shipped an implemented-not-proven capability, which is exactly the status that does not advance the walked count. Moving it would require the live validation §7 says has not happened.

13. Observations — noticed, deliberately not acted on

  • make docker-push also tags and pushes :latest, which the project's own rules forbid. I used make docker followed by an explicit docker push …:0.106.0 instead, so no :latest was moved. The Makefile target is a loaded gun for anyone who runs the documented command; not changed here because it is outside this task's scope.
  • internal/monitor/storage_fill_test.go is not gofmt-clean, and was already so at HEAD — confirmed by stashing my changes and re-running gofmt -l. Not touched; it is not mine and fixing it would put unrelated churn in this diff.
  • The two checkers are now ~95% identical in their reachability half. A shared helper is the obvious next move and was deliberately not done here, per the prompt: they have different sources, different error taxonomies (one has a sentinel, one does not) and different snapshot types, and the existing code keeps them separate on purpose. Worth revisiting if a third box checker appears.
  • The first ArgoCD sync reported Synced/Healthy at the PREVIOUS revision (ab2262c) while the pod was still ContainerCreating. Waiting and re-reading gave c03f629 and the correct image. A sync status sampled too early is not the deploy's verdict — the running image tag is.
  • alerting.box_unreachable_windows is in no deployed config file, by design, so the live hub is running on the compiled default. If the operator wants to tune it, the key has to be added to the hub ConfigMap first.