12 KiB
REPORT — the hub says something when it loses sight of the off-site endpoints (2026-08-18)
Shipped: hub v0.106.0, deployed and verified. Both box checkers now carry a second, independent reachability signal with paired all-clears. The fill logic is untouched. Part 6 was done, not dropped.
NOT proven live — see §7. No real or constructed outage has exercised the emit path.
1. Confirmed baselines
| item | value |
|---|---|
felhom.eu main @ start |
78a244bf0f… — matches the prompt's anchor |
| hub version in → out | v0.105.0 → v0.106.0 (read from hub/CHANGELOG.md head) |
scripts/ version in → out |
due_checks_gate.py v1.0.0 → v1.0.1 |
| highest R in use at start | R-343, so R-339 / R-340 free as specified |
Had the three target files moved? No. Verified by hash before editing:
61a16462756099d2fa60dd0c50aeac8c internal/monitor/pbsdr_box.go
d12b3e665d729a0291d2397895f23d1f internal/monitor/offsite_box.go
702d17487ffb4ac17a9d18050bb546b7 internal/notify/dispatcher.go
All landmarks in §5 of the prompt resolved as described; nothing was stale.
2. Files created / modified
Created: hub/internal/monitor/box_reachability_test.go,
hub/internal/notify/dispatcher_box_reachability_test.go, REPORT-hub-blindness.md.
Modified: hub/internal/monitor/pbsdr_box.go, hub/internal/monitor/offsite_box.go,
hub/internal/notify/dispatcher.go, hub/cmd/hub/main.go, hub/CHANGELOG.md,
hub/internal/monitor/{pbsdr_box_test.go,offsite_box_test.go} (new constructor arg),
manifests/hub.yaml, CONTEXT.md, STATUS.md, documentation/backlog/OPEN-ITEMS.md,
documentation/architecture/00-capability-map.md, scripts/due_checks_gate.py,
scripts/instructions_gate.py, scripts/test_due_checks_gate.py,
scripts/test_instructions_gate.py, scripts/CHANGELOG.md.
3. Commits pushed to main
| commit | contents |
|---|---|
ab2262c91c1e976ae3b983e9df6b45dabfcb9d23 |
the code, tests, and register/doc edits |
c03f629d43ed3d175e2c8561486cadc9a432640f |
manifests/hub.yaml 0.105.0 → 0.106.0 — the change that actually deploys |
| (Part 6 commit — see §10) | the today-override announcement + scripts/CHANGELOG.md |
4. Tests and the three red-proofs
All named tests pass. Groups A–F in internal/monitor/box_reachability_test.go, Group G in
internal/notify/dispatcher_box_reachability_test.go:
| test | result |
|---|---|
TestPBSDRBox_Unreachable_SustainedOutage (A) |
PASS |
TestPBSDRBox_Unreachable_BlipBelowThreshold (B) |
PASS |
TestPBSDRBox_Unreachable_Recovery (C) |
PASS |
TestPBSDRBox_UsageUnsupported_IsNotBlindness (D) |
PASS |
TestPBSDRBox_BornBlind_StillReports (E) |
PASS |
TestOffsiteBox_Unreachable_AndRecovery (F) |
PASS |
TestPBSDRBox_ZeroCapacitySuccess_ClearsBlindness (§8 truth table) |
PASS |
TestBoxRecovery_ReachesTheOperatorDespiteInfoSeverity (G) |
PASS |
TestBoxUnreachable_ReachesTheOperatorOnItsOwnSeverity (G) |
PASS |
TestBoxRecovery_PairedWithTheCorrectDownType (G) |
PASS |
None of the three red-proofs passed on the first attempt — each turned its test red, and each did so for the reason under test, which I checked in the message rather than in the count.
Red-proof 1 — threshold 3 → 1. Group B seen failing:
box_reachability_test.go:132: two failed windows emitted [pbsdr_box_unreachable pbsdr_box_unreachable], want silence below the threshold
The message names the premature events, not an incidental error. Reverted (defaultBoxUnreachableWindows = 3 restored).
Red-proof 2 — remove the ErrUsageUnsupported counter guard (deleted its early return so the
branch falls through). Group D seen failing:
box_reachability_test.go:213: ErrUsageUnsupported emitted [pbsdr_box_unreachable × 8] — an expected pre-update condition must never alert
The message names the unexpected event type, as the prompt required — not merely a count. Reverted.
Red-proof 3 — remove pbsdr_box_recovered from recoveredPairedDownTypes. Group G seen failing,
and the first failure is the end-to-end mail assertion, which is what proves the test exercises
the wiring rather than the map:
dispatcher_box_reachability_test.go:41: pbsdr_box_recovered: operator mails = 0, want 1 — the all-clear must reach the operator; 0 means the recoveredPairedDownTypes entry is missing and "info" was dropped by the severity gate
Reverted. grep -rn MUTATED internal/ returns nothing.
5. Test count
go test ./... — 21 packages, all green (18 with tests, 3 with none). internal/monitor gained 7
tests; internal/notify gained 3. go build ./... and go vet ./... clean.
Repo gates: 10/10 OK, rc=0.
6. Deployed version and the wiring evidence
ArgoCD app "felhom": Synced Healthy rev=c03f629d43ed3d175e2c8561486cadc9a432640f
pod: hub-654bbc8fbc-9wld9 1/1 Running
running image: gitea.dooplex.hu/admin/felhom-hub:0.106.0
The required post-deploy check — both constructor log lines carrying the threshold:
19:29:07 [INFO] Offsite pool-box checker initialized: box=611714 fill warn=80% crit=90%,
oversub warn=2.00x, unreachable after 3 consecutive failed reads, refresh 15m0s
19:29:07 [INFO] PBS-DR box checker initialized: fill warn=80% crit=90%,
unreachable after 3 consecutive failed reads, refresh 15m0s
Two lines, both carrying the threshold — the parameter reached both checkers. Their absence would
have meant the config was inert however green the tests were. Note this also exercised the
absent-key path: box_unreachable_windows is deliberately not in any deployed config, so both
checkers fell back to the documented default of 3, which is what the log shows.
7. NOT yet live-validated — explicitly
No real or constructed endpoint outage has exercised the emit path end to end. Everything in §4
is an injected fake with a scripted error and an injected clock. What is proven: the checkers emit the
right events with the right details, and the dispatcher routes both new *_recovered types to a real
operator mail. What is not proven: that a genuine ep0 or Hetzner failure produces those errors in
the shape the checkers expect.
A real outage cannot be manufactured without making ep0 or the Hetzner API unreachable, and ep0 is Tier 2 protected — that was not done. The constructed-outage option, named but not performed: point the tenantsync client at a blackholed address on a scratch hub instance and let three windows elapse.
8. Teardown
This run provisioned nothing. No VM, no container beyond the hub's own rolling deployment, no
drill target, no scratch guest. All three layers N/A. ep0, both demo boxes and the drill VM were
untouched, as were the agent's credential-consume and self-heal paths.
9. Register rows
R-339 opened and marked SHIPPED (hub v0.106.0), with PROVEN-LIVE explicitly still owed and an instruction not to close it on the unit tests.
R-340 opened, READY (M) — the honest boundary: the hub's ep0 read is the usage op, which rides
the local API daemon, and the 2026-08-18 incident explicitly cleared that daemon while the HTTPS
proxy on 8007 was wedged. R-339's check would have shown green for all 9 h 37 m of the outage that
motivated it. Overlap with R-336's remaining half is noted so whichever runs second reuses the
first's evidence rather than re-measuring a protected machine.
R-336's next-step cell corrected. The replacement text, verbatim:
CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist. This cell used to read "PVE storage status is the prime suspect, and its interval is tunable". The first half is right and the second half is false.
pvestatdstats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (pvesm set <id> --disable 1) around the backup window, and that is substantially more than a tuning knob: it collides withfelhom-agent/internal/pbsdr/manager.go's health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. Doc-only correction — no agent code was changed. The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3.
10. Part 6 — DONE, not dropped
Both gates now announce the FELHOM_GATE_TODAY override loudly before any verdict, and
instructions_gate.py no longer swallows a malformed one — it used to fall through to the real
date in silence while due_checks_gate.py already exited 2 on the same input, so one variable had two
gates disagreeing about what a mistake means. Both exit 2 now.
Tests extended in both suites (42 and 73 assertions, all green). Red-proof: the announcement
was deleted from due_checks_gate.py and its two assertions were seen failing —
P6: valid override is announced and P6: the announcement says the real date is being ignored —
then reverted.
11. Gate and CI status
python3 scripts/repo_gates.py → rc=0, all ten gates OK, including due-checks.
The due-checks gate did NOT refuse this push. R-341's first check comes due 2026-08-19 UTC and
this work ran on 2026-08-18 (17:18–19:30 UTC), so the gate reported "2 dated check(s) pending, none
due yet" throughout. No row was cleared, no date edited, no --no-verify used. Every push today
went through the armed hook.
CI, by run ID — all three pushes green:
| run | head_sha | conclusion |
|---|---|---|
| 357 | ab2262c91 |
success — the code, tests and register edits |
| 358 | c03f629d4 |
success — the manifest bump that deployed it |
| 359 | 104ef34f5 |
success — Part 6 and this report |
(Run 356 on 78a244bf0, the baseline, was also green — so these greens are attributable to this
work rather than inherited from a red baseline. Per §13 of the prompt, a red run here would have been
mine.)
12. unproven.py --summary
where felhom stands — 55 claims, verified_on 2026-08-09
walked 23
partial 14 (6 cite evidence, 8 prose only)
built 14 (0 cite evidence, 14 prose only)
missing 4 (0 cite evidence, 4 prose only)
NOT WALKED: 32 of 55
No number moved. Correct: this shipped an implemented-not-proven capability, which is exactly the status that does not advance the walked count. Moving it would require the live validation §7 says has not happened.
13. Observations — noticed, deliberately not acted on
make docker-pushalso tags and pushes:latest, which the project's own rules forbid. I usedmake dockerfollowed by an explicitdocker push …:0.106.0instead, so no:latestwas moved. The Makefile target is a loaded gun for anyone who runs the documented command; not changed here because it is outside this task's scope.internal/monitor/storage_fill_test.gois not gofmt-clean, and was already so atHEAD— confirmed by stashing my changes and re-runninggofmt -l. Not touched; it is not mine and fixing it would put unrelated churn in this diff.- The two checkers are now ~95% identical in their reachability half. A shared helper is the obvious next move and was deliberately not done here, per the prompt: they have different sources, different error taxonomies (one has a sentinel, one does not) and different snapshot types, and the existing code keeps them separate on purpose. Worth revisiting if a third box checker appears.
- The first ArgoCD sync reported
Synced/Healthyat the PREVIOUS revision (ab2262c) while the pod was stillContainerCreating. Waiting and re-reading gavec03f629and the correct image. A sync status sampled too early is not the deploy's verdict — the running image tag is. alerting.box_unreachable_windowsis in no deployed config file, by design, so the live hub is running on the compiled default. If the operator wants to tune it, the key has to be added to the hub ConfigMap first.