# REPORT — the hub says something when it loses sight of the off-site endpoints (2026-08-18) **Shipped: hub v0.106.0, deployed and verified.** Both box checkers now carry a second, independent **reachability** signal with paired all-clears. The fill logic is untouched. **Part 6 was done, not dropped.** **NOT proven live** — see §7. No real or constructed outage has exercised the emit path. --- ## 1. Confirmed baselines | item | value | |---|---| | felhom.eu `main` @ start | `78a244bf0f…` — **matches the prompt's anchor** | | hub version in → out | **v0.105.0 → v0.106.0** (read from `hub/CHANGELOG.md` head) | | `scripts/` version in → out | `due_checks_gate.py v1.0.0` → **v1.0.1** | | highest R in use at start | R-343, so **R-339 / R-340 free** as specified | **Had the three target files moved?** No. Verified by hash before editing: ``` 61a16462756099d2fa60dd0c50aeac8c internal/monitor/pbsdr_box.go d12b3e665d729a0291d2397895f23d1f internal/monitor/offsite_box.go 702d17487ffb4ac17a9d18050bb546b7 internal/notify/dispatcher.go ``` All landmarks in §5 of the prompt resolved as described; nothing was stale. ## 2. Files created / modified **Created:** `hub/internal/monitor/box_reachability_test.go`, `hub/internal/notify/dispatcher_box_reachability_test.go`, `REPORT-hub-blindness.md`. **Modified:** `hub/internal/monitor/pbsdr_box.go`, `hub/internal/monitor/offsite_box.go`, `hub/internal/notify/dispatcher.go`, `hub/cmd/hub/main.go`, `hub/CHANGELOG.md`, `hub/internal/monitor/{pbsdr_box_test.go,offsite_box_test.go}` (new constructor arg), `manifests/hub.yaml`, `CONTEXT.md`, `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`, `documentation/architecture/00-capability-map.md`, `scripts/due_checks_gate.py`, `scripts/instructions_gate.py`, `scripts/test_due_checks_gate.py`, `scripts/test_instructions_gate.py`, `scripts/CHANGELOG.md`. ## 3. Commits pushed to `main` | commit | contents | |---|---| | `ab2262c91c1e976ae3b983e9df6b45dabfcb9d23` | the code, tests, and register/doc edits | | `c03f629d43ed3d175e2c8561486cadc9a432640f` | `manifests/hub.yaml` 0.105.0 → 0.106.0 — the change that actually deploys | | *(Part 6 commit — see §10)* | the today-override announcement + `scripts/CHANGELOG.md` | ## 4. Tests and the three red-proofs **All named tests pass.** Groups A–F in `internal/monitor/box_reachability_test.go`, Group G in `internal/notify/dispatcher_box_reachability_test.go`: | test | result | |---|---| | `TestPBSDRBox_Unreachable_SustainedOutage` (A) | PASS | | `TestPBSDRBox_Unreachable_BlipBelowThreshold` (B) | PASS | | `TestPBSDRBox_Unreachable_Recovery` (C) | PASS | | `TestPBSDRBox_UsageUnsupported_IsNotBlindness` (D) | PASS | | `TestPBSDRBox_BornBlind_StillReports` (E) | PASS | | `TestOffsiteBox_Unreachable_AndRecovery` (F) | PASS | | `TestPBSDRBox_ZeroCapacitySuccess_ClearsBlindness` (§8 truth table) | PASS | | `TestBoxRecovery_ReachesTheOperatorDespiteInfoSeverity` (G) | PASS | | `TestBoxUnreachable_ReachesTheOperatorOnItsOwnSeverity` (G) | PASS | | `TestBoxRecovery_PairedWithTheCorrectDownType` (G) | PASS | **None of the three red-proofs passed on the first attempt** — each turned its test red, and each did so **for the reason under test**, which I checked in the message rather than in the count. **Red-proof 1 — threshold 3 → 1.** Group B seen failing: > `box_reachability_test.go:132: two failed windows emitted [pbsdr_box_unreachable pbsdr_box_unreachable], want silence below the threshold` The message names the premature events, not an incidental error. **Reverted** (`defaultBoxUnreachableWindows = 3` restored). **Red-proof 2 — remove the `ErrUsageUnsupported` counter guard** (deleted its early return so the branch falls through). Group D seen failing: > `box_reachability_test.go:213: ErrUsageUnsupported emitted [pbsdr_box_unreachable × 8] — an expected pre-update condition must never alert` The message names the **unexpected event type**, as the prompt required — not merely a count. **Reverted.** **Red-proof 3 — remove `pbsdr_box_recovered` from `recoveredPairedDownTypes`.** Group G seen failing, and the first failure is the **end-to-end mail assertion**, which is what proves the test exercises the wiring rather than the map: > `dispatcher_box_reachability_test.go:41: pbsdr_box_recovered: operator mails = 0, want 1 — the all-clear must reach the operator; 0 means the recoveredPairedDownTypes entry is missing and "info" was dropped by the severity gate` **Reverted.** `grep -rn MUTATED internal/` returns nothing. ## 5. Test count `go test ./...` — **21 packages, all green** (18 with tests, 3 with none). `internal/monitor` gained 7 tests; `internal/notify` gained 3. `go build ./...` and `go vet ./...` clean. Repo gates: **10/10 OK, rc=0**. ## 6. Deployed version and the wiring evidence ``` ArgoCD app "felhom": Synced Healthy rev=c03f629d43ed3d175e2c8561486cadc9a432640f pod: hub-654bbc8fbc-9wld9 1/1 Running running image: gitea.dooplex.hu/admin/felhom-hub:0.106.0 ``` **The required post-deploy check — both constructor log lines carrying the threshold:** ``` 19:29:07 [INFO] Offsite pool-box checker initialized: box=611714 fill warn=80% crit=90%, oversub warn=2.00x, unreachable after 3 consecutive failed reads, refresh 15m0s 19:29:07 [INFO] PBS-DR box checker initialized: fill warn=80% crit=90%, unreachable after 3 consecutive failed reads, refresh 15m0s ``` **Two lines, both carrying the threshold — the parameter reached both checkers.** Their absence would have meant the config was inert however green the tests were. Note this also exercised the **absent-key** path: `box_unreachable_windows` is deliberately not in any deployed config, so both checkers fell back to the documented default of 3, which is what the log shows. ## 7. NOT yet live-validated — explicitly **No real or constructed endpoint outage has exercised the emit path end to end.** Everything in §4 is an injected fake with a scripted error and an injected clock. What is proven: the checkers emit the right events with the right details, and the dispatcher routes both new `*_recovered` types to a real operator mail. What is **not** proven: that a genuine ep0 or Hetzner failure produces those errors in the shape the checkers expect. A real outage cannot be manufactured without making ep0 or the Hetzner API unreachable, and **ep0 is Tier 2 protected — that was not done.** The constructed-outage option, named but not performed: point the tenantsync client at a blackholed address on a **scratch** hub instance and let three windows elapse. ## 8. Teardown **This run provisioned nothing.** No VM, no container beyond the hub's own rolling deployment, no drill target, no scratch guest. All three layers N/A. `ep0`, both demo boxes and the drill VM were untouched, as were the agent's credential-consume and self-heal paths. ## 9. Register rows **R-339 opened and marked SHIPPED** (hub v0.106.0), with PROVEN-LIVE explicitly still owed and an instruction not to close it on the unit tests. **R-340 opened, READY (M)** — the honest boundary: the hub's ep0 read is the `usage` op, which rides the **local API daemon**, and the 2026-08-18 incident explicitly cleared that daemon while the HTTPS proxy on 8007 was wedged. **R-339's check would have shown green for all 9 h 37 m of the outage that motivated it.** Overlap with R-336's remaining half is noted so whichever runs second reuses the first's evidence rather than re-measuring a protected machine. **R-336's next-step cell corrected. The replacement text, verbatim:** > **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to > read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is > right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second > cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is > no knob to turn down. The only lever PVE actually offers is disabling the storage entry > (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a > tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an > inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a > design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports > reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** > The remaining step is unchanged: cut the poll rate by whatever means survives that question, then > confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. ## 10. Part 6 — DONE, not dropped Both gates now announce the `FELHOM_GATE_TODAY` override loudly before any verdict, and **`instructions_gate.py` no longer swallows a malformed one** — it used to fall through to the real date in silence while `due_checks_gate.py` already exited 2 on the same input, so one variable had two gates disagreeing about what a mistake means. Both exit 2 now. Tests extended in both suites (**42** and **73** assertions, all green). Red-proof: the announcement was deleted from `due_checks_gate.py` and its two assertions were seen failing — `P6: valid override is announced` and `P6: the announcement says the real date is being ignored` — then reverted. ## 11. Gate and CI status `python3 scripts/repo_gates.py` → **rc=0, all ten gates OK**, including `due-checks`. **The due-checks gate did NOT refuse this push.** R-341's first check comes due 2026-08-19 UTC and this work ran on 2026-08-18 (17:18–19:30 UTC), so the gate reported *"2 dated check(s) pending, none due yet"* throughout. **No row was cleared, no date edited, no `--no-verify` used.** Every push today went through the armed hook. **CI, by run ID — all three pushes green:** | run | head_sha | conclusion | |---|---|---| | **357** | `ab2262c91` | success — the code, tests and register edits | | **358** | `c03f629d4` | success — the manifest bump that deployed it | | **359** | `104ef34f5` | success — Part 6 and this report | (Run 356 on `78a244bf0`, the baseline, was also green — so these greens are attributable to this work rather than inherited from a red baseline. Per §13 of the prompt, a red run here would have been mine.) ## 12. `unproven.py --summary` ``` where felhom stands — 55 claims, verified_on 2026-08-09 walked 23 partial 14 (6 cite evidence, 8 prose only) built 14 (0 cite evidence, 14 prose only) missing 4 (0 cite evidence, 4 prose only) NOT WALKED: 32 of 55 ``` **No number moved.** Correct: this shipped an implemented-not-proven capability, which is exactly the status that does not advance the walked count. Moving it would require the live validation §7 says has not happened. ## 13. Observations — noticed, deliberately not acted on - **`make docker-push` also tags and pushes `:latest`**, which the project's own rules forbid. I used `make docker` followed by an explicit `docker push …:0.106.0` instead, so no `:latest` was moved. The Makefile target is a loaded gun for anyone who runs the documented command; not changed here because it is outside this task's scope. - **`internal/monitor/storage_fill_test.go` is not gofmt-clean, and was already so at `HEAD`** — confirmed by stashing my changes and re-running `gofmt -l`. Not touched; it is not mine and fixing it would put unrelated churn in this diff. - **The two checkers are now ~95% identical in their reachability half.** A shared helper is the obvious next move and was deliberately not done here, per the prompt: they have different sources, different error taxonomies (one has a sentinel, one does not) and different snapshot types, and the existing code keeps them separate on purpose. Worth revisiting if a third box checker appears. - **The first ArgoCD sync reported `Synced/Healthy` at the PREVIOUS revision** (`ab2262c`) while the pod was still `ContainerCreating`. Waiting and re-reading gave `c03f629` and the correct image. A sync status sampled too early is not the deploy's verdict — the running image tag is. - **`alerting.box_unreachable_windows` is in no deployed config file**, by design, so the live hub is running on the compiled default. If the operator wants to tune it, the key has to be added to the hub ConfigMap first.