Files
felhom.eu/REPORT-hub-blindness.md
T
2026-08-18 19:33:08 +02:00

235 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — the hub says something when it loses sight of the off-site endpoints (2026-08-18)
**Shipped: hub v0.106.0, deployed and verified.** Both box checkers now carry a second, independent
**reachability** signal with paired all-clears. The fill logic is untouched. **Part 6 was done, not
dropped.**
**NOT proven live** — see §7. No real or constructed outage has exercised the emit path.
---
## 1. Confirmed baselines
| item | value |
|---|---|
| felhom.eu `main` @ start | `78a244bf0f…` — **matches the prompt's anchor** |
| hub version in → out | **v0.105.0 → v0.106.0** (read from `hub/CHANGELOG.md` head) |
| `scripts/` version in → out | `due_checks_gate.py v1.0.0` → **v1.0.1** |
| highest R in use at start | R-343, so **R-339 / R-340 free** as specified |
**Had the three target files moved?** No. Verified by hash before editing:
```
61a16462756099d2fa60dd0c50aeac8c internal/monitor/pbsdr_box.go
d12b3e665d729a0291d2397895f23d1f internal/monitor/offsite_box.go
702d17487ffb4ac17a9d18050bb546b7 internal/notify/dispatcher.go
```
All landmarks in §5 of the prompt resolved as described; nothing was stale.
## 2. Files created / modified
**Created:** `hub/internal/monitor/box_reachability_test.go`,
`hub/internal/notify/dispatcher_box_reachability_test.go`, `REPORT-hub-blindness.md`.
**Modified:** `hub/internal/monitor/pbsdr_box.go`, `hub/internal/monitor/offsite_box.go`,
`hub/internal/notify/dispatcher.go`, `hub/cmd/hub/main.go`, `hub/CHANGELOG.md`,
`hub/internal/monitor/{pbsdr_box_test.go,offsite_box_test.go}` (new constructor arg),
`manifests/hub.yaml`, `CONTEXT.md`, `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md`,
`documentation/architecture/00-capability-map.md`, `scripts/due_checks_gate.py`,
`scripts/instructions_gate.py`, `scripts/test_due_checks_gate.py`,
`scripts/test_instructions_gate.py`, `scripts/CHANGELOG.md`.
## 3. Commits pushed to `main`
| commit | contents |
|---|---|
| `ab2262c91c1e976ae3b983e9df6b45dabfcb9d23` | the code, tests, and register/doc edits |
| `c03f629d43ed3d175e2c8561486cadc9a432640f` | `manifests/hub.yaml` 0.105.0 → 0.106.0 — the change that actually deploys |
| *(Part 6 commit — see §10)* | the today-override announcement + `scripts/CHANGELOG.md` |
## 4. Tests and the three red-proofs
**All named tests pass.** Groups A–F in `internal/monitor/box_reachability_test.go`, Group G in
`internal/notify/dispatcher_box_reachability_test.go`:
| test | result |
|---|---|
| `TestPBSDRBox_Unreachable_SustainedOutage` (A) | PASS |
| `TestPBSDRBox_Unreachable_BlipBelowThreshold` (B) | PASS |
| `TestPBSDRBox_Unreachable_Recovery` (C) | PASS |
| `TestPBSDRBox_UsageUnsupported_IsNotBlindness` (D) | PASS |
| `TestPBSDRBox_BornBlind_StillReports` (E) | PASS |
| `TestOffsiteBox_Unreachable_AndRecovery` (F) | PASS |
| `TestPBSDRBox_ZeroCapacitySuccess_ClearsBlindness` (§8 truth table) | PASS |
| `TestBoxRecovery_ReachesTheOperatorDespiteInfoSeverity` (G) | PASS |
| `TestBoxUnreachable_ReachesTheOperatorOnItsOwnSeverity` (G) | PASS |
| `TestBoxRecovery_PairedWithTheCorrectDownType` (G) | PASS |
**None of the three red-proofs passed on the first attempt** — each turned its test red, and each did
so **for the reason under test**, which I checked in the message rather than in the count.
**Red-proof 1 — threshold 3 → 1.** Group B seen failing:
> `box_reachability_test.go:132: two failed windows emitted [pbsdr_box_unreachable pbsdr_box_unreachable], want silence below the threshold`
The message names the premature events, not an incidental error. **Reverted** (`defaultBoxUnreachableWindows = 3` restored).
**Red-proof 2 — remove the `ErrUsageUnsupported` counter guard** (deleted its early return so the
branch falls through). Group D seen failing:
> `box_reachability_test.go:213: ErrUsageUnsupported emitted [pbsdr_box_unreachable × 8] — an expected pre-update condition must never alert`
The message names the **unexpected event type**, as the prompt required — not merely a count.
**Reverted.**
**Red-proof 3 — remove `pbsdr_box_recovered` from `recoveredPairedDownTypes`.** Group G seen failing,
and the first failure is the **end-to-end mail assertion**, which is what proves the test exercises
the wiring rather than the map:
> `dispatcher_box_reachability_test.go:41: pbsdr_box_recovered: operator mails = 0, want 1 — the all-clear must reach the operator; 0 means the recoveredPairedDownTypes entry is missing and "info" was dropped by the severity gate`
**Reverted.** `grep -rn MUTATED internal/` returns nothing.
## 5. Test count
`go test ./...` — **21 packages, all green** (18 with tests, 3 with none). `internal/monitor` gained 7
tests; `internal/notify` gained 3. `go build ./...` and `go vet ./...` clean.
Repo gates: **10/10 OK, rc=0**.
## 6. Deployed version and the wiring evidence
```
ArgoCD app "felhom": Synced Healthy rev=c03f629d43ed3d175e2c8561486cadc9a432640f
pod: hub-654bbc8fbc-9wld9 1/1 Running
running image: gitea.dooplex.hu/admin/felhom-hub:0.106.0
```
**The required post-deploy check — both constructor log lines carrying the threshold:**
```
19:29:07 [INFO] Offsite pool-box checker initialized: box=611714 fill warn=80% crit=90%,
oversub warn=2.00x, unreachable after 3 consecutive failed reads, refresh 15m0s
19:29:07 [INFO] PBS-DR box checker initialized: fill warn=80% crit=90%,
unreachable after 3 consecutive failed reads, refresh 15m0s
```
**Two lines, both carrying the threshold — the parameter reached both checkers.** Their absence would
have meant the config was inert however green the tests were. Note this also exercised the
**absent-key** path: `box_unreachable_windows` is deliberately not in any deployed config, so both
checkers fell back to the documented default of 3, which is what the log shows.
## 7. NOT yet live-validated — explicitly
**No real or constructed endpoint outage has exercised the emit path end to end.** Everything in §4
is an injected fake with a scripted error and an injected clock. What is proven: the checkers emit the
right events with the right details, and the dispatcher routes both new `*_recovered` types to a real
operator mail. What is **not** proven: that a genuine ep0 or Hetzner failure produces those errors in
the shape the checkers expect.
A real outage cannot be manufactured without making ep0 or the Hetzner API unreachable, and **ep0 is
Tier 2 protected — that was not done.** The constructed-outage option, named but not performed: point
the tenantsync client at a blackholed address on a **scratch** hub instance and let three windows
elapse.
## 8. Teardown
**This run provisioned nothing.** No VM, no container beyond the hub's own rolling deployment, no
drill target, no scratch guest. All three layers N/A. `ep0`, both demo boxes and the drill VM were
untouched, as were the agent's credential-consume and self-heal paths.
## 9. Register rows
**R-339 opened and marked SHIPPED** (hub v0.106.0), with PROVEN-LIVE explicitly still owed and an
instruction not to close it on the unit tests.
**R-340 opened, READY (M)** — the honest boundary: the hub's ep0 read is the `usage` op, which rides
the **local API daemon**, and the 2026-08-18 incident explicitly cleared that daemon while the HTTPS
proxy on 8007 was wedged. **R-339's check would have shown green for all 9 h 37 m of the outage that
motivated it.** Overlap with R-336's remaining half is noted so whichever runs second reuses the
first's evidence rather than re-measuring a protected machine.
**R-336's next-step cell corrected. The replacement text, verbatim:**
> **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to
> read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is
> right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second
> cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is
> no knob to turn down. The only lever PVE actually offers is disabling the storage entry
> (`pvesm set <id> --disable 1`) around the backup window, and that is **substantially more than a
> tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an
> inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a
> design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports
> reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.**
> The remaining step is unchanged: cut the poll rate by whatever means survives that question, then
> confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3.
## 10. Part 6 — DONE, not dropped
Both gates now announce the `FELHOM_GATE_TODAY` override loudly before any verdict, and
**`instructions_gate.py` no longer swallows a malformed one** — it used to fall through to the real
date in silence while `due_checks_gate.py` already exited 2 on the same input, so one variable had two
gates disagreeing about what a mistake means. Both exit 2 now.
Tests extended in both suites (**42** and **73** assertions, all green). Red-proof: the announcement
was deleted from `due_checks_gate.py` and its two assertions were seen failing —
`P6: valid override is announced` and `P6: the announcement says the real date is being ignored` —
then reverted.
## 11. Gate and CI status
`python3 scripts/repo_gates.py` → **rc=0, all ten gates OK**, including `due-checks`.
**The due-checks gate did NOT refuse this push.** R-341's first check comes due 2026-08-19 UTC and
this work ran on 2026-08-18 (17:18–19:30 UTC), so the gate reported *"2 dated check(s) pending, none
due yet"* throughout. **No row was cleared, no date edited, no `--no-verify` used.** Every push today
went through the armed hook.
**CI, by run ID — all three pushes green:**
| run | head_sha | conclusion |
|---|---|---|
| **357** | `ab2262c91` | success — the code, tests and register edits |
| **358** | `c03f629d4` | success — the manifest bump that deployed it |
| **359** | `104ef34f5` | success — Part 6 and this report |
(Run 356 on `78a244bf0`, the baseline, was also green — so these greens are attributable to this
work rather than inherited from a red baseline. Per §13 of the prompt, a red run here would have been
mine.)
## 12. `unproven.py --summary`
```
where felhom stands — 55 claims, verified_on 2026-08-09
walked 23
partial 14 (6 cite evidence, 8 prose only)
built 14 (0 cite evidence, 14 prose only)
missing 4 (0 cite evidence, 4 prose only)
NOT WALKED: 32 of 55
```
**No number moved.** Correct: this shipped an implemented-not-proven capability, which is exactly the
status that does not advance the walked count. Moving it would require the live validation §7 says
has not happened.
## 13. Observations — noticed, deliberately not acted on
- **`make docker-push` also tags and pushes `:latest`**, which the project's own rules forbid. I used
`make docker` followed by an explicit `docker push …:0.106.0` instead, so no `:latest` was moved.
The Makefile target is a loaded gun for anyone who runs the documented command; not changed here
because it is outside this task's scope.
- **`internal/monitor/storage_fill_test.go` is not gofmt-clean, and was already so at `HEAD`** —
confirmed by stashing my changes and re-running `gofmt -l`. Not touched; it is not mine and fixing
it would put unrelated churn in this diff.
- **The two checkers are now ~95% identical in their reachability half.** A shared helper is the
obvious next move and was deliberately not done here, per the prompt: they have different sources,
different error taxonomies (one has a sentinel, one does not) and different snapshot types, and the
existing code keeps them separate on purpose. Worth revisiting if a third box checker appears.
- **The first ArgoCD sync reported `Synced/Healthy` at the PREVIOUS revision** (`ab2262c`) while the
pod was still `ContainerCreating`. Waiting and re-reading gave `c03f629` and the correct image. A
sync status sampled too early is not the deploy's verdict — the running image tag is.
- **`alerting.box_unreachable_windows` is in no deployed config file**, by design, so the live hub is
running on the compiled default. If the operator wants to tune it, the key has to be added to the
hub ConfigMap first.