diff --git a/REPORT.md b/REPORT.md index 12bf8df..408f4b5 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,51 +1,268 @@ -# REPORT — R-50 island control plane: Phase A artifacts (felhom.eu) (2026-07-25) +# REPORT — R-81: "no signal" is not "bad signal" (hub v0.75.0) (2026-07-26) -Phase A of the R-50 implementation (spike GO). Fresh **appliance** installs are now born on the island -(F1-immune). Pairs with agent **v0.96.0** (which attaches the guest's island net1). byo unchanged. +Class fix for the third instance of one bug: **absence of a signal treated as evidence of failure.** +Input: `documentation/audits/DIAG-backup-missed-2026-07-26.md` (`add5b9b`). -## Changed here -- `scripts/felhom-host-install.sh` → **v1.19.0**: `ensure_island_bridge()` (portless `vmbr9`, - `169.254.253.1/30`, idempotent); appliance default binds `local_api` on the island - (`169.254.253.1:8443`), writes `island_bridge`/`island_guest_addr`, and pins `lan_resolver.host_ip` - to the LAN IP explicitly (Finding-1 — else moving the bind silently kills LAN DNS). `--no-island` - opt-out; explicit `--bridge-ip` overrides. Gate `hostinstall_gates.py` ALL PASS. -- `hub/internal/web/configs.go` → `hostInstallVersion` 1.16.0 → **1.19.0** (F-1 sync; clears prior - drift). Hub CHANGELOG → v0.73.2. Web render test green. +--- -## Determination (A0/A3) — recorded -- **A0 = config-only for the bootstrap endpoint** (`main.go` sets it from `listen_addr`); the version - bump lives in the agent's net1-attach. The operator chose **agent-bringup** (v0.96.0) over golden-bake - to eliminate the golden-coupling failure mode. -- **A3 = healer verify-only** — eth0-only (`parseMode` dev-scoped); red-proof test in the agent. -- **net1 deviation from the spike inventory** (it placed net1 in host-install, but host-install doesn't - create the guest — the agent does) resolved by attaching net1 in the agent at bring-up. +## 1. Confirmed baselines -## Coupling / deploy order (important) -An island host-install REQUIRES agent ≥ 0.96.0. **Vouch 0.96.0 before island installs go live**, else a -fresh install writes island config an older agent ignores → guest has no island NIC → broken control -plane. This is the one hard ordering constraint. +| Repo | Start `main` | Version | → Shipped | +|---|---|---|---| +| felhom.eu (hub) | `add5b9b` | v0.74.0 | **v0.75.0** (`f5a5e2b` code+docs, `88b41ec` manifest pin) | +| felhom-agent | `dfd5d73` | v0.96.0 | unchanged — **N/A this task** (cause-fix is R-84) | +| felhom-controller | after R-77 | v0.173.0 | unchanged | -## Status of the phases -- **Phase A (artifacts): DONE.** Agent v0.96.0 committed (`dfd5d73`), **published** (sha `af938601…`) and - **VOUCHED** in the hub Day-0 manifest (agent 0.96.0; golden 0.161.0 / min_agent 0.93.0 / wrapper - preserved). host-install v1.19.0 + hub sync committed (`485321f`). Golden unchanged (net1 comes from the - agent). Live site serves 1.19.0. -- **Phase B (migration): B0+B1+A4 DONE; B2/B3 pending operator go.** - - **B0** `RUNBOOK-island-migration.md` written (`4650493`). - - **B1** drill migration ran verbatim → island `/storage` **HTTP 200**, LAN DNS held on the LAN IP - (Finding-1 pin), apps healthy, hub reports 0.96.0. No rollback. - - **A4** the agent auto-attached `net1` on the island at provision (scratch guest from golden, **zero - manual edits**) → the provisioning path is proven live (spike method caveat CLOSED). F1 replay + cold - reboot were proven on the same drill (spike P6/P7). *(A literal ISO wipe-reinstall was not run — the - surgical provision proves the R-50-specific behavior; the ISO/first-boot/day-0 glue is unchanged.)* - - **B2 demo-hp — DONE (2026-07-25).** Agent 0.95.0→0.96.0, vmbr9, guest 9201 net1, island bind, DNS pinned - to `.87`. Verify: island `/storage` **200**, LAN DNS OK, **7/7 apps unchanged** (cloudflared/traefik uptime - 42h = never restarted), public URL **302**, hub reports 0.96.0. No rollback. - - **B3 demo-felhom — DONE (2026-07-25).** Same procedure; DNS pinned to `.162`. Verify: island `/storage` - **200**, LAN DNS OK, **15/15 apps unchanged**, public URL **302**, hub reports 0.96.0. No rollback. - - Both boxes' `.bak-0.95.0` agent + `/etc/network/interfaces.pre-island.bak` retained for rollback. -- **Phase 5 (closing docs): DONE.** Capability-map "site/network change" row → **PROVEN-LIVE**; ROADMAP R-50 → - **SHIPPED (fleet-migrated)**; **R-74** added (island on Peti's cluster); nodes.md updated (both boxes island, - agent 0.96.0). **Firewall (runbook §8):** on both boxes the island bind already moved the listener off the - LAN (nothing on `:8443`) — the LAN surface is closed by the bind; the portless-bridge narrowing is moot. -- **Phase C (Peti cluster): parked as R-74** — its own supervised runbook (SDN/bridge parity), coordinated with Peti. +Clean-tree gate passed before the build: `git status --porcelain` empty, `HEAD == origin/main`. + +--- + +## 2. Phase-0 — the anchor probe + +**Question:** does the hub retain host-report HISTORY, or only the latest report? + +**Answer: HISTORY IS RETAINED — 90 days, indexed and queryable by customer + time.** + +- `host_reports` is an append-only table (`SaveHostReport` INSERTs; `GetLatestHostReportJSON` only + ever read the newest row). +- Retention: `cfg.Retention.MaxDays`, **90** both as the code default (`cmd/hub/main.go:631,654`) and + as the live value in `manifests/hub.yaml` (`retention.max_days: 90`). `Prune` deletes + `host_reports` older than that on the 04:30 job. +- Index `idx_host_reports_customer ON host_reports(customer_id, received_at DESC)` — the exact access + pattern needed, already present. +- Live confirmation on the production DB: 664 / 436 / 96 retained rows for demo-felhom / demo-hp / + drill-r50, spanning 7.63 / 4.64 / 0.91 days. + +**Branch taken: the free anchor.** No new persisted state, no `last_backup_seen` column, no agent +change. The agent's store is *point-in-time*; the hub has *memory*. Judging backup evidence across +the retained window is both the cheapest option and semantically the right question — **"when did I +last SEE evidence of a backup?"** rather than "what does this one report happen to say?". + +--- + +## 3. Files modified + +| File | Change | +|---|---| +| `hub/internal/monitor/deadline.go` | Three-valued verdict; hub-history fold-in; anchored absence; distinct reason strings; `newestBackupEvidence`; landmine comment on `backupStaleAfter`; new `backupEvidenceLookback` | +| `hub/internal/store/store.go` | `HostReportRow`, `GetHostReportsSince`, `GetFirstHostReportAt`, `SetHostReportsReceivedAtForTest` | +| `hub/internal/monitor/deadline_test.go` | Signature update; note explaining the zero-evidence rows | +| `hub/internal/monitor/deadline_anchor_test.go` | **NEW** — 15 tests, scenarios A–E + the named boundary contract | +| `hub/CHANGELOG.md` | v0.75.0 | +| `CONTEXT.md` | New **Standing rulings** section (S-1, S-2) + the v0.75.0 entry | +| `REUSE.md` | New canonical-helper row; new dangerous-lookalike row | +| `documentation/backlog/ROADMAP.md` | R-80 closed, R-81 shipped, **R-82 / R-83 / R-84 added** | +| `documentation/architecture/00-capability-map.md` | Note row — no status flips | +| `manifests/hub.yaml` | image pin 0.74.0 → **0.75.0** | + +Commits: **`f5a5e2b`** (code + docs), **`88b41ec`** (manifest pin). Both on `main`, pushed. + +--- + +## 4. The fix + +`assessBackupFreshness` returned `{missed bool, reason string}` and collapsed *absence of records* +into *failure*. It now returns a three-valued verdict: + +``` +verdictOK positive evidence of a recent backup +verdictUnknown no evidence yet, and the anchored window has not elapsed → silent, LOGGED +verdictMissed positive evidence of a problem → alarm +``` + +Three inputs, all injected — the function stays **pure**: + +1. the latest report (unchanged parsing), +2. `backupEvidence.newestSeen` — the newest backup evidence across a bounded 7-day window of retained + reports (`GetHostReportsSince` + `newestBackupEvidence`), +3. `backupEvidence.firstReportAt` — first contact (`GetFirstHostReportAt`), the absence anchor. + +Absence is graded against the **existing** `backupStaleAfter` (26 h), exactly as v0.73.0 reused +offsite `staleAfter` — **no new knob**. A zero anchor fails toward visibility (the v0.73.0 +legacy-shape precedent). The window scan early-exits at the first sufficiently-fresh evidence, so the +healthy path reads one row; only the genuinely-broken path walks the lookback. + +`CheckBackupDeadlines` logs each deferral (`verdict UNKNOWN`) and the summary line gained a +`backup unknown (deferred)` counter — a quiet check must never be indistinguishable from one that +did not run. + +The invariant is written at the head of `assessBackupFreshness`, naming all three instances +(v0.12.0, v0.73.0, R-81), and pinned by a boundary test whose name states what it protects. + +--- + +## 5. Test results + +**Full gate green:** `go build ./... && go vet ./... && go test ./...` — all rc=0, every package `ok`. + +**Test count: 493 → 508 (+15).** (Measured with `git grep -hE '^func Test'` at `HEAD` vs the working +tree — a first attempt via `go test -list` after a `git stash` gave a bogus 434 because the stash +broke compilation and silently dropped the whole package. Recorded because that is exactly the +exit-code-lie class this project has been bitten by; likewise `go vet ... | head` reported rc=0 while +vet was actually failing.) + +All 21 tests in `internal/monitor` covering this area pass, including the 6 pre-existing ones. + +### The three companion red-proofs — applied, observed, restored + +**A — Scenario A (the 07-26 case must not alarm).** Removed the `ev.haveSeen` fold-in (pre-R-81 +shape: judge the latest report alone): + +``` +--- FAIL: TestBackupFreshness_AgentRestartBlindWindow_NoAlarm (0.00s) + deadline_anchor_test.go:61: 07-26 shape must NOT alarm; got verdict=2 reason="newest backup is 176h0m0s old (limit 26h0m0s)" +--- FAIL: TestCheckBackupDeadlines_RestartBlindWindow_NoEvent (0.03s) + deadline_anchor_test.go:388: the 07-26 restart shape must NOT raise expected_backup_missed; got [expected_backup_missed] +``` + +The reason string is **verbatim the message demo-felhom actually sent to the customer channel** that +morning. Restored. + +**B — Scenario B (a genuinely dead box must still alarm).** Applied the *naive* fix — absence always +returns UNKNOWN, never a fault: + +``` +--- FAIL: TestBackupFreshness_NoEvidenceBeyondAnchor_Alarms (0.00s) + deadline_anchor_test.go:97: a box with NO backup evidence for 240h MUST alarm; got verdict=1 reason="no backup evidence yet, but only watching for 240h0m0s" +--- FAIL: TestBackupFreshness_Contract_AbsenceIsUnknownUntilAnchorElapses/just_outside_the_window + deadline_anchor_test.go:156: CONTRACT VIOLATED: absence beyond the window MUST alarm (watched=26h1m0s, limit=26h0m0s); got verdict=1 … +--- FAIL: TestBackupFreshness_Contract_AbsenceIsUnknownUntilAnchorElapses/long_past_the_window + deadline_anchor_test.go:156: CONTRACT VIOLATED: absence beyond the window MUST alarm (watched=720h0m0s, limit=26h0m0s); got verdict=1 … +--- FAIL: TestCheckBackupDeadlines_NeverBackedUpBeyondAnchor_Alarms (0.03s) + deadline_anchor_test.go:411: a host with no backup for 120h MUST alarm; got [] +``` + +**This is the proof that matters most.** A suite that only pinned Scenario A would have passed +against this over-suppressed implementation — and over-suppression is strictly worse than the bug it +replaces. Restored. + +**C — Scenario C (a fresh box is not born failing).** Restored the literal pre-fix branch +(`if !havePBS && !haveVzdump { return missed }`): + +``` +--- FAIL: TestBackupFreshness_EmptyArraysWithinGrace_Unknown (0.00s) + deadline_anchor_test.go:74: absence inside the anchored grace must NOT alarm; got reason="no PBS snapshot or successful backup in the latest host-report" +--- FAIL: …/newborn,_1_minute + deadline_anchor_test.go:159: CONTRACT VIOLATED: no evidence + no elapsed window must NOT alarm (watched=1m0s, limit=26h0m0s); got reason="no PBS snapshot or successful backup in the latest host-report" +--- FAIL: …/newborn,_1_hour (watched=1h0m0s) +--- FAIL: …/just_inside_the_window (watched=25h59m0s) +--- FAIL: …/exactly_at_the_window (watched=26h0m0s) +--- FAIL: TestCheckBackupDeadlines_NewbornHost_NoEvent (0.03s) + deadline_anchor_test.go:426: a newborn host must not alarm; got "No fresh verified backup: no PBS snapshot or successful backup in the latest host-report" +``` + +Restored; full suite green after each. + +### Scenario D and E + +- **D** — `TestBackupFreshness_ExistingBehavioursUnchanged` pins all three pre-existing outcomes + **including their exact reason strings** (`"newest backup is 30h0m0s old (limit 26h0m0s)"`, + `"newest PBS snapshot failed verification"`, `"latest host-report could not be parsed"`), each with + fresh window evidence present so the result cannot be an artefact of the new input. Plus + `TestBackupFreshness_WindowEvidenceDoesNotRescueFailedVerify` — the anchor must not suppress an + integrity fault. +- **E** — `TestBackupFreshness_ReasonStringsAreDistinct` fails if any two of the six failure modes + ever produce the same message. + +--- + +## 6. Replay of the 2026-07-26 shape — **actual reports, not a fixture** + +Two replays, both against the real thing. + +**(a) Real report payloads** extracted from the production hub DB (every retained report in the +7-day lookback at or before the check instant), fed through `newestBackupEvidence` + +`assessBackupFreshness` at `now = 2026-07-26 03:00:00 UTC`: + +``` +demo-felhom rows=600 first=2026-07-18T16:30:51Z windowEvidence=2026-07-25T06:30:14Z -> OK +demo-hp rows=417 first=2026-07-21T16:24:48Z windowEvidence=2026-07-25T10:23:31Z -> OK +drill-r50 rows= 77 first=2026-07-25T09:50:55Z windowEvidence=none -> UNKNOWN + "no backup evidence yet, but only watching for 17h0m0s (grace 26h0m0s since first + contact 2026-07-25T09:50:55Z) — newborn host, not a fault" +``` + +**(b) Through the real store queries** — a copy of the live `hub.db` opened with `store.New`, driving +`GetHostReportsSince` / `GetFirstHostReportAt` / `GetLatestHostReportJSON` for real. This was worth +doing separately: the window query compares `received_at` as a **SQLite datetime string**, which is a +genuine correctness risk that a hand-built fixture would not have exercised. Identical verdicts, plus: + +``` +peti-felhom NO host-report → deadline check skips the backup half entirely +``` + +**All three would now be silent. Zero events.** Both harnesses were throwaway and are deleted; the +tree is clean. Only the DB *copy* was ever opened writable — the live `/data/hub.db` was not touched. + +Note what the verdicts say: demo-felhom and demo-hp are **OK** (real evidence was found, not +suppressed), and drill-r50 is **UNKNOWN** (correctly deferred as a newborn, and it will alarm on its +own if it is still backup-less after the grace). None of the three is silenced by a blanket rule. + +--- + +## 7. Deployment + +| Step | Result | +|---|---| +| Image build + push | `gitea.dooplex.hu/admin/felhom-hub:0.75.0`, 25 MB, digest `sha256:a2ebaa50…10453` | +| Manifest pin | `manifests/hub.yaml` → `0.75.0` (commit `88b41ec`) — **GitOps only**, no `kubectl set image` | +| ArgoCD | hard-refresh + deliberate sync → **Synced / Healthy** | +| Rollout | `deployment "hub" successfully rolled out` | +| Pod | `hub-56946d5cb4-kwwn7` 1/1 Running | +| **Running image (pod status, not spec)** | `felhom-hub:0.75.0` @ `sha256:a2ebaa50…10453` — **matches the pushed digest** | +| Startup log | `[INFO] felhom-hub 0.75.0 starting` … `[INFO] deadline-check: next run at 2026-07-27 05:00 CEST (in 17h14m19s)` | + +The pin landed and matches — checked deliberately, since the DIAG found `hub.yaml` pinned at +`0.73.1` while the CHANGELOG read `0.73.2`. + +--- + +## 8. NOT yet live-validated — explicit + +1. **The 05:00 CEST deadline check has not fired under v0.75.0.** Next run 2026-07-27 05:00 CEST. All + evidence above is unit-proven + replayed against real data; the live firing is unobserved. +2. **The deferred-UNKNOWN INFO line has not appeared in a production log** — proven by unit test only. + drill-r50's 26 h grace expires 2026-07-26 ~11:50 UTC, so by 03:00 on 07-27 it will instead + **alarm** — correctly, and that will be the first real test of Scenario B in production. +3. **No genuinely-dead box has been observed alarming live.** The B path is unit-proven only. +4. **The cold path (full 7-day lookback scan) has not run in production.** Cost is bounded and the + healthy path early-exits at one row, but the scan-heavy branch is untimed live. +5. **demo-hp and drill-r50 host-level state remains uncollected** (no SSH key; break-glass not used) — + unchanged from the DIAG. + +--- + +## 9. Observations — recorded, not acted on + +1. **`parseBackupTime`'s silent `continue` is a latent member of the same class.** An unparseable + timestamp is dropped without a trace, so a wire-format change on the agent side would present as + "no backup" — the exact 07-26 failure, with no log line to catch it. Left untouched per scope + (the agent emits clean RFC3339 `Z`, confirmed in the DIAG). Worth a WARN log when R-84 touches the + agent side of this contract. +2. **The customer-facing Hungarian copy is unchanged.** `notify/templates.go:106` — *"A mai + biztonsági mentés nem készült el a határidőig!"* — still overstates scope (it reads as *all* + backups failed; this check only covers the host/PBS tier). Not in this task's Parts 1–3. The DIAG + proposed it as fix #4; it needs a `felhom-ui-design` copy pass and the confirm gate. +3. **`GetFirstHostReportAt` returns the prune horizon, not true first contact, for hosts older than + 90 days.** Accepted and documented at the function: it only makes the anchor *more* conservative + for long-lived hosts (the window has elapsed either way) and never shortens a newborn's grace. +4. **The `backupStaleAfter` landmine is now written into the constant's comment** and carried as a + named dependency in R-82. It will alarm on a healthy weekly PBS snapshot six days in seven. +5. **`retention.max_days: 90` is now load-bearing for a monitor**, not just for storage. Shortening it + below `backupEvidenceLookback` would silently narrow the evidence window. Not currently guarded. +6. **The R-80 real finding stands untouched:** the PBS/offsite-DR tier has no schedule at all — + demo-felhom still holds exactly one PBS snapshot (2026-07-18), demo-hp zero. R-81 fixes how the hub + *reads* backup state; it does nothing about the DR tier being empty. **That is R-82, and it is the + more serious of the two.** +7. The customer-channel suppression the DIAG recommended (removing `expected_backup_missed` from + demo-felhom's `enabled_events`) was **not** applied — still an operator decision, and now much less + pressing since the false-positive path is closed. + +--- + +## 10. Scope discipline + +Not started, recorded only: **R-82** (backup target split — the real finding), **R-83** (ratify or +retire `07-backup-architecture.md`), **R-84** (persist the agent's backup store — the cause-fix). +No branches; every commit staged with explicit paths (never `git add -A`, per the shared-worktree +rule); no `Co-Authored-By`.