diff --git a/CONTEXT.md b/CONTEXT.md index bf85d4f3..0fa2641a 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -15,6 +15,31 @@ > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## Box REACHABILITY is a separate signal from box FILL — and the cadence difference is deliberate (2026-08-18, R-339, hub v0.106.0) + +Both off-site checkers now carry two independent signals, and conflating them is the mistake to avoid: + +- **FILL** — how full is the store. Escalation-only (one emit on a band rise), and **a degraded or + missing reading drives NO band transition**, because missing ≠ 0%. That rule is unchanged by + v0.106.0 and must stay unchanged: it is what stops a dead endpoint from reading as an empty one. +- **REACHABILITY** — can the hub see the store at all. Counted in consecutive failed fetch windows, + emitted past a threshold (default 3 ≈ 30–45 min) with a paired all-clear. + +**Why reachability REPEATS rather than escalating once.** `emitFill`/`emitOversub` fire once on a band +rise and stay silent while the condition persists, which is right for a capacity trend. Applied to +reachability it would produce exactly ONE mail at roughly minute 30 of a nine-hour outage — and one +mail is missable, which is the whole failure being fixed. So the unreachable event fires on every +failed window past the threshold and relies on the dispatcher's 1-hour per-type operator cooldown to +become an hourly "still blind" heartbeat. **If you find yourself "fixing" this back to the band shape, +this paragraph is why not.** The cost is a `suppressed` notification row every ~15 min during an +outage — honest bookkeeping. + +**Two boundaries worth holding in mind.** `ErrUsageUnsupported` is *not* blindness — an ep0 on an old +tenantsync answers "no such op", which means we reached it; the counter is not advanced, or the alert +would fire for days on a healthy pre-update box. And the reachability read rides ep0's **local API +daemon**, not the HTTPS proxy on 8007 — so it would have shown green throughout the 2026-08-18 outage, +which is R-340 and is the honest limit of what this watches. + ## The managed floor now tracks the vouched golden — and that is rule 2 ARRIVING, not an exception (2026-08-18, R-343) Since 2026-08-18 12:36:58Z the global managed controller floor diff --git a/STATUS.md b/STATUS.md index 886614f8..c5f5ad63 100644 --- a/STATUS.md +++ b/STATUS.md @@ -43,6 +43,16 @@ record with no machine** — created 13 August, no host, no backups, nothing to ## Shipped +- **The system now tells you when it cannot see the off-site copies** (R-339). Until today, a + completely dead off-site store and a perfectly healthy one looked **identical** to you — the checks + only ever watched how full a store was getting, and a failed reading was written to a log nobody + reads. That is why Monday's nine-and-a-half-hour outage reached you only by accident, through the + weekly backup that happened to fall inside it. After about half an hour of being unable to see a + store you now get a mail, repeated hourly while it lasts, and one all-clear when it comes back. + **Caveat worth knowing:** this watches whether the machine answers at all — it would *not* have + caught Monday's exact fault, which was one service wedged while the machine stayed healthy. That + second check is written down as the next step. + - **The auto-update floor is current again** (R-343). You raised it to today's version this afternoon. It was **not** left behind by accident — our own rule says the floor is raised *last*, after the image is vouched, because it acts within seconds. It did: **one machine updated itself diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 63a7b66d..1dc54327 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -157,5 +157,6 @@ | Per-customer offsite fill + staleness + freeze lever | hub v0.41 | **IMPLEMENTED** | `OffsiteChecker` (`hub/internal/monitor/offsite.go`): fill 90/95% vs soft quota, staleness >48h | No live-fired leg: `CAMPAIGN-offsite-overnight-2026-07-10` recorded no quota/fill/staleness emails, and the freeze write-block was **inconclusive** (only the Hetzner `readonly:true` API op succeeded). Demoted | | **Box-level Storage Box aggregate (total fill, Σ quotas, oversubscription alert)** | hub v0.64.0 (R-5) | **IMPLEMENTED** (data pipeline PROVEN-LIVE) | `monitor.OffsiteBoxChecker` — fetch-throttled Hetzner GET (1/15 min), fill (used/`storage_box_type.size`, 80/90%) + oversubscription (Σ shared+enabled ConfigJSON quotas / capacity, 2.0×), escalation-only operator alert on the customer-less `"pool-box"` scope; Offsite-tab panel + dashboard tile. **Phase-0-pinned** live shape (box 611714) + **live-computed in-cluster:** `0.2% full (2.6 GB of 1.00 TB), Σ shared quota 150 GB, oversub 0.15x`. Tests + 4 red-proofs; hub v0.64.0 REPORT | Two open legs: the **UI render** is unit-verified only (hub UI password-gated → no screenshot); the **alert emails** are unit + red-proof verified but NOT fired live (real pool nominal — a live-fire emails Viktor). Thresholds pending Viktor's ruling (named config keys). READ-ONLY (GET) | | **Operator sees PBS DR datastore fill at a glance (Offsite "PBS DR" tab + dashboard gauge)** | hub v0.65.0 + tenantsync v1.2.0 (R-5) | **IMPLEMENTED** (data pipeline PROVEN-LIVE) | The PBS DR datastore (`felhom-offsite` on ep0) fill — NOT a Hetzner box. **Option A:** a read-only `usage` op on the `felhom-tenantsync` ep0 forced command (twin of `fingerprint`; `df` on the datastore path — no customer_id, no admin token, NO mutation), polled by `monitor.PBSDRBoxChecker` (OffsiteBoxChecker clone; 15-min throttle; states ok/unavailable/degraded; fill 80/90% on the `"pbsdr-box"` operator scope). `/offsite` split into Restic + PBS DR tabs; two dashboard gauges. **Graceful: hub deploy ⟂ ep0 update** (ep0 ≤ v1.1.0 → gauge "n/a" until updated). **Phase-0-pinned** (`df` on ep0 PBS 4.2.3) + **live-computed in-cluster** (ep0 updated to v1.2.0 this session): `19.1% full (7.1 GB of 37.2 GB)`. 10 Go tests + a bash harness + 3 red-proofs; hub v0.65.0 REPORT | Open legs: **UI render** unit-verified only (hub UI password-gated); the **fill alert email** is unit + red-proof verified, NOT fired live (datastore nominal at 19%). Separate PBS threshold keys (default 80/90); no oversubscription (namespaces, not quotas). READ-ONLY | +| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) | | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | \ No newline at end of file diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index ff8a5ccf..cafc0040 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -636,7 +636,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-333** | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | | **R-334** | **CLOSED 2026-08-18 — golden 0.216.0 baked, published and VOUCHED; CI green by run id.** ~~WAIVER + open item: controller v0.215.0 is released and deployed, and NO golden carries it.~~ Convicted by `golden_currency_gate.py` on the 2026-08-14 push: newest released controller **0.215.0**, newest golden bake **0.214.0** (`documentation/tests/golden-0.214.0-2026-08-12`). **A machine installed right now receives 0.214.0** — i.e. a brand-new box would ship WITHOUT the R-328 severity fix and would keep emailing nobody about a failing disk. The running fleet is unaffected (demo-hp guest 9201 is on 0.215.0 and healthy); this is purely the day-0 install path. **Not baked in this session deliberately:** the task scoped deployment to demo-hp only, and the second half of the fix — vouching the bake in the hub's day-0 artifact manifest — is **operator-password-gated, so CC cannot complete it**; a baked-but-unvouched golden is worse than none. **The push was made with `git push --no-verify` and it is stated here and in the session report**, per `.claude/rules/gates.md` — the gate has no waiver parser, so recording a waiver does not clear it. **STILL OPEN and now one version WIDER, 2026-08-18:** the newest released controller is **0.216.0** (v0.215.0's own follow-up fix, R-335) and the newest bake is still **0.214.0**, so a new install now misses *two* releases. Re-convicted on this date's documentation-only push, which was likewise made with `--no-verify`; the gate reads `felhom-controller/CHANGELOG.md` and `documentation/tests/golden-*`, **neither of which that session touched** — the conviction is inherited, not caused | **CLOSED 2026-08-18.** Baked from `RUNBOOK-manual-build.md` §4.0+§4.1 in the DooPlex drill VM and published: **`GOLDEN_VERSION=0.216.0`**, **`GOLDEN_SHA256=ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b`**, 656,970,239 bytes at `…/generic/felhom-golden/0.216.0/golden.tar.zst`. **The published bytes were verified, not just the script's print** — the artifact was downloaded back out of Gitea and hashed, and it matches. **Vouched by the operator, all THREE fields together**, confirmed by reading the hub's own store rather than the save: `artifact_golden_version=0.216.0`, `artifact_agent_version=0.129.0`, `artifact_min_agent=0.129.0` (2026-08-18 11:00:59–11:01:00), and the hub's recorded sha256 matches the downloaded artifact. The R-216 shape was checked on the machine: `MinAgent` 0.129.0 is **equal to**, not above, the newest **published** agent. **`golden_currency_gate.py` rc=0 and `repo_gates.py --fast` rc=0 — all nine gates — and CI is GREEN BY RUN ID: run **353**, `head_sha 7d81681d6`, conclusion `success`** (the two prior runs 351/352 on this same afternoon were red on exactly this row, which is the contrast). That push needed **no `--no-verify`** — the first of the day that did not. Evidence: `documentation/tests/golden-0.216.0-2026-08-18/`, report `REPORT-golden-0.216.0.md`. **Closed with the run id quoted deliberately**: this row was re-confirmed once and widened once, and closing it on a local green a third time would have left the same ambiguity | — | Bake a golden on **0.216.0** per `runbooks/RUNBOOK-manual-build.md` §4.1, then vouch it — a THREE-field change (`golden_version` + `agent_version` + `min_agent`). Until then every NEW install lacks the severity fix | CC bakes; **Viktor vouches** | | **R-335** | **One physical disk was walked TWICE per run, and the second walk sustained it against itself.** Found on live hardware ~2h after the v0.215.0 deploy, **by noticing the release's own positive observable disagreed with its own persisted artefact**: the hourly check logged *"3 disk(s) evaluated"* while `disk-health-state.json` held **two** records. Cause: demo-hp's `c11-scratch` and `felhom-backup` are the same physical NVMe (`/dev/nvme0n1`) and resolve to the same `diskKey`. **Not cosmetic** — `RunDiskHealthCheck` writes a disk's new record before the next entry reads it, so the SECOND copy consumed the FIRST copy's write as its prior: the disk **sustained against itself and reached Hiba on a FIRST sighting**, defeating truth-table row 6 — the exact rule separating a one-hour benign excursion from a false critical — and would have emitted **two identical events** for one drive. **Latent, not active, on demo-hp** (all three entries healthy, zero counters), but any aliased disk developing a single pending sector would have gone straight to Hiba. **This is the shape standing rule 3 warns about: an absent alarm was not evidence — the two artefacts had to be read AGAINST each other** | **CLOSED — controller v0.216.0, 2026-08-14.** Each `diskKey` is evaluated once per run; both entries stay marked `seen` so neither looks like a disappeared disk, and the card still renders both storage rows (the dedup is about state and alerts, not display). Pinned by `TestDiskCheck_SameDiskTwiceIsEvaluatedOnce`; companion red-proof run and reverted — deleting the guard makes the first sighting emit `Kind:2` (Hiba-from-sectors) at 8 sectors | — | — | CC | -| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | Find what polls `felhom-pbs` this hard (PVE storage status is the prime suspect, and its interval is tunable) and cut it; then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 | CC | +| **R-336** | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 | CC | | **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | | **R-338** | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-341** | **Does the fd slope change after the PBS 4.2.5 upgrade? — two dated checks, and the answer is expected to be NO.** ep0 was upgraded 4.2.2-1 → 4.2.5-1 on 2026-08-18 09:51Z on the operator's ruling, **for rehearsal value, not as a fix**: the full changelog range was read (128 lines, all three entries) and swept for connection-handling vocabulary, and it contains **no mechanism** by which descriptor reaping would change — the single keyword hit was `S3 … honor the node's proxy settings`, HTTP-proxy config for S3, not the PBS proxy daemon. **The 32-minute post-upgrade window is indistinguishable from the before window** (+5 fd/1919 s = 225/day vs +4 fd/1885 s = 183/day; the two differ by ONE descriptor and Poisson uncertainty on such counts is ±2, so both are consistent with one unchanged rate — the higher after-figure is noise, not a regression). **Thirty minutes cannot settle it in either direction and this row exists so nobody pretends it did.** New `t0` = **fd 17 at 2026-08-18 09:51:22Z, proxy PID 551655**; before-rate to beat = **183–200/day**. **Interpretation fixed in advance** (`evidence-ep0-pbs-upgrade-2026-08-18/stop1-ruling.txt`, written before any numbers existed): unchanged = EXPECTED, not a failed upgrade; changed = a SURPRISE needing explanation, not a confirmation | **WATCHING — NEW 2026-08-18** | elapsed time only | **Two dated checks, both CC:** **+24 h — 2026-08-19 ~10:00Z** and **+7 d — 2026-08-25 ~10:00Z**. Command (the incident's own positive observable): `ssh root@ 'PID=$(systemctl show proxmox-backup-proxy -p MainPID --value); ls /proc/$PID/fd \| wc -l; ss -lnt "( sport = :8007 )"; ss -tn state all "( sport = :8007 )" \| awk "NR>1{print \$1}" \| sort \| uniq -c'`. **Record the ESTAB/CLOSE-WAIT split, not just the total** — the split is what says which leak it is. If PID ≠ 551655 the window is void: something restarted the proxy and the count began again | CC on both dates | @@ -644,6 +644,9 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-342** | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-343** | **The managed controller floor was raised 0.214.0 → 0.216.0 — and it was NOT the no-op it was expected to be: it moved a live box nine seconds later.** Raised by the operator 2026-08-18 **12:36:58Z**, in a **separate save after** the artifact vouch. **Read back from the store, not the form** (`hub_settings.min_controller_version`, `GetGlobalMinControllerVersion` — `hub/internal/store/store.go:1751`): `min_controller_version = 0.216.0`, updated 12:36:58. **WHY IT HAD BEEN BEHIND — this was NOT drift, and describing it as "two releases behind" without this context reads as a defect it was not.** `publish-train-rules.md` rule 1 is *manifest before floor*, and rule 2 requires the floor field to be filled **LAST, in a separate save**, because the DB row overrides the env floor and **acts immediately on the next report cycle**. That rule was earned: on the 2026-07-11 publish train the floor was saved together with the manifest, acted at once, and pushed controller 0.113.0 onto Peti's box **~9 minutes ahead of agent 0.81.0** — the exact forbidden skew, benign only because that box had no NAS shares. R-120's row records the same deliberate choice (*"Floor untouched per publish-train rule 2"*). So the floor sitting at 0.214.0 was **policy being followed**, not neglect. **What it was functionally while it sat there:** not a live problem — every reporting box was at or above it — but a **safety net set two versions low**. The floor is what drags a box forward if it ever falls behind (restored from an old backup, reinstalled, long offline), and at 0.214.0 it would have pulled such a box only to two versions back, missing R-328's severity fix and R-335's follow-up. **THE MEASURED BLAST RADIUS — five reads, and the third and fifth are the findings.** **(1) Floor:** `0.216.0` @ 12:36:58Z, from the store. **(2) Per-customer overrides:** **zero** — all five `customer_configs` rows carry an empty `min_controller_version`, so nothing hides behind a lower override and the global applies to everyone. **(3) Every box's controller version:** `demo-felhom` **0.216.0**, `demo-hp` **0.216.0** — both AT the floor **now**; `drill-r50` **0.213.0** (status `blocked`, last report 2026-08-12, powered off/reverted) and `peti-felhom` **0.115.0** (host row DELETED) are **below** it but are **not reporting boxes**. **(4) Directives/holds:** no `managed floor HELD` line exists; the hub logged `[INFO] Global controller-version floor set to "0.216.0"`. **(5) THE FINDING — a controller DID auto-update after the raise.** `demo-felhom` had been on **0.214.0** since 2026-08-12 16:44 and the hub recorded `controller_updated — Controller frissítve: 0.214.0 → 0.216.0` at **12:37:07Z**, then `controller_started (0.216.0)` at 12:37:12Z — **nine seconds after the raise**, exactly the immediate action rule 2 documents. `demo-hp` was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move. **No error, warning or critical event followed** — the update completed and the controller came back up. **So the change was real, not inert: "every reporting box is at or above the floor" is true BECAUSE of the raise, not independently of it.** **Why it is safe by construction**, cited rather than asserted: `ResolveManagedFloor` (`hub/internal/store/store.go:2068`) sets `Held` and clears the floor entirely when `Floor > GoldenVersion` (the R-216 shape) — floor 0.216.0 **equals** golden 0.216.0, so that guard does not trip — and holds per-box when the box's agent is below the manifest's `MinAgent`, or unknown, or unparseable; both boxes report agent 0.129.0 against `MinAgent` 0.129.0, so the floor was served rather than held. That second guard remains armed for any box reporting with an old agent. **No ISO rebuild is required:** the golden is fetched at first boot from the hub's manifest, which is 0.216.0 — at the floor, not below it — and rule 5's `assert_golden_ge_floor` is a **build-time** gate for *future* builds (`scripts/iso/build-felhom-iso.sh:77`, called at **:267**) which **fails open with a warning** when its inputs are absent (`:78-82`: `if [[ -z "$golden" \| \| -z "$floor" ]]` → `log_warn "… UNENFORCED …"` → `return 0`), both confirmed in the script. **`peti-felhom` was NOT contacted and needs no contact** — from the PETI row: its host row was deleted 2026-07-15 and *"a report from a deleted host 401s and is not persisted"*, so it cannot receive a floor directive at all and the raise cannot reach it | **OPEN — NEW 2026-08-18.** Deliberately NOT closed: the task's closing condition was *all five reads clean, no directive served*, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later | — | **Confirm the 0.216.0 update on `demo-felhom` is healthy in normal operation** (it reported and restarted clean, but it has not yet run a full backup cycle on 0.216.0 at the time of writing), then close. Separately: `drill-r50` at 0.213.0 will be dragged to 0.216.0 by this floor if it is ever booted and reports — that is the floor doing its job, noted so it is not read as a surprise | CC | +| **R-339** | **The hub was SILENT when it lost sight of the off-site stores — and a 9 h 37 m outage proved it.** Both box checkers (`OffsiteBoxChecker` over the Hetzner API, `PBSDRBoxChecker` over ep0's `usage` op) held their last snapshot and returned quietly on a failed fetch. That is **correct for a fill signal** — a missing reading must never be mistaken for 0%, which is why degraded data drives no band transition — but the consequence was that a completely dead off-site endpoint and a healthy one were **indistinguishable on the operator channel**. During the 2026-08-18 ep0 incident the hub said nothing for the entire outage; the only mails came from the boxes' own backup failures, and **only because the WEEKLY offsite run happened to fall inside the window**. Two days earlier, nothing would have fired at all | **SHIPPED — hub v0.106.0, 2026-08-18.** Reachability is now a second, independent signal: consecutive failed fetch windows counted per checker, `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default **3 windows (≈30–45 min)**, with paired `*_recovered` all-clears wired into `recoveredPairedDownTypes` — necessary because both recoveries are severity `info` and `severityNotifies` drops `info`. Threshold tunable via `alerting.box_unreachable_windows`. **The fill logic is untouched**: no threshold, throttle, band or escalate-once behaviour changed. Evidence: `internal/monitor/box_reachability_test.go` (Scenarios A–F) + `internal/notify/dispatcher_box_reachability_test.go` (the cross-package wiring, asserting an actual operator mail), plus three companion red-proofs each seen failing with a message naming the right cause | — | **PROVEN-LIVE still owed.** No real or constructed outage has exercised the emit path end to end, and one cannot be manufactured without making ep0 or the Hetzner API unreachable — ep0 is Tier 2 protected, so that is forbidden. The honest route is a constructed outage against a scratch hub instance with the tenantsync client pointed at a blackholed address. **Do not close this row on the unit tests** | CC | +| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** | CC | +