From c8da8e043905ca6395490a5337ca63f0717b3b8c Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 21:09:04 +0200 Subject: [PATCH] =?UTF-8?q?burn-down=20night:=20rulings=20=C2=A71=20record?= =?UTF-8?q?ed=20(09=20=C2=A73=20decisions=20128-130);=20R-888,=20R-337,=20?= =?UTF-8?q?R-375=20closed=20by=20ruling;=20R-126/R-856=20rulings=20in=20th?= =?UTF-8?q?eir=20rows;=20night=20log=20(199=20->=20196)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../architecture/09-update-architecture.md | 13 ++++++++++++ .../night-burndown-2026-10-05/NIGHT-LOG.md | 21 +++++++++++++++++++ documentation/backlog/CLOSED-ITEMS.md | 10 +++++++++ documentation/backlog/OPEN-ITEMS.md | 9 ++------ scripts/wire_contract_gate.py | 6 +++--- 5 files changed, 49 insertions(+), 10 deletions(-) create mode 100644 documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 8e5dc7ee..f9565507 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -905,6 +905,19 @@ its length, and both fixes cost something the household would notice — operato 127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer. *Operator ruling 2026-10-05.* (R-861) +### 2026-10-05 (~21:00) — three rulings, the reviewer's picks given to the operator (recorded before the work; the burn-down night) + +128. **R-126 — a `.fab` export onto a network drive.** **Refuse an export WITHOUT a password to any network drive; with + a password it is allowed.** Rejected: the row's other option, filtering network drives out of the export + destination list. *Ruling 2026-10-05 ~21:00, the burn-down night brief §1.* +129. **R-856 — a crash restart reached the household twice.** **After a crash boot, the controller's app mails wait the + same grace period as after a normal start.** The hub's "restarted after an unexpected stop" line stays the one + message about the crash. *Ruling 2026-10-05 ~21:00, the burn-down night brief §1.* `08` §5. +130. **R-888, R-337, R-375 — closed.** R-888: the hub needs neither `stacks.deployed` nor `storage.decommissioned` today + (no hub page reads them; the fields stay on the wire and stay allow-listed in `scripts/wire_contract_gate.py`). + R-337: resolved by itself and never seen again. R-375: a note with no defect behind it (nothing in the product + reads a PBS target's size). *Ruling 2026-10-05 ~21:00, the burn-down night brief §1.* + ### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief) 100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) — diff --git a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md new file mode 100644 index 00000000..cbda53ab --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md @@ -0,0 +1,21 @@ +# Night log — the burn-down night, 2026-10-05 → 06 + +Brief: `drills/NIGHT-burndown-2026-10-05.md` (workspace root). One line per row touched: +row · result (fixed / closed-stale / failed / moved to A or D / needs-live) · minutes · commit. + +**Baselines read from live source at 21:03 CEST:** felhom.eu `30650cad6e` · controller `6f1ba1fe43` (v0.297.0) · agent +`208fac8027` (v0.147.0) · catalog `4828dc754d` · hub v0.137.0 (CHANGELOG head) · register **199** (19 P2, 95 P3, 85 P4). + +**How the night is worked:** six helper sessions, one per repo lane, each in its own git worktree and branch cut from +`origin/main` (controller ×2, hub, installer, agent, catalog). Helpers commit locally, never push. The lead reviews, +cherry-picks onto `main`, writes CHANGELOG, closes rows, pushes and watches CI. + +## Rows + +| Row | Result | Min | Commit | +|---|---|---|---| +| R-888 | closed — ruling (decision 130) | 5 | (this commit) | +| R-337 | closed — ruling (decision 130) | 2 | (this commit) | +| R-375 | closed — ruling (decision 130) | 2 | (this commit) | +| R-126 | ruling recorded (decision 128); build → helper ctrl-a | 2 | (this commit) | +| R-856 | ruling recorded (decision 129); build → helper ctrl-a | 2 | (this commit) | diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 20de06c1..c020d13b 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,16 @@ --- +## 2026-10-05 (night) — the burn-down night: rulings §1 + +The full text of every row below: `git show 30650cad:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-375** | **A PBS datastore signal was noted and explicitly not filed.** (P4) | CLOSED 2026-10-05 — RULED (decision 130, the burn-down night §1): a note with no defect behind it | Nothing in the product reads a PBS target's Total/Used/Avail: the space preflight skips a PBS target (felhom-agent `internal/backup/runner.go:265`, burn-down round 2 check). `09` §3 decision 130. | +| **R-337** | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** (P4) | CLOSED 2026-10-05 — RULED (decision 130, the burn-down night §1): resolved by itself, never seen again | The lag cleared on its own on 2026-08-18 (the row). The refresh path is the job goroutine after the runner returns (burn-down round 2 check: felhom-agent `internal/localapi/server.go:885`, `pickLatestBackup` `:1304-1318`). Not seen again since 2026-08-18. `09` §3 decision 130. | +| **R-888** | **Two fields the controller reports are never read by the hub: `stacks.deployed` and `storage.decommissioned`.** (P4) | CLOSED 2026-10-05 — RULED (decision 130, the burn-down night §1): the hub needs neither field today | No hub surface decodes `stacks.deployed` or `storage.decommissioned` (grep of `hub/internal`, 2026-10-05 21:40: no reader). The fields stay on the wire and stay allow-listed in `scripts/wire_contract_gate.py` with this reason; a surface that needs one later decodes it then. `09` §3 decision 130. | + ## 2026-10-05 (late night) — burn-down round 2: the operator's answer (43 accepted), small rows fixed with releases The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN-ITEMS.md` (the commit before each closing commit; the burn-down evidence table is `audits/burndown-2026-10-05/`). diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 30008a6c..40bb9019 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -5,7 +5,6 @@ page keeps only what is **open**, and it is the file to read first. Root `REPORT and **overwritten** — nothing durable may live only there; a session that must not clobber it writes a non-overwritten `REPORT-.md` sibling instead (`CLAUDE.md:82-87`), of which 14 now exist. - ## How a row is filed (2026-10-03 — the triage; operator may reverse the scale and the list) **One table shape:** `| ID | Category | Sev | What | State | Blocked on | Next action | Owner |`. @@ -113,7 +112,6 @@ match what the reader sees is how an instrument stops being believed (R-421). It **nothing was proven on the day this line was drawn** and a stopping line that moves a status is a stopping line that lies. - ## Install & onboarding — 12 rows (P3 8, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | @@ -204,7 +202,6 @@ stopping line that lies. | **R-246** | Backup & restore | P4 | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | | **R-279** | Backup & restore | P4 | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC | | **R-336** | Backup & restore | P4 | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 **SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it.** `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **The leak is OURS, and the poll rate is not what feeds it.** Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by **`felhom-agent`** on the boxes — 194 on each, `ss -tnp` naming a single PID per box, and **zero** held by `pvestatd` or `proxmox-backup-client`. Confirmed independently from ep0's access log over the same 46.18 h window: `libwww-perl` (pvestatd) **81,192 requests -> 0 descriptors**, `proxmox-backup-client` **80,061 requests -> 0 descriptors**, `Go-http-client/1.1` (the agent) **387 `/snapshots` calls -> 388 sockets — one per call, within one**. So **162,404 requests, 99.5% of the traffic, produce 0% of the leak.** **Mechanism, named from source:** `felhom-agent/internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` — a composite literal, so `IdleConnTimeout` is the zero value = **no limit** (`http.DefaultTransport` sets 90 s; a literal does not inherit it) — and `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh client every cycle**, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; `CloseIdleConnections`/`IdleConnTimeout`/`MaxIdleConns` appear **nowhere** in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h `DefaultVerifyCadence` (7.7 cycles) = 192.4 predicted vs **194 observed per box**. **CONSEQUENCE — RE-RANK.** The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") **would have produced a null result and read as a failed fix.** Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and **zero** descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a **scaling/cost item, not the leak fix**, and the leak fix is **R-344**. **Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED** — Phase C is held at STOP 1 with its prediction pre-registered in `evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt`. Do not record a proportionality verdict here until that window has run. **RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX.** That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: *cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing.* Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — **ours, in `felhom-agent`, R-344** — would have been further from being found, not closer. **Measured 2026-08-20:** `pvestatd` (`libwww-perl`) and `proxmox-backup-client` made **162,404 requests** in a 46 h window and leaked **zero** descriptors; the agent made 811 and leaked **388**. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its **baseline of 17**, with the poll rate completely unchanged — 85,000/day before and after. **WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits:** ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so **at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33**. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — `pvestatd` stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with `felhom-agent/internal/pbsdr/manager.go`'s health model. **NEW ACCEPTANCE CRITERION, since the old one is void:** the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. | CC | -| **R-375** | Backup & restore | P4 | **A PBS datastore signal was noted and explicitly not filed.** `audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:171`: *"Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here."* Recorded so the note has a number and stops depending on someone re-reading that report. **Age when filed: 4 days.** | **OPEN — LOW** | — | Confirm the cause on ep0 the next time it is touched; it is a read-only check. | CC | | **R-526** | Backup & restore | P4 | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** **Re-ranked 2026-10-03: P3->P4: operator teardown op on a protected box; adopt path covers the need.** | — | — | operator | | **R-541** | Backup & restore | P4 | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** **Re-ranked 2026-10-03: P3->P4: needs a second pool box or an outgrown customer first; operator-only and far off (0.3% full).** | — | — | CC | | **R-570** | Backup & restore | P4 | **[P3-LOW] The off-site stale-note display still has a Hungarian-text fallback, for boxes that have not run off-site since v0.251.0.** OPENED 2026-09-17 by R-553's fix: `offboxWarningDisplay` (controller/internal/web/handlers.go) decides on `LastWarningKind`, but a box upgraded to 0.251.0 carries the PERSISTED old sentence with no kind until its next off-site run rewrites it, so the substring test survives under `kind == ""`. **Close when every fleet box has completed one off-site run on ≥ 0.251.0** (the hub's reports carry the controller version; the off-site anchor is `offbox.last_success`), then delete the fallback, its constant and its legacy test rows. **Hard dependency: localisation slice 2 (R-557) must NOT translate the producer `"Sikeres — nincs mentésre jelölt alkalmazás"` (controller/internal/backup/offbox.go) until this row closes** — translating it while the fallback is load-bearing strands exactly those boxes. | **WATCHING - rank P3-LOW; owner: operator (the fleet condition), CC (the deletion)** **Re-ranked 2026-10-03: P3->P4: cleanup waiting on a fleet condition; no defect today.** | — | — | operator | @@ -233,7 +230,7 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | | **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | -| **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | +| **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | **READY — RULED 2026-10-05 ~21:00 (`09` §3 decision 128): refuse an export without a password to any network drive; with a password it is allowed. Being built (burn-down night).** | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | @@ -293,8 +290,7 @@ stopping line that lies. | **R-177** | Monitoring & notifications | P4 | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | | **R-285** | Monitoring & notifications | P4 | **A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere.** During the 2026-08-09 rehearsal the hub sent, all `status: sent` to the operator channel: `host_stale` 08:58 UTC, `node_stale` 09:00, **`host_down` 09:28 (error)**, **`node_down` 09:30 (error)**, `host_leaf_changed` 09:31, `host_recovered` 09:31, `node_recovered` 09:34, `offsite_delivery_stuck` 09:34 — eight operator mails for work that was deliberate, attended and announced. **This is the OPPOSITE gap from the one R-281 filed:** the alarms are not missing, they are indiscriminate. `host_stale` at 30 min and `host_down` at 60 min (`monitor/host_staleness.go:22-23`, `downAfter = 2 * threshold`) cannot distinguish a wiped-on-purpose box from a dead one, and `host_leaf_changed` firing on a reinstall is correct-but-expected. **Note the interaction with the mute used on 2026-08-09 evening:** blocking a customer silences everything, so today the only two settings are *page me for planned work* and *tell me nothing at all*. **What is owed is a middle:** a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | **READY (M) — NEW 2026-08-09** | — | The evidence is the operator's mailbox plus `events`/`notification_log` for 2026-08-09 | CC | -| **R-337** | Monitoring & notifications | P4 | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | -| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator | +| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — RULED 2026-10-05 ~21:00 (`09` §3 decision 129): after a crash boot, the controller's app mails wait the same grace period as after a normal start. Being built (burn-down night).** | — | Build decision 129 | CC | | **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC | | **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. **Checked from source 2026-10-05 (burn-down round 2):** homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager re | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | | **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. **Checked from source 2026-10-05 (burn-down round 2):** Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | @@ -315,7 +311,6 @@ stopping line that lies. | **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator | | **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator | | **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC | -| **R-888** | Hub & operator | P4 | **Two fields the controller reports are never read by the hub: `stacks.deployed` and `storage.decommissioned`.** Surfaced 2026-10-05 when the wire-contract gate stopped counting a tag named only in a comment (R-555): the hub decodes no deployed-app list and no per-drive decommissioned marker from the box report. Allow-listed in `scripts/wire_contract_gate.py` so the gate stays green. **Needs a decision, not a fix:** does any hub surface (customer page, drive view) need either? If yes, decode and show it; if no, record why. | **OPEN** | — | Operator: say whether the hub needs either field | operator | ## Business & legal — 6 rows (P2 4, P4 2) diff --git a/scripts/wire_contract_gate.py b/scripts/wire_contract_gate.py index 14130350..01aac46d 100644 --- a/scripts/wire_contract_gate.py +++ b/scripts/wire_contract_gate.py @@ -299,11 +299,11 @@ ALLOWLIST = { "R-555 / R-331: no producer since slice 8C; the Backup card reads offsite.repo_size_bytes " "instead (hub/internal/web/backup_card.go)."), (_CH, "stacks.deployed"): ( - "R-555: surfaced 2026-10-05 — the hub decodes no deployed-app list from the report; whether a hub " - "surface needs it is not decided. Filed as R-888 for the operator to decide."), + "R-555: surfaced 2026-10-05 — the hub decodes no deployed-app list from the report; no hub " + "surface needs it (R-888 closed by ruling, 09 §3 decision 130, 2026-10-05)."), (_CH, "storage.decommissioned"): ( "R-555: surfaced 2026-10-05 — the hub decodes no per-drive decommissioned marker from the report; " - "whether a hub surface needs it is not decided. Filed as R-888 for the operator to decide."), + "no hub surface needs it (R-888 closed by ruling, 09 §3 decision 130, 2026-10-05)."), } # A node whose IMMEDIATE CHILDREN are still checked but whose DEEPER descendants are not, because the