diff --git a/CONTEXT.md b/CONTEXT.md index 4c237035..12bd9572 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,14 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-08 (evening) — the decision sheet D1–D10 built on main, unreleased (`09` §3 185–194).** Controller (`3f84f82`): +> operator actions (`internal/report/opactions.go`, closed list), dashboard sessions on disk as sha256 + password +> fingerprint (`internal/web/session_store.go`), six health keys, the R-893 hold (`restore_mixed`) on every failure after +> a definition/volume moved. Hub: `operator_actions` table + reply field, wait-channel presence + the ticked fast delete +> (`intent/presence.go`), health mail without the raw note, Cloudflare `/zones` reach check (zone == domain), R-435 +> window credits (capped, spent once). Catalog `31e9651` (wger 256M). Ship hub + controller the same day; live proofs +> listed in `REPORT-day3-2026-10-08.md`. Register 132 → 136 (+4 another session's). + > **2026-10-08 (morning) — the first real kernel night read back.** demo-felhom: staged exactly the told 7.0.14-20 at > 04:39:04, healthy 04:40:38, default now 7.0.14-20, apps ~1.5 min away, no alarm. demo-hp: no whole-guest backup that > night (yesterday's 08:49 press + 24 h cadence > window end) → no step; R-899 filed. Reply-To proven by the operator's diff --git a/REPORT-day3-2026-10-08.md b/REPORT-day3-2026-10-08.md new file mode 100644 index 00000000..18eb25ca --- /dev/null +++ b/REPORT-day3-2026-10-08.md @@ -0,0 +1,88 @@ +# REPORT — 2026-10-08 (evening): the operator's ten answers (D1–D10) built — code and tests, no delivery + +(`REPORT-day3-…` because other sessions write in this clone today.) Rulings recorded first: `09` §3 decisions 185–194 +(commit `4510dd7d`), and in each row. + +| Part | Result | +|---|---| +| **A** — operator actions (D1; R-314, R-279) | **Done on main.** Controller `internal/report/opactions.go` + `scheduler.RunNow`; hub table `operator_actions`, four buttons, reply field `operator_actions`, result field `operator_action_results`. Closed list both sides (`TestOpActions_ClosedList`, `TestOperatorActions_ClosedList`); unknown → `refused`, nothing called; once per id; cross-customer result ignored. Fixed on the way: `ExtendAbandon` could move a deletion earlier or edit the box date in the hub phase; `StopAbandon` reported success while the hub's deletion stayed pending; reply read limit 4 → 64 KiB. Review fix: a `run_job` „done" says it ran, not what it found. | +| **B** — delete a switched-off box at once (D2; R-30) | **Done on hub main.** Presence from the wait channel (connected / not connected since T / unknown), shown on the host page; the delete goes ahead at once only with the tick AND ≥ 360 s not connected, re-checked at the POST. 7 mutations red-proved. | +| **C** — health mail + dashboard (D3; R-79) | **Done on main, both halves.** Hub: the household's health mail drops the raw note and points to the dashboard (operator mail unchanged). Controller: the last six health producers carry a key; wire text unchanged. | +| **D** — signed in across a restart (D4; R-35) | **Done on controller main.** Only sha256(cookie) on disk (0600). Review fixes: rows written under another password are dropped at load; a failed revoking save removes the file. 8 tests. | +| **E** — Cloudflare key check (D6; R-138) | **Done on hub main.** Fake Cloudflare in tests only; the token is never logged. Review fixes: the token's one zone must BE the customer's domain (a parent zone was accepted, and checking it against today's customers was order-dependent). Limit filed: R-913 (read scope ≠ write scope). | +| **F** — one lost off-site copy is an error (D7; R-435) | **Done on hub main.** Review fixes: each window explains at most its hub-set cap; each window explains one fall only; a window stuck open explains nothing; unreadable windows fall back to the half-rule. `08` §6.5 written (count from the box, net count). | +| **G** — failed restore keeps the app stopped (D8; R-893) | **Done on controller main.** Review fixes: the hold now covers every failure after the definition or a volume moved (not only a failed replay); the hold is persisted before the stop. R-893 stays open, NARROWED to the second half („put back exactly as it was") — no new row needed. | +| **H** — D5, D10 | **R-717 closed** (D5). **wger `mem_request` 256M pushed** (catalog `31e9651`, CI 1559 success), after a hub read showed no box runs wger (app-telemetry page 12:22 UTC: `wger` 0; controls `paperless`, `bentopdf` present). | +| R-554, R-462 | **Not done.** R-554 is still stopped on a customer promise (decision 2 below). R-462 not started (no time left after the review fixes). | + +**Rows: 132 before → 136 after. Opened 1 (R-913). Closed 1 (R-717).** The other +4 are another session's (R-909, R-910, +R-911, R-912: dashboard layout and logos). State changes: R-314, R-279, R-30, R-79, R-35, R-138, R-435 → VERIFY (built, ships tomorrow, +closes after the live proof); R-893 → NARROWED. + +**Commits.** Controller: `ee8a526`…`d3e17e9` (cherry-picked from the helper branches onto main), `3f84f82` (review +fixes). Hub/felhom.eu: `4510dd7d` (rulings), then the integration push (this report's commit). Catalog: `31e9651`. + +## Security review — what it found and what happened + +Automatic commit reviews plus two read-only review helpers. Fixed and red-proved the same evening (10): Cloudflare +parent zone; order-dependence of the zone check; R-435 fail-open on unreadable windows; a box's `count_after` +widening a window; one window explaining several falls; a stale open window; D4 revocation lost on a failed save +(two shapes); D8 earlier failure branches started the app on a mix; the hold written after the stop; `run_job` +„done" read as a pass. **Written as limits, not fixed:** D2 — presence is the guest controller's, so a live host +with a stopped guest reads „not connected" (the cost accepted with decision 186; noted on R-30); D8 — placed files +alone still restart after a successful rollback (option C as ruled; option A is next); D6 — tokens saved before the +check were never checked (R-138 closes only after each is checked once); D6 — read scope ≠ write scope (R-913). + +## Waits for tomorrow's releases + +- **hub:** the morning's items (R-243, R-901 one-year deletion, R-304 mail, R-415 guard) + D1 hub half, D2, D3 mail, D6, D7. +- **controller:** the morning's R-899/R-304 + the afternoon's R-304 mail, R-298 button + another session's dashboard + layout fixes + D1, D3 dashboard, D4, D8. MinAgent unchanged (0.131.0). **Release the hub and the controller the + same day** (D1's wire fields; an older hub sends no actions, an older controller ignores them and the rows expire). +- **agent:** the morning's R-899 + ring label + R-304 424 (nothing new tonight). +- **catalog:** already pushed (wger 256M; reaches a box at its next sync). +- **DooPlex:** the `ep0-copy` clean-up job (D9) — dry run, then daily — after the releases are read back. + +## Live proofs tomorrow's session must run (scratch 9202 only, after the releases) + +1. **D1:** from the hub's host page, `run_job fill-watch` → the hub's result row; control from another channel: the + controller log pull's „checked N filesystem(s)" line. `offsite_backup_now` → result; control: the off-site snapshot + list on 9202's backup page. `abandon_stop` / `abandon_extend` only if 9202 has a countdown. +2. **D2:** stop guest 9202; time until the host page says „not connected"; control: the ingress log's last + `/api/v1/wait` line from 9202. **Do NOT tick-delete demo-hp** — the fast delete removes the Proxmox host; prove it + only through `GET /hosts//delete-impact` (`off_tick_required`). Evidence off before 9202 starts again. +3. **D3:** a `health_critical` on 9202 → the household mail in the catch-all has no `{`, has the dashboard line; the + operator mail still has the note. The dashboard banner in English for an English household. +4. **D4:** sign in on 9202, restart the controller, the next page needs no login; then change the password → the old + cookie is refused after the next restart. +5. **D6:** re-save each customer's config once with its current token (operator present): each must pass; a refusal + keeps the old token. Then close R-138. +6. **D7:** a hand-granted window on 9202 removes N → no alarm; control: the hub's own window row. +7. **D8:** design slice 0 on 9202 with a throwaway app (version N snapshot, update to N+1, restore with a truncated + dump) → the app held, the page sentence, the operator line. Evidence off before teardown. +8. **D9:** install the DooPlex job, dry run first, then daily. + +## Machines + +demo-hp, demo-felhom, Tester 1: not touched. 9202 and the bench: not touched. Hub: read only (three page reads with +the operator password: `/apps`, `/hosts`, `/configs`). DooPlex: no change (helper work ran in git worktrees under +`/mnt/5_hdd/felhom.eu/worktrees/`; no Docker command). No release, no deploy, no reboot, no prune. Provisioned nothing. + +## Small fixes without a row + +`scripts/test_gate_decoys.py` crashed in a linked git worktree (lock path); controller `TestR650_NoBareDockerExec` +raced a vanishing temp file; the three D1 abandon fixes above. + +Instruction files: no edit. + +## Decisions for the operator + +1. **R-913 — how sure must the Cloudflare key check be?** The hub can see what a pasted key may READ, not what it may + CHANGE. A key made by hand with „read one zone, change all zones" would pass. **Pick: (a) you always make the key + with the one recipe** („Edit zone DNS", Zone Resources: Include → Specific zone); it costs nothing. (b) The hub + makes the key itself: safest, but the hub then holds a key that can make keys. *If you do nothing:* the check stays + as built, and the limit stays written down. +2. **R-554 — what should the household's recovery note say once the old setup wizard is deleted?** Today the note on + the box tells the household to run a setup script on port 8081 to restore. **Pick: the note says „Call Felhom; we + restore your box from its backup"**, and CC deletes the wizard. *If you do nothing:* the wizard stays, unused, and a + box whose first start cannot reach the hub shows it to the household. diff --git a/REUSE.md b/REUSE.md index dfbd3203..29025dd9 100644 --- a/REUSE.md +++ b/REUSE.md @@ -71,6 +71,8 @@ | Symbol | File | Short signature | Use for | Gotchas | |---|---|---|---|---| +| `store.WindowCreditsBetween` | hub/internal/store/offsite_keys.go | `(customerID, from, to) ([]WindowCredit, unknown)` | R-435: what the hub's clean-up windows may explain | Each window ≤ its `max_remove`; the CALLER spends each window once (`OffsiteChecker.usedWindows`); unknown → the half-rule, never „explained" | +| `cloudflare.TokenZones` / `ZoneEquals` | hub/internal/cloudflare/reach.go | `(ctx, base, token) (names, total, err)` | R-138: what zones a pasted token can see | Every failure wraps `ErrReachUnknown` (refuse the save); the token never appears in an error; `/zones` is READ scope, not write | | `offsitekeys.Registrar` (`Install` / `Confirm` / `Audit` / `OpenWindow` / `CloseWindow` / `MoveAside`) | hub/internal/offsitekeys/offsitekeys.go | `(ctx, Target, password, …)` | EVERY write to a sub-account's `.ssh/authorized_keys` and every repo move-aside | **The only writer of that file, and the only deleter on a sub-account (`DeleteSetAside`: `.orphaned-*` only, decision 74).** `read()` is read-only (R-827). Uses the provider's port-23 restricted shell (`dd of=` takes stdin, `mv` overwrites, `test` does NOT exist — measured); an unpinned line is a deletion route and is dropped on every install; the window line goes FIRST (first match wins). Never `rm`. | | `offsitekeys.Service` (`RegisterKey`, `ConfirmKey`, `AuditAll`, `OpenWindowFor`, `CloseWindowFor`, `SweepExpiredWindows`) | hub/internal/offsitekeys/service.go | — | Binding the registrar to the store, descriptor and operator events | The box-facing API (`/api/v1/offsite/register-key…`) answers with NO credential — pinned by `TestOffsiteKeyEndpoints_AuthAndNoPasswordInAnyResponse`. | | `(*Store).SaveOneTimeSecret` / `OffsitePassword` / `SealLegacyOffsiteSecrets` | hub/internal/store/offsite_seal.go | — | Storing / reading the sub-account password | **Sealed AES-256-GCM; no key → refused (fail-closed).** Under `go test` every store gets a fixed key (`testing.Testing()`); production needs `OFFSITE_SECRET_KEY`. Never serve the value to a box. | @@ -80,6 +82,8 @@ | Symbol | File | Short signature | Use for | Gotchas | |---|---|---|---|---| +| `store.CreateOperatorAction` / `PendingOperatorActions` / `RecordOperatorActionResult` / `ValidateOperatorAction` | hub/internal/store/opactions.go | rows of `operator_actions` | Operator actions carried in the report reply (decision 185); POST `/hosts/{id}/operator-action` → `handleOperatorAction` | `ValidateOperatorAction` IS the closed list (`TestOperatorActions_ClosedList`); a result is matched on (customer, id) — never on id alone | +| `intent.Hub.MarkWaitStart` / `MarkWaitEnd` / `Presence` | hub/internal/intent/presence.go | `(customerID, now)` | Box presence from the wait channel (R-30); the delete-host fast path | In memory: a hub younger than 333 s answers `unknown`, which never permits the fast delete; `hostOffDeleteAfter` ≥ the window (pinned) | | `(*Server).hostDetailData` | hub/internal/web/hosts.go (~L282) | `(host *store.Host, r) map[string]interface{}` | The ONE view-model builder for the shared `host_detail_body` sub-template (standalone `/hosts/{id}` + customer Host tab) | Booleans/counts only for DR/escrow; carries `Deletable` (= status != "ok") which gates the danger-zone card. Never add a secret field. | | `parseHostAddresses` + `(*Server).hostNetwork` / `hostNetworkView` (v0.85.0) | hub/internal/web/hosts.go | `(reportJSON) []hostAddressView` · `(host, reportJSON) hostNetworkView` | The host page's Network card: every routable address the box holds + its WireGuard allocation | Needs agent **>= 0.119.0** (`minAgentForAddresses`); below it the wire has no `addresses` key and the card renders **UNKNOWN, never "no addresses"** — an absent signal is not a negative result. WireGuard is TWO facts: the hub's allocation (`GetWGPeerForHost`, authoritative) AND whether the box confirms holding it — the allocation alone cannot distinguish a live tunnel from a peer that was never applied. The WG row is split out by comparing against the ALLOCATION, never by matching the interface name `wg-felhom`, which is a unit name that can change. | | `(*Store).GetHostRecoveryMeta` + `(*Server).handleHostRevealRecoveryCredential` | hub/internal/store/host_recovery.go · hub/internal/web/hosts.go | `(hostID) (*HostRecoveryMeta, error)` · `POST /hosts/{id}/reveal-recovery-credential` | The break-glass console credential, split into a RENDER half and a RETRIEVE half (v0.84.0) | **Use `GetHostRecoveryMeta` on any page-render path** — its struct and its `SELECT` both omit the `secret` column, so it cannot leak one; `GetHostRecoveryCredential` (which does select it) belongs only to the two retrieval handlers. The reveal is POST so the ServeHTTP-level CSRF check applies and no secret is reachable by URL; it writes ONE `recovery_credential_revealed` event via `SaveEvent` and calls NO dispatcher (the `handleRequestLogTail` shape). `api/handler.go handleAdminGetRecoveryCredential` (global key) is the independent fallback for when the UI is down — never route the UI through it. Secret at rest is plaintext → R-133. | diff --git a/STATUS.md b/STATUS.md index 3cdb9e2a..c16243e1 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,9 +2,30 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-08 (afternoon): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller -0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 129. Reports: -`REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon).** +**Updated 2026-10-08 (evening): hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller +0.303.0 (nothing delivered today — tonight is the second kernel night). The open-items list is at 136. Reports: +`REPORT-day-2026-10-08.md` (morning), `REPORT-day2-2026-10-08.md` (afternoon), `REPORT-day3-2026-10-08.md` (evening).** + +## Evening (2026-10-08): your ten answers (D1–D10) built — they ship tomorrow + +- **Buttons in the hub** (D1): run an off-site backup now, run one of four checks now, stop or extend a household's + deletion countdown. The list is fixed. No button deletes anything or shortens a countdown. +- **A switched-off box can be deleted at once** (D2), after you tick „I checked: the box is off" and the hub has heard + nothing from it for 6 minutes. The host page shows when the box was last connected. +- **The household's „system health" mail is short** (D3): no technical note; it points to the dashboard. The dashboard + shows every health warning in the household's language. +- **The household stays signed in** when the box restarts or updates (D4). The disk keeps only a fingerprint. +- **opengist's item is closed** (D5): the address block is enough. +- **The hub checks a pasted Cloudflare key** (D6): it must reach exactly that customer's own domain. +- **One lost off-site copy outside a clean-up window = an error mail to you** (D7). +- **A failed restore of one app keeps the app stopped** (D8) when its version or its data volumes already moved, and + the page says it needs our help. „Put back exactly as it was" is the next step (on the list). +- **wger needs 256 MB to install** (D10). Live in the catalog; no box runs wger. +- **D9 (the DooPlex clean-up job) is tomorrow**, after the releases are read back. +- Security checks read every change. Ten problems were fixed and tested the same evening. Two stay as written limits + (one is the cost you accepted with D2; files alone wait for D8's next step). One needs you (decision 1 below). + +**Needs you:** two decisions at the end of `REPORT-day3-2026-10-08.md`. ## Afternoon (2026-10-08): your four answers built, and one sheet of decisions @@ -16,7 +37,7 @@ - **Cloudflare:** on the list as a later item. - **Nine stuck items have a one-page design.** Two were no longer true and are closed; one lost item is back on the list. -### The decision sheet — answer „all as picked", or name the numbers you change +### The decision sheet — ANSWERED 2026-10-08 14:16, all as picked (`09` §3 decisions 185–194) | # | Question | Pick | Cost | If you do nothing | |---|---|---|---|---| diff --git a/documentation/architecture/01-topology-and-trust.md b/documentation/architecture/01-topology-and-trust.md index 26ad9b20..50975275 100644 --- a/documentation/architecture/01-topology-and-trust.md +++ b/documentation/architecture/01-topology-and-trust.md @@ -213,6 +213,14 @@ app. Controller down → the gated app answers an error, never the app. Measured on Cloudflare) is register row R-494, P3, not blocking. Measured reason this was ruled now: the 2026-09-14 first-hour drill used a `*.felhom.eu` customer domain with no tunnel, and the dashboard link in the setup-code mail did not resolve. +- **A pasted Cloudflare API token must reach exactly the customer's own zone** (R-138 option C, `09` §3 decision 190; + hub main 2026-10-08, unreleased). On save the hub asks Cloudflare `GET /zones` with the token and keeps it only when + the token sees exactly ONE zone and that zone IS the customer's domain (a zone above it is refused: a sibling + customer added later would be reachable). Cloudflare unreachable → the save fails and the old token stays. The token + is never logged. **Limit:** `/zones` lists what the token can READ; a hand-built token with Zone:Read on one zone and + DNS:Edit on all zones would pass — the token wizard's single „Specific zone" scope does not build that, and a + customer token cannot read its own policies. Tokens saved before this check are checked the first time each config + is saved with a new token or domain. - **Tunnel placement: INSIDE the guest** (corrected 2026-10-01, R-754 — the operator's brief of that evening: the build is right, correct the document). `cloudflared` is a container the CONTROLLER renders and keeps up (`internal/infra`, `EnsureBaseStack`, a protected stack), with the tunnel token from `controller.yaml`. *This page used to say it ran on diff --git a/documentation/architecture/04-control-plane-authorization.md b/documentation/architecture/04-control-plane-authorization.md index cd39ec76..97b696fa 100644 --- a/documentation/architecture/04-control-plane-authorization.md +++ b/documentation/architecture/04-control-plane-authorization.md @@ -162,6 +162,16 @@ op the agent verifies** (same pipeline, §2.3) — never unauthenticated config. queue — then the agent polls, verifies, executes, and audits. One command + passphrase, from the desk. **Never** a site visit. +### 6.1 Operator actions in the report reply [DESIGN, R-314/R-279 — `09` §3 decision 185, controller + hub main 2026-10-08, unreleased] + +Routine, non-destructive operator requests need no signature (`03` §4 asks for one only to destroy or overwrite the +only copy). The hub stores a row, bumps the box's intent, and the report reply carries `operator_actions:[{id,action,arg}]` +until the next report answers `operator_action_results:[{id,outcome,message}]`. The list is CLOSED on both sides: +`offsite_backup_now`, `abandon_stop`, `abandon_extend` (1–30 days; never earlier than the current date; refused once +the deletion is the hub's), `run_job` (`fill-watch`, `offsite-integrity`, `offsite-proof`, `disk-health-check`). +Anything else is refused at the hub's POST and again on the box. No action deletes data, starts a countdown or +shortens one (`TestOpActions_ClosedList`, `TestOperatorActions_ClosedList`). Unanswered after 24 h: expired. + ## 7. Hardware readiness (Viktor's "build the foundation now") Software `ssh-ed25519` now; a FIDO2 `sk-ssh-ed25519@openssh.com` key later is a **no-op on the diff --git a/documentation/architecture/05-hub-architecture.md b/documentation/architecture/05-hub-architecture.md index fe8f56f6..635fdefe 100644 --- a/documentation/architecture/05-hub-architecture.md +++ b/documentation/architecture/05-hub-architecture.md @@ -95,6 +95,16 @@ Evolves the existing staleness checker (60s **cadence**, a **configured** thresh than waiting for a guest report to go stale. - **Guest-report recency = secondary** app-level signal. +**Box presence from the wait channel [DESIGN, R-30 — `09` §3 decision 186, hub main 2026-10-08, unreleased].** The +hub records, in memory, each controller wait (`GET /api/v1/wait`) per customer: **connected** (a hold is open, or one +started < 333 s ago = 243 s cadence + 90 s grace), **not connected since T**, or **unknown** (the hub started < 333 s +ago). The host page shows it. Alerting is NOT changed. **„Delete host" goes ahead at once** — instead of waiting for +the report clock — only when the host is online by its report, the operator ticked „I checked: the box is off", AND +presence has been „not connected" for ≥ 360 s, re-checked at the POST; unknown never permits it. The delete writes an +INFO line and one `host_deleted_box_off` event. Wrong case: a box whose controller crashed while its host runs — its +agent is locked out until re-enrolled; no household data is touched. `hub/internal/intent/presence.go`, +`hub/internal/web/r30_presence_delete_test.go`. + **Backup-deadline checker:** today it is *event-based* — it scans for `backup_completed`/`backup_failed` events since local midnight and alerts if none. Two changes: (1) **mechanism** — move it to a field check on `host_reports`' last-backup-per-target (cleaner now that backup state arrives in the host diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index f9a089a7..fa464642 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -651,7 +651,13 @@ was added. **Known limit, not fixed (R-893):** after a failed OFF-SITE replay, t removed stay, and when the snapshot's older definition was written, the rollback branch does not put the newer one back — the older app starts on rolled-back data. An order change cannot fix it: the only undo is a logical dump and the volume it should land in was replaced. Option B (a loader that rebuilds instead of overlays) or a pre-restore volume -copy would; both are larger than this ruling. +copy would; both are larger than this ruling. **`[DESIGN]` 2026-10-08 (R-893 option C, `09` §3 decision 192; controller +main, unreleased): the mixed start is gone where the version or a volume moved.** Any failure after the snapshot's +definition is written or a named volume replaced — a failed placement, volume replay, DB-only start or replay — writes +the live definition back and HOLDS the app stopped (`restore_mixed`, operator-cleared); the household reads that it +needs our help; the operator gets a `backup_run_failures` line. The hold is persisted before the stop. Placed files +ALONE (no version change, no volume replaced) still restart after a successful rollback (`TestR379_ScenarioA`): that +mix, and the undo of everything, is option A — „put back exactly as it was", the next slice on R-893. **[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.** Recorded here rather than only in a closed register row, because a decision that survives only inside @@ -1327,7 +1333,7 @@ crosses the line — **R-158**. | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** **`[BETA-DEFERRED]`** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | | 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** **`[BETA-DEFERRED]`** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) **`[FACT]` 2026-10-03 (decision 70): ep0's datastore now has a nightly copy on DooPlex** (PBS pull-sync `ep0-felhom-offsite`, `remove-vanished false`, weekly verify, failures mailed to the operator); first pull 201 s, 12 GB, 4 of 4 snapshots, matching ep0 namespace for namespace. Restore route: `runbooks/ep0-datastore-copy.md` — never walked. `audits/offsite-lock-build-2026-10-03/partF/`. **`[FACT]` 2026-10-04: the restore route from the DooPlex copy was WALKED** on demo-hp (scratch VMID): listed in 2 s, restored in 186 s (15 GB logical), data read by `pct mount`; the restored config carries `onboot: 1` and the host's real drive binds — strip them on a restore beside the original (R-834). The copy now keeps 8 weekly copies (decision 71). `runbooks/ep0-datastore-copy.md` route 1. | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** **`[BETA-DEFERRED]`** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. **`[FACT]` 2026-10-03 (decisions 68–69, hub v0.127.0, controller v0.289.1): THE BOX CAN NO LONGER DELETE ITS RESTIC HISTORY.** Both demo boxes run on an append-only key the hub pinned; a delete from the box is refused (`403 Forbidden`, measured on each, counts 13→13 and 100→100); the box never receives the sub-account password (the old endpoint answers 410, measured with demo-hp's own key); the hub reads every key file daily. Retention runs only in a hub-opened window behind a fake-snapshot guard — mechanics proven live (window 1 on demo-hp: opened, guard refused, closed in 3 s, operator mailed); **a real prune inside a window is NOT yet proven** (R-824). Residual: an add-only attacker can still fill the quota and plant past-dated snapshots (R-822). `audits/offsite-lock-build-2026-10-03/`. **`[FACT]` 2026-10-04 (controller v0.290.0, hub v0.128.0):** the guard now skips same-day-superseded young snapshots instead of refusing; window 2 on demo-hp ran without refusal and removed nothing (127→127 — every candidate was a young same-day copy); weekly windows are ON fleet-wide. **A window that actually removes snapshots has still not been observed** (first candidates age past 8 days around 2026-10-11). The household's set-aside deletion is the hub's after 7 days (decision 74), proven live on a planted scratch dir. **`[FACT]` 2026-10-05 (controller v0.294.0, R-867 CLOSED): A WINDOW THAT REMOVES SNAPSHOTS IS OBSERVED, on both demo boxes.** The guard's line is now built from the policy's own constants (keep-daily 7 CALENDAR days, `09` decision 104) — the 8-day age line sat inside the keep window and refused every honest window (demo-felhom 2026-10-05 02:15 UTC, an error mail). One window each, opened by the operator's one-shot grant and run by the household's "run now": demo-felhom 16 → 14, demo-hp 145 → 127 — the removed snapshots are exactly those the policy predicted; the hub's rows read `pruned` with the same counts, no drop event, no mail, the key files clean. The UNATTENDED weekly window has not yet been observed removing (next due ~2026-10-11/12). Residual unchanged: R-822 (past-dated gap-fills steer older keeps). `audits/night-fixes-2026-10-05/partA/`. | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** **`[BETA-DEFERRED]`** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. **`[FACT]` 2026-10-03 (decisions 68–69, hub v0.127.0, controller v0.289.1): THE BOX CAN NO LONGER DELETE ITS RESTIC HISTORY.** Both demo boxes run on an append-only key the hub pinned; a delete from the box is refused (`403 Forbidden`, measured on each, counts 13→13 and 100→100); the box never receives the sub-account password (the old endpoint answers 410, measured with demo-hp's own key); the hub reads every key file daily. Retention runs only in a hub-opened window behind a fake-snapshot guard — mechanics proven live (window 1 on demo-hp: opened, guard refused, closed in 3 s, operator mailed); **a real prune inside a window is NOT yet proven** (R-824). Residual: an add-only attacker can still fill the quota and plant past-dated snapshots (R-822). `audits/offsite-lock-build-2026-10-03/`. **`[FACT]` 2026-10-04 (controller v0.290.0, hub v0.128.0):** the guard now skips same-day-superseded young snapshots instead of refusing; window 2 on demo-hp ran without refusal and removed nothing (127→127 — every candidate was a young same-day copy); weekly windows are ON fleet-wide. **A window that actually removes snapshots has still not been observed** (first candidates age past 8 days around 2026-10-11). The household's set-aside deletion is the hub's after 7 days (decision 74), proven live on a planted scratch dir. **`[FACT]` 2026-10-05 (controller v0.294.0, R-867 CLOSED): A WINDOW THAT REMOVES SNAPSHOTS IS OBSERVED, on both demo boxes.** The guard's line is now built from the policy's own constants (keep-daily 7 CALENDAR days, `09` decision 104) — the 8-day age line sat inside the keep window and refused every honest window (demo-felhom 2026-10-05 02:15 UTC, an error mail). One window each, opened by the operator's one-shot grant and run by the household's "run now": demo-felhom 16 → 14, demo-hp 145 → 127 — the removed snapshots are exactly those the policy predicted; the hub's rows read `pruned` with the same counts, no drop event, no mail, the key files clean. The UNATTENDED weekly window has not yet been observed removing (next due ~2026-10-11/12). Residual unchanged: R-822 (past-dated gap-fills steer older keeps). `audits/night-fixes-2026-10-05/partA/`. **`[DESIGN]` 2026-10-08 (R-435, `09` §3 decision 191; hub main, unreleased): on a pinned tier even ONE snapshot that disappears outside a hub-opened window raises an error mail** — each window explains at most its hub-set cap; NAS tiers keep the half-rule; the count is the box's own and net (`08` §6.5). **2026-10-08 (R-893, decision 192; controller main, unreleased):** a failed off-site replay of one app after a moved definition or a replaced volume now HOLDS the app stopped for support instead of starting it on mixed data (§6.3); „put back exactly as it was" is the next slice. | | 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** **`[BETA-DEFERRED]`** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | | 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** **`[BETA-DEFERRED]`** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | | 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** **`[BETA-DEFERRED]`** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D **`[FACT]` 2026-10-03:** the PBS (whole-guest) off-site copy now also exists OFF Hetzner — DooPlex pulls ep0's datastore nightly (decision 70). The restic tier is still Hetzner-only. | diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index 01f5a9af..e4d9e2b5 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -387,6 +387,19 @@ Design home: `07` §6.1.1 (`09` decisions 109–110; CC decisions 115–116, *op --- +## 6.5 One lost off-site copy on a pinned tier [DESIGN, R-435 — `09` §3 decision 191, hub main 2026-10-08, unreleased] + +On a **pinned** off-site tier (the hub holds a confirmed append-only key, decision 69) the only legitimate way the +snapshot count can fall is a clean-up window the hub opened (decision 68). So the hub raises +`offsite_snapshots_dropped` (**error**) when the count falls by even ONE more than those windows explain since the +previous trustworthy report. Each window explains at most its own hub-set `max_remove` — a box's `count_after` cannot +widen it, and a window closed by timeout or still open explains exactly that cap; each window explains one fall only; +a window stuck open past its deadline explains nothing; when the windows cannot be read, the half-rule decides (fail +closed). Non-pinned (NAS) tiers keep the more-than-half rule. **Limits:** the count comes from the box (R-895's +caveat holds); it is a NET count — new snapshots between two reports hide the same number of deletions; the „window +spent" memory is in-process, so after a hub restart one window can explain one more fall. Pinned by +`hub/internal/monitor/r435_pinned_drop_test.go` and `hub/internal/store/r435_windows_between_test.go`. + ## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0] **"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.** diff --git a/documentation/architecture/10-localisation.md b/documentation/architecture/10-localisation.md index 1ede3b73..c0ab4b44 100644 --- a/documentation/architecture/10-localisation.md +++ b/documentation/architecture/10-localisation.md @@ -376,6 +376,13 @@ harmless (`TestFlashKeyRoundTrip`). The same shape one layer in: an **alert bann background health cycle and read minutes later, so `Alert` carries `MessageKey` + `MessageArgs` and `GetAlerts(lang)` renders on the way out. +**[FACT] 2026-10-08 (R-79, `09` §3 decision 187; controller main, unreleased): every health issue and warning now +carries its key** — the last six producers (Docker unreachable, protected container down, storage unavailable / not +separate / usage high / almost full) joined the seven resource ones; the wire text is unchanged byte for byte. The +Docker error inside its sentence stays English (Docker writes it). **The household's health mail no longer carries the +raw details note** (hub main, unreleased): it says what the headline says and points to the dashboard; the operator's +mail keeps the note (`hub/internal/notify/r79_health_mail_test.go`). + **[DESIGN] Word order is Go's explicit argument index, not a second placeholder syntax.** The plan proposed a named-parameter (`{{.Name}}`) form for multi-parameter Go messages. English reorders with `%[2]s`, which `fmt` already understands, so the Hungarian value stays **the format string the code diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 5ded6f9b..e7e0978b 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -154,12 +154,12 @@ stopping line that lies. | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-08 — owner Viktor.** (a) DONE with the operator's yes in chat: `notify_failure` now mails admin@felhom.eu through Resend; proven by one test mail that reached the inbox (`audits/day-2026-10-08/r232/`; no backup was started). (b) partly: the hub database leaves DooPlex nightly to ep0 (R-173); everything else stays on the box. (c)–(h) unchanged. **READY** for the rest | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — the operator mail is BUILT on main 2026-10-08 (decision 183; controller `a40729a` + hub, ships with the next releases).** The honesty fix is on main too (agent `91b9405`, controller `75b3b39`). Left: the runbook „open a retained package for a household" (design option C, second slice). Design `audits/day-2026-10-08/design-R-304.md`. | R-198, R-199, R-224, R-241 | Release controller + hub; write the retained-package runbook from the 2026-08-12 drill §4; then close | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC | -| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-893.md` — re-verified; and the screen's „your data is back as it was" (`err.backup.db_restore_failed_rolled_back`) is false for files, other volumes and the app version. Question D8 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D8 (`09` §3 decision 192):** yes, both, in that order — first the app stays stopped for support (option C), next „put back exactly as it was" (option A). | **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | +| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-893.md` — re-verified; and the screen's „your data is back as it was" (`err.backup.db_restore_failed_rolled_back`) is false for files, other volumes and the app version. Question D8 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D8 (`09` §3 decision 192):** yes, both, in that order — first the app stays stopped for support (option C), next „put back exactly as it was" (option A). | **NARROWED 2026-10-08 — the first half (D8 option C, the app held stopped) is built on controller main and ships tomorrow; what stays open is the second half, „put back exactly as it was” (option A: a pre-restore copy of the app's volumes and placed files, a fit check, disk for one copy; the slice-0 measurement on 9202 first).** **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | | **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** **2026-10-07 07:58: kept open for later (`09` §3 decision 166).** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** **2026-10-06: leg (a) PUSHED** to the live catalog (`c265b37`). **2026-10-06 (burn-down night, later): leg (b) NEEDS A DESIGN.** `07` §7.4 sets no direction; refusing the restore without the DB password, or `ALTER USER` after it, each change restore behaviour on customer data. | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-231** | Backup & restore | P3 | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | — | — | operator | -| **R-314** | Backup & restore | P3 | **`StopAbandon` has no web route — a customer who telephones is served by a command line.** `--abandon-stop` exists on the controller binary (`cmd/controller/main.go:86`) and refuses rather than silently no-opping, which is right. But there is no handler: the only in-product way to cancel a countdown is the customer finding their recovery code (Scenario G cancels it at the moment the code proves they still have it). **An operator who is telephoned instead has to reach a shell on the customer's machine.** Used this session on the operator's ruling, container stopped first so the running controller could not overwrite `settings.json` from memory — a sequencing subtlety that is itself an argument for a route **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-314-279-177.md` — the box's own 14-day countdown has no operator door (the hub's 7-day phase does, decision 74); one door for R-314 + R-279: a fixed list of operator actions carried in the report reply. Question D1 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D1 (`09` §3 decision 185):** yes — option A, a fixed list of operator actions carried in the box's report reply. | **READY (S) — NEW 2026-08-12, RANK 3** **2026-10-05 (burn-down night): NEEDS A DESIGN.** There is no operator door into a RUNNING controller (its HTTP surface is the household's session; the CLI runs in a separate process and cannot safely trigger a job in the running one). Same design as R-177 and R-279. | R-241, R-307 | An operator-authenticated POST that calls the same `StopAbandon`, so the telephone path and the code path converge | CC | +| **R-314** | Backup & restore | P3 | **`StopAbandon` has no web route — a customer who telephones is served by a command line.** `--abandon-stop` exists on the controller binary (`cmd/controller/main.go:86`) and refuses rather than silently no-opping, which is right. But there is no handler: the only in-product way to cancel a countdown is the customer finding their recovery code (Scenario G cancels it at the moment the code proves they still have it). **An operator who is telephoned instead has to reach a shell on the customer's machine.** Used this session on the operator's ruling, container stopped first so the running controller could not overwrite `settings.json` from memory — a sequencing subtlety that is itself an argument for a route **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-314-279-177.md` — the box's own 14-day countdown has no operator door (the hub's 7-day phase does, decision 74); one door for R-314 + R-279: a fixed list of operator actions carried in the report reply. Question D1 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D1 (`09` §3 decision 185):** yes — option A, a fixed list of operator actions carried in the box's report reply. | **VERIFY — built on main 2026-10-08 (D1, decision 185), ships with tomorrow's controller + hub releases; closes after the live proof on 9202** (`abandon_stop` / `abandon_extend` through the hub's buttons). **READY (S) — NEW 2026-08-12, RANK 3** **2026-10-05 (burn-down night): NEEDS A DESIGN.** There is no operator door into a RUNNING controller (its HTTP surface is the household's session; the CLI runs in a separate process and cannot safely trigger a job in the running one). Same design as R-177 and R-279. | R-241, R-307 | An operator-authenticated POST that calls the same `StopAbandon`, so the telephone path and the code path converge | CC | | **R-401** | Backup & restore | P3 | **Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store.** Controller v0.228.0 (R-399) made `--read-data-subset=100%` the default for every box. The whole justification is a single data point: `demo-hp`, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% **39.2 s** — four seconds. Re-proven live 2026-08-31 at 38.7 s. **It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA.** A 50 GB store is ~370x the data and this curve says nothing about it. **Nothing was invented from that one point** — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. **`readDataSubsetRe` already accepts `n/m`**, so a rotating schedule (`1/7` on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. **THE TRIGGER IS AN EVENT, NOT A DATE:** the slow-check WARN from v0.228.0 firing on any box (`integritySlowNoticeThreshold`, 5 min) — that line names the duration, the depth and this row. **WHAT HAPPENS IF NOBODY ACTS:** every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. **Whoever acts must also revisit `integrityCheckTimeout` (30 min)**, which is now the number a large store meets first. | **OPEN — WATCHING** | — | When the WARN fires: measure the curve on that store, then choose between a rotation (`n/m`), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC | | **R-409** | Backup & restore | P3 | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC | | **R-433** | Backup & restore | P3 | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. **-- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628).** **Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box?** Hetzner: *"With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation"*, citing `docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots`; and separately *"A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken."* **So the Storage BOX documentation governs, not the Storage Share FAQ** - which is exactly the trap `provider-questions-2026-09-01.md` warned the reader about, and the answer came back on the right side of it. **File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it**, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). **HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on `u629488`. That is now the measurement this row needs, and it is cheap.** **Q2 - is `--append-only` enforced by Hetzner, or taken from what the client sends?** Hetzner: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, with `fluix.one/blog/hetzner-restic-append-only/`. **So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's `authorized_keys`** - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. **This is the answer R-95 has been blocked on** and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. **NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof.** What is owed is a real forced-command key on `u629488`, a restic `forget --prune` through it that is REFUSED, and a backup through it that still succeeds. | **BLOCKED** — **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | — | — | CC + operator | @@ -170,7 +170,7 @@ stopping line that lies. | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | | **R-246** | Backup & restore | P4 | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | -| **R-279** | Backup & restore | P4 | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-314-279-177.md` — still true: the only on-demand off-site run is the household's dashboard route; same door as R-314. Question D1 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D1 (`09` §3 decision 185):** yes — `offsite_backup_now` is one of the operator actions in the report reply. | **READY (XS) — NEW 2026-08-09** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the same missing operator door as R-314 and R-177. | — | Same shape as R-177; solve both together | CC | +| **R-279** | Backup & restore | P4 | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-314-279-177.md` — still true: the only on-demand off-site run is the household's dashboard route; same door as R-314. Question D1 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D1 (`09` §3 decision 185):** yes — `offsite_backup_now` is one of the operator actions in the report reply. | **VERIFY — built on main 2026-10-08 (D1, decision 185): `offsite_backup_now` in the report reply; ships tomorrow; closes after the live proof on 9202.** **READY (XS) — NEW 2026-08-09** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the same missing operator door as R-314 and R-177. | — | Same shape as R-177; solve both together | CC | | **R-336** | Backup & restore | P4 | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** **2026-10-06 night: no hub half** — the pollers are pvestatd and proxmox-backup-client on the boxes; the lever (disabling the storage entry) collides with the agent's pbsdr health model; still a design question. | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 **SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it.** `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **The leak is OURS, and the poll rate is not what feeds it.** Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by **`felhom-agent`** on the boxes — 194 on each, `ss -tnp` naming a single PID per box, and **zero** held by `pvestatd` or `proxmox-backup-client`. Confirmed independently from ep0's access log over the same 46.18 h window: `libwww-perl` (pvestatd) **81,192 requests -> 0 descriptors**, `proxmox-backup-client` **80,061 requests -> 0 descriptors**, `Go-http-client/1.1` (the agent) **387 `/snapshots` calls -> 388 sockets — one per call, within one**. So **162,404 requests, 99.5% of the traffic, produce 0% of the leak.** **Mechanism, named from source:** `felhom-agent/internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` — a composite literal, so `IdleConnTimeout` is the zero value = **no limit** (`http.DefaultTransport` sets 90 s; a literal does not inherit it) — and `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh client every cycle**, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; `CloseIdleConnections`/`IdleConnTimeout`/`MaxIdleConns` appear **nowhere** in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h `DefaultVerifyCadence` (7.7 cycles) = 192.4 predicted vs **194 observed per box**. **CONSEQUENCE — RE-RANK.** The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") **would have produced a null result and read as a failed fix.** Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and **zero** descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a **scaling/cost item, not the leak fix**, and the leak fix is **R-344**. **Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED** — Phase C is held at STOP 1 with its prediction pre-registered in `evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt`. Do not record a proportionality verdict here until that window has run. **RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX.** That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: *cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing.* Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — **ours, in `felhom-agent`, R-344** — would have been further from being found, not closer. **Measured 2026-08-20:** `pvestatd` (`libwww-perl`) and `proxmox-backup-client` made **162,404 requests** in a 46 h window and leaked **zero** descriptors; the agent made 811 and leaked **388**. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its **baseline of 17**, with the poll rate completely unchanged — 85,000/day before and after. **WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits:** ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so **at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33**. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — `pvestatd` stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with `felhom-agent/internal/pbsdr/manager.go`'s health model. **NEW ACCEPTANCE CRITERION, since the old one is void:** the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. | CC | | **R-526** | Backup & restore | P4 | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** **Re-ranked 2026-10-03: P3->P4: operator teardown op on a protected box; adopt path covers the need.** | — | — | operator | | **R-541** | Backup & restore | P4 | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** **Re-ranked 2026-10-03: P3->P4: needs a second pool box or an outgrown customer first; operator-only and far off (0.3% full).** | — | — | CC | @@ -192,14 +192,14 @@ stopping line that lies. | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -## Security & access — 14 rows (P2 1, P3 11, P4 2) +## Security & access — 15 rows (P2 1, P3 11, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. **2026-10-07 07:58: `09` §3 decision 165 — (a) A1 yes before the first paying customer; (b) B3 accept + B2 hygiene in the same bundle; (c) C2 accept.** **2026-10-07: (a) A1 and (b) B2 DELIVERED (agent 0.151.0 + bundle on demo-hp, demo-felhom, Tester 1; probe 68/68).** Live on demo-hp: no `tee` grant left in `sudo -l -U felhom-agent`; the verb `felhom-priv-apply ^controller-image [0-9]+$` is the route; a hand-fed `docker.io/library/alpine:latest` → `REFUSED [I1]` rc 3, the guest's image file unchanged; the old `pct exec … tee` asks for a password; felhom-op's pct lines anchored (`audits/day-2026-10-07/C/C-live-demo-hp.txt`). **LEFT:** one managed controller swap seen through the verb — no newer controller existed today; the next controller release shows it. (c) accepted (decision 165). | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | -| **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-138.md` — the shared-zone premise is gone (own domain per customer, `01` §7), but nothing enforces it and an account-wide token has the same reach; option B (hub refuses a duplicate or nested domain) needs no decision and also covers R-415. Question D6 on STATUS's decision sheet. **-- 2026-10-08 (afternoon):** option B BUILT on hub main (the duplicate/nested/felhom.eu domain guard, R-415 closed with it). Left: option C (check the key's reach with Cloudflare) — question D6. **-- 2026-10-08 14:16 operator ruling D6 (`09` §3 decision 190):** yes — option C, the hub checks with Cloudflare that a pasted key reaches only the customer's own zone. | READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | +| **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-138.md` — the shared-zone premise is gone (own domain per customer, `01` §7), but nothing enforces it and an account-wide token has the same reach; option B (hub refuses a duplicate or nested domain) needs no decision and also covers R-415. Question D6 on STATUS's decision sheet. **-- 2026-10-08 (afternoon):** option B BUILT on hub main (the duplicate/nested/felhom.eu domain guard, R-415 closed with it). Left: option C (check the key's reach with Cloudflare) — question D6. **-- 2026-10-08 14:16 operator ruling D6 (`09` §3 decision 190):** yes — option C, the hub checks with Cloudflare that a pasted key reaches only the customer's own zone. | **VERIFY — built on hub main 2026-10-08 (D6, decision 190), ships tomorrow; closes after each stored customer token is checked once (tokens saved before the check were never checked); R-913 holds the read-scope limit.** READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`28a5203`; the catalog clone stores no credentials; the token is supplied per fetch; R-615's repo comparison ignores credentials on both sides, so a token never re-clones (`TestR616_TokenSetSameRepoNoRecloneOriginClean`) and a credentialed origin is cleaned at the next pull. The operator's Gitea admin token rotation (the row's second half) is still owed). Ships with the next controller release; close after delivery. **2026-10-06: DELIVERED** in controller v0.298.0 (the clone stores no credentials). Left: the operator's Gitea admin token rotation. | — | — | CC + operator | @@ -211,6 +211,7 @@ stopping line that lies. | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | | **R-904** | Security & access | P4 | **Cloudflare can read every household's app traffic; replacing it with our own relay is a later item.** Facts (reviewer discussion 2026-10-08; `01`): app traffic and the dashboard reach the box through the Cloudflare Tunnel (`01` §5 trust table, rows end-user ↔ apps and customer ↔ controller UI; §7), and the tunnel's public end is Cloudflare's edge, where TLS ends — so Cloudflare can technically read that traffic (the FAQ says so since 2026-10-08, R-900; „TLS ends at the edge" is not written in `01` — add it there). What Cloudflare gives today, free: inbound reach with no router setup, the CGNAT answer (`01` §4, §7); certificates (`01` §7, the free tier covers one level below a zone); the geo-WAF the hub enforces (`01` §5 last row, §7); flood protection (not in `01`). The alternative named: our own EU relay over WireGuard with TLS passthrough by SNI, certificates on the box, the geo-block on the relay. Its costs: one more machine the operator keeps up, and a single point of reach for every box; weaker flood protection; about a week of work after a spike. Operator ruling 2026-10-08 09:07 (`09` §3 decision 184): a later item. | **DEFERRED — after the first customers (operator ruling 2026-10-08 09:07)** | — | A spike after the first customers (the relay's reach, cost and flood behaviour, measured) | operator | +| **R-913** | Security & access | P4 | **The Cloudflare token check (R-138, decision 190) reads what a token can SEE, not what it can WRITE.** FOUND 2026-10-08 by the security review of the R-138 build: `GET /zones` lists zones the token can read; a hand-built token with Zone:Read on the customer's zone and DNS:Edit on ALL zones would pass the check and could still change every household's DNS. The token wizard's single „Specific zone" scope does not build such a token, and a customer token cannot read its own policies (that needs „API Tokens Read"). Written as a limit in `01` §7. **Options:** (a) the operator mints every customer token from one recipe and the hub checks nothing more; (b) the hub mints the token itself with the operator's account token (one zone, DNS:Edit) — a new privileged credential on the hub; (c) a negative probe: with the pasted token, try to READ the DNS records of another customer's zone by its id (the hub knows the ids) and refuse on success — catches the read side only. | **OPEN — needs the operator** (a is free today; b is a design) | operator: which of a/b/c | Pick a/b/c; if nothing: the check stands as built and the limit stays written in `01` §7 | operator | ## Box system & updates — 6 rows (P2 1, P3 5) @@ -219,7 +220,7 @@ stopping line that lies. | **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). **2026-10-07 evening (decision 176): the night run moved to TONIGHT 7→8.** Pre-night fixes delivered (agent v0.153.0 — ring 0 stages exactly the told kernel, R-898 closed; controller v0.303.0 — the first capture after a boot waits for the drive, R-897 closed; hub v0.143.1 — Reply-To on household mails). Armed at 19:05: demo-felhom due 7.0.14-20 (household told 18:12), demo-hp due 7.0.14-22 (told 18:38; demo-hp's day-old apt lists refreshed by a report-only pass with its updates switched off for ~1 min). **WHERE IT STOPPED: read back the night on 2026-10-08 morning** (each box on its kernel as default, the apps back, minutes down, no false backup failure), then approve the set (the two boxes ran DIFFERENT kernels — the button waits until both boot the same one) and Part E4. **2026-10-08 read-back (`audits/kernel-night-2026-10-07/readback/`): demo-felhom PASSED its first real kernel night** — whole-guest backup 04:35, OS + Proxmox steps 04:37–04:38, kernel staged 04:39:04 as EXACTLY the told 7.0.14-20 (not the newer 22 — R-898's fix held), restart 04:39:08, new kernel judged healthy 04:40:38 (38 s after the agent started), now the default; apps away ~1.5 min; no operator alarm, no false backup failure (demo-felhom has no drive apps, so R-897's fix was not exercised there); crash guard 0. **demo-hp took NO step** — no whole-guest backup that night (R-899): it stays on 7.0.14-20, still due 7.0.14-22; a new household mail is allowed after 14:38 today, so it should run the night 2026-10-08→09. Reply-To proven: the operator's reply to the test mail went to admin@felhom.eu (Gmail 2026-10-07 18:29 UTC). **WHERE IT STOPPED:** read back demo-hp the morning of 2026-10-09; the approval button waits until both demo boxes booted the SAME kernel (now 20 vs due 22) — likely a step to 22 on demo-felhom later; then Part E4. | — | — | CC | | **R-899** | Box system & updates | P3 | **A daytime whole-guest backup moves the 24 h cadence past the night window, so the next night has no whole-guest backup — and with it no OS leg and no kernel step.** SEEN 2026-10-08: demo-hp's last whole-guest backup was a morning press at 2026-10-07 08:49:53; due again 2026-10-08 08:49, after the window [04:30, 08:30) closed → no backup, no night leg; the household had been told „tonight" (18:38) and nothing happened (harmless: no change). The kernel lane recovers by itself (a new mail after 20 h, step the next night, inside the 3-mail limit), but every daytime press costs one night. Fix direction (a design, not taken): the cadence counts nights, not 24 h (due when the last whole-guest backup is older than ~20 h at the window's start), or a press does not reset the night's cadence. `audits/kernel-night-2026-10-07/readback/demo-hp-why-no-step.txt` | **VERIFY — fixed on main 2026-10-08, ships with tomorrow's controller + agent releases** (operator ruling 2026-10-08 07:12, option A, `09` §3 decision 177: every night takes its own). Controller `6d07ca2`: a durable ledger (last successful press / last successful scheduled run per tier); a tier the agent calls „not due" because of a press is still due inside tonight's window, never by the safety valve (`TestR899_*`, red-proved). Agent `4c69c25`: a press arrives as `trigger=manual` and runs no OS leg; the after-boot kernel reports carry the saved ring instead of „ring 1". The household mail case („tonight" for a night whose backup cannot start because of a press) is removed by the rule; no hub change. `07` §6.1. | — | Release controller + agent tomorrow; read back the first night after a daytime press; then close | CC | | **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator | -| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-35.md` — wider than the row: every controller update and crash restart also logs the household out; the family gate already keeps its sessions on disk. Question D4 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D4 (`09` §3 decision 188):** yes — option C, sessions kept across restarts, only a fingerprint on disk. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | +| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-35.md` — wider than the row: every controller update and crash restart also logs the household out; the family gate already keeps its sessions on disk. Question D4 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D4 (`09` §3 decision 188):** yes — option C, sessions kept across restarts, only a fingerprint on disk. | **VERIFY — built on controller main 2026-10-08 (D4, decision 188), ships tomorrow; closes after the 9202 check (sign in, restart the controller, the next page needs no login).** **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | | **R-242** | Box system & updates | P3 | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | — | — | operator | | **R-468** | Box system & updates | P3 | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | — | — | CC | @@ -230,12 +231,12 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-243** | Monitoring & notifications | P2 | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **VERIFY — built on hub main 2026-10-08 (`b119301c`), ships with tomorrow's hub release.** New operator-only `offsite_escrow_pending` (warning): off-site ON, escrow not done, no successful off-site run for 7 days (CC picked 7 days, `09` §3 decision 179 — operator may reverse); weekly re-send, persisted, clears with an info line. `08` §6.3. **Expected to fire for Tester 2 on the first sweep after the deploy** (hub read 2026-10-08: off-site on, escrow `pending`, never a success). | — | Deploy hub tomorrow; confirm the Tester 2 mail arrives (positive control); then close | operator | | **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. **2026-10-06: STOPPED BEFORE ANY CODE — the measurement contradicted the design** (decision 155 said A then C). On scratch 9202, Docker 29.8.2, the `OOMKilled` flag was TRUE in all four shapes tried: a child process killed while the container kept running (`oom_kill` 0 → 3, flag true) and three main-process kills (exit 137, flag true); an `oom` event each time. The false flags of 2026-09-15 were on Docker 29.8.0. Cost read for the record: one exec read of 21 containers on demo-hp = 1.6 s. `audits/design-build-2026-10-06/`C/. **2026-10-06 18:24: operator ruling, option A (decision 157):** nothing is built; the row stays open as a WATCH for a box whose Docker reports a false flag. **Added:** a Docker engine set may not be approved until it reports a memory kill correctly on the boxes that ran it (built in this row's next session). **2026-10-06 (night): the Docker-approval memory-kill check is BUILT** — hub v0.140.0 (LIVE: the approval waits for a passing `oom_check` on every ring-0 box; a failed, errored or missing one blocks) and the agent's wrapper (`felhom-agent` `acccb66`, unreleased — ships as v0.150.0 after the 2026-10-07 read-back). Proven by hand on demo-hp's guest: `OOMKilled=true`, exit 137, the `oom` event — seen only with an events window ending after the run (the wrapper waits 2 s). `audits/readback-2026-10-07/`F/. | — | Release agent v0.150.0 + bundle and deliver it; then the row is a WATCH only | CC | -| **R-79** | Monitoring & notifications | P3 | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-79.md` — half stale: the dashboard seam was built by R-516; six producers still show raw text (option A, no decision — CC builds); the household's health mail carries the raw details JSON (option B needs a word). Question D3 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D3 (`09` §3 decision 187):** yes — the household's health mail drops the raw note and points to the dashboard; the dashboard shows each issue in the household's language. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the issues and warnings travel to the hub as sentences; changing them is the two-repo spike the row itself names. | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC | +| **R-79** | Monitoring & notifications | P3 | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-79.md` — half stale: the dashboard seam was built by R-516; six producers still show raw text (option A, no decision — CC builds); the household's health mail carries the raw details JSON (option B needs a word). Question D3 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D3 (`09` §3 decision 187):** yes — the household's health mail drops the raw note and points to the dashboard; the dashboard shows each issue in the household's language. | **VERIFY — both halves built on main 2026-10-08 (D3, decision 187): controller keys for the last six producers, hub mail without the raw note; ships tomorrow; closes after a health mail is read on a Tier-0 box.** **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the issues and warnings travel to the hub as sentences; changing them is the two-repo spike the row itself names. | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC | | **R-211** | Monitoring & notifications | P3 | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC | | **R-333** | Monitoring & notifications | P3 | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | | **R-340** | Monitoring & notifications | P3 | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC | | **R-388** | Monitoring & notifications | P3 | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor | -| **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-435.md` — the box's key can no longer delete (decisions 68–69); deletion through any other credential still needs >50 % to alarm. Question D7 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D7 (`09` §3 decision 191):** yes — option C, an error mail when even one snapshot disappears outside a hub-opened clean-up window (pinned tiers). | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Covering one-app deletions needs per-app snapshot counts from the controller and a per-tag threshold measured on real counts — two repos, a mechanism nobody has measured. | — | — | CC | +| **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-435.md` — the box's key can no longer delete (decisions 68–69); deletion through any other credential still needs >50 % to alarm. Question D7 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D7 (`09` §3 decision 191):** yes — option C, an error mail when even one snapshot disappears outside a hub-opened clean-up window (pinned tiers). | **VERIFY — built on hub main 2026-10-08 (D7, decision 191), ships tomorrow; closes after the 9202 proof (a hand-granted window removes N, no alarm; the hub's window row as the control).** **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Covering one-app deletions needs per-app snapshot counts from the controller and a per-tag threshold measured on real counts — two repos, a mechanism nobody has measured. | — | — | CC | | **R-521** | Monitoring & notifications | P3 | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`800b32c` — controller half: apps a lost drive stopped no longer mail app_start_failed one by one; test + red-proof). Ships with the next controller release; close after delivery. Left: the hub's cooldown outliving a recovery (F6/F7) and whether a household gets a mail for a lost drive (operator). **2026-10-06: controller half DELIVERED** (v0.298.0). | — | — | CC + operator | | **R-724** | Monitoring & notifications | P3 | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** **2026-10-06 night: confirmed not a controller-only fix** — the only host-network container is samba (`infra/samba.go:123`) and the agent's local API has no guest-network route; the fix is a new agent route (e.g. the PVE lxc interfaces read) + a controller reader — a two-repo design. | — | — | CC | | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | @@ -250,7 +251,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-173** | Hub & operator | P2 | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **NARROWED 2026-10-05 (evening) — option A IN FORCE** (`09` decision 125). The hub writes a nightly `VACUUM INTO` snapshot at 02:00 (hub v0.136.0, `05` §16.3, keep 2; the volume grew to 2 Gi); DooPlex checks it (`integrity_check`, size, ≥1 host, ≤26 h old), encrypts it with a key ep0 never sees and pushes it at 02:30 to ep0's `operator` namespace with a write-only token; a read-only token restore-tests it every Sunday 04:30 (and refuses a readable console password); `HubDBBackupStale`/`HubDBRestoreTestStale` alarm on success-only timestamps, `absent()` included. Both keys are off DooPlex (operator, 2026-10-05). PVC label fixed (`enabled`). Proven live: first push 7 s, restore test, token limits, a key rebuilt from the paper `data` field decrypts, runbook §3 steps 1–3 (4/4 console passwords open with the saved seal key, 0/4 with a random one). `audits/hub-db-offsite-2026-10-05/` | — | **LEFT:** runbook §3 steps 4–5 (the copy into a live PVC) are not exercised — they need the hub down; do them at the next planned hub maintenance or a DR drill on a scratch k3s. Close then. | CC | -| **R-30** | Hub & operator | P3 | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-30.md` — the line is 45 min now (live ConfigMap); the wait channel reconnects every 241–243 s (DooPlex ingress log, read-only); slice 1 (show live-link presence on the host page) needs no decision. Question D2 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D2 (`09` §3 decision 186):** yes — „delete host" goes ahead at once after „I checked: the box is off" when the hub has had no connection for ~6 min. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk.** **2026-10-06 night: stopped — a design:** presence from the wait channel is not a direct signal (nginx holds the box's connection; a powered-off box is seen only when no new wait begins — about 240 s + grace, unmeasured), and the same presence gates host-delete before RESET (a destructive path). Needs a measured presence rule and the operator's word on the delete guard. | — | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | CC | +| **R-30** | Hub & operator | P3 | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-30.md` — the line is 45 min now (live ConfigMap); the wait channel reconnects every 241–243 s (DooPlex ingress log, read-only); slice 1 (show live-link presence on the host page) needs no decision. Question D2 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D2 (`09` §3 decision 186):** yes — „delete host" goes ahead at once after „I checked: the box is off" when the hub has had no connection for ~6 min. **-- 2026-10-08 security review of the build:** presence comes from the in-guest controller's wait channel, so a running host whose guest is stopped (or whose controller is in its 30-min crash-loop pause, or after an ingress outage while the controller backs off up to 5 min) reads „not connected" — and the tick then deletes a live host, whose agent is locked out until re-enrolled. That is the cost the operator accepted with decision 186 (no household data is touched); the host page says „Box connection", not „host off". Option C (the WireGuard handshake as host-level presence) would remove it. | **VERIFY — built on hub main 2026-10-08 (D2, decision 186: presence + delete at once after the tick), ships with tomorrow's hub release; closes after the 9202 measurement (stop 9202, time until „not connected”, control: the ingress log's last wait line).** **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk.** **2026-10-06 night: stopped — a design:** presence from the wait channel is not a direct signal (nginx holds the box's connection; a powered-off box is seen only when no new wait begins — about 240 s + grace, unmeasured), and the same presence gates host-delete before RESET (a destructive path). Needs a measured presence rule and the operator's word on the delete guard. | — | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | CC | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** **2026-10-06 night: the race half fixed on felhom.eu main (hub, unreleased):** a second Save while the first still provisions is refused with 409 („already running — wait about a minute, then reload; do not save again"); nothing saved, nothing created; per customer, in memory. `TestProvision_R31_*`, red-proof `audits/night-burndown-2026-10-06/hub/R-31-red.txt`. LEFT: the async save with a status card (the escrow-card idiom). | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index 777704c2..75c307c3 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,3 +1,37 @@ +## Unreleased (2026-10-08, evening) — the operator's decision sheet: operator actions (D1), box presence + delete a switched-off box at once (D2), the household's health mail points to the dashboard (D3), the Cloudflare token reach check (D6), one lost off-site snapshot is an error (D7) — ships with tomorrow's hub release + +Operator rulings 2026-10-08 14:16, `09` §3 decisions 185–194. **Operator action on deploy: none.** The controller of the +same day carries D1's other half; an older controller ignores `operator_actions` (rows then expire after 24 h). + +- **D1 — operator actions (R-314, R-279).** Table `operator_actions(id, customer_id, action, arg, requested_at, + requested_by, done_at, outcome, message)`; four buttons on the host page (POST `/hosts/{id}/operator-action`, + operator auth + CSRF); the row is carried in the report reply (`operator_actions`) until the box answers + (`operator_action_results`); a result naming another customer's id is ignored; unanswered after 24 h → `expired`; a + customer RESET → `cancelled`. CLOSED list, validated at the POST (400, no row): `offsite_backup_now`, `abandon_stop`, + `abandon_extend` (1–30), `run_job` (`fill-watch`, `offsite-integrity`, `offsite-proof`, `disk-health-check`) + (`TestOperatorActions_ClosedList`). Who pressed = the channel and address (one password, no user names), logged at + INFO at the press and at the result; each result is a stored `operator_action` event (never a customer mail). New + wire-gate root for the reply field. +- **D2 — box presence from the wait channel (R-30, slices 1 and 2).** `intent.Hub` records wait starts/ends per + customer: connected / not connected since T / unknown (hub younger than 333 s). The host page shows „Box connection". + „Delete host" goes ahead at once when the host is online by report, the operator ticked „I checked: the box is off", + and presence has been „not connected" ≥ 360 s — re-checked at the POST; otherwise 409 as before. INFO audit line + one + `host_deleted_box_off` event (warning, hub-minted). `TestPresence_*`, `TestR30_*` (red-proved, 7 mutations). +- **D3 — the household's health mail (R-79 option B).** `health_degraded` / `health_critical` / `health_recovered` + customer mails drop the raw details note and end with „A részleteket a vezérlőpultodon látod." / „You can see the + details on your dashboard." (`mail.customer.line.health_dashboard`). The operator's mail keeps the note; other event + types unchanged. 6 mail goldens regenerated. `TestR79_*`. +- **D6 — the Cloudflare token reach check (R-138 option C).** On create/edit, a new token (or the same token on a new + domain) is checked with Cloudflare `GET /zones`: saved only when it sees exactly one zone and that zone IS the + customer's domain. Cloudflare down → refused, the old token stays. Never logged. **Security review fixes (same day):** + a parent zone that could hold a sibling customer is refused, order-independently (strict equality). Tests with a fake + Cloudflare, `TestR138_*`, `TestZoneEquals`, `TestTokenZones_*`. Limit written in `01` §7 (listing = read scope). +- **D7 — one lost off-site snapshot is an error (R-435 option C).** Pinned tiers: any fall the hub's own clean-up + windows do not explain raises `offsite_snapshots_dropped` (error, „outside any clean-up window the hub opened"). + **Security review fixes (same day):** each window explains at most its hub-set `max_remove` (a box's `count_after` + cannot widen it); each window explains one fall only; a window stuck open past its deadline explains nothing; + unreadable windows fall back to the half-rule. NAS tiers keep the half-rule. `TestR435_*` (store + monitor). `08` §6.5. + ## Unreleased (2026-10-08) — an alarm when a box never backs up off-site because its escrow is pending (R-243; `09` §3 decision 179); a deleted customer's audit rows go after 1 year (R-901; decision 181); the operator's older-recovery-package mail (R-304; decision 183); no duplicate or nested customer domain (R-415, R-138 option B) — ships with tomorrow's hub release **Operator action on deploy: none.** Expect ONE `offsite_escrow_pending` mail for **Tester 2** on the first sweep after diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 0160105f..66ce22fd 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,10 @@ +## gates — the decoy suite runs in a linked git worktree (2026-10-08, fixed without a row) + +- `test_gate_decoys.py` took its lock at `ROOT/.git/decoy-suite.lock`; in a `git worktree add` copy `.git` is a FILE, so + the suite crashed (`NotADirectoryError`) and `repo_gates.py` went red for every helper working in a worktree. The lock + now lives in `git rev-parse --absolute-git-dir` (per worktree — each worktree mutates its own files), falling back to + `ROOT/.git`. Measured: the full `repo_gates.py --fast` run green inside `worktrees/hub-integ`. + ## gates — site gate 20: the one picture viewer (2026-10-08) - Every `data-gallery` opener must be an `` to an existing file under `website/assets/`; a page with openers must diff --git a/scripts/test_gate_decoys.py b/scripts/test_gate_decoys.py index 3ae62375..460085f9 100644 --- a/scripts/test_gate_decoys.py +++ b/scripts/test_gate_decoys.py @@ -39,7 +39,15 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) # already mutated, and B's "restore" then wrote A's decoy back for good: website/en/apps.html was left without its # Radicale logo, one step from a push. An exclusive lock makes a second run WAIT for the first. import fcntl -_LOCK = open(os.path.join(ROOT, ".git", "decoy-suite.lock"), "w") +# The lock lives in THIS worktree's git dir: in a linked worktree (`git worktree add`) `.git` is a FILE, and the old +# `ROOT/.git/decoy-suite.lock` crashed with NotADirectoryError (2026-10-08). Each worktree mutates its own files, so a +# per-worktree lock is the right scope. +try: + _GITDIR = subprocess.run(["git", "rev-parse", "--absolute-git-dir"], cwd=ROOT, capture_output=True, text=True, + check=True).stdout.strip() +except (OSError, subprocess.CalledProcessError): + _GITDIR = os.path.join(ROOT, ".git") +_LOCK = open(os.path.join(_GITDIR, "decoy-suite.lock"), "w") fcntl.flock(_LOCK, fcntl.LOCK_EX) # ── WHAT THIS FILE COVERS ────────────────────────────────────────────────────────────────────────