diff --git a/documentation/audits/day-2026-10-08/design-R-304.md b/documentation/audits/day-2026-10-08/design-R-304.md index 3b51c696..eddbe0ad 100644 --- a/documentation/audits/day-2026-10-08/design-R-304.md +++ b/documentation/audits/day-2026-10-08/design-R-304.md @@ -1,6 +1,6 @@ # R-304 — the household's old recovery code and the retained packages: a one-page design (2026-10-08) -**Status:** design only. Nothing here is built except today's honesty fix (below). Customer data and promises are the +**Status:** option C's first slice BUILT 2026-10-08 afternoon on the operator's ruling (09:04, `09` §3 decision 183) — see the end. Ships with the next controller + hub releases. Customer data and promises are the operator's: they are the two questions at the end. ## Where it stands (read in source today, not from the row) @@ -48,3 +48,12 @@ today's fix turns from a silent „wrong code" into an honest „not all checked the honest route is „contact support" (A/C). 2. **Do you want C's first slice built (an operator mail when a household's code opens — or may open — an old package)?** *If you do nothing:* nothing is built; you learn of such a household only when they write to you. + +## Built 2026-10-08 (afternoon) — option C, first slice + +- **Controller** (`a40729a`): on a 422 (opens an older package) or a 424 (may open one), `recovery_older_package` + (warning) to the hub — the class and the package's date, never the code; at most once per day per box + (`/recovery-older-mail.day`). `TestR304_OlderPackageMail_*`. +- **Hub** (same day): the type is allowlisted and operator-only. `TestR304_RecoveryOlderPackageIsAllowlistedAndOperatorOnly`. +- **Not built:** the runbook „open a retained package for a household" (the drill's §4 written down) — next slice. +- **Question 1 (the promise wording) was not answered:** the wording stays. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index ffe8d013..b3a4c3c3 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -150,7 +150,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-08 — owner Viktor.** (a) DONE with the operator's yes in chat: `notify_failure` now mails admin@felhom.eu through Resend; proven by one test mail that reached the inbox (`audits/day-2026-10-08/r232/`; no backup was started). (b) partly: the hub database leaves DooPlex nightly to ep0 (R-173); everything else stays on the box. (c)–(h) unchanged. **READY** for the rest | — | — | operator | -| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — ruled 2026-10-08 09:04 (`09` §3 decision 183): the operator gets a mail when a household's code opens, or may open, an older escrow package (design option C).** The honesty fix is on main (agent `91b9405`, controller `75b3b39`, ships tomorrow); design `audits/day-2026-10-08/design-R-304.md`. The promise-wording question (design Q1) was not answered: the wording stays as it is. | R-198, R-199, R-224, R-241 | Build the operator mail (controller event + hub allowlist and operator-only), ships with the next releases | CC | +| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — the operator mail is BUILT on main 2026-10-08 (decision 183; controller `a40729a` + hub, ships with the next releases).** The honesty fix is on main too (agent `91b9405`, controller `75b3b39`). Left: the runbook „open a retained package for a household" (design option C, second slice). Design `audits/day-2026-10-08/design-R-304.md`. | R-198, R-199, R-224, R-241 | Release controller + hub; write the retained-package runbook from the 2026-08-12 drill §4; then close | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC | | **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | | **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** **2026-10-07 07:58: kept open for later (`09` §3 decision 166).** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC | diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index 90476542..644dd43b 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,9 +1,13 @@ -## Unreleased (2026-10-08) — an alarm when a box never backs up off-site because its escrow is pending (R-243; `09` §3 decision 179); a deleted customer's audit rows go after 1 year (R-901; decision 181) — ships with tomorrow's hub release +## Unreleased (2026-10-08) — an alarm when a box never backs up off-site because its escrow is pending (R-243; `09` §3 decision 179); a deleted customer's audit rows go after 1 year (R-901; decision 181); the operator's older-recovery-package mail (R-304; decision 183) — ships with tomorrow's hub release **Operator action on deploy: none.** Expect ONE `offsite_escrow_pending` mail for **Tester 2** on the first sweep after the deploy: its latest report (2026-10-04) says off-site ON, escrow `pending`, no successful run ever — the state the operator believes it is in (decision 170). +- **R-304 option C (decision 183):** new event type `recovery_older_package` (sent by the controller of the same day when + a household's code opens, or may open, an older sealed escrow package): in `allowedEventTypes` and + `notify.operatorOnlyEvents`, no household text. Test `TestR304_RecoveryOlderPackageIsAllowlistedAndOperatorOnly` + (red-proved: without the allowlist line „the controller's push would 400"). - **R-901 (operator ruling 2026-10-08 09:04, decision 181):** after a customer is DELETED, its `events` and `notification_log` rows are deleted 1 year after the deletion — a new daily step in `pruneAll` (`store.PruneDeletedCustomerAudit`). A deletion is a `customer_resets` journal row with leg `customer_delete` = ok and diff --git a/hub/internal/api/handler.go b/hub/internal/api/handler.go index e6260553..9907dc31 100644 --- a/hub/internal/api/handler.go +++ b/hub/internal/api/handler.go @@ -2169,6 +2169,9 @@ var allowedEventTypes = map[string]bool{ // R-243 — the hub raises these itself (monitor/offsite_escrow_pending.go); allowlisted like the line above. "offsite_escrow_pending": true, "offsite_escrow_pending_cleared": true, + // R-304 option C — controller (same day) sends it on an unlock that opens / may open an older package. Operator-only + // (notify.operatorOnlyEvents); NO customerMessages entry. + "recovery_older_package": true, // controller v0.289.0 (decision 69): the customer-chosen deletion of set-aside history is deferred // to the operator — the box's append-only key cannot delete. Operator-only (notify.operatorOnlyEvents). "offbox_abandon_deferred": true, diff --git a/hub/internal/api/r304_older_package_event_test.go b/hub/internal/api/r304_older_package_event_test.go new file mode 100644 index 00000000..efcf644f --- /dev/null +++ b/hub/internal/api/r304_older_package_event_test.go @@ -0,0 +1,19 @@ +package api + +import ( + "testing" + + "gitea.dooplex.hu/admin/felhom-hub/internal/notify" +) + +// R-304 option C (decision 183): the controller's recovery_older_package must be accepted (or POST /event 400s — R-77) +// and must reach only the operator. RED-PROOF: drop it from allowedEventTypes → "must be in allowedEventTypes". +func TestR304_RecoveryOlderPackageIsAllowlistedAndOperatorOnly(t *testing.T) { + et := "recovery_older_package" + if !allowedEventTypes[et] { + t.Fatalf("%s must be in allowedEventTypes — the controller's push would 400", et) + } + if !notify.IsOperatorOnly(et) { + t.Fatalf("%s must be operator-only — reopening old history is the operator's (R-312)", et) + } +} diff --git a/hub/internal/notify/dispatcher.go b/hub/internal/notify/dispatcher.go index 24f296e4..f998701f 100644 --- a/hub/internal/notify/dispatcher.go +++ b/hub/internal/notify/dispatcher.go @@ -707,6 +707,9 @@ var operatorOnlyEvents = map[string]bool{ // the reminder on every page; this is the operator's „they have not acted" line. Listed in the SAME commit. "offsite_escrow_pending": true, "offsite_escrow_pending_cleared": true, + // R-304 option C (2026-10-08, decision 183): a household's recovery code opens — or may open — an OLDER sealed + // escrow package; reopening old history is operator-only (R-312), so only the operator is told. Same commit. + "recovery_older_package": true, // v0.127.0 (decisions 68–69, R-820/R-822). The off-site key registrar, its daily check and the // clean-up window: custody facts about key lines and fingerprints — the household can take no // action on any of them. Listed in the SAME commit that mints them. diff --git a/hub/internal/notify/r243_operator_only_test.go b/hub/internal/notify/r243_operator_only_test.go index 00a85c82..40d16eea 100644 --- a/hub/internal/notify/r243_operator_only_test.go +++ b/hub/internal/notify/r243_operator_only_test.go @@ -6,7 +6,8 @@ import "testing" // on every page, and a missing customerMessages entry would mail them raw operator English. Red-proof: delete either // line from operatorOnlyEvents and this fails naming it. func TestR243_EscrowPendingIsOperatorOnly(t *testing.T) { - for _, e := range []string{"offsite_escrow_pending", "offsite_escrow_pending_cleared"} { + // R-304 option C's recovery_older_package is pinned here too (same reason: operator-only, no household text). + for _, e := range []string{"offsite_escrow_pending", "offsite_escrow_pending_cleared", "recovery_older_package"} { if !operatorOnlyEvents[e] { t.Errorf("%s is not operator-only", e) }