diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 4e48ed57..5f31c996 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -110,7 +110,7 @@ fixed file, delivered by the signed config bundle): | `FELHOM_DNSMASQ` | the LAN split-horizon resolver | drop-ins via `felhom-priv-apply dnsmasq` (only `bind-interfaces`, `no-resolv`, `listen-address`, `server`, `local`, `address`); `rm` one exact name; exact `pct exec` reads | — | | `FELHOM_GUESTHOOK` | the pre-start self-heal hook | the hook is a FIXED bundle file; the agent only checks it (`SnippetReady`) and registers it; exact vmid/slot | — | | `FELHOM_INTERMEDIARY` | the shared drive parent + live drive binds | boot script + unit are FIXED bundle files (the agent only enables the unit); one-segment drive names (no leading dot, no `..`) | — | -| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen (pinned-registry) image ref and restart the guest's bootstrap — the household's data, not host root | +| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen image ref and restart the guest's bootstrap — **any image from any registry**, because sudo cannot see the `tee` content on stdin and only the `inspect` line is pinned (corrected 2026-10-06 night, `audits/night-burndown-2026-10-06/design-R-861.md`) — the household's data, not host root | | `FELHOM_STALELOCK` / `FELHOM_SCRATCH_TEARDOWN` | stale-lock clear; failed restore-test scratch | exact vmid; the scratch band `99000[0-9]` was already exact | — | | `FELHOM_WG` | the off-site tunnel | conf via `felhom-priv-apply wg` (only the keys `renderConf` writes; no `PostUp`/`PreUp`/`DNS`/`Table`; `/32` only) | — | | `FELHOM_SELFUPDATE` | commit / rollback of the A/B flip | **`apply` removed**: the flip runs only inside `felhom-os-apply agent_update`, after the operator signature, host, window and nonce are checked as root and the staged bytes are hashed ONCE and copied to a root-owned dir (`/var/lib/felhom-os-apply/agent-update/`); the wrapper accepts only that dir | — | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index cb639555..9269b27d 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -243,7 +243,7 @@ It carries guest sizing, `pve_storage[]`, and an app half with per-app `storage_ the identity bundle's shape is `{tunnel_token, pbs_token, wg_private_key, restic_repo_password}` (`felhom-agent/internal/escrow/identity.go:26-39`). -**[DESIGN] Off-site deletion custody (decisions 68–69, 2026-10-03) — BUILT hub v0.127.0 / controller v0.289.1, live on both demo boxes.** The restic repository password stays on the box only. The box's off-site key is append-only (pinned in `authorized_keys`, written by the hub — the key registrar); the box never receives the sub-account password, which the hub stores encrypted at rest. Old snapshots are pruned by the box itself, only in a weekly window the hub opens, behind a fake-snapshot guard (R-822); until that ships nothing prunes. A Felhom-side pruner holding repository passwords is rejected. +**[DESIGN] Off-site deletion custody (decisions 68–69, 2026-10-03) — BUILT hub v0.127.0 / controller v0.289.1, live on both demo boxes.** The restic repository password stays on the box only. The box's off-site key is append-only (pinned in `authorized_keys`, written by the hub — the key registrar); the box never receives the sub-account password, which the hub stores encrypted at rest. Old snapshots are pruned by the box itself, only in a weekly window the hub opens, behind a fake-snapshot guard (R-822); the window has run live on both demo boxes since 2026-10-05 (R-95, closed). The guard's residual — past-dated fakes steering the keeps, and a window check that trusts the box's own counts — is R-822 and R-895. A Felhom-side pruner holding repository passwords is rejected. **[FACT] Three parts of the Recipe are empty or wrong on the live fleet**, and they are exactly the parts a host-loss recovery would read (INV Part D2.3): `hosts.dr_record_json` is `{}` on all three diff --git a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md index abc528d5..4d4df665 100644 --- a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md @@ -17,3 +17,12 @@ no reboot. |---|---|---|---| | R-892 | **§3 stopped** — the shell has no SSH agent; DooPlex's two own keys (`id_rsa`, `id_ed25519`) are refused by `root@192.168.0.154`, direct and through demo-hp. The key the operator copied at 20:20 came from the agent of the operator's own login, so it is not a key this shell holds. Not tried, by the brief: another key, the vaulted password, a disk edit. Next try needs: `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` run **from DooPlex as kisfenyo** (it then asks for the box's root password once). Evidence `s3/s3-ssh-refused.txt` | 5 | — | | R-894 | **fixed on agent main** `74b5eae` (ships as v0.150.0 with the memory-kill check). Saved copy on disk, read only on an unreadable storage; fresh → not due, old → due, none → due/unknown; a storage that answers wins. 4 red-proofs `s4/`. `07` §6.1 updated | 50 | agent `74b5eae` | +| (§1.2 fence) | **The app list per box could not be read in full.** demo-hp and demo-felhom read directly (`docker ps`, read only); the hub's Apps page has no per-box list; reading a copy of the hub database was **refused by the session's permission check** — stopped there, no workaround, the copy deleted at once. **So every catalog change tonight goes to branch `night-held-2026-10-06`, none to `main`** (the fence holds without the list). | 10 | — | +| R-173 | **skipped — waits for an event.** Its only LEFT item is runbook §3 steps 4–5 (copy into a live PVC), which need the hub down — the next planned hub maintenance or a scratch-k3s DR drill. Read only tonight; nothing to do | 2 | — | +| R-861 | **design written** (`design-R-861.md`): pick (a) close before the first paying customer, (b)/(c) accept; `03` §3.1 sentence corrected ("pinned-registry" was false) | 35 | (this batch) | +| R-812 | **design written** (`design-R-812.md`): pick a Proxmox-packages slow lane (no kernel), per-set approval; kernel by hand until R-836 is measured. Read-only on demo-hp/demo-felhom (one temp file the helper made in /tmp on demo-hp was deleted at once) | 30 | (this batch) | +| R-105 | **design written** (`design-R-105.md`): both empty fields have no live writer/reader — pick: retire them and correct `05`/`06` (operator word) | 25 | (this batch) | +| R-32 | **design written** (`design-R-32.md`): pick A — purge through the sub-account's own login before deleting it; byte view waits on one read-only measurement | 40 | (this batch) | +| R-822 | **design written** (`design-R-822.md`): pick — accept the residual and close (operator's word); option B filed as **R-895** (P2) | — | (this batch) | +| R-895 | **OPENED** — the hub's clean-up-window check trusts the box's own snapshot counts (needs a design + a read-only measurement) | 5 | (this batch) | +| (07 §6.5 area) | stale sentence fixed: „until that ships nothing prunes" — the window runs live since 2026-10-05 | 2 | (this batch) | diff --git a/documentation/audits/night-burndown-2026-10-06/design-R-105.md b/documentation/audits/night-burndown-2026-10-06/design-R-105.md new file mode 100644 index 00000000..067cb6ba --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/design-R-105.md @@ -0,0 +1,65 @@ +# R-105 — the two "empty DR records" — design proposal (burn-down night 2026-10-06, no code) + +Baselines read: felhom.eu `8e2dc204` (hub v0.140.0 live), felhom-agent `74b5eae`. Architecture: `05-hub-architecture.md` +§9 and §11 (the "slim DR record"), `06-offsite-connectivity.md` §3.5 (the escrow directive). + +## 1. The problem, and what was measured +The row says two hub-held records are `{}` on every box: `hosts.dr_record_json` and `host_escrow.directive_json` +(the third, `dr_recipe…drives`, was fixed 2026-07-28 and is not re-measured tonight). + +**Tonight's reading is in SOURCE, not in the database** — reading a copy of the hub database was refused by the +session's permission check, so no value was read. Source proves more than a reading would: for these two fields the +empty value is the only value the shipped code can produce. The hub's host pages (read through the operator UI) show +"DR Recipe present / Key Escrow present" on demo-hp, demo-felhom and Tester 1, and "none / none" on Tester 2. + +## 2. What the code does today (read in source) +- **`hosts.dr_record_json` has no writer and no reader.** Created by hub v0.7.0 (`7c0c7545`) with default `'{}'` + (`hub/internal/store/store.go:360`); scanned into `Host.DRRecordJSON` (`store.go:2795`, `:2867`) and read by + nothing else in the hub, the agent or the controller (repo-wide grep). The `05` §9 "slim DR record" was never built. +- **`host_escrow.directive_json` is written only by a by-hand flag.** The agent sends a directive only when the + operator runs `--selftest=escrow-create -directive ` (`felhom-agent/cmd/felhom-agent/main.go:201`, + `:2767-2771`, `:3132-3135`). The production ceremony (the customer's escrow wizard) runs ONE fixed argv with no + `-directive` (`felhom-agent/configs/felhom-agent.sudoers:284`). When that upload carries an identity blob, the hub + stores the missing directive as `'{}'` (`hub/internal/api/handler.go:1293-1305`, `store.go:3612`) — so every + wizard escrow writes `{}`, and also overwrites the one directive made by hand on 2026-07-04 (`06` §3.5). +- **Nothing reads the directive's fields.** The hub serves it only from `/re-enroll` and `/restore-directive` + (`hub/internal/api/dr.go:101`, `:155`); no agent or controller code calls either route (grep). +- **The DR path that IS built does not need them.** The agent's restore plan reads the desired-state + `restore_directive` plus the DR recipe (`felhom-agent/internal/dr/plan.go:45`, `:104`); the recipe carries the PBS + repo id and namespace (`internal/hub/dr_recipe.go` `DRPBSCoord`), the hub holds the endpoint's PBS fingerprint + (`hub/internal/tenantsync/client.go:136`), and the wrapped key is `host_escrow.blob`. + +So the row's two thirds are not a fault in a running path. They are a **design that was half-built and then +bypassed**, and two architecture documents still describe it as if it existed. + +## 3. Options +**A. Retire both, and correct the documents.** Drop the `DRRecordJSON` scan field; stop storing the directive +(the column stays, read as `{}`); `05` §9/§11 and `06` §3.5 say where each fact really lives (recipe, tenantsync, +escrow blob). Cost: ~45 min, hub only, no box change. Can go wrong: if a future re-enroll client wants the +directive, it must be rebuilt — the routes stay and serve `{}`. + +**B. Populate the directive from the ceremony.** The agent fills `{pbs repo id, namespace, endpoint fingerprint}` +in the wizard ceremony. Cost: agent + sudoers argv change (the argv is pinned byte-for-byte) + a bundle delivery; +~2 h and a release. Can go wrong: a second copy of facts the recipe already carries, which can disagree with it — +the R-106 namespace defect was exactly such a disagreement. + +**C. Do nothing more; close R-105 with this evidence.** Cost: nothing. The two documents keep describing a record +that does not exist, which is how this row was filed in the first place. + +## 4. The pick — PROPOSAL for the operator, not a decision +**Option A.** One source per fact; the recipe is reported every cycle and was fixed to match the backup (R-106). +It removes a design promise in `05` §9, so it needs the operator's word (a design decision is not a defect). +Until then the row can move to WAITING-ON-OPERATOR: there is no data at risk — the empty fields have no reader. + +## 5. First slice and its proof +- Red test first: a hub test that the escrow PUT with an identity blob and no directive leaves `directive_json` + unchanged from a hand-made value (today it overwrites with `{}` — fails); then decide by the option. +- A test pinning "no reader": an AST/grep test in the hub that `DRRecordJSON` is not referenced outside the store + (so a new reader cannot appear without this design being revisited). +- Live proof: none needed on a box (hub-only); after deploy, the host page still shows "DR Recipe present". + +## 6. Open questions for the operator +1. Retire the "slim DR record" of `05` §9 (Option A)? If you do nothing: the fields stay empty and unread, the + documents stay wrong, nothing breaks. +2. Tester 2's host page shows no DR recipe and no key escrow. Expected for a box that never did the escrow wizard — + is that the case? If you do nothing: a Tester 2 host loss has no hub-held recovery material. diff --git a/documentation/audits/night-burndown-2026-10-06/design-R-32.md b/documentation/audits/night-burndown-2026-10-06/design-R-32.md new file mode 100644 index 00000000..18bc9be4 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/design-R-32.md @@ -0,0 +1,85 @@ +# R-32 — RESET leaves the household's off-site ciphertext behind — design proposal (burn-down night 2026-10-06, no code) + +Baselines read: felhom.eu `8e2dc204` (hub v0.140.0), felhom-controller `5e7522e023` (v0.301.0). Architecture: +`07-backup-architecture.md` §246 (off-site deletion custody, decisions 68–69); `06-offsite-connectivity.md`; the +row's own three-part ruling from the 2026-07-21 rehearsal. + +## 1. The problem +RESET is meant to destroy a household's off-site copy. On the shared pool box it deletes the Hetzner SUB-ACCOUNT. A +sub-account is a login, not the data: its home directory stays. Re-enabling off-site for the same customer creates a +sub-account with the SAME home (`felhom-`) over the old ciphertext, whose key that RESET destroyed. Measured +on the pool box the night of 2026-07-21: **49 MB attributed** (2 snapshots, 48.7 MiB) against **1.4 GB + 3.0 MB +unattributed** in two `.orphaned-*` folders (`restic-and-pool.txt`, R-32). The orphan card then appeared — the +rehearsal's S7 had said in advance that it would be a finding. + +## 2. What the code does today (read in source) +- RESET's off-site leg calls `s.offsite.Deprovision` (`hub/internal/web/customer_reset.go:217-225`) and journals + `hetzner: ok`, logging "repo data destroyed" (`:225`). +- `Deprovision`, shared tier: `DeleteSubaccount` per labelled sub-account, nothing else + (`hub/internal/offsite/offsite.go:303-326`). Its doc comment says "The offsite repo DATA dies with the + sub-account/box" (`:276-280`) — **true for the dedicated tier (the box is deleted, `:283-300`), false for the shared + one.** A comment that asserts an invariant the code does not provide. +- The home directory is fixed per customer: `HomeDirectory: "felhom-" + customerID` (`offsite.go:349`). So a new + lifecycle lands on the old folder. +- The hub already deletes off-site data in ONE place, through the sub-account's own password login (port 23): + `Registrar.DeleteSetAside` (`hub/internal/offsitekeys/offsitekeys.go:424-445`) — `rm -rf` of a `.orphaned-…` + folder only, refusing anything else (`IsSetAsidePath`, `:448-459`), after the household's 7-day abandonment delay + (decision 74, `service.go` „Decision 74"). +- The move-aside for a reinstall WITHOUT RESET: the box asks the hub to rename the old repo to `.orphaned-` + (`felhom-controller/controller/internal/backup/offbox.go:329-345`). The ruling keeps this — custody survives there. +- The operator's Restic tab shows the pool box's totals from the provider API and, per customer, the usage the BOX + reports (`hub/internal/web/offsite_box.go:121-185`). Bytes that no box reports (an `.orphaned-*` folder, an old + lifecycle) are visible only in the pool total, not per customer. + +## 3. Options +**A. Purge through the sub-account, then delete it.** In `Deprovision` (shared), before `DeleteSubaccount`: log in with +the hub's stored password for that sub-account (it already does this for `authorized_keys` and `DeleteSetAside`), +remove the live repo and every `.orphaned-*`, check the home holds no repository left, then delete the +sub-account. A purge that fails stops the leg (`hetzner: failed`, re-run resumes) — the sub-account is NOT deleted, +because after that only the main account can reach the folder. +- Costs: hub only, small; one new registrar method (`PurgeRepos`) beside `DeleteSetAside`, same refusals style. +- No new credential: the main-account password the ruling named is not needed. +- Can go wrong: the stored password no longer works (rotated, or the hub DB restored from an older copy) → the leg + fails loudly and the operator must decide (main-account clean-up by hand). Deletes household data — that is RESET's + purpose, and the existing RESET ack covers it per the ruling; still the operator's word on the route. + +**B. Purge with the pool box's MAIN account (the ruling's words).** The hub gets the main-account SFTP password. +- Costs: a new credential that reaches every household's folder on the pool box. A hub compromise then deletes every + household's history in one step. Bigger blast radius than A for the same result. +- Only advantage: also reaches folders of sub-accounts deleted BEFORE this fix (old lifecycles). + +**C. A new home folder per lifecycle (`felhom--`), no purge.** The new sub-account never sees old ciphertext, +so the orphan card cannot lie. +- Costs: small. But nothing is ever deleted: the ruling's part (1) is not met, and dead ciphertext fills the pool box + for ever, unseen unless part (3) is built. + +**Part (3), the byte view, for every option:** per customer, the bytes in its folder (all repos, set-aside copies +included) beside the bytes its box attributes. Needs a per-folder size from the provider. **Not measured:** whether the +password login on port 23 answers `du -s` (the shell has `rm`, `mv`, `dd` — memory `storagebox-subaccount-shell…`). So +not buildable tonight: it rests on a mechanism nobody has measured, and tonight's fence forbids touching the Storage Box. + +## 4. The pick — PROPOSAL for the operator, not a decision +**A, with part (3) after one read-only measurement.** It meets the ruling's part (1) with no new credential, keeps the +move-aside guard for reinstall-without-RESET (part 2 — untouched), and makes the journal's "repo data destroyed" true. +Old lifecycles' folders (if any are left on the pool box) are a one-time clean-up by hand — the operator's (question 2). + +## 5. First slice and its proof +- Build (hub): `offsitekeys.Registrar.PurgeRepos(ctx, t, pw)` — removes `t.RepoPath` and every `IsSetAsidePath` match, + nothing else; then lists and refuses success if any remains. `offsite.Provisioner.Deprovision` (shared) calls it + through a seam BEFORE `DeleteSubaccount`. Fix the doc comment at `offsite.go:276-280`. +- Red test first (must FAIL today): a fake provider API and a fake shell; RESET's `Deprovision` for a shared customer. + Assert the shell saw `rm -rf ` before the API saw `DeleteSubaccount`. Today it fails: no shell call at all. +- Second red test: the purge fails → `Deprovision` returns an error and `DeleteSubaccount` was NOT called (else the + folder becomes unreachable). +- Keep green: `DeleteSetAside`'s refusals; the dedicated tier unchanged; RESET's journal re-run. +- Live proof on a SCRATCH customer only (never a household): enable off-site, one run from scratch box 9202, RESET, + re-enable. Positive observable: the new sub-account's home lists no `restic`/`.orphaned-*` folder, and no orphan card + appears. Control from a different channel: the provider's pool-box `stats` size before and after (coarse, API). + Evidence off the machine before teardown. + +## 6. Open questions for the operator +1. RESET deletes the household's off-site folder through the sub-account's own login before removing it (option A), + instead of through the pool box's main account — agree? If you do nothing: RESET keeps leaving the ciphertext; a + re-enabled customer sees an orphan card again. +2. Folders left by RESETs done before this fix (the 2026-07-21 measurement found 1.4 GB): clean them up once by hand + with the main account, or leave them? If you do nothing: they stay and use pool-box space; nothing reads them. diff --git a/documentation/audits/night-burndown-2026-10-06/design-R-812.md b/documentation/audits/night-burndown-2026-10-06/design-R-812.md new file mode 100644 index 00000000..647958bc --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/design-R-812.md @@ -0,0 +1,108 @@ +# R-812 — what is left of OS updates: the Proxmox packages and the kernel — design proposal (burn-down night 2026-10-06, no code) + +Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `11-os-updates.md` (§1, §3, §5.2, §5.5, §5.6, +§5.9, §7.1, §8). Related rows: R-836 (the kernel's fallback), R-808 (the intention, `ROADMAP.md`). + +## 1. The problem +`11` §8 steps 1–5 are BUILT: the guest's Debian, the host's Debian, the fleet view, the Docker engine, the root files +(R-840, closed). Step 6 is not: **the host's Proxmox packages and its kernel are never updated on any box.** The host +fast lane takes only Debian-origin packages, so everything from the Proxmox repository waits for ever — including +packages with ordinary names (`zfsutils-linux`, `shim-signed`, `corosync`, `ceph-common`, `amd64-microcode`, `11` C3). + +Measured tonight (read only; `apt list --upgradable`, `pveversion`, `dpkg -l`, `ls /boot`; no `apt update` — the package +lists are from 2026-10-05, the last daily refresh): +- **demo-hp:** `pve-manager` 9.2.2 → 9.2.21 pending; **77 packages pending, all from the Proxmox repository** (the Debian + lane has taken the rest): `qemu-server` 9.1.15 → 9.2.10, `pve-container` 6.1.10 → 6.1.14, `proxmox-backup-client` + 4.2.0 → 4.2.7, `zfsutils-linux` 2.4.2 → 2.4.4, `shim-signed` 1.48 → 1.51, `pve-firmware`, `corosync`, `ceph-*`, `frr`, + `amd64-microcode`, `libpve-*`. Kernel: runs 7.0.14-20 (installed BY HAND in the 2026-10-04 spike), 7.0.2-6 kept. +- **demo-felhom:** `pve-manager` 9.2.2; **78 pending**; runs kernel **7.0.2-6, the install-time kernel** — 7.0.14-20 is in + the repository and not installed. +- So a box keeps its install-time Proxmox and kernel. Proxmox security fixes (for example in `pve-manager`'s web API, in + `qemu-server`, in the Secure Boot loader `shim-signed`) never arrive. The spike counted 79 Proxmox packages "not + covered" on the host (`11` §7.1). + +## 2. What the code does today (read in source) +- The fast-lane origin rule: `felhom-agent/configs/felhom-os-apply:63` `FAST_ORIGINS = ("Debian", "Debian-Security")`; + a plan package of any other origin is refused **R2** (`:491-492`, and in the simulation `:1164` `origin_ok`). +- The host's name rule on top: `:91-92` `HOST_SLOW_RE` (`linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`, + `firmware-*`, `grub*`, `shim*`, `systemd-boot*`, `*-microcode`, `efibootmgr`) → refusal **R14** (`:493-494`, `:1234-1235`); + the pending selection skips them (`:1136`). +- Layers and lanes: `:437-442` — `guest`/`host` are lane `fast` only; `docker` is the only `slow` layer. There is no + host slow layer. +- Tests pinning this: `configs/test_felhom_os_apply.py:417` `test_R2_non_debian_origin_in_the_plan`, `:422` + `test_R2_non_debian_origin_in_the_simulation`, `:584`/`:590` `test_R14_kernel_package_*`, `:639` + `test_pending_fast_skips_proxmox_docker_and_kernel` (a pending `pve-manager` 9.2.21 from „Proxmox Debian Repository" + is NOT taken). +- What a Proxmox step restarts (spike `partH/H5-proxmox-simulate.txt`, read from the maintainer scripts, not run): + `pve-manager` → pvedaemon, pveproxy, pvescheduler, pvestatd, spiceproxy; `pve-cluster` → pmxcfs (`/etc/pve` gone for + seconds — the agent's API calls fail meanwhile); `qemu-server`, `pve-firewall`, `pve-ha-*`, `corosync`, `chrony`, + `zfs-zed`. **No guest is stopped by these scripts** (read, not measured). One new package (`proxmox-firewall-data`). +- The kernel (`11` §5.6, R-836, measured on demo-hp with the operator's word): installing a kernel makes it the GRUB + default at once; `--next-boot` and `grub-reboot` are NOT one-shots here (`/boot` is ext4 on LVM; GRUB cannot write its + env block); a kernel that hangs before userspace stays the default on every power cycle. `sp5100_tco` exists on + demo-hp, blacklisted, never armed. demo-felhom (Intel N100) has no watchdog measured. + +## 3. Options +**A. A host slow lane for the Proxmox USERSPACE packages, no kernel, no reboot.** +- What: a new layer `pve` (lane `slow`) in the wrapper: origin „Proxmox Debian Repository" only, still never a name in + `HOST_SLOW_RE` (kernel, boot, firmware, microcode stay out), no removals, no new package except a named allow-list (the + `proxmox-firewall-data` kind). Approval like the Docker engine set (`11` §5.8): the operator approves one SET per box + generation on the System page after ring 0 ran it; ring 1 only by signed job. Runs in the night leg after the host + Debian step, under the same heavy-op gate. Health = the host rule (`HostHealthVerdict`) + `pveversion` reports the + new version. +- Cost: wrapper layer + refusals + tests; hub candidate/approval/System page (copies the Docker set's code); ~1.5 days. +- Can go wrong: pmxcfs restart while the agent writes `/etc/pve` (the gate already excludes backups/restore-tests; the + agent's own reconcile must wait too); a Proxmox point release that needs a newer kernel (`proxmox-ve` depends on + `proxmox-default-kernel` — the plan must not pull the kernel in: the simulation refusal R14 already catches it); + undo is only „install the previous version" (Proxmox keeps 30–66 versions, `11` C2 — good). +- Brings the 77–78 pending packages, minus the boot-chain ones, to every box. Leaves `shim-signed`, `pve-firmware`, + microcode and the kernel out. + +**B. The kernel lane, with a fallback that works on these hosts — measured first.** +- What: install the kernel (plus `shim-signed`, `pve-firmware`, microcode, which also only act at boot), keep the old + one as the permanent default, boot the new one ONCE, and make it the default only from userspace after a healthy boot. + The one-shot needs one of: (1) **UEFI `BootNext`** with a second boot entry whose GRUB config defaults to the new kernel + — the firmware clears BootNext itself on use, so a hang falls back on the next power cycle; (2) a GRUB env block on + the ESP (vfat, writable by GRUB); (3) arming a hardware watchdog early (`sp5100_tco` on demo-hp) so a hang power-cycles + into the old default. None is measured. +- Cost: a spike with ≥4 reboots per box on both demo hosts (Secure Boot ON on demo-hp, OFF on demo-felhom; AMD vs Intel), + then the lane: ~3–5 days. Every customer box restarts all its apps once per kernel (≈1–2 min at night). +- Can go wrong: firmware that ignores `BootNext` (seen on consumer boards); Secure Boot refusing a second entry; a + hang with no watchdog still needs a person — it is a dead box at a household. Telling households that the box may + restart at night is a **promise to users** (`11` §5.7) — the operator's. + +**C. The kernel only by an operator-present maintenance step.** +- What: no automatic kernel lane. The System page shows „kernel behind / reboot needed"; the operator runs the step per + box by a signed job at a time they choose, ready to power-cycle (or ask the household to). +- Cost: small (a signed job and a runbook), but a person per box per kernel; does not scale past a handful of boxes. +- Can go wrong: kernels are never installed because nobody schedules them — the state today with a button. + +## 4. The pick — PROPOSAL for the operator, not a decision +**A now, C for the kernel until B is measured.** A closes the biggest gap (78 Proxmox packages, the web API and +qemu/lxc tooling) with no reboot and the same approval shape the operator already uses for Docker. The kernel stays a +separate, operator-scheduled act (C) until R-836's fallback is measured on both demo hosts; then B replaces C. + +**R-812 should split.** Close R-812 (P2) when A ships to ring 0 and ring 1, with R-836 (kernel fallback, P3) carrying +the kernel. The kernel's urgency then becomes its own question: a P2 row „kernel security fixes reach a box" only if the +operator wants it before the first paying customer (question 2). + +## 5. First slice and its proof +- Build: wrapper layer `pve` lane `slow` — origin „Proxmox Debian Repository" only; R14 still applies; R2 for any other + origin; refusal for a removal or an unlisted new package; appliance only (R12's root-owned install record). +- Red test first (must FAIL on today's code): a host plan with `pve-manager 9.2.21` origin „Proxmox Debian Repository", + layer `pve`, lane `slow` → today refused R12 („layer is not guest, host or docker"); after: installed. Companion + tests that must stay refused: the same plan carrying `proxmox-kernel-7.0` (R14) or `shim-signed` (R14); a Debian + package in a `pve` plan (R2); `test_pending_fast_skips_proxmox_docker_and_kernel` unchanged. +- Live proof on **demo-felhom** (ring 0, Tier 0, its Proxmox is the install-time one): one approved set of the pending + Proxmox packages, no kernel. Positive observable: `pveversion` reads 9.2.21; the host health rule passes; the guest + kept running (container `StartedAt` unchanged). Control from a different channel: the hub's System page row for the + box, and an app on the guest answering 200 every 5 s through the step. No reboot needed, so no operator word for a + reboot; the operator's word is the approval click, as for Docker. Evidence off the box at the end. + +## 6. Open questions for the operator +1. **May Proxmox's own packages (no kernel) update by night, after you approve each set on the System page, as the + Docker engine does?** Pick: yes (option A). If you do nothing: every box keeps its install-time Proxmox; 78 packages + behind on the demo boxes today, and growing. +2. **Must kernel security fixes reach customer boxes before the first paying customer?** If yes, B is a ~1-week arc with + reboots on both demo boxes (your word before each). Pick: no — the kernel by your hand (C) until B is measured. If you + do nothing: kernels stay at the install-time version; a kernel hole stays open until you run the step by hand. diff --git a/documentation/audits/night-burndown-2026-10-06/design-R-822.md b/documentation/audits/night-burndown-2026-10-06/design-R-822.md new file mode 100644 index 00000000..9ea3fe25 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/design-R-822.md @@ -0,0 +1,68 @@ +# R-822 — fake snapshots steering the off-site retention — design proposal (burn-down night 2026-10-06, no code) + +Baselines read: felhom.eu `8e2dc204`, felhom-controller `5e7522e023` (v0.301.0), hub v0.140.0. Architecture: +`07-backup-architecture.md` §246 (off-site deletion custody, decisions 68–69) and threat row 10 (§1298); +`audits/offsite-append-only-2026-10-03/DESIGN.md` §3 Option 1. + +## 1. The problem +The box's off-site key can only ADD (decision 69). Someone holding that key can add snapshots with any date and make the +box's own honest pruner delete the real ones. Measured in the lab 2026-10-03 (restic 0.14.0, rclone `--append-only`): +13 empty future-dated snapshots made the policy `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` select all 3 real +snapshots (`audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt`). The guard (controller v0.289.0, +tuned v0.290.0 and v0.294.0) refuses that shape. **What is left (read in source tonight, not measured):** a fake dated +in the PAST, inside a week or month that already holds a real keep, and newer than that real snapshot, wins the bucket. +The real snapshot then falls out of the plan as an ordinary old removal. The guard cannot tell it from honest retention. + +## 2. What the code does today (read in source) +- The guard runs on the box before any `forget`, inside a hub window (`felhom-controller/controller/internal/backup/offbox_window.go:21-38`, `offsiteGuard` `:131-158`). It refuses: a future date (`:133-135`); a date after the window opened (`:136-138`); a plan that removes a snapshot younger than 7 calendar days and not superseded the same day (`:141-149`); a plan above `MaxRemove` (`:152-154`). +- A YOUNG real snapshot superseded the same day by a newer one is only EXCLUDED, then removed in a later window once old (`:142-145`, R-824). So a same-day fake, planted after the night run, also wins the daily bucket — later. +- `MaxRemove` = half of the count, at least 5 (`felhom.eu/hub/internal/offsitekeys/service.go:240-246`). So the old real history goes in about two windows. +- The hub's after-window check compares counts the BOX reports: `countBefore` comes with the box's request (`service.go:284`), `CountAfter` with its close report (`service.go:343`). Fakes added between windows keep the count level, so the drop alarm (`EventWindowDrop`, `:350`) does not fire. +- **The same exposure already exists without any fake:** during a window the box's own key may delete (`service.go:303`, the deleting line). A box that is broken into at window time can delete everything and report any counts. Decision 68 accepted that, with the count check as the backstop. + +**Result, inferred from the above:** an attacker who held the box's key once can, over about two weekly windows, shrink +the real off-site history to the last 7 days. Nobody is alarmed. It needs the attacker to plant fakes; it does not need +the attacker to still be there. + +## 3. Options +**A. Accept the residual and close the row.** Record it in `07` §246 as a stated limit of decision 68. +- Costs: nothing to build. +- Can go wrong: the slow, silent loss above. It is the same class of loss decision 68 already accepts for a broken-into + box in a window (box-reported counts), so A adds no NEW weakness — it names one. + +**B. The hub keeps its own snapshot list (no repository password needed).** The hub already logs in to each sub-account +daily (the `authorized_keys` audit). It also lists `/snapshots/`: file names (= snapshot ids) and their upload +times on the provider. The hub then: (1) counts before and after each window ITSELF, not from the box; (2) alarms when +snapshot files appear that no box run explains (a run reports its time; a file uploaded outside a run, or more files +than runs, is the injection signature); (3) refuses to open a window while such files exist. +- Costs: hub only, about one day. One more SFTP listing per customer per day. Custody unchanged (07 §8a): the hub sees + file names and sizes, never contents. +- Can go wrong: a manual run, a catch-up run or a crash retry must be counted as "a run", or false alarms. Not measured: + whether the provider's restricted shell gives upload times (`ls -l` on port 23) — measure on the scratch sub-account. +- It also closes the box-trusted count in decision 68's backstop — worth having on its own. + +**C. A stricter box guard (refuse when a keep is empty or tiny).** Rejected: a fake can point at a real snapshot's tree +with another date, at no cost. The guard runs on the box that the attacker held. + +## 4. The pick — PROPOSAL for the operator, not a decision +**A now, B as the next build.** Close R-822 as a stated, bounded limit of decision 68 (the loss is the last-7-days +floor, at ≤ half the snapshots per weekly window), and file B as its own row: "the hub trusts the box's own count before +and after a window". B is the real defence for both the fake-snapshot case and the broken-into-box-in-a-window case. It +needs a measurement first, so it is not tonight's work. + +## 5. First slice and its proof (for B) +- Measure first, on the scratch sub-account: `ls -l /snapshots/` through the password login on port 23 — does it + list names and upload times? (Read only.) +- Build: hub `offsitekeys` gets `ListSnapshots(ctx, target, pw) ([]SnapshotFile, error)`; `OpenWindowFor` and + `CloseWindowFor` take the counts from it, not from the box (the box's numbers are logged beside). +- Red test first (must FAIL today): a fake registrar whose listing holds 17 files before and 3 after; the box reports + 17 → 15. Assert `EventWindowDrop` fires. Today it does not (the hub uses the box's 15). +- Live proof on scratch 9202's sub-account: open a one-shot window; the hub's own before/after counts appear in its + log and match `restic snapshots` run on the box. Control from a different channel: the provider's listing read by + hand over SFTP. + +## 6. Open questions for the operator +1. Accept the residual (an attacker who once held a box's key can shrink its off-site history to 7 days over about two + weeks, silently) and close R-822? If you do nothing: the row stays open; nothing changes on the boxes. +2. Build B (the hub counts the snapshots itself, about a day of work, after one read-only measurement)? If you do + nothing: the window's alarm keeps trusting the numbers the box sends. diff --git a/documentation/audits/night-burndown-2026-10-06/design-R-861.md b/documentation/audits/night-burndown-2026-10-06/design-R-861.md new file mode 100644 index 00000000..332d8ffb --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/design-R-861.md @@ -0,0 +1,94 @@ +# R-861 — the three sudoers leftovers — design proposal (burn-down night 2026-10-06, no code) + +Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `03-host-agent.md` §3.1 (the group table) +and §11, `11-os-updates.md` §5.4.2 (the config bundle), `09` §3 decision 122. Read only; nothing ran on a box. + +## 1. The problem +After agent v0.146.1 a compromised agent PROCESS (user `felhom-agent`) no longer becomes host root: measured 2026-10-05, +29 attack lines refused, `sudo -l` 93/93 on both demo boxes (`audits/hub-safety-2026-10-05/partF/`). Three paths were +left open on purpose and named in `03` §3.1. The row asks the operator: accept them, or close them before the first +paying customer? The question for each is: **what does an attacker who already owns the agent process gain through it, +beyond what the agent already has?** + +What the agent already has (the baseline): its Proxmox token holds `VM.Backup`, `VM.Allocate`, `VM.Config.*`, +`VM.PowerMgmt`, `Pool.Allocate` on `/pool/felhom` (`felhom.eu/scripts/felhom-host-install.sh:309`). So it can already +back up a customer guest and restore it into a scratch guest (the restore-test does exactly that), stop and start felhom +guests, and it writes the guest's `bootstrap.json` itself. **The household's data is already in its reach, offline.** + +## 2. What the code does today (read in source) + +**(a) The controller image ref.** `FELHOM_CONTROLLERSWAP` allows `pct exec -- tee /etc/felhom-controller-image` +(`felhom-agent/configs/felhom-agent.sudoers:122`). The ref goes in on STDIN (`internal/localapi/controllerswap.go:150`). +The regex `controllerImageRe` (`controllerswap.go:36`) is the agent's OWN check — a compromised agent skips it, and sudo +cannot see stdin. The guest's bootstrap unit then runs `docker run … "$IMAGE"` with the docker socket, the read-only +bootstrap dir and `/mnt` (`configs/build-golden.sh:316,355-362`). +**Correction to `03` §3.1:** the table says *"a chosen (pinned-registry) image ref"*. That is not true: only the +`docker image inspect` line is pinned (`sudoers:119`); the `tee` content is free, so the bootstrap pulls and runs **any +image from any registry**. Gain over the baseline: a LIVE foothold in the guest with the docker socket (guest root, the +household's running apps and its LAN), and it survives agent restarts until the next swap. Not host root. + +**(b) The felhom-op SSH key.** `felhom-priv-apply sshd-key` installs one plain key line (no `command=`/`from=`; +`configs/felhom-priv-apply:60-61`) from a file the agent staged; the key itself comes from the hub, unsigned. A +compromised agent can therefore put its own key on `felhom-op`. `felhom-op`'s sudo (`configs/felhom-op.sudoers`) is +scoped: restart wg/agent/sshd, `pct list`, and `pct start|stop|unlock [0-9]*` — **on any guest of the host**, not only +the felhom pool. Gain over the baseline: power control of NON-felhom guests on a BYO host (the household's own other +VMs), and an interactive login. Not host root. (Side note: these `pct` lines still use the `*` glob, which matches +spaces — the R-861 shape 1; no harmful `pct start/stop` option is known, so this is hygiene, not a hole.) + +**(c) The escrow ceremony.** The root child (`FELHOM_ESCROW`, `sudoers:283-284`) returns the recovery code R on the +agent's stdout pipe (`internal/localapi/escrow_ceremony.go:19-29,93`); the agent holds R in memory for one claim +(`:140-147`). The agent's own hub key may read this box's escrow blob (`hub/internal/api/handler.go:252,1357-1372`, +self-scoped). So a compromised agent can learn R, fetch the blob and unwrap this box's PBS encryption key (and the +identity bundle). Gain over the baseline: **off-box** decryption of this box's off-site archives, which lasts after the +compromise is cleaned up, until the key is rotated. One box only; not root. + +## 3. Options + +**(a)** +- **A1. A checking wrapper verb.** `felhom-priv-apply controller-image `: as root, read stdin, require + `^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$`, then write the guest file. The sudoers `tee` line + is removed. Same mechanism as decision 122 (b); delivered by the signed config bundle. Cost: ~1–2 h (one verb, its + Python tests, the Go call path, `TestManifestCoveredBySudoers`, the capability probe). What can go wrong: the old + agent binary still calls `tee` → the bundle must follow the agent update (the usual order). Leaves: a chosen OLD + controller version from our registry (a downgrade) — still possible. +- **A2. Check in the guest.** The bootstrap script refuses a ref outside the pattern. Cost: a golden change, and + existing guests keep the old script until re-baked — slow to reach the fleet. +- **A3. Accept.** Write the corrected sentence in `03` §3.1. + +**(b)** +- **B1. Sign the key.** The operator signs the felhom-op key with the operator key; `felhom-os-apply` verifies as root. + Cost: medium; every key rotation needs the operator's offline signature. +- **B2. Narrow felhom-op's sudo** to anchored regexes (`^start [0-9]+$` …) — hygiene only; it does not stop the key swap. + Cost: ~30 min, rides the bundle. +- **B3. Accept** (felhom-op is not root; the gain is power control of guests). + +**(c)** +- **C1. R bypasses the agent.** The root child writes R straight into the guest (to the controller), so the agent never + sees it. Cost: a ceremony redesign across agent and controller; touches the escrow promise to the household. +- **C2. Accept** — the ceremony is designed so the box handles K once; the residual is "a compromised agent can read + this one box's backups off-site", which `03` §3.1 already names. + +## 4. The pick — PROPOSAL for the operator, not a decision +- **(a) A1, before the first paying customer.** It is the only one of the three that gives a live, persistent foothold + next to the household's running apps, and the fix uses a mechanism already built and measured. Fix the `03` sentence + in the same commit. +- **(b) B3 now, B2 as hygiene** with the next bundle; B1 later if the OOB door is ever opened wider. +- **(c) C2.** C1 changes an escrow promise and is a redesign — not before the first customer. + +## 5. First slice and its proof (for A1) +- Build: verb `controller-image` in `configs/felhom-priv-apply` (stdin ≤ 256 bytes, the regex, a numeric vmid, then + `pct exec -- tee` as root); sudoers: remove `tee /etc/felhom-controller-image$`, add + `/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$`; `controllerswap.go:150` calls the verb. +- Red test first: `configs/test_felhom_priv_apply.py` — `controller-image 9201` with stdin + `docker.io/library/alpine:latest` must exit 3; with `gitea.dooplex.hu/admin/felhom-controller:0.301.0` exit 0. And + `TestSudoersRefusesTheR861Injections` gains `pct exec 9201 -- tee /etc/felhom-controller-image` as a REFUSED line — + it fails on today's sudoers. +- Live proof on scratch 9202 after a bundle there (not tonight): `sudo -l -U felhom-agent` lists no `tee`; a managed + controller update still swaps (positive observable: the new controller version in `docker ps` AND the hub's host + report — two channels); a hand-fed `alpine` ref is refused (journal tag `felhom-priv-apply`). + +## 6. Open questions for the operator +1. **(a)** Close the free image ref before the first paying customer (A1, ~1–2 h, ships with a bundle)? If you do + nothing: a compromised agent can run any container next to the household's apps, with the docker socket. +2. **(b)+(c)** Accept both for the first customers (felhom-op is not root; the escrow residual is one box's off-site + backups)? If you do nothing: they stay open, named in `03` §3.1, as today. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index c2b40cde..839da317 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -146,19 +146,20 @@ stopping line that lies. | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | -## Backup & restore — 33 rows (P2 7, P3 12, P4 14) +## Backup & restore — 34 rows (P2 8, P3 12, P4 14) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | -| **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC | +| **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-32.md`) — pick A: purge the repo + set-aside copies through the sub-account's own password login BEFORE deleting it (no main-account credential); the „repo data destroyed" comment at `hub/internal/offsite/offsite.go:276-280` is false for the shared tier; part (3) waits on a read-only measurement (`du` on port 23). Operator: agree to route A. | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | +| **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: TRACED in source — not a fault in a running path.** `hosts.dr_record_json` has no writer and no reader; `host_escrow.directive_json` is filled only by the by-hand `-directive` flag (the wizard's fixed argv carries none, so every wizard escrow stores `{}`), and its only routes (`/re-enroll`, `/restore-directive`) have no client. The built DR path reads the recipe, tenantsync and the escrow blob. Design with a pick (retire both and correct `05` §9/§11, `06` §3.5 — needs the operator's word): `audits/night-burndown-2026-10-06/design-R-105.md`. The `drives` third was not re-measured (the hub-DB read was refused). | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC | | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. | — | Read back the 2026-10-07 night (local tier) and the ~2026-10-08 night (both tiers on demo-hp); measure one press (Part D) | CC | | **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | | **R-894** | Backup & restore | P3 | **After an agent restart, an UNREADABLE off-site storage makes the off-site tier look DUE, so the box asks for a copy that cannot be made.** MEASURED 2026-10-05 on demo-hp, read 2026-10-06 (`audits/readback-2026-10-07/C/`): the last off-site copy was 2026-10-01 20:15Z (ep0's own listing, verify `ok`), so the 7-day tier was NOT due; the agent had restarted at 04:57 local; at 06:25 `GET …/storage/felhom-pbs/content` answered 500 *Can't connect to 10.77.0.1:8007*; `newestArchiveOn` returned `unknown` and fell back to the in-memory record (`internal/localapi/server.go`), which a restart empties (`internal/backup/store.go` — memory only, R-348); so the tier read DUE, the controller requested it, and vzdump failed (*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks (`07` §6.4); whether it did that night is NOT KNOWN (the controller's log was lost to a later restart). The hub got `whole_guest_backup_failed` (error); its operator mail then failed (fixed in the hub this session: a failed operator mail is retried). The code's rule is deliberate: an unreadable storage must not suppress a backup. | **OPEN — filed 2026-10-06** **2026-10-06 night: fixed on agent main `74b5eae` (the row's first option — the newest success per tier on disk, read only when the storage cannot be read; `07` §6.1); ships with agent v0.150.0, closes when delivered.** | a design: keep the newest success per tier on disk (as `RestoreTestState` does) so the fallback is the last known copy, not "never"; or report an unreachable storage so no app is stopped for it | Decide the fallback; build it in the agent with a test that restarts the agent and then cannot read the storage; measure once | CC | -| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | +| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-822.md`) — the past-dated residual can shrink real history to the last 7 days in ~2 windows, unalarmed, because the hub's window check trusts the box's own counts (`hub/internal/offsitekeys/service.go:284,343`); the same box-trusted count already sits under decision 68. Pick: accept + close (operator's word); the hub's own snapshot count is the new row R-895. | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | +| **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** **2026-10-06: leg (a) PUSHED** to the live catalog (`c265b37`). **2026-10-06 (burn-down night, later): leg (b) NEEDS A DESIGN.** `07` §7.4 sets no direction; refusing the restore without the DB password, or `ALTER USER` after it, each change restore behaviour on customer data. | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-231** | Backup & restore | P3 | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | — | — | operator | @@ -201,7 +202,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | -| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | +| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | @@ -220,7 +221,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator | +| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-812.md`) — what is left is `11` §8 step 6: the host's Proxmox-origin packages and its kernel are never updated (read tonight: demo-hp 77 and demo-felhom 78 Proxmox packages pending on the 2026-10-05 lists, `pve-manager` 9.2.2 → 9.2.21; demo-felhom still runs its install-time kernel 7.0.2-6). Pick: a `pve` slow lane for the Proxmox packages without the kernel, approved per set like the Docker engine; the kernel stays an operator-run step until R-836's fallback is measured. Proposed split: close R-812 when that lane ships; R-836 carries the kernel. Two operator questions in the design. | — | — | CC + operator | | **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |