Night 2026-10-06: designs R-105, R-812, R-861, R-32, R-822 (rows point at them); R-895 filed; 03 §3.1 and 07 corrected
gates / gates (push) Successful in 2m44s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 20:44:06 +02:00
parent 8e2dc2049f
commit ed8b10330c
9 changed files with 438 additions and 8 deletions
@@ -17,3 +17,12 @@ no reboot.
|---|---|---|---|
| R-892 | **§3 stopped** — the shell has no SSH agent; DooPlex's two own keys (`id_rsa`, `id_ed25519`) are refused by `root@192.168.0.154`, direct and through demo-hp. The key the operator copied at 20:20 came from the agent of the operator's own login, so it is not a key this shell holds. Not tried, by the brief: another key, the vaulted password, a disk edit. Next try needs: `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` run **from DooPlex as kisfenyo** (it then asks for the box's root password once). Evidence `s3/s3-ssh-refused.txt` | 5 | — |
| R-894 | **fixed on agent main** `74b5eae` (ships as v0.150.0 with the memory-kill check). Saved copy on disk, read only on an unreadable storage; fresh → not due, old → due, none → due/unknown; a storage that answers wins. 4 red-proofs `s4/`. `07` §6.1 updated | 50 | agent `74b5eae` |
| (§1.2 fence) | **The app list per box could not be read in full.** demo-hp and demo-felhom read directly (`docker ps`, read only); the hub's Apps page has no per-box list; reading a copy of the hub database was **refused by the session's permission check** — stopped there, no workaround, the copy deleted at once. **So every catalog change tonight goes to branch `night-held-2026-10-06`, none to `main`** (the fence holds without the list). | 10 | — |
| R-173 | **skipped — waits for an event.** Its only LEFT item is runbook §3 steps 4–5 (copy into a live PVC), which need the hub down — the next planned hub maintenance or a scratch-k3s DR drill. Read only tonight; nothing to do | 2 | — |
| R-861 | **design written** (`design-R-861.md`): pick (a) close before the first paying customer, (b)/(c) accept; `03` §3.1 sentence corrected ("pinned-registry" was false) | 35 | (this batch) |
| R-812 | **design written** (`design-R-812.md`): pick a Proxmox-packages slow lane (no kernel), per-set approval; kernel by hand until R-836 is measured. Read-only on demo-hp/demo-felhom (one temp file the helper made in /tmp on demo-hp was deleted at once) | 30 | (this batch) |
| R-105 | **design written** (`design-R-105.md`): both empty fields have no live writer/reader — pick: retire them and correct `05`/`06` (operator word) | 25 | (this batch) |
| R-32 | **design written** (`design-R-32.md`): pick A — purge through the sub-account's own login before deleting it; byte view waits on one read-only measurement | 40 | (this batch) |
| R-822 | **design written** (`design-R-822.md`): pick — accept the residual and close (operator's word); option B filed as **R-895** (P2) | — | (this batch) |
| R-895 | **OPENED** — the hub's clean-up-window check trusts the box's own snapshot counts (needs a design + a read-only measurement) | 5 | (this batch) |
| (07 §6.5 area) | stale sentence fixed: „until that ships nothing prunes" — the window runs live since 2026-10-05 | 2 | (this batch) |
@@ -0,0 +1,65 @@
# R-105 — the two "empty DR records" — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom.eu `8e2dc204` (hub v0.140.0 live), felhom-agent `74b5eae`. Architecture: `05-hub-architecture.md`
§9 and §11 (the "slim DR record"), `06-offsite-connectivity.md` §3.5 (the escrow directive).
## 1. The problem, and what was measured
The row says two hub-held records are `{}` on every box: `hosts.dr_record_json` and `host_escrow.directive_json`
(the third, `dr_recipe…drives`, was fixed 2026-07-28 and is not re-measured tonight).
**Tonight's reading is in SOURCE, not in the database** — reading a copy of the hub database was refused by the
session's permission check, so no value was read. Source proves more than a reading would: for these two fields the
empty value is the only value the shipped code can produce. The hub's host pages (read through the operator UI) show
"DR Recipe present / Key Escrow present" on demo-hp, demo-felhom and Tester 1, and "none / none" on Tester 2.
## 2. What the code does today (read in source)
- **`hosts.dr_record_json` has no writer and no reader.** Created by hub v0.7.0 (`7c0c7545`) with default `'{}'`
(`hub/internal/store/store.go:360`); scanned into `Host.DRRecordJSON` (`store.go:2795`, `:2867`) and read by
nothing else in the hub, the agent or the controller (repo-wide grep). The `05` §9 "slim DR record" was never built.
- **`host_escrow.directive_json` is written only by a by-hand flag.** The agent sends a directive only when the
operator runs `--selftest=escrow-create -directive <file>` (`felhom-agent/cmd/felhom-agent/main.go:201`,
`:2767-2771`, `:3132-3135`). The production ceremony (the customer's escrow wizard) runs ONE fixed argv with no
`-directive` (`felhom-agent/configs/felhom-agent.sudoers:284`). When that upload carries an identity blob, the hub
stores the missing directive as `'{}'` (`hub/internal/api/handler.go:1293-1305`, `store.go:3612`) — so every
wizard escrow writes `{}`, and also overwrites the one directive made by hand on 2026-07-04 (`06` §3.5).
- **Nothing reads the directive's fields.** The hub serves it only from `/re-enroll` and `/restore-directive`
(`hub/internal/api/dr.go:101`, `:155`); no agent or controller code calls either route (grep).
- **The DR path that IS built does not need them.** The agent's restore plan reads the desired-state
`restore_directive` plus the DR recipe (`felhom-agent/internal/dr/plan.go:45`, `:104`); the recipe carries the PBS
repo id and namespace (`internal/hub/dr_recipe.go` `DRPBSCoord`), the hub holds the endpoint's PBS fingerprint
(`hub/internal/tenantsync/client.go:136`), and the wrapped key is `host_escrow.blob`.
So the row's two thirds are not a fault in a running path. They are a **design that was half-built and then
bypassed**, and two architecture documents still describe it as if it existed.
## 3. Options
**A. Retire both, and correct the documents.** Drop the `DRRecordJSON` scan field; stop storing the directive
(the column stays, read as `{}`); `05` §9/§11 and `06` §3.5 say where each fact really lives (recipe, tenantsync,
escrow blob). Cost: ~45 min, hub only, no box change. Can go wrong: if a future re-enroll client wants the
directive, it must be rebuilt — the routes stay and serve `{}`.
**B. Populate the directive from the ceremony.** The agent fills `{pbs repo id, namespace, endpoint fingerprint}`
in the wizard ceremony. Cost: agent + sudoers argv change (the argv is pinned byte-for-byte) + a bundle delivery;
~2 h and a release. Can go wrong: a second copy of facts the recipe already carries, which can disagree with it —
the R-106 namespace defect was exactly such a disagreement.
**C. Do nothing more; close R-105 with this evidence.** Cost: nothing. The two documents keep describing a record
that does not exist, which is how this row was filed in the first place.
## 4. The pick — PROPOSAL for the operator, not a decision
**Option A.** One source per fact; the recipe is reported every cycle and was fixed to match the backup (R-106).
It removes a design promise in `05` §9, so it needs the operator's word (a design decision is not a defect).
Until then the row can move to WAITING-ON-OPERATOR: there is no data at risk — the empty fields have no reader.
## 5. First slice and its proof
- Red test first: a hub test that the escrow PUT with an identity blob and no directive leaves `directive_json`
unchanged from a hand-made value (today it overwrites with `{}` — fails); then decide by the option.
- A test pinning "no reader": an AST/grep test in the hub that `DRRecordJSON` is not referenced outside the store
(so a new reader cannot appear without this design being revisited).
- Live proof: none needed on a box (hub-only); after deploy, the host page still shows "DR Recipe present".
## 6. Open questions for the operator
1. Retire the "slim DR record" of `05` §9 (Option A)? If you do nothing: the fields stay empty and unread, the
documents stay wrong, nothing breaks.
2. Tester 2's host page shows no DR recipe and no key escrow. Expected for a box that never did the escrow wizard —
is that the case? If you do nothing: a Tester 2 host loss has no hub-held recovery material.
@@ -0,0 +1,85 @@
# R-32 — RESET leaves the household's off-site ciphertext behind — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom.eu `8e2dc204` (hub v0.140.0), felhom-controller `5e7522e023` (v0.301.0). Architecture:
`07-backup-architecture.md` §246 (off-site deletion custody, decisions 68–69); `06-offsite-connectivity.md`; the
row's own three-part ruling from the 2026-07-21 rehearsal.
## 1. The problem
RESET is meant to destroy a household's off-site copy. On the shared pool box it deletes the Hetzner SUB-ACCOUNT. A
sub-account is a login, not the data: its home directory stays. Re-enabling off-site for the same customer creates a
sub-account with the SAME home (`felhom-<customer id>`) over the old ciphertext, whose key that RESET destroyed. Measured
on the pool box the night of 2026-07-21: **49 MB attributed** (2 snapshots, 48.7 MiB) against **1.4 GB + 3.0 MB
unattributed** in two `.orphaned-*` folders (`restic-and-pool.txt`, R-32). The orphan card then appeared — the
rehearsal's S7 had said in advance that it would be a finding.
## 2. What the code does today (read in source)
- RESET's off-site leg calls `s.offsite.Deprovision` (`hub/internal/web/customer_reset.go:217-225`) and journals
`hetzner: ok`, logging "repo data destroyed" (`:225`).
- `Deprovision`, shared tier: `DeleteSubaccount` per labelled sub-account, nothing else
(`hub/internal/offsite/offsite.go:303-326`). Its doc comment says "The offsite repo DATA dies with the
sub-account/box" (`:276-280`) — **true for the dedicated tier (the box is deleted, `:283-300`), false for the shared
one.** A comment that asserts an invariant the code does not provide.
- The home directory is fixed per customer: `HomeDirectory: "felhom-" + customerID` (`offsite.go:349`). So a new
lifecycle lands on the old folder.
- The hub already deletes off-site data in ONE place, through the sub-account's own password login (port 23):
`Registrar.DeleteSetAside` (`hub/internal/offsitekeys/offsitekeys.go:424-445`) — `rm -rf` of a `<repo>.orphaned-…`
folder only, refusing anything else (`IsSetAsidePath`, `:448-459`), after the household's 7-day abandonment delay
(decision 74, `service.go` „Decision 74").
- The move-aside for a reinstall WITHOUT RESET: the box asks the hub to rename the old repo to `<repo>.orphaned-<date>`
(`felhom-controller/controller/internal/backup/offbox.go:329-345`). The ruling keeps this — custody survives there.
- The operator's Restic tab shows the pool box's totals from the provider API and, per customer, the usage the BOX
reports (`hub/internal/web/offsite_box.go:121-185`). Bytes that no box reports (an `.orphaned-*` folder, an old
lifecycle) are visible only in the pool total, not per customer.
## 3. Options
**A. Purge through the sub-account, then delete it.** In `Deprovision` (shared), before `DeleteSubaccount`: log in with
the hub's stored password for that sub-account (it already does this for `authorized_keys` and `DeleteSetAside`),
remove the live repo and every `<repo>.orphaned-*`, check the home holds no repository left, then delete the
sub-account. A purge that fails stops the leg (`hetzner: failed`, re-run resumes) — the sub-account is NOT deleted,
because after that only the main account can reach the folder.
- Costs: hub only, small; one new registrar method (`PurgeRepos`) beside `DeleteSetAside`, same refusals style.
- No new credential: the main-account password the ruling named is not needed.
- Can go wrong: the stored password no longer works (rotated, or the hub DB restored from an older copy) → the leg
fails loudly and the operator must decide (main-account clean-up by hand). Deletes household data — that is RESET's
purpose, and the existing RESET ack covers it per the ruling; still the operator's word on the route.
**B. Purge with the pool box's MAIN account (the ruling's words).** The hub gets the main-account SFTP password.
- Costs: a new credential that reaches every household's folder on the pool box. A hub compromise then deletes every
household's history in one step. Bigger blast radius than A for the same result.
- Only advantage: also reaches folders of sub-accounts deleted BEFORE this fix (old lifecycles).
**C. A new home folder per lifecycle (`felhom-<id>-<n>`), no purge.** The new sub-account never sees old ciphertext,
so the orphan card cannot lie.
- Costs: small. But nothing is ever deleted: the ruling's part (1) is not met, and dead ciphertext fills the pool box
for ever, unseen unless part (3) is built.
**Part (3), the byte view, for every option:** per customer, the bytes in its folder (all repos, set-aside copies
included) beside the bytes its box attributes. Needs a per-folder size from the provider. **Not measured:** whether the
password login on port 23 answers `du -s` (the shell has `rm`, `mv`, `dd` — memory `storagebox-subaccount-shell…`). So
not buildable tonight: it rests on a mechanism nobody has measured, and tonight's fence forbids touching the Storage Box.
## 4. The pick — PROPOSAL for the operator, not a decision
**A, with part (3) after one read-only measurement.** It meets the ruling's part (1) with no new credential, keeps the
move-aside guard for reinstall-without-RESET (part 2 — untouched), and makes the journal's "repo data destroyed" true.
Old lifecycles' folders (if any are left on the pool box) are a one-time clean-up by hand — the operator's (question 2).
## 5. First slice and its proof
- Build (hub): `offsitekeys.Registrar.PurgeRepos(ctx, t, pw)` — removes `t.RepoPath` and every `IsSetAsidePath` match,
nothing else; then lists and refuses success if any remains. `offsite.Provisioner.Deprovision` (shared) calls it
through a seam BEFORE `DeleteSubaccount`. Fix the doc comment at `offsite.go:276-280`.
- Red test first (must FAIL today): a fake provider API and a fake shell; RESET's `Deprovision` for a shared customer.
Assert the shell saw `rm -rf <repo>` before the API saw `DeleteSubaccount`. Today it fails: no shell call at all.
- Second red test: the purge fails → `Deprovision` returns an error and `DeleteSubaccount` was NOT called (else the
folder becomes unreachable).
- Keep green: `DeleteSetAside`'s refusals; the dedicated tier unchanged; RESET's journal re-run.
- Live proof on a SCRATCH customer only (never a household): enable off-site, one run from scratch box 9202, RESET,
re-enable. Positive observable: the new sub-account's home lists no `restic`/`.orphaned-*` folder, and no orphan card
appears. Control from a different channel: the provider's pool-box `stats` size before and after (coarse, API).
Evidence off the machine before teardown.
## 6. Open questions for the operator
1. RESET deletes the household's off-site folder through the sub-account's own login before removing it (option A),
instead of through the pool box's main account — agree? If you do nothing: RESET keeps leaving the ciphertext; a
re-enabled customer sees an orphan card again.
2. Folders left by RESETs done before this fix (the 2026-07-21 measurement found 1.4 GB): clean them up once by hand
with the main account, or leave them? If you do nothing: they stay and use pool-box space; nothing reads them.
@@ -0,0 +1,108 @@
# R-812 — what is left of OS updates: the Proxmox packages and the kernel — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `11-os-updates.md` (§1, §3, §5.2, §5.5, §5.6,
§5.9, §7.1, §8). Related rows: R-836 (the kernel's fallback), R-808 (the intention, `ROADMAP.md`).
## 1. The problem
`11` §8 steps 1–5 are BUILT: the guest's Debian, the host's Debian, the fleet view, the Docker engine, the root files
(R-840, closed). Step 6 is not: **the host's Proxmox packages and its kernel are never updated on any box.** The host
fast lane takes only Debian-origin packages, so everything from the Proxmox repository waits for ever — including
packages with ordinary names (`zfsutils-linux`, `shim-signed`, `corosync`, `ceph-common`, `amd64-microcode`, `11` C3).
Measured tonight (read only; `apt list --upgradable`, `pveversion`, `dpkg -l`, `ls /boot`; no `apt update` — the package
lists are from 2026-10-05, the last daily refresh):
- **demo-hp:** `pve-manager` 9.2.2 → 9.2.21 pending; **77 packages pending, all from the Proxmox repository** (the Debian
lane has taken the rest): `qemu-server` 9.1.15 → 9.2.10, `pve-container` 6.1.10 → 6.1.14, `proxmox-backup-client`
4.2.0 → 4.2.7, `zfsutils-linux` 2.4.2 → 2.4.4, `shim-signed` 1.48 → 1.51, `pve-firmware`, `corosync`, `ceph-*`, `frr`,
`amd64-microcode`, `libpve-*`. Kernel: runs 7.0.14-20 (installed BY HAND in the 2026-10-04 spike), 7.0.2-6 kept.
- **demo-felhom:** `pve-manager` 9.2.2; **78 pending**; runs kernel **7.0.2-6, the install-time kernel** — 7.0.14-20 is in
the repository and not installed.
- So a box keeps its install-time Proxmox and kernel. Proxmox security fixes (for example in `pve-manager`'s web API, in
`qemu-server`, in the Secure Boot loader `shim-signed`) never arrive. The spike counted 79 Proxmox packages "not
covered" on the host (`11` §7.1).
## 2. What the code does today (read in source)
- The fast-lane origin rule: `felhom-agent/configs/felhom-os-apply:63` `FAST_ORIGINS = ("Debian", "Debian-Security")`;
a plan package of any other origin is refused **R2** (`:491-492`, and in the simulation `:1164` `origin_ok`).
- The host's name rule on top: `:91-92` `HOST_SLOW_RE` (`linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`,
`firmware-*`, `grub*`, `shim*`, `systemd-boot*`, `*-microcode`, `efibootmgr`) → refusal **R14** (`:493-494`, `:1234-1235`);
the pending selection skips them (`:1136`).
- Layers and lanes: `:437-442` — `guest`/`host` are lane `fast` only; `docker` is the only `slow` layer. There is no
host slow layer.
- Tests pinning this: `configs/test_felhom_os_apply.py:417` `test_R2_non_debian_origin_in_the_plan`, `:422`
`test_R2_non_debian_origin_in_the_simulation`, `:584`/`:590` `test_R14_kernel_package_*`, `:639`
`test_pending_fast_skips_proxmox_docker_and_kernel` (a pending `pve-manager` 9.2.21 from „Proxmox Debian Repository"
is NOT taken).
- What a Proxmox step restarts (spike `partH/H5-proxmox-simulate.txt`, read from the maintainer scripts, not run):
`pve-manager` → pvedaemon, pveproxy, pvescheduler, pvestatd, spiceproxy; `pve-cluster` → pmxcfs (`/etc/pve` gone for
seconds — the agent's API calls fail meanwhile); `qemu-server`, `pve-firewall`, `pve-ha-*`, `corosync`, `chrony`,
`zfs-zed`. **No guest is stopped by these scripts** (read, not measured). One new package (`proxmox-firewall-data`).
- The kernel (`11` §5.6, R-836, measured on demo-hp with the operator's word): installing a kernel makes it the GRUB
default at once; `--next-boot` and `grub-reboot` are NOT one-shots here (`/boot` is ext4 on LVM; GRUB cannot write its
env block); a kernel that hangs before userspace stays the default on every power cycle. `sp5100_tco` exists on
demo-hp, blacklisted, never armed. demo-felhom (Intel N100) has no watchdog measured.
## 3. Options
**A. A host slow lane for the Proxmox USERSPACE packages, no kernel, no reboot.**
- What: a new layer `pve` (lane `slow`) in the wrapper: origin „Proxmox Debian Repository" only, still never a name in
`HOST_SLOW_RE` (kernel, boot, firmware, microcode stay out), no removals, no new package except a named allow-list (the
`proxmox-firewall-data` kind). Approval like the Docker engine set (`11` §5.8): the operator approves one SET per box
generation on the System page after ring 0 ran it; ring 1 only by signed job. Runs in the night leg after the host
Debian step, under the same heavy-op gate. Health = the host rule (`HostHealthVerdict`) + `pveversion` reports the
new version.
- Cost: wrapper layer + refusals + tests; hub candidate/approval/System page (copies the Docker set's code); ~1.5 days.
- Can go wrong: pmxcfs restart while the agent writes `/etc/pve` (the gate already excludes backups/restore-tests; the
agent's own reconcile must wait too); a Proxmox point release that needs a newer kernel (`proxmox-ve` depends on
`proxmox-default-kernel` — the plan must not pull the kernel in: the simulation refusal R14 already catches it);
undo is only „install the previous version" (Proxmox keeps 30–66 versions, `11` C2 — good).
- Brings the 77–78 pending packages, minus the boot-chain ones, to every box. Leaves `shim-signed`, `pve-firmware`,
microcode and the kernel out.
**B. The kernel lane, with a fallback that works on these hosts — measured first.**
- What: install the kernel (plus `shim-signed`, `pve-firmware`, microcode, which also only act at boot), keep the old
one as the permanent default, boot the new one ONCE, and make it the default only from userspace after a healthy boot.
The one-shot needs one of: (1) **UEFI `BootNext`** with a second boot entry whose GRUB config defaults to the new kernel
— the firmware clears BootNext itself on use, so a hang falls back on the next power cycle; (2) a GRUB env block on
the ESP (vfat, writable by GRUB); (3) arming a hardware watchdog early (`sp5100_tco` on demo-hp) so a hang power-cycles
into the old default. None is measured.
- Cost: a spike with ≥4 reboots per box on both demo hosts (Secure Boot ON on demo-hp, OFF on demo-felhom; AMD vs Intel),
then the lane: ~3–5 days. Every customer box restarts all its apps once per kernel (≈1–2 min at night).
- Can go wrong: firmware that ignores `BootNext` (seen on consumer boards); Secure Boot refusing a second entry; a
hang with no watchdog still needs a person — it is a dead box at a household. Telling households that the box may
restart at night is a **promise to users** (`11` §5.7) — the operator's.
**C. The kernel only by an operator-present maintenance step.**
- What: no automatic kernel lane. The System page shows „kernel behind / reboot needed"; the operator runs the step per
box by a signed job at a time they choose, ready to power-cycle (or ask the household to).
- Cost: small (a signed job and a runbook), but a person per box per kernel; does not scale past a handful of boxes.
- Can go wrong: kernels are never installed because nobody schedules them — the state today with a button.
## 4. The pick — PROPOSAL for the operator, not a decision
**A now, C for the kernel until B is measured.** A closes the biggest gap (78 Proxmox packages, the web API and
qemu/lxc tooling) with no reboot and the same approval shape the operator already uses for Docker. The kernel stays a
separate, operator-scheduled act (C) until R-836's fallback is measured on both demo hosts; then B replaces C.
**R-812 should split.** Close R-812 (P2) when A ships to ring 0 and ring 1, with R-836 (kernel fallback, P3) carrying
the kernel. The kernel's urgency then becomes its own question: a P2 row „kernel security fixes reach a box" only if the
operator wants it before the first paying customer (question 2).
## 5. First slice and its proof
- Build: wrapper layer `pve` lane `slow` — origin „Proxmox Debian Repository" only; R14 still applies; R2 for any other
origin; refusal for a removal or an unlisted new package; appliance only (R12's root-owned install record).
- Red test first (must FAIL on today's code): a host plan with `pve-manager 9.2.21` origin „Proxmox Debian Repository",
layer `pve`, lane `slow` → today refused R12 („layer is not guest, host or docker"); after: installed. Companion
tests that must stay refused: the same plan carrying `proxmox-kernel-7.0` (R14) or `shim-signed` (R14); a Debian
package in a `pve` plan (R2); `test_pending_fast_skips_proxmox_docker_and_kernel` unchanged.
- Live proof on **demo-felhom** (ring 0, Tier 0, its Proxmox is the install-time one): one approved set of the pending
Proxmox packages, no kernel. Positive observable: `pveversion` reads 9.2.21; the host health rule passes; the guest
kept running (container `StartedAt` unchanged). Control from a different channel: the hub's System page row for the
box, and an app on the guest answering 200 every 5 s through the step. No reboot needed, so no operator word for a
reboot; the operator's word is the approval click, as for Docker. Evidence off the box at the end.
## 6. Open questions for the operator
1. **May Proxmox's own packages (no kernel) update by night, after you approve each set on the System page, as the
Docker engine does?** Pick: yes (option A). If you do nothing: every box keeps its install-time Proxmox; 78 packages
behind on the demo boxes today, and growing.
2. **Must kernel security fixes reach customer boxes before the first paying customer?** If yes, B is a ~1-week arc with
reboots on both demo boxes (your word before each). Pick: no — the kernel by your hand (C) until B is measured. If you
do nothing: kernels stay at the install-time version; a kernel hole stays open until you run the step by hand.
@@ -0,0 +1,68 @@
# R-822 — fake snapshots steering the off-site retention — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom.eu `8e2dc204`, felhom-controller `5e7522e023` (v0.301.0), hub v0.140.0. Architecture:
`07-backup-architecture.md` §246 (off-site deletion custody, decisions 68–69) and threat row 10 (§1298);
`audits/offsite-append-only-2026-10-03/DESIGN.md` §3 Option 1.
## 1. The problem
The box's off-site key can only ADD (decision 69). Someone holding that key can add snapshots with any date and make the
box's own honest pruner delete the real ones. Measured in the lab 2026-10-03 (restic 0.14.0, rclone `--append-only`):
13 empty future-dated snapshots made the policy `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` select all 3 real
snapshots (`audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt`). The guard (controller v0.289.0,
tuned v0.290.0 and v0.294.0) refuses that shape. **What is left (read in source tonight, not measured):** a fake dated
in the PAST, inside a week or month that already holds a real keep, and newer than that real snapshot, wins the bucket.
The real snapshot then falls out of the plan as an ordinary old removal. The guard cannot tell it from honest retention.
## 2. What the code does today (read in source)
- The guard runs on the box before any `forget`, inside a hub window (`felhom-controller/controller/internal/backup/offbox_window.go:21-38`, `offsiteGuard` `:131-158`). It refuses: a future date (`:133-135`); a date after the window opened (`:136-138`); a plan that removes a snapshot younger than 7 calendar days and not superseded the same day (`:141-149`); a plan above `MaxRemove` (`:152-154`).
- A YOUNG real snapshot superseded the same day by a newer one is only EXCLUDED, then removed in a later window once old (`:142-145`, R-824). So a same-day fake, planted after the night run, also wins the daily bucket — later.
- `MaxRemove` = half of the count, at least 5 (`felhom.eu/hub/internal/offsitekeys/service.go:240-246`). So the old real history goes in about two windows.
- The hub's after-window check compares counts the BOX reports: `countBefore` comes with the box's request (`service.go:284`), `CountAfter` with its close report (`service.go:343`). Fakes added between windows keep the count level, so the drop alarm (`EventWindowDrop`, `:350`) does not fire.
- **The same exposure already exists without any fake:** during a window the box's own key may delete (`service.go:303`, the deleting line). A box that is broken into at window time can delete everything and report any counts. Decision 68 accepted that, with the count check as the backstop.
**Result, inferred from the above:** an attacker who held the box's key once can, over about two weekly windows, shrink
the real off-site history to the last 7 days. Nobody is alarmed. It needs the attacker to plant fakes; it does not need
the attacker to still be there.
## 3. Options
**A. Accept the residual and close the row.** Record it in `07` §246 as a stated limit of decision 68.
- Costs: nothing to build.
- Can go wrong: the slow, silent loss above. It is the same class of loss decision 68 already accepts for a broken-into
box in a window (box-reported counts), so A adds no NEW weakness — it names one.
**B. The hub keeps its own snapshot list (no repository password needed).** The hub already logs in to each sub-account
daily (the `authorized_keys` audit). It also lists `<repo>/snapshots/`: file names (= snapshot ids) and their upload
times on the provider. The hub then: (1) counts before and after each window ITSELF, not from the box; (2) alarms when
snapshot files appear that no box run explains (a run reports its time; a file uploaded outside a run, or more files
than runs, is the injection signature); (3) refuses to open a window while such files exist.
- Costs: hub only, about one day. One more SFTP listing per customer per day. Custody unchanged (07 §8a): the hub sees
file names and sizes, never contents.
- Can go wrong: a manual run, a catch-up run or a crash retry must be counted as "a run", or false alarms. Not measured:
whether the provider's restricted shell gives upload times (`ls -l` on port 23) — measure on the scratch sub-account.
- It also closes the box-trusted count in decision 68's backstop — worth having on its own.
**C. A stricter box guard (refuse when a keep is empty or tiny).** Rejected: a fake can point at a real snapshot's tree
with another date, at no cost. The guard runs on the box that the attacker held.
## 4. The pick — PROPOSAL for the operator, not a decision
**A now, B as the next build.** Close R-822 as a stated, bounded limit of decision 68 (the loss is the last-7-days
floor, at ≤ half the snapshots per weekly window), and file B as its own row: "the hub trusts the box's own count before
and after a window". B is the real defence for both the fake-snapshot case and the broken-into-box-in-a-window case. It
needs a measurement first, so it is not tonight's work.
## 5. First slice and its proof (for B)
- Measure first, on the scratch sub-account: `ls -l <repo>/snapshots/` through the password login on port 23 — does it
list names and upload times? (Read only.)
- Build: hub `offsitekeys` gets `ListSnapshots(ctx, target, pw) ([]SnapshotFile, error)`; `OpenWindowFor` and
`CloseWindowFor` take the counts from it, not from the box (the box's numbers are logged beside).
- Red test first (must FAIL today): a fake registrar whose listing holds 17 files before and 3 after; the box reports
17 → 15. Assert `EventWindowDrop` fires. Today it does not (the hub uses the box's 15).
- Live proof on scratch 9202's sub-account: open a one-shot window; the hub's own before/after counts appear in its
log and match `restic snapshots` run on the box. Control from a different channel: the provider's listing read by
hand over SFTP.
## 6. Open questions for the operator
1. Accept the residual (an attacker who once held a box's key can shrink its off-site history to 7 days over about two
weeks, silently) and close R-822? If you do nothing: the row stays open; nothing changes on the boxes.
2. Build B (the hub counts the snapshots itself, about a day of work, after one read-only measurement)? If you do
nothing: the window's alarm keeps trusting the numbers the box sends.
@@ -0,0 +1,94 @@
# R-861 — the three sudoers leftovers — design proposal (burn-down night 2026-10-06, no code)
Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `03-host-agent.md` §3.1 (the group table)
and §11, `11-os-updates.md` §5.4.2 (the config bundle), `09` §3 decision 122. Read only; nothing ran on a box.
## 1. The problem
After agent v0.146.1 a compromised agent PROCESS (user `felhom-agent`) no longer becomes host root: measured 2026-10-05,
29 attack lines refused, `sudo -l` 93/93 on both demo boxes (`audits/hub-safety-2026-10-05/partF/`). Three paths were
left open on purpose and named in `03` §3.1. The row asks the operator: accept them, or close them before the first
paying customer? The question for each is: **what does an attacker who already owns the agent process gain through it,
beyond what the agent already has?**
What the agent already has (the baseline): its Proxmox token holds `VM.Backup`, `VM.Allocate`, `VM.Config.*`,
`VM.PowerMgmt`, `Pool.Allocate` on `/pool/felhom` (`felhom.eu/scripts/felhom-host-install.sh:309`). So it can already
back up a customer guest and restore it into a scratch guest (the restore-test does exactly that), stop and start felhom
guests, and it writes the guest's `bootstrap.json` itself. **The household's data is already in its reach, offline.**
## 2. What the code does today (read in source)
**(a) The controller image ref.** `FELHOM_CONTROLLERSWAP` allows `pct exec <vmid> -- tee /etc/felhom-controller-image`
(`felhom-agent/configs/felhom-agent.sudoers:122`). The ref goes in on STDIN (`internal/localapi/controllerswap.go:150`).
The regex `controllerImageRe` (`controllerswap.go:36`) is the agent's OWN check — a compromised agent skips it, and sudo
cannot see stdin. The guest's bootstrap unit then runs `docker run … "$IMAGE"` with the docker socket, the read-only
bootstrap dir and `/mnt` (`configs/build-golden.sh:316,355-362`).
**Correction to `03` §3.1:** the table says *"a chosen (pinned-registry) image ref"*. That is not true: only the
`docker image inspect` line is pinned (`sudoers:119`); the `tee` content is free, so the bootstrap pulls and runs **any
image from any registry**. Gain over the baseline: a LIVE foothold in the guest with the docker socket (guest root, the
household's running apps and its LAN), and it survives agent restarts until the next swap. Not host root.
**(b) The felhom-op SSH key.** `felhom-priv-apply sshd-key` installs one plain key line (no `command=`/`from=`;
`configs/felhom-priv-apply:60-61`) from a file the agent staged; the key itself comes from the hub, unsigned. A
compromised agent can therefore put its own key on `felhom-op`. `felhom-op`'s sudo (`configs/felhom-op.sudoers`) is
scoped: restart wg/agent/sshd, `pct list`, and `pct start|stop|unlock [0-9]*` — **on any guest of the host**, not only
the felhom pool. Gain over the baseline: power control of NON-felhom guests on a BYO host (the household's own other
VMs), and an interactive login. Not host root. (Side note: these `pct` lines still use the `*` glob, which matches
spaces — the R-861 shape 1; no harmful `pct start/stop` option is known, so this is hygiene, not a hole.)
**(c) The escrow ceremony.** The root child (`FELHOM_ESCROW`, `sudoers:283-284`) returns the recovery code R on the
agent's stdout pipe (`internal/localapi/escrow_ceremony.go:19-29,93`); the agent holds R in memory for one claim
(`:140-147`). The agent's own hub key may read this box's escrow blob (`hub/internal/api/handler.go:252,1357-1372`,
self-scoped). So a compromised agent can learn R, fetch the blob and unwrap this box's PBS encryption key (and the
identity bundle). Gain over the baseline: **off-box** decryption of this box's off-site archives, which lasts after the
compromise is cleaned up, until the key is rotated. One box only; not root.
## 3. Options
**(a)**
- **A1. A checking wrapper verb.** `felhom-priv-apply controller-image <vmid>`: as root, read stdin, require
`^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$`, then write the guest file. The sudoers `tee` line
is removed. Same mechanism as decision 122 (b); delivered by the signed config bundle. Cost: ~1–2 h (one verb, its
Python tests, the Go call path, `TestManifestCoveredBySudoers`, the capability probe). What can go wrong: the old
agent binary still calls `tee` → the bundle must follow the agent update (the usual order). Leaves: a chosen OLD
controller version from our registry (a downgrade) — still possible.
- **A2. Check in the guest.** The bootstrap script refuses a ref outside the pattern. Cost: a golden change, and
existing guests keep the old script until re-baked — slow to reach the fleet.
- **A3. Accept.** Write the corrected sentence in `03` §3.1.
**(b)**
- **B1. Sign the key.** The operator signs the felhom-op key with the operator key; `felhom-os-apply` verifies as root.
Cost: medium; every key rotation needs the operator's offline signature.
- **B2. Narrow felhom-op's sudo** to anchored regexes (`^start [0-9]+$` …) — hygiene only; it does not stop the key swap.
Cost: ~30 min, rides the bundle.
- **B3. Accept** (felhom-op is not root; the gain is power control of guests).
**(c)**
- **C1. R bypasses the agent.** The root child writes R straight into the guest (to the controller), so the agent never
sees it. Cost: a ceremony redesign across agent and controller; touches the escrow promise to the household.
- **C2. Accept** — the ceremony is designed so the box handles K once; the residual is "a compromised agent can read
this one box's backups off-site", which `03` §3.1 already names.
## 4. The pick — PROPOSAL for the operator, not a decision
- **(a) A1, before the first paying customer.** It is the only one of the three that gives a live, persistent foothold
next to the household's running apps, and the fix uses a mechanism already built and measured. Fix the `03` sentence
in the same commit.
- **(b) B3 now, B2 as hygiene** with the next bundle; B1 later if the OOB door is ever opened wider.
- **(c) C2.** C1 changes an escrow promise and is a redesign — not before the first customer.
## 5. First slice and its proof (for A1)
- Build: verb `controller-image` in `configs/felhom-priv-apply` (stdin ≤ 256 bytes, the regex, a numeric vmid, then
`pct exec <vmid> -- tee` as root); sudoers: remove `tee /etc/felhom-controller-image$`, add
`/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$`; `controllerswap.go:150` calls the verb.
- Red test first: `configs/test_felhom_priv_apply.py` — `controller-image 9201` with stdin
`docker.io/library/alpine:latest` must exit 3; with `gitea.dooplex.hu/admin/felhom-controller:0.301.0` exit 0. And
`TestSudoersRefusesTheR861Injections` gains `pct exec 9201 -- tee /etc/felhom-controller-image` as a REFUSED line —
it fails on today's sudoers.
- Live proof on scratch 9202 after a bundle there (not tonight): `sudo -l -U felhom-agent` lists no `tee`; a managed
controller update still swaps (positive observable: the new controller version in `docker ps` AND the hub's host
report — two channels); a hand-fed `alpine` ref is refused (journal tag `felhom-priv-apply`).
## 6. Open questions for the operator
1. **(a)** Close the free image ref before the first paying customer (A1, ~1–2 h, ships with a bundle)? If you do
nothing: a compromised agent can run any container next to the household's apps, with the docker socket.
2. **(b)+(c)** Accept both for the first customers (felhom-op is not root; the escrow residual is one box's off-site
backups)? If you do nothing: they stay open, named in `03` §3.1, as today.