Night 2026-10-06: designs R-105, R-812, R-861, R-32, R-822 (rows point at them); R-895 filed; 03 §3.1 and 07 corrected
gates / gates (push) Successful in 2m44s
gates / gates (push) Successful in 2m44s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -17,3 +17,12 @@ no reboot.
|
||||
|---|---|---|---|
|
||||
| R-892 | **§3 stopped** — the shell has no SSH agent; DooPlex's two own keys (`id_rsa`, `id_ed25519`) are refused by `root@192.168.0.154`, direct and through demo-hp. The key the operator copied at 20:20 came from the agent of the operator's own login, so it is not a key this shell holds. Not tried, by the brief: another key, the vaulted password, a disk edit. Next try needs: `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` run **from DooPlex as kisfenyo** (it then asks for the box's root password once). Evidence `s3/s3-ssh-refused.txt` | 5 | — |
|
||||
| R-894 | **fixed on agent main** `74b5eae` (ships as v0.150.0 with the memory-kill check). Saved copy on disk, read only on an unreadable storage; fresh → not due, old → due, none → due/unknown; a storage that answers wins. 4 red-proofs `s4/`. `07` §6.1 updated | 50 | agent `74b5eae` |
|
||||
| (§1.2 fence) | **The app list per box could not be read in full.** demo-hp and demo-felhom read directly (`docker ps`, read only); the hub's Apps page has no per-box list; reading a copy of the hub database was **refused by the session's permission check** — stopped there, no workaround, the copy deleted at once. **So every catalog change tonight goes to branch `night-held-2026-10-06`, none to `main`** (the fence holds without the list). | 10 | — |
|
||||
| R-173 | **skipped — waits for an event.** Its only LEFT item is runbook §3 steps 4–5 (copy into a live PVC), which need the hub down — the next planned hub maintenance or a scratch-k3s DR drill. Read only tonight; nothing to do | 2 | — |
|
||||
| R-861 | **design written** (`design-R-861.md`): pick (a) close before the first paying customer, (b)/(c) accept; `03` §3.1 sentence corrected ("pinned-registry" was false) | 35 | (this batch) |
|
||||
| R-812 | **design written** (`design-R-812.md`): pick a Proxmox-packages slow lane (no kernel), per-set approval; kernel by hand until R-836 is measured. Read-only on demo-hp/demo-felhom (one temp file the helper made in /tmp on demo-hp was deleted at once) | 30 | (this batch) |
|
||||
| R-105 | **design written** (`design-R-105.md`): both empty fields have no live writer/reader — pick: retire them and correct `05`/`06` (operator word) | 25 | (this batch) |
|
||||
| R-32 | **design written** (`design-R-32.md`): pick A — purge through the sub-account's own login before deleting it; byte view waits on one read-only measurement | 40 | (this batch) |
|
||||
| R-822 | **design written** (`design-R-822.md`): pick — accept the residual and close (operator's word); option B filed as **R-895** (P2) | — | (this batch) |
|
||||
| R-895 | **OPENED** — the hub's clean-up-window check trusts the box's own snapshot counts (needs a design + a read-only measurement) | 5 | (this batch) |
|
||||
| (07 §6.5 area) | stale sentence fixed: „until that ships nothing prunes" — the window runs live since 2026-10-05 | 2 | (this batch) |
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
# R-105 — the two "empty DR records" — design proposal (burn-down night 2026-10-06, no code)
|
||||
|
||||
Baselines read: felhom.eu `8e2dc204` (hub v0.140.0 live), felhom-agent `74b5eae`. Architecture: `05-hub-architecture.md`
|
||||
§9 and §11 (the "slim DR record"), `06-offsite-connectivity.md` §3.5 (the escrow directive).
|
||||
|
||||
## 1. The problem, and what was measured
|
||||
The row says two hub-held records are `{}` on every box: `hosts.dr_record_json` and `host_escrow.directive_json`
|
||||
(the third, `dr_recipe…drives`, was fixed 2026-07-28 and is not re-measured tonight).
|
||||
|
||||
**Tonight's reading is in SOURCE, not in the database** — reading a copy of the hub database was refused by the
|
||||
session's permission check, so no value was read. Source proves more than a reading would: for these two fields the
|
||||
empty value is the only value the shipped code can produce. The hub's host pages (read through the operator UI) show
|
||||
"DR Recipe present / Key Escrow present" on demo-hp, demo-felhom and Tester 1, and "none / none" on Tester 2.
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- **`hosts.dr_record_json` has no writer and no reader.** Created by hub v0.7.0 (`7c0c7545`) with default `'{}'`
|
||||
(`hub/internal/store/store.go:360`); scanned into `Host.DRRecordJSON` (`store.go:2795`, `:2867`) and read by
|
||||
nothing else in the hub, the agent or the controller (repo-wide grep). The `05` §9 "slim DR record" was never built.
|
||||
- **`host_escrow.directive_json` is written only by a by-hand flag.** The agent sends a directive only when the
|
||||
operator runs `--selftest=escrow-create -directive <file>` (`felhom-agent/cmd/felhom-agent/main.go:201`,
|
||||
`:2767-2771`, `:3132-3135`). The production ceremony (the customer's escrow wizard) runs ONE fixed argv with no
|
||||
`-directive` (`felhom-agent/configs/felhom-agent.sudoers:284`). When that upload carries an identity blob, the hub
|
||||
stores the missing directive as `'{}'` (`hub/internal/api/handler.go:1293-1305`, `store.go:3612`) — so every
|
||||
wizard escrow writes `{}`, and also overwrites the one directive made by hand on 2026-07-04 (`06` §3.5).
|
||||
- **Nothing reads the directive's fields.** The hub serves it only from `/re-enroll` and `/restore-directive`
|
||||
(`hub/internal/api/dr.go:101`, `:155`); no agent or controller code calls either route (grep).
|
||||
- **The DR path that IS built does not need them.** The agent's restore plan reads the desired-state
|
||||
`restore_directive` plus the DR recipe (`felhom-agent/internal/dr/plan.go:45`, `:104`); the recipe carries the PBS
|
||||
repo id and namespace (`internal/hub/dr_recipe.go` `DRPBSCoord`), the hub holds the endpoint's PBS fingerprint
|
||||
(`hub/internal/tenantsync/client.go:136`), and the wrapped key is `host_escrow.blob`.
|
||||
|
||||
So the row's two thirds are not a fault in a running path. They are a **design that was half-built and then
|
||||
bypassed**, and two architecture documents still describe it as if it existed.
|
||||
|
||||
## 3. Options
|
||||
**A. Retire both, and correct the documents.** Drop the `DRRecordJSON` scan field; stop storing the directive
|
||||
(the column stays, read as `{}`); `05` §9/§11 and `06` §3.5 say where each fact really lives (recipe, tenantsync,
|
||||
escrow blob). Cost: ~45 min, hub only, no box change. Can go wrong: if a future re-enroll client wants the
|
||||
directive, it must be rebuilt — the routes stay and serve `{}`.
|
||||
|
||||
**B. Populate the directive from the ceremony.** The agent fills `{pbs repo id, namespace, endpoint fingerprint}`
|
||||
in the wizard ceremony. Cost: agent + sudoers argv change (the argv is pinned byte-for-byte) + a bundle delivery;
|
||||
~2 h and a release. Can go wrong: a second copy of facts the recipe already carries, which can disagree with it —
|
||||
the R-106 namespace defect was exactly such a disagreement.
|
||||
|
||||
**C. Do nothing more; close R-105 with this evidence.** Cost: nothing. The two documents keep describing a record
|
||||
that does not exist, which is how this row was filed in the first place.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
**Option A.** One source per fact; the recipe is reported every cycle and was fixed to match the backup (R-106).
|
||||
It removes a design promise in `05` §9, so it needs the operator's word (a design decision is not a defect).
|
||||
Until then the row can move to WAITING-ON-OPERATOR: there is no data at risk — the empty fields have no reader.
|
||||
|
||||
## 5. First slice and its proof
|
||||
- Red test first: a hub test that the escrow PUT with an identity blob and no directive leaves `directive_json`
|
||||
unchanged from a hand-made value (today it overwrites with `{}` — fails); then decide by the option.
|
||||
- A test pinning "no reader": an AST/grep test in the hub that `DRRecordJSON` is not referenced outside the store
|
||||
(so a new reader cannot appear without this design being revisited).
|
||||
- Live proof: none needed on a box (hub-only); after deploy, the host page still shows "DR Recipe present".
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. Retire the "slim DR record" of `05` §9 (Option A)? If you do nothing: the fields stay empty and unread, the
|
||||
documents stay wrong, nothing breaks.
|
||||
2. Tester 2's host page shows no DR recipe and no key escrow. Expected for a box that never did the escrow wizard —
|
||||
is that the case? If you do nothing: a Tester 2 host loss has no hub-held recovery material.
|
||||
@@ -0,0 +1,85 @@
|
||||
# R-32 — RESET leaves the household's off-site ciphertext behind — design proposal (burn-down night 2026-10-06, no code)
|
||||
|
||||
Baselines read: felhom.eu `8e2dc204` (hub v0.140.0), felhom-controller `5e7522e023` (v0.301.0). Architecture:
|
||||
`07-backup-architecture.md` §246 (off-site deletion custody, decisions 68–69); `06-offsite-connectivity.md`; the
|
||||
row's own three-part ruling from the 2026-07-21 rehearsal.
|
||||
|
||||
## 1. The problem
|
||||
RESET is meant to destroy a household's off-site copy. On the shared pool box it deletes the Hetzner SUB-ACCOUNT. A
|
||||
sub-account is a login, not the data: its home directory stays. Re-enabling off-site for the same customer creates a
|
||||
sub-account with the SAME home (`felhom-<customer id>`) over the old ciphertext, whose key that RESET destroyed. Measured
|
||||
on the pool box the night of 2026-07-21: **49 MB attributed** (2 snapshots, 48.7 MiB) against **1.4 GB + 3.0 MB
|
||||
unattributed** in two `.orphaned-*` folders (`restic-and-pool.txt`, R-32). The orphan card then appeared — the
|
||||
rehearsal's S7 had said in advance that it would be a finding.
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- RESET's off-site leg calls `s.offsite.Deprovision` (`hub/internal/web/customer_reset.go:217-225`) and journals
|
||||
`hetzner: ok`, logging "repo data destroyed" (`:225`).
|
||||
- `Deprovision`, shared tier: `DeleteSubaccount` per labelled sub-account, nothing else
|
||||
(`hub/internal/offsite/offsite.go:303-326`). Its doc comment says "The offsite repo DATA dies with the
|
||||
sub-account/box" (`:276-280`) — **true for the dedicated tier (the box is deleted, `:283-300`), false for the shared
|
||||
one.** A comment that asserts an invariant the code does not provide.
|
||||
- The home directory is fixed per customer: `HomeDirectory: "felhom-" + customerID` (`offsite.go:349`). So a new
|
||||
lifecycle lands on the old folder.
|
||||
- The hub already deletes off-site data in ONE place, through the sub-account's own password login (port 23):
|
||||
`Registrar.DeleteSetAside` (`hub/internal/offsitekeys/offsitekeys.go:424-445`) — `rm -rf` of a `<repo>.orphaned-…`
|
||||
folder only, refusing anything else (`IsSetAsidePath`, `:448-459`), after the household's 7-day abandonment delay
|
||||
(decision 74, `service.go` „Decision 74").
|
||||
- The move-aside for a reinstall WITHOUT RESET: the box asks the hub to rename the old repo to `<repo>.orphaned-<date>`
|
||||
(`felhom-controller/controller/internal/backup/offbox.go:329-345`). The ruling keeps this — custody survives there.
|
||||
- The operator's Restic tab shows the pool box's totals from the provider API and, per customer, the usage the BOX
|
||||
reports (`hub/internal/web/offsite_box.go:121-185`). Bytes that no box reports (an `.orphaned-*` folder, an old
|
||||
lifecycle) are visible only in the pool total, not per customer.
|
||||
|
||||
## 3. Options
|
||||
**A. Purge through the sub-account, then delete it.** In `Deprovision` (shared), before `DeleteSubaccount`: log in with
|
||||
the hub's stored password for that sub-account (it already does this for `authorized_keys` and `DeleteSetAside`),
|
||||
remove the live repo and every `<repo>.orphaned-*`, check the home holds no repository left, then delete the
|
||||
sub-account. A purge that fails stops the leg (`hetzner: failed`, re-run resumes) — the sub-account is NOT deleted,
|
||||
because after that only the main account can reach the folder.
|
||||
- Costs: hub only, small; one new registrar method (`PurgeRepos`) beside `DeleteSetAside`, same refusals style.
|
||||
- No new credential: the main-account password the ruling named is not needed.
|
||||
- Can go wrong: the stored password no longer works (rotated, or the hub DB restored from an older copy) → the leg
|
||||
fails loudly and the operator must decide (main-account clean-up by hand). Deletes household data — that is RESET's
|
||||
purpose, and the existing RESET ack covers it per the ruling; still the operator's word on the route.
|
||||
|
||||
**B. Purge with the pool box's MAIN account (the ruling's words).** The hub gets the main-account SFTP password.
|
||||
- Costs: a new credential that reaches every household's folder on the pool box. A hub compromise then deletes every
|
||||
household's history in one step. Bigger blast radius than A for the same result.
|
||||
- Only advantage: also reaches folders of sub-accounts deleted BEFORE this fix (old lifecycles).
|
||||
|
||||
**C. A new home folder per lifecycle (`felhom-<id>-<n>`), no purge.** The new sub-account never sees old ciphertext,
|
||||
so the orphan card cannot lie.
|
||||
- Costs: small. But nothing is ever deleted: the ruling's part (1) is not met, and dead ciphertext fills the pool box
|
||||
for ever, unseen unless part (3) is built.
|
||||
|
||||
**Part (3), the byte view, for every option:** per customer, the bytes in its folder (all repos, set-aside copies
|
||||
included) beside the bytes its box attributes. Needs a per-folder size from the provider. **Not measured:** whether the
|
||||
password login on port 23 answers `du -s` (the shell has `rm`, `mv`, `dd` — memory `storagebox-subaccount-shell…`). So
|
||||
not buildable tonight: it rests on a mechanism nobody has measured, and tonight's fence forbids touching the Storage Box.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
**A, with part (3) after one read-only measurement.** It meets the ruling's part (1) with no new credential, keeps the
|
||||
move-aside guard for reinstall-without-RESET (part 2 — untouched), and makes the journal's "repo data destroyed" true.
|
||||
Old lifecycles' folders (if any are left on the pool box) are a one-time clean-up by hand — the operator's (question 2).
|
||||
|
||||
## 5. First slice and its proof
|
||||
- Build (hub): `offsitekeys.Registrar.PurgeRepos(ctx, t, pw)` — removes `t.RepoPath` and every `IsSetAsidePath` match,
|
||||
nothing else; then lists and refuses success if any remains. `offsite.Provisioner.Deprovision` (shared) calls it
|
||||
through a seam BEFORE `DeleteSubaccount`. Fix the doc comment at `offsite.go:276-280`.
|
||||
- Red test first (must FAIL today): a fake provider API and a fake shell; RESET's `Deprovision` for a shared customer.
|
||||
Assert the shell saw `rm -rf <repo>` before the API saw `DeleteSubaccount`. Today it fails: no shell call at all.
|
||||
- Second red test: the purge fails → `Deprovision` returns an error and `DeleteSubaccount` was NOT called (else the
|
||||
folder becomes unreachable).
|
||||
- Keep green: `DeleteSetAside`'s refusals; the dedicated tier unchanged; RESET's journal re-run.
|
||||
- Live proof on a SCRATCH customer only (never a household): enable off-site, one run from scratch box 9202, RESET,
|
||||
re-enable. Positive observable: the new sub-account's home lists no `restic`/`.orphaned-*` folder, and no orphan card
|
||||
appears. Control from a different channel: the provider's pool-box `stats` size before and after (coarse, API).
|
||||
Evidence off the machine before teardown.
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. RESET deletes the household's off-site folder through the sub-account's own login before removing it (option A),
|
||||
instead of through the pool box's main account — agree? If you do nothing: RESET keeps leaving the ciphertext; a
|
||||
re-enabled customer sees an orphan card again.
|
||||
2. Folders left by RESETs done before this fix (the 2026-07-21 measurement found 1.4 GB): clean them up once by hand
|
||||
with the main account, or leave them? If you do nothing: they stay and use pool-box space; nothing reads them.
|
||||
@@ -0,0 +1,108 @@
|
||||
# R-812 — what is left of OS updates: the Proxmox packages and the kernel — design proposal (burn-down night 2026-10-06, no code)
|
||||
|
||||
Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `11-os-updates.md` (§1, §3, §5.2, §5.5, §5.6,
|
||||
§5.9, §7.1, §8). Related rows: R-836 (the kernel's fallback), R-808 (the intention, `ROADMAP.md`).
|
||||
|
||||
## 1. The problem
|
||||
`11` §8 steps 1–5 are BUILT: the guest's Debian, the host's Debian, the fleet view, the Docker engine, the root files
|
||||
(R-840, closed). Step 6 is not: **the host's Proxmox packages and its kernel are never updated on any box.** The host
|
||||
fast lane takes only Debian-origin packages, so everything from the Proxmox repository waits for ever — including
|
||||
packages with ordinary names (`zfsutils-linux`, `shim-signed`, `corosync`, `ceph-common`, `amd64-microcode`, `11` C3).
|
||||
|
||||
Measured tonight (read only; `apt list --upgradable`, `pveversion`, `dpkg -l`, `ls /boot`; no `apt update` — the package
|
||||
lists are from 2026-10-05, the last daily refresh):
|
||||
- **demo-hp:** `pve-manager` 9.2.2 → 9.2.21 pending; **77 packages pending, all from the Proxmox repository** (the Debian
|
||||
lane has taken the rest): `qemu-server` 9.1.15 → 9.2.10, `pve-container` 6.1.10 → 6.1.14, `proxmox-backup-client`
|
||||
4.2.0 → 4.2.7, `zfsutils-linux` 2.4.2 → 2.4.4, `shim-signed` 1.48 → 1.51, `pve-firmware`, `corosync`, `ceph-*`, `frr`,
|
||||
`amd64-microcode`, `libpve-*`. Kernel: runs 7.0.14-20 (installed BY HAND in the 2026-10-04 spike), 7.0.2-6 kept.
|
||||
- **demo-felhom:** `pve-manager` 9.2.2; **78 pending**; runs kernel **7.0.2-6, the install-time kernel** — 7.0.14-20 is in
|
||||
the repository and not installed.
|
||||
- So a box keeps its install-time Proxmox and kernel. Proxmox security fixes (for example in `pve-manager`'s web API, in
|
||||
`qemu-server`, in the Secure Boot loader `shim-signed`) never arrive. The spike counted 79 Proxmox packages "not
|
||||
covered" on the host (`11` §7.1).
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- The fast-lane origin rule: `felhom-agent/configs/felhom-os-apply:63` `FAST_ORIGINS = ("Debian", "Debian-Security")`;
|
||||
a plan package of any other origin is refused **R2** (`:491-492`, and in the simulation `:1164` `origin_ok`).
|
||||
- The host's name rule on top: `:91-92` `HOST_SLOW_RE` (`linux-*`, `proxmox-kernel*`, `pve-kernel*`, `pve-firmware`,
|
||||
`firmware-*`, `grub*`, `shim*`, `systemd-boot*`, `*-microcode`, `efibootmgr`) → refusal **R14** (`:493-494`, `:1234-1235`);
|
||||
the pending selection skips them (`:1136`).
|
||||
- Layers and lanes: `:437-442` — `guest`/`host` are lane `fast` only; `docker` is the only `slow` layer. There is no
|
||||
host slow layer.
|
||||
- Tests pinning this: `configs/test_felhom_os_apply.py:417` `test_R2_non_debian_origin_in_the_plan`, `:422`
|
||||
`test_R2_non_debian_origin_in_the_simulation`, `:584`/`:590` `test_R14_kernel_package_*`, `:639`
|
||||
`test_pending_fast_skips_proxmox_docker_and_kernel` (a pending `pve-manager` 9.2.21 from „Proxmox Debian Repository"
|
||||
is NOT taken).
|
||||
- What a Proxmox step restarts (spike `partH/H5-proxmox-simulate.txt`, read from the maintainer scripts, not run):
|
||||
`pve-manager` → pvedaemon, pveproxy, pvescheduler, pvestatd, spiceproxy; `pve-cluster` → pmxcfs (`/etc/pve` gone for
|
||||
seconds — the agent's API calls fail meanwhile); `qemu-server`, `pve-firewall`, `pve-ha-*`, `corosync`, `chrony`,
|
||||
`zfs-zed`. **No guest is stopped by these scripts** (read, not measured). One new package (`proxmox-firewall-data`).
|
||||
- The kernel (`11` §5.6, R-836, measured on demo-hp with the operator's word): installing a kernel makes it the GRUB
|
||||
default at once; `--next-boot` and `grub-reboot` are NOT one-shots here (`/boot` is ext4 on LVM; GRUB cannot write its
|
||||
env block); a kernel that hangs before userspace stays the default on every power cycle. `sp5100_tco` exists on
|
||||
demo-hp, blacklisted, never armed. demo-felhom (Intel N100) has no watchdog measured.
|
||||
|
||||
## 3. Options
|
||||
**A. A host slow lane for the Proxmox USERSPACE packages, no kernel, no reboot.**
|
||||
- What: a new layer `pve` (lane `slow`) in the wrapper: origin „Proxmox Debian Repository" only, still never a name in
|
||||
`HOST_SLOW_RE` (kernel, boot, firmware, microcode stay out), no removals, no new package except a named allow-list (the
|
||||
`proxmox-firewall-data` kind). Approval like the Docker engine set (`11` §5.8): the operator approves one SET per box
|
||||
generation on the System page after ring 0 ran it; ring 1 only by signed job. Runs in the night leg after the host
|
||||
Debian step, under the same heavy-op gate. Health = the host rule (`HostHealthVerdict`) + `pveversion` reports the
|
||||
new version.
|
||||
- Cost: wrapper layer + refusals + tests; hub candidate/approval/System page (copies the Docker set's code); ~1.5 days.
|
||||
- Can go wrong: pmxcfs restart while the agent writes `/etc/pve` (the gate already excludes backups/restore-tests; the
|
||||
agent's own reconcile must wait too); a Proxmox point release that needs a newer kernel (`proxmox-ve` depends on
|
||||
`proxmox-default-kernel` — the plan must not pull the kernel in: the simulation refusal R14 already catches it);
|
||||
undo is only „install the previous version" (Proxmox keeps 30–66 versions, `11` C2 — good).
|
||||
- Brings the 77–78 pending packages, minus the boot-chain ones, to every box. Leaves `shim-signed`, `pve-firmware`,
|
||||
microcode and the kernel out.
|
||||
|
||||
**B. The kernel lane, with a fallback that works on these hosts — measured first.**
|
||||
- What: install the kernel (plus `shim-signed`, `pve-firmware`, microcode, which also only act at boot), keep the old
|
||||
one as the permanent default, boot the new one ONCE, and make it the default only from userspace after a healthy boot.
|
||||
The one-shot needs one of: (1) **UEFI `BootNext`** with a second boot entry whose GRUB config defaults to the new kernel
|
||||
— the firmware clears BootNext itself on use, so a hang falls back on the next power cycle; (2) a GRUB env block on
|
||||
the ESP (vfat, writable by GRUB); (3) arming a hardware watchdog early (`sp5100_tco` on demo-hp) so a hang power-cycles
|
||||
into the old default. None is measured.
|
||||
- Cost: a spike with ≥4 reboots per box on both demo hosts (Secure Boot ON on demo-hp, OFF on demo-felhom; AMD vs Intel),
|
||||
then the lane: ~3–5 days. Every customer box restarts all its apps once per kernel (≈1–2 min at night).
|
||||
- Can go wrong: firmware that ignores `BootNext` (seen on consumer boards); Secure Boot refusing a second entry; a
|
||||
hang with no watchdog still needs a person — it is a dead box at a household. Telling households that the box may
|
||||
restart at night is a **promise to users** (`11` §5.7) — the operator's.
|
||||
|
||||
**C. The kernel only by an operator-present maintenance step.**
|
||||
- What: no automatic kernel lane. The System page shows „kernel behind / reboot needed"; the operator runs the step per
|
||||
box by a signed job at a time they choose, ready to power-cycle (or ask the household to).
|
||||
- Cost: small (a signed job and a runbook), but a person per box per kernel; does not scale past a handful of boxes.
|
||||
- Can go wrong: kernels are never installed because nobody schedules them — the state today with a button.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
**A now, C for the kernel until B is measured.** A closes the biggest gap (78 Proxmox packages, the web API and
|
||||
qemu/lxc tooling) with no reboot and the same approval shape the operator already uses for Docker. The kernel stays a
|
||||
separate, operator-scheduled act (C) until R-836's fallback is measured on both demo hosts; then B replaces C.
|
||||
|
||||
**R-812 should split.** Close R-812 (P2) when A ships to ring 0 and ring 1, with R-836 (kernel fallback, P3) carrying
|
||||
the kernel. The kernel's urgency then becomes its own question: a P2 row „kernel security fixes reach a box" only if the
|
||||
operator wants it before the first paying customer (question 2).
|
||||
|
||||
## 5. First slice and its proof
|
||||
- Build: wrapper layer `pve` lane `slow` — origin „Proxmox Debian Repository" only; R14 still applies; R2 for any other
|
||||
origin; refusal for a removal or an unlisted new package; appliance only (R12's root-owned install record).
|
||||
- Red test first (must FAIL on today's code): a host plan with `pve-manager 9.2.21` origin „Proxmox Debian Repository",
|
||||
layer `pve`, lane `slow` → today refused R12 („layer is not guest, host or docker"); after: installed. Companion
|
||||
tests that must stay refused: the same plan carrying `proxmox-kernel-7.0` (R14) or `shim-signed` (R14); a Debian
|
||||
package in a `pve` plan (R2); `test_pending_fast_skips_proxmox_docker_and_kernel` unchanged.
|
||||
- Live proof on **demo-felhom** (ring 0, Tier 0, its Proxmox is the install-time one): one approved set of the pending
|
||||
Proxmox packages, no kernel. Positive observable: `pveversion` reads 9.2.21; the host health rule passes; the guest
|
||||
kept running (container `StartedAt` unchanged). Control from a different channel: the hub's System page row for the
|
||||
box, and an app on the guest answering 200 every 5 s through the step. No reboot needed, so no operator word for a
|
||||
reboot; the operator's word is the approval click, as for Docker. Evidence off the box at the end.
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. **May Proxmox's own packages (no kernel) update by night, after you approve each set on the System page, as the
|
||||
Docker engine does?** Pick: yes (option A). If you do nothing: every box keeps its install-time Proxmox; 78 packages
|
||||
behind on the demo boxes today, and growing.
|
||||
2. **Must kernel security fixes reach customer boxes before the first paying customer?** If yes, B is a ~1-week arc with
|
||||
reboots on both demo boxes (your word before each). Pick: no — the kernel by your hand (C) until B is measured. If you
|
||||
do nothing: kernels stay at the install-time version; a kernel hole stays open until you run the step by hand.
|
||||
@@ -0,0 +1,68 @@
|
||||
# R-822 — fake snapshots steering the off-site retention — design proposal (burn-down night 2026-10-06, no code)
|
||||
|
||||
Baselines read: felhom.eu `8e2dc204`, felhom-controller `5e7522e023` (v0.301.0), hub v0.140.0. Architecture:
|
||||
`07-backup-architecture.md` §246 (off-site deletion custody, decisions 68–69) and threat row 10 (§1298);
|
||||
`audits/offsite-append-only-2026-10-03/DESIGN.md` §3 Option 1.
|
||||
|
||||
## 1. The problem
|
||||
The box's off-site key can only ADD (decision 69). Someone holding that key can add snapshots with any date and make the
|
||||
box's own honest pruner delete the real ones. Measured in the lab 2026-10-03 (restic 0.14.0, rclone `--append-only`):
|
||||
13 empty future-dated snapshots made the policy `--keep-daily 7 --keep-weekly 4 --keep-monthly 6` select all 3 real
|
||||
snapshots (`audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt`). The guard (controller v0.289.0,
|
||||
tuned v0.290.0 and v0.294.0) refuses that shape. **What is left (read in source tonight, not measured):** a fake dated
|
||||
in the PAST, inside a week or month that already holds a real keep, and newer than that real snapshot, wins the bucket.
|
||||
The real snapshot then falls out of the plan as an ordinary old removal. The guard cannot tell it from honest retention.
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- The guard runs on the box before any `forget`, inside a hub window (`felhom-controller/controller/internal/backup/offbox_window.go:21-38`, `offsiteGuard` `:131-158`). It refuses: a future date (`:133-135`); a date after the window opened (`:136-138`); a plan that removes a snapshot younger than 7 calendar days and not superseded the same day (`:141-149`); a plan above `MaxRemove` (`:152-154`).
|
||||
- A YOUNG real snapshot superseded the same day by a newer one is only EXCLUDED, then removed in a later window once old (`:142-145`, R-824). So a same-day fake, planted after the night run, also wins the daily bucket — later.
|
||||
- `MaxRemove` = half of the count, at least 5 (`felhom.eu/hub/internal/offsitekeys/service.go:240-246`). So the old real history goes in about two windows.
|
||||
- The hub's after-window check compares counts the BOX reports: `countBefore` comes with the box's request (`service.go:284`), `CountAfter` with its close report (`service.go:343`). Fakes added between windows keep the count level, so the drop alarm (`EventWindowDrop`, `:350`) does not fire.
|
||||
- **The same exposure already exists without any fake:** during a window the box's own key may delete (`service.go:303`, the deleting line). A box that is broken into at window time can delete everything and report any counts. Decision 68 accepted that, with the count check as the backstop.
|
||||
|
||||
**Result, inferred from the above:** an attacker who held the box's key once can, over about two weekly windows, shrink
|
||||
the real off-site history to the last 7 days. Nobody is alarmed. It needs the attacker to plant fakes; it does not need
|
||||
the attacker to still be there.
|
||||
|
||||
## 3. Options
|
||||
**A. Accept the residual and close the row.** Record it in `07` §246 as a stated limit of decision 68.
|
||||
- Costs: nothing to build.
|
||||
- Can go wrong: the slow, silent loss above. It is the same class of loss decision 68 already accepts for a broken-into
|
||||
box in a window (box-reported counts), so A adds no NEW weakness — it names one.
|
||||
|
||||
**B. The hub keeps its own snapshot list (no repository password needed).** The hub already logs in to each sub-account
|
||||
daily (the `authorized_keys` audit). It also lists `<repo>/snapshots/`: file names (= snapshot ids) and their upload
|
||||
times on the provider. The hub then: (1) counts before and after each window ITSELF, not from the box; (2) alarms when
|
||||
snapshot files appear that no box run explains (a run reports its time; a file uploaded outside a run, or more files
|
||||
than runs, is the injection signature); (3) refuses to open a window while such files exist.
|
||||
- Costs: hub only, about one day. One more SFTP listing per customer per day. Custody unchanged (07 §8a): the hub sees
|
||||
file names and sizes, never contents.
|
||||
- Can go wrong: a manual run, a catch-up run or a crash retry must be counted as "a run", or false alarms. Not measured:
|
||||
whether the provider's restricted shell gives upload times (`ls -l` on port 23) — measure on the scratch sub-account.
|
||||
- It also closes the box-trusted count in decision 68's backstop — worth having on its own.
|
||||
|
||||
**C. A stricter box guard (refuse when a keep is empty or tiny).** Rejected: a fake can point at a real snapshot's tree
|
||||
with another date, at no cost. The guard runs on the box that the attacker held.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
**A now, B as the next build.** Close R-822 as a stated, bounded limit of decision 68 (the loss is the last-7-days
|
||||
floor, at ≤ half the snapshots per weekly window), and file B as its own row: "the hub trusts the box's own count before
|
||||
and after a window". B is the real defence for both the fake-snapshot case and the broken-into-box-in-a-window case. It
|
||||
needs a measurement first, so it is not tonight's work.
|
||||
|
||||
## 5. First slice and its proof (for B)
|
||||
- Measure first, on the scratch sub-account: `ls -l <repo>/snapshots/` through the password login on port 23 — does it
|
||||
list names and upload times? (Read only.)
|
||||
- Build: hub `offsitekeys` gets `ListSnapshots(ctx, target, pw) ([]SnapshotFile, error)`; `OpenWindowFor` and
|
||||
`CloseWindowFor` take the counts from it, not from the box (the box's numbers are logged beside).
|
||||
- Red test first (must FAIL today): a fake registrar whose listing holds 17 files before and 3 after; the box reports
|
||||
17 → 15. Assert `EventWindowDrop` fires. Today it does not (the hub uses the box's 15).
|
||||
- Live proof on scratch 9202's sub-account: open a one-shot window; the hub's own before/after counts appear in its
|
||||
log and match `restic snapshots` run on the box. Control from a different channel: the provider's listing read by
|
||||
hand over SFTP.
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. Accept the residual (an attacker who once held a box's key can shrink its off-site history to 7 days over about two
|
||||
weeks, silently) and close R-822? If you do nothing: the row stays open; nothing changes on the boxes.
|
||||
2. Build B (the hub counts the snapshots itself, about a day of work, after one read-only measurement)? If you do
|
||||
nothing: the window's alarm keeps trusting the numbers the box sends.
|
||||
@@ -0,0 +1,94 @@
|
||||
# R-861 — the three sudoers leftovers — design proposal (burn-down night 2026-10-06, no code)
|
||||
|
||||
Baselines read: felhom-agent `74b5eae`, felhom.eu `8e2dc204`. Architecture: `03-host-agent.md` §3.1 (the group table)
|
||||
and §11, `11-os-updates.md` §5.4.2 (the config bundle), `09` §3 decision 122. Read only; nothing ran on a box.
|
||||
|
||||
## 1. The problem
|
||||
After agent v0.146.1 a compromised agent PROCESS (user `felhom-agent`) no longer becomes host root: measured 2026-10-05,
|
||||
29 attack lines refused, `sudo -l` 93/93 on both demo boxes (`audits/hub-safety-2026-10-05/partF/`). Three paths were
|
||||
left open on purpose and named in `03` §3.1. The row asks the operator: accept them, or close them before the first
|
||||
paying customer? The question for each is: **what does an attacker who already owns the agent process gain through it,
|
||||
beyond what the agent already has?**
|
||||
|
||||
What the agent already has (the baseline): its Proxmox token holds `VM.Backup`, `VM.Allocate`, `VM.Config.*`,
|
||||
`VM.PowerMgmt`, `Pool.Allocate` on `/pool/felhom` (`felhom.eu/scripts/felhom-host-install.sh:309`). So it can already
|
||||
back up a customer guest and restore it into a scratch guest (the restore-test does exactly that), stop and start felhom
|
||||
guests, and it writes the guest's `bootstrap.json` itself. **The household's data is already in its reach, offline.**
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
|
||||
**(a) The controller image ref.** `FELHOM_CONTROLLERSWAP` allows `pct exec <vmid> -- tee /etc/felhom-controller-image`
|
||||
(`felhom-agent/configs/felhom-agent.sudoers:122`). The ref goes in on STDIN (`internal/localapi/controllerswap.go:150`).
|
||||
The regex `controllerImageRe` (`controllerswap.go:36`) is the agent's OWN check — a compromised agent skips it, and sudo
|
||||
cannot see stdin. The guest's bootstrap unit then runs `docker run … "$IMAGE"` with the docker socket, the read-only
|
||||
bootstrap dir and `/mnt` (`configs/build-golden.sh:316,355-362`).
|
||||
**Correction to `03` §3.1:** the table says *"a chosen (pinned-registry) image ref"*. That is not true: only the
|
||||
`docker image inspect` line is pinned (`sudoers:119`); the `tee` content is free, so the bootstrap pulls and runs **any
|
||||
image from any registry**. Gain over the baseline: a LIVE foothold in the guest with the docker socket (guest root, the
|
||||
household's running apps and its LAN), and it survives agent restarts until the next swap. Not host root.
|
||||
|
||||
**(b) The felhom-op SSH key.** `felhom-priv-apply sshd-key` installs one plain key line (no `command=`/`from=`;
|
||||
`configs/felhom-priv-apply:60-61`) from a file the agent staged; the key itself comes from the hub, unsigned. A
|
||||
compromised agent can therefore put its own key on `felhom-op`. `felhom-op`'s sudo (`configs/felhom-op.sudoers`) is
|
||||
scoped: restart wg/agent/sshd, `pct list`, and `pct start|stop|unlock [0-9]*` — **on any guest of the host**, not only
|
||||
the felhom pool. Gain over the baseline: power control of NON-felhom guests on a BYO host (the household's own other
|
||||
VMs), and an interactive login. Not host root. (Side note: these `pct` lines still use the `*` glob, which matches
|
||||
spaces — the R-861 shape 1; no harmful `pct start/stop` option is known, so this is hygiene, not a hole.)
|
||||
|
||||
**(c) The escrow ceremony.** The root child (`FELHOM_ESCROW`, `sudoers:283-284`) returns the recovery code R on the
|
||||
agent's stdout pipe (`internal/localapi/escrow_ceremony.go:19-29,93`); the agent holds R in memory for one claim
|
||||
(`:140-147`). The agent's own hub key may read this box's escrow blob (`hub/internal/api/handler.go:252,1357-1372`,
|
||||
self-scoped). So a compromised agent can learn R, fetch the blob and unwrap this box's PBS encryption key (and the
|
||||
identity bundle). Gain over the baseline: **off-box** decryption of this box's off-site archives, which lasts after the
|
||||
compromise is cleaned up, until the key is rotated. One box only; not root.
|
||||
|
||||
## 3. Options
|
||||
|
||||
**(a)**
|
||||
- **A1. A checking wrapper verb.** `felhom-priv-apply controller-image <vmid>`: as root, read stdin, require
|
||||
`^gitea\.dooplex\.hu/admin/felhom-controller:[0-9]+\.[0-9]+\.[0-9]+$`, then write the guest file. The sudoers `tee` line
|
||||
is removed. Same mechanism as decision 122 (b); delivered by the signed config bundle. Cost: ~1–2 h (one verb, its
|
||||
Python tests, the Go call path, `TestManifestCoveredBySudoers`, the capability probe). What can go wrong: the old
|
||||
agent binary still calls `tee` → the bundle must follow the agent update (the usual order). Leaves: a chosen OLD
|
||||
controller version from our registry (a downgrade) — still possible.
|
||||
- **A2. Check in the guest.** The bootstrap script refuses a ref outside the pattern. Cost: a golden change, and
|
||||
existing guests keep the old script until re-baked — slow to reach the fleet.
|
||||
- **A3. Accept.** Write the corrected sentence in `03` §3.1.
|
||||
|
||||
**(b)**
|
||||
- **B1. Sign the key.** The operator signs the felhom-op key with the operator key; `felhom-os-apply` verifies as root.
|
||||
Cost: medium; every key rotation needs the operator's offline signature.
|
||||
- **B2. Narrow felhom-op's sudo** to anchored regexes (`^start [0-9]+$` …) — hygiene only; it does not stop the key swap.
|
||||
Cost: ~30 min, rides the bundle.
|
||||
- **B3. Accept** (felhom-op is not root; the gain is power control of guests).
|
||||
|
||||
**(c)**
|
||||
- **C1. R bypasses the agent.** The root child writes R straight into the guest (to the controller), so the agent never
|
||||
sees it. Cost: a ceremony redesign across agent and controller; touches the escrow promise to the household.
|
||||
- **C2. Accept** — the ceremony is designed so the box handles K once; the residual is "a compromised agent can read
|
||||
this one box's backups off-site", which `03` §3.1 already names.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
- **(a) A1, before the first paying customer.** It is the only one of the three that gives a live, persistent foothold
|
||||
next to the household's running apps, and the fix uses a mechanism already built and measured. Fix the `03` sentence
|
||||
in the same commit.
|
||||
- **(b) B3 now, B2 as hygiene** with the next bundle; B1 later if the OOB door is ever opened wider.
|
||||
- **(c) C2.** C1 changes an escrow promise and is a redesign — not before the first customer.
|
||||
|
||||
## 5. First slice and its proof (for A1)
|
||||
- Build: verb `controller-image` in `configs/felhom-priv-apply` (stdin ≤ 256 bytes, the regex, a numeric vmid, then
|
||||
`pct exec <vmid> -- tee` as root); sudoers: remove `tee /etc/felhom-controller-image$`, add
|
||||
`/usr/local/sbin/felhom-priv-apply ^controller-image [0-9]+$`; `controllerswap.go:150` calls the verb.
|
||||
- Red test first: `configs/test_felhom_priv_apply.py` — `controller-image 9201` with stdin
|
||||
`docker.io/library/alpine:latest` must exit 3; with `gitea.dooplex.hu/admin/felhom-controller:0.301.0` exit 0. And
|
||||
`TestSudoersRefusesTheR861Injections` gains `pct exec 9201 -- tee /etc/felhom-controller-image` as a REFUSED line —
|
||||
it fails on today's sudoers.
|
||||
- Live proof on scratch 9202 after a bundle there (not tonight): `sudo -l -U felhom-agent` lists no `tee`; a managed
|
||||
controller update still swaps (positive observable: the new controller version in `docker ps` AND the hub's host
|
||||
report — two channels); a hand-fed `alpine` ref is refused (journal tag `felhom-priv-apply`).
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. **(a)** Close the free image ref before the first paying customer (A1, ~1–2 h, ships with a bundle)? If you do
|
||||
nothing: a compromised agent can run any container next to the household's apps, with the docker socket.
|
||||
2. **(b)+(c)** Accept both for the first customers (felhom-op is not root; the escrow residual is one box's off-site
|
||||
backups)? If you do nothing: they stay open, named in `03` §3.1, as today.
|
||||
Reference in New Issue
Block a user