hub-safety session: R-135/R-133/R-604/R-530/R-508/R-509/R-880 closed, R-861/R-173/R-518/R-519 narrowed, R-879/R-881 opened (336 → 332); 03 §3.1, 05 §16, golden 0.296.0, the hub-DB off-site plan, STATUS
gates / gates (push) Successful in 32s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 13:47:57 +02:00
parent 2b30733b0d
commit 0826e41b31
44 changed files with 1511 additions and 31 deletions
@@ -196,10 +196,11 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 |
| A second login step for the dashboard (a TOTP code or a passkey) | — | **MISSING** | — | One password, one bcrypt hash (`controller/internal/web/auth.go:37-44`) → R-811 (added 2026-10-03) |
| The household can leave Felhom, or outlive it — the box runs without the hub, the household owns its domain, tunnel and off-site account, and can export everything | — | **MISSING** (as a written answer) | the LOST-hub half only: `_recovery-inventory-2026-07-28.md` §D2.4, `07` §8 row 11b | Leaving and hand-over are answered nowhere → R-810 (spike, added 2026-10-03) |
| **The host agent cannot reach root without the operator key: exact sudo patterns, no agent-written file installed where root reads it without a content check, the agent binary only by an operator-signed update checked as root** | agent **v0.146.1** (R-861; delivered by a step bundle, R-880) | **PROVEN-LIVE on both demo boxes (2026-10-05) — with three named residuals** | `audits/hub-safety-2026-10-05/partF/` (real sudo: 23 of 29 attacks allowed before, 0 after; 64 capability commands allowed; `sudo -l` 93/93 on demo-hp and demo-felhom after the bundle; capability probe 67/67; a staged unit over `/etc/sudoers.d` refused live) | `03` §3.1: the controller-swap image ref (guest-scoped), the felhom-op SSH key (hub-delivered, unsigned; felhom-op's sudo is scoped), the escrow ceremony relays R → R-861 |
| WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` |
| The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe |
| The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing |
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → **R-133** |
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. **Sealed at rest since hub v0.135.0 (R-133 CLOSED):** the off-site seal and key; live 2026-10-05: 4 legacy rows sealed at start-up, 0 left plain, and the demo-hp reveal still minted a PVE ticket (HTTP 200; a wrong password 401) — `audits/hub-safety-2026-10-05/partB/`. A database backup now needs `OFFSITE_SECRET_KEY` too → R-173 |
## F. Notifications & monitoring
@@ -233,6 +234,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
+55 -1
View File
@@ -72,6 +72,59 @@ Explicitly does **not**:
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
### 3.1 The agent's admin commands, group by group — and how each is narrowed (R-861, agent v0.146.1) `[DESIGN — 2026-10-05, CC; operator may reverse]`
**The question.** Can a compromised agent PROCESS (running as the `felhom-agent` user) become root on its host without
the operator's key? Read on 2026-10-04 (R-861) and measured 2026-10-05 with the real sudo 1.9.16 in a throwaway
container: **with the v0.145.0 sudoers, yes — 23 of 29 attack command lines were allowed.** Two shapes did it:
1. **A glob in the arguments.** Sudo's `*` in arguments also matches spaces, so one grant smuggled extra options:
`pct set [0-9]* -onboot 1` allowed `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest);
`mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` allowed a `..` path onto `/etc/sudoers.d`; `nft add element …
*` allowed `; flush ruleset`.
2. **A file the agent wrote, installed where root reads it.** A `.mount` unit (bind any directory over `/etc`), a
dnsmasq drop-in (`dhcp-script=` runs as root), the wg-quick config (`PostUp=` runs as root), the OOB sshd config
(`AuthorizedKeysFile` + `StrictModes no`), the guest pre-start hook (Proxmox runs it as root), the shared-parent boot
script, and the agent BINARY itself (the escrow ceremony and the guest hook run it as root; `apply` took a sha the
agent passed).
**The rule after v0.146.1** (the R-861 fix direction: each becomes a root-owned wrapper that checks its own input, or a
fixed file, delivered by the signed config bundle):
- Every varying argument list is a **sudo regular expression** (`^…$`): one value per slot, a fixed character set, no
`..`, no extra argument. Literal lines stay literal.
- **No file the agent wrote is installed where root reads it.** Either the content is FIXED and comes with the signed
bundle (the hook, the shared parent), or a root wrapper checks the CONTENT against the agent's own renderers before
installing it (`felhom-priv-apply`), or the operator's signature is checked as root (`felhom-os-apply agent_update`).
- The pins: `TestManifestCoveredBySudoers` (every command the agent runs is still allowed), `TestSudoersRefusesTheR861
Injections` (the 29 attacks are not), the real-sudo run of both (`audits/hub-safety-2026-10-05/partF/
sudo-container-proof.txt`), and `sudo -l -U felhom-agent` on both demo boxes after the bundle.
| Group | What it is for | How it is narrowed (v0.146.1) | Left open |
|---|---|---|---|
| `FELHOM_MOUNT` | fs-UUID mount units for enrolled drives | install only via `felhom-priv-apply unit <name>`: `[Unit]` only Description + `After=local-fs-pre.target`, `Where=` `/mnt/<name>` or `/mnt/felhom-drives/<name>` and equal to the unit name, `What=` a UUID or a network source, no `bind`/`suid`/`dev`, no continuation lines; systemctl verbs on `mnt-…\.mount` only | — |
| `FELHOM_NETMOUNT` | NAS automount pairs, re-arm, clean-up | same checker; a network share must carry `nosuid,nodev` (the agent now renders them); `rm`/`rmdir`/`reset-failed` one exact name | — |
| `FELHOM_DISK` | SMART, thin-pool, PV and pool reads | exact device / LV patterns (no extra options such as `smartctl -s off`, `lvs --config`) | read-only |
| `FELHOM_PROVISION` | bootstrap config mount, autostart | exact `mpN` spec (`…/guests/<vmid>/bootstrap,mp=/…[,ro=1]`) — no smuggled `--dev0` | — |
| `FELHOM_FORMAT` | data-bearing probe + guarded mkfs | one device path, no space; `mkfs` only through `felhom-mkfs-guarded` (its own root checks) with `ext4`/`xfs` | blkid/lsblk read any `/dev` path (read-only) |
| `FELHOM_DNSMASQ` | the LAN split-horizon resolver | drop-ins via `felhom-priv-apply dnsmasq` (only `bind-interfaces`, `no-resolv`, `listen-address`, `server`, `local`, `address`); `rm` one exact name; exact `pct exec` reads | — |
| `FELHOM_GUESTHOOK` | the pre-start self-heal hook | the hook is a FIXED bundle file; the agent only checks it (`SnippetReady`) and registers it; exact vmid/slot | — |
| `FELHOM_INTERMEDIARY` | the shared drive parent + live drive binds | boot script + unit are FIXED bundle files (the agent only enables the unit); one-segment drive names (no leading dot, no `..`) | — |
| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen (pinned-registry) image ref and restart the guest's bootstrap — the household's data, not host root |
| `FELHOM_STALELOCK` / `FELHOM_SCRATCH_TEARDOWN` | stale-lock clear; failed restore-test scratch | exact vmid; the scratch band `99000[0-9]` was already exact | — |
| `FELHOM_WG` | the off-site tunnel | conf via `felhom-priv-apply wg` (only the keys `renderConf` writes; no `PostUp`/`PreUp`/`DNS`/`Table`; `/32` only) | — |
| `FELHOM_SELFUPDATE` | commit / rollback of the A/B flip | **`apply` removed**: the flip runs only inside `felhom-os-apply agent_update`, after the operator signature, host, window and nonce are checked as root and the staged bytes are hashed ONCE and copied to a root-owned dir (`/var/lib/felhom-os-apply/agent-update/`); the wrapper accepts only that dir | — |
| `FELHOM_SSHD` | the out-of-band operator sshd | config via `felhom-priv-apply sshd-config` (the ONE template, only the Port varies, never 22); the felhom-op key via `sshd-key` (one plain key, no `command=`/`from=` options) | felhom-op's key itself is hub-delivered, not signed: a compromised agent can install its own key for **felhom-op** — whose sudo is scoped (`felhom-op.sudoers`), not root |
| `FELHOM_OOB` | the OOB firewall sets | `add element` takes exactly `{ <ip>[/n] }` or `{ <port> }` — no chained command | — |
| `FELHOM_PBSDR` / `FELHOM_BACKUPTARGET` | PBS DR entry; whole-system backup target | unchanged: the arguments stay coarse, and the root wrappers (`felhom-pbs-apply`, `felhom-backup-target-apply`) are the gate (fixed verbs, own validation) | coarse argv into a checking wrapper |
| `FELHOM_ESCROW` | the recovery-code ceremony (runs the agent binary as root) | the binary is only ever an operator-signed one (`FELHOM_SELFUPDATE`); as root it pins the PVE secret dir and the WG state dir, refuses a storage id that is a path, and reads its two staged files by walking the path with `openat(O_NOFOLLOW)` (no symlink anywhere) | **by design the agent relays R**, so a compromised agent can still learn this box's PBS key through the ceremony — not root, but the backup key |
| `FELHOM_SELFHEAL` / `FELHOM_GUESTNET` / `FELHOM_OSAPPLY` | networking restart; guest DHCP watchdog; OS updates | exact; `felhom-os-apply --plan …` stays the glob line on purpose — the bundle's own self-check reads that exact text, and the wrapper refuses any other plan path (R1) | — |
**What this does not change.** The operator key (`/etc/felhom/operator-signers`, root-owned, never a bundle path) stays
the one trust root; the agent's API token is untouched; a box gets the new rule only through the signed
`agent_config_update` (the order: signed `agent_update` first, then the bundle — after the bundle, an agent below
0.146.0 cannot update itself on that box).
## 4. Control model — reconcile + signed destructive ops
Two channels, split by **reversibility**, not by transport.
@@ -655,7 +708,8 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea
runs as root at guest start (and `pct reboot` is granted); `FELHOM_INTERMEDIARY` installs a script and a systemd unit
that run as root at boot; `FELHOM_ESCROW` runs the agent binary as root, and `FELHOM_SELFUPDATE apply` accepts a sha
the agent itself passes. So a compromised agent PROCESS is root on its host; the root-owned trust files (decision 93,
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants.
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. **Narrowed in agent
v0.146.1 — §3.1 lists every group, the rule now, and what stays open.**
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
@@ -130,6 +130,17 @@ R-216. Either way the box's reported agent must meet the chosen MinAgent, else t
the Hosts page and logged once per change as `managed floor SERVED`. The rules the operator follows:
`runbooks/publish-train-rules.md` rule 1.
**Boxes left behind (hub v0.135.0, R-604 + R-530).** A per-customer floor wins over the global one, so a global
raise does not move a box whose OWN floor is lower — and `managed floor SERVED` is logged once per change, so that box
was silent (demo-hp missed four raises, 2026-09-21). Now the raise logs one line per such customer and sends ONE
operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set
(`customer_configs.min_controller_set_at`; "unknown" for one set before v0.135.0), and the System page's "Version
floors" table lists the global floor, every per-customer floor with its age, and which ones the global cannot move.
Agents are a separate train (they update only by a per-box signed job, R-530's ruling): the System page shows each
box's agent against the vouched one ("0.142.0 → 0.145.0 (since …)", amber, red after the wait), and a box behind the
vouched agent for 7 days raises `agent_behind` (warning, operator-only; `OS_ALARM_AGENT_BEHIND_AFTER`; the clock starts
when the hub first sees the box behind). `[DESIGN — CC 2026-10-05, operator may reverse; `09` decision 119]`
## 6. Authorization — signed-op queue + editing flow
Implements Part 4's gate on the hub side. The hub holds **no signing key**.
@@ -403,3 +414,35 @@ count appeared was the bind page's passphrase hint, and its English half is now
phrase you received from your operator during setup") — because "five words" stops being true for an
English household, and was already wrong for one whose passphrase predates this release. The
Hungarian „öt szó" is correct and unchanged.
## 16. The operator surface's own safety [DESIGN — hub v0.135.0, CC 2026-10-05, operator may reverse]
### 16.1 Form protection (CSRF) on both login paths (R-135)
The operator logs in two ways: a browser session (`hub_session` cookie + a per-session token on every form) and HTTP
Basic for scripts. Until v0.135.0 a state-changing request with NO cookie skipped the token check — on the reasoning
that it must be a script. It need not be: a browser caches Basic credentials per origin and resends them on a
cross-site form POST (SameSite does not govern the Authorization header). Now a request without a session passes
only with Basic credentials AND the header `X-Felhom-Operator` (any value; scripts send `cli`). A page on another site
cannot add a custom header without a CORS preflight, which the hub never answers. The gate sits in `ServeHTTP` before
the route switch, so it covers every route at once (38 state-changing routes + an unknown path in
`r135_csrf_test.go`); `/login` and the public `/bind/<token>` stay exempt (no operator session to ride; the bind token
is the capability). The other choice — dropping browser-usable Basic auth entirely — was not taken: CC's headless runs
and the runbooks drive the hub with Basic auth, and the header costs them one flag.
### 16.2 Secrets at rest in `hub.db` (R-133, R-821)
Two columns are SEALED (AES-256-GCM, `enc:v1:` + nonce, one key: `OFFSITE_SECRET_KEY` from `Secret/offsite-secret-key`,
never in the database or git): the off-site sub-account passwords (`one_time_secrets.value`, v0.127.0) and, since
v0.135.0, each box's break-glass console password (`host_recovery.secret`) — the same helpers, not a second scheme.
Legacy rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, measured live: 4 rows). No key → a save is
refused; a wrong key → a reveal is a 500 with nothing in the body or the log. Both retrieval paths (the operator page and
the global-key API) open through `GetHostRecoveryCredential`, so the break-glass route still works with the UI down.
**What a copy of `hub.db` still holds readable** (R-879): each box's hub API key, each household's owner passphrase
(`customer_configs.retrieval_password`) and API key, and the PBS-DR token values (`host_pbs_secrets.value`, kept after
use). **And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only
with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export
on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md`
(R-173 — awaiting the operator's decision).
@@ -133,6 +133,16 @@ startup a record still marked running becomes a failed, interrupted result („A
restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the
precedent is kept for what it was written for.
**A backup RUN cut off by a power cut or a restart is said (R-519, controller v0.296.0, `09` decision 123).** The
same shape for the app-data run: `appdata-run.json` beside the restore record, written at both ends of a run. A start
that finds it still running turns it into a notice on /backups and /backups/apps („A legutóbbi mentés (…) megszakadt,
mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every step OK. Measured BIGNIGHT F2 (2026-09-14):
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the
operator (R-519 narrowed).*
### Lane 2 — the operator: guest and host recovery
Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator
@@ -661,6 +671,11 @@ successes only. After an agent restart the success is read back from the tier's
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both
measurements (≈6 min / 9 apps, ≈8 min / 12 apps), say "minutes, not seconds", and that the off-site copy in the same
run makes it longer.
### 6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, `09` §3 decision 36)
@@ -327,6 +327,8 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
breaks a `not_running` run.
@@ -334,9 +336,9 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
what the box cannot heal — a running container with no connection (wrong token, blocked network).
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`,
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended,
operator may reverse* (`11` §8.3).
next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`,
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`,
`OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119).
---
@@ -853,6 +853,34 @@ its length, and both fixes cost something the household would notice — operato
a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs
one call (pinned by a test). agent v0.145.0.
### 2026-10-05 (afternoon) — decided by CC — operator may reverse (the hub-safety / R-861 brief)
119. **Boxes left behind (R-604, R-530).** Options for the agent alarm's wait: 3 days (a box off for a long weekend
alarms), 7 days (one week, the same wait as the bundle-behind alarm, R-840), 14 days. **Chosen 7 days**
(`OS_ALARM_AGENT_BEHIND_AFTER`), counted from when the hub first sees the box behind. A global floor raise names,
in one operator mail, every box whose own LOWER floor it cannot move. hub v0.135.0, `05` §5.
120. **How the Basic-auth operator path is protected from cross-site POSTs (R-135).** Options: (a) drop browser-usable
Basic auth (CC's headless runs and every runbook POST break); (b) require a custom header on a cookie-less
state change (a browser cannot add one cross-site without a CORS preflight the hub never answers; scripts add one
flag). **Chosen (b)**, header `X-Felhom-Operator`. hub v0.135.0, `05` §16.1.
121. **The console password's seal (R-133).** Options: (a) a second key and scheme for `host_recovery`; (b) the off-site
seal and key already in force (R-821). **Chosen (b)** — the brief asked for the existing pattern, and one key is
one custody question. Consequence named: a database backup needs this key off DooPlex too (R-173). `05` §16.2.
122. **How the agent's admin commands are narrowed (R-861).** Options per group: (a) a root wrapper per group with its
own argv; (b) exact sudo regex patterns for every varying argument + ONE content checker for every agent-written
file root reads + fixed bundle files where the content never varies + the signed update verified as root by the
existing `felhom-os-apply`. **Chosen (b)** — fewest new root programs, and the checker is pinned to the agent's
own renderers by contract tests. Delivery order: signed `agent_update` first, then the bundle. agent v0.146.1,
`03` §3.1.
123. **How long a cut-off backup run is said on the page (R-519).** Options: until the next run of any kind (a failed
run would clear the warning), until the next run that ends with every step OK, or until dismissed. **Chosen: until
a run ends with every step OK.** controller v0.296.0.
124. **How a bundle that adds a path reaches a box (R-880).** An installed `felhom-os-apply` refuses any path not in its
own table (R16). Options: (a) copy the new wrapper onto each box by hand as root (does not scale, leaves the signed
route); (b) a STEP bundle — the box's current bundle with only `felhom-os-apply` replaced, published as
`<ver>-step1` — then the release's bundle, both by signed jobs. **Chosen (b)**, `felhom-agent/scripts/build-step-bundle.py`;
delivered to demo-hp, demo-felhom and Tester 1 on 2026-10-05. `11` §5.4.2 rule unchanged.
### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief)
100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) —
@@ -0,0 +1,8 @@
== Part A live, hub 0.135.0, 2026-10-05T09:15:34Z, ClusterIP, Basic auth (password from the credentials file, not printed)
POST /configuration/global-floor, Basic, NO header, Origin evil (empty form): 403
POST /no-such-route, Basic, NO header: 403
POST /no-such-route, Basic + X-Felhom-Operator: cli (passes the gate → router 404): 404
POST /no-such-route, header but NO credentials: 401
GET /system, Basic, no header: 200
2026/10/05 11:15:34 [WARN] CSRF rejected: POST /configuration/global-floor from 10.42.0.1:36344
2026/10/05 11:15:34 [WARN] CSRF rejected: POST /no-such-route from 10.42.0.1:41323
@@ -0,0 +1,50 @@
# Hub state-changing routes and how each is protected (hub v0.135.0, R-135)
Every route below is reached through `RequireAuth` → `ServeHTTP`; the CSRF gate is the first thing `ServeHTTP` does for any
method other than GET/HEAD/OPTIONS, before the route switch — so the protection is the same for every route, and an
unknown path is refused by the gate before it can 404.
| Route (representative path) | Protected how | Test |
|---|---|---|
| `POST /configuration` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /apps/demo/reset-telemetry` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /apps/demo/dismiss-issues` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/endpoints` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/endpoints/1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /appliances/1/bind` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /appliances/1/discard` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /hosts/h1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /hosts/h1/reveal-recovery-credential` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /hosts/h1/request-logs` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/block` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/selfbind-link` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/unblock` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/geo/disable` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/create-config` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/request-log-tail` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/new` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configuration/global-floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configuration/artifacts` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configuration/password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/edit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/offsite-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/claim-resend` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/pbsdr-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/offsite-freeze` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/regen-password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/reset` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/remove-unpinned/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/abandon-cancel/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/window-grant/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/windows-enabled` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/key-audit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/ring/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/enabled/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/approve-now` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/approve-docker` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /no-such-route` | the gate, before routing (not a route) | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /login` | exempt (no session to ride; a wrong password is 401) | — |
| `POST /bind/<token>` | exempt (public self-bind; the e-mailed URL token is the capability, rate-limited) | existing `selfbind_test.go` |
| `GET` routes | not gated by design (a GET must not change state). **Not audited in this session** for a GET that writes — the route switch sends POST-only actions to handlers that check `MethodPost`, but the GET renderers were not read line by line | `TestR135_GetIsNotGated` |
@@ -0,0 +1,6 @@
== Part B live, 2026-10-05T09:16:00Z: live hub.db copied to scratch, only prefix + length selected, copy shredded after
Tester-2-be8404|enc:v1:|87|2026-10-04 16:07:15
demo-felhom-8363b5|enc:v1:|87|2026-07-18 16:30:41
demo-hp-bb76ea|enc:v1:|87|2026-07-21 16:24:27
tester-1-d70be4|enc:v1:|87|2026-10-04 19:40:27
rows NOT sealed: 0
@@ -0,0 +1,6 @@
== Part B live reveal, 2026-10-05T09:16:15Z: POST /hosts/demo-hp-bb76ea/reveal-recovery-credential (Basic + X-Felhom-Operator), body to a 0600 scratch file, shredded after
reveal HTTP 200
username root@pam password length 32 set_at 2026-07-21T16:24:27Z
the revealed password logs in to demo-hp's Proxmox API (POST /api2/json/access/ticket, root@pam): HTTP 200
control, a wrong password: HTTP 401
2026/10/05 11:16:15 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars)
@@ -0,0 +1,42 @@
== Part C readings on DooPlex, READ ONLY, 2026-10-05T09:17:24Z
-- where the hub database lives
hub-data pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 1Gi longhorn
-rw-r--r-- 1 root root 374534144 Oct 5 11:14 hub.db
-rw-r--r-- 1 root root 32768 Oct 5 11:16 hub.db-shm
-rw-r--r-- 1 root root 313152 Oct 5 11:16 hub.db-wal
973.4M 373.4M 584.1M 39% /data
-- the exclusion label: PVC (git, manifests/hub.yaml:47, commit 868e8465 2026-02-16 'updated hub yaml', no reason given) vs the live Longhorn Volume
PVC label: disabled
Volume label: enabled
-- recurring jobs
backup-daily backup 0 4 * * * 1 [default]
backup-weekly backup 0 5 * * 0 1 [default]
-- backups of the hub volume (Longhorn backupstore)
2026-10-04T03:05:01Z Completed 708837376
2026-10-05T02:06:26Z Completed 713031680
-- backup target
nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc?nfsOptions=soft,timeo=330,retrans=3 true
-- which disk holds the target
/dev/sda1
/dev/sdb1
-- DooPlex's own backup service
Mon 2026-10-05 11:30:00 CEST 12min Mon 2026-10-05 11:15:00 CEST 2min 25s ago backup-freshness.timer backup-freshness.service
Tue 2026-10-06 03:19:15 CEST 16h Mon 2026-10-05 03:15:21 CEST 8h ago dooplex-backup.timer dooplex-backup.service
Result=success
ExecMainStatus=0
-- what tells anyone when a backup fails
80:export NOTIFY_ON_FAILURE="true"
81:# export NOTIFY_WEBHOOK_URL="https://your-webhook-url"
137: if [ "${NOTIFY_ON_FAILURE}" = "true" ] && [ -n "${NOTIFY_WEBHOOK_URL}" ]; then
140: "${NOTIFY_WEBHOOK_URL}" || true
prometheus rule backup-freshness-alerts.yml MinecraftBackupStale
prometheus rule backup-freshness-alerts.yml BackupFreshnessExporterDead
prometheus rule longhorn-alerts.yml LonghornVolumeSpaceCritical
prometheus rule longhorn-alerts.yml LonghornVolumeSpaceWarning
prometheus rule longhorn-alerts.yml LonghornVolumeDegraded
prometheus rule longhorn-alerts.yml LonghornNodeStoragePressure
(no rule watches a Longhorn BACKUP's success or age, nor dooplex-backup.service; the only backup-freshness rule is MinecraftBackupStale)
-- does anything leave DooPlex for the hub DB? the off-site route that exists today: ep0 PBS reached through felhom-ep0-pbs-tunnel (pull only, ep0 -> DooPlex)
active
/usr/bin/proxmox-backup-client
/usr/bin/sqlite3
@@ -0,0 +1,9 @@
== Part D live: GET /system on hub 0.135.0 (Basic auth); extracted, no tokens
Version floors: Version floors Global controller floor: 0.292.0 · vouched agent: 0.145.0 Customer Own floor Set Global floor moves it? demo-felhom Demo Ügyfél 0.295.0 unknown no — its own floor applies (at or above the global) demo-hp Demo HP 0.295.0 unknown no — its own floor applies (at or above the global) tester-1 Tester 1 0.295.0 unknown no — its own floor applies (at or above the global)
Agent cell Tester-2-be8404: [('c-warn', '0.142.0 → 0.145.0 (since 2026-10-05)', '3 minor releases behind — sign an agent_update for this box')]
Agent cell demo-felhom-8363b5: []
Agent cell demo-hp-bb76ea: []
Agent cell tester-1-d70be4: []
<td title="current (vouched 0.145.0)">0.145.0</td>
<td title="current (vouched 0.145.0)">0.145.0</td>
<td title="current (vouched 0.145.0)">0.145.0</td>
@@ -0,0 +1,28 @@
== R-518 measure, demo-hp guest 9201, controller 0.295.0, 2026-10-05: POST /api/guest-backup/trigger (the button's call) at 09:19:05Z
2026/10/05 09:19:07 backup_handlers.go:349: [INFO] [web] manual whole-guest backup triggered (quiesce loop)
2026/10/05 09:19:07 quiesce.go:427: [INFO] [quiesce] manual backup requested — quiescing now
2026/10/05 09:19:08 quiesce.go:517: [INFO] [quiesce] backup due on 2 tier(s) — quiescing 9 stack(s): [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin]
2026/10/05 09:19:29 quiesce.go:566: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 started — polling
2026/10/05 09:23:44 quiesce.go:210: [INFO] [quiesce] a backup cycle is already running — skipping this scheduled check
2026/10/05 09:24:09 quiesce.go:643: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 done — next tier may start (app still quiesced)
2026/10/05 09:24:09 quiesce.go:307: [INFO] [quiesce] tier felhom-pbs is BUSY — the agent refused the backup because a concurrent heavy operation holds it. This is contention, NOT a failure: the tier stays due and retries in 15m0s (contended for 0s)
2026/10/05 09:24:09 quiesce.go:504: [INFO] [quiesce] unquiescing (last tier is busy — deferring to a later cycle): restarting 9 stack(s)
-- container StartedAt after the backup (the apps the quiesce stopped):
2026-10-05T09:24:10.019684173Z adventurelog-postgres
2026-10-05T09:24:10.200085281Z adventurelog-frontend
2026-10-05T09:24:15.705236204Z adventurelog
2026-10-05T09:24:16.331987016Z bentopdf
2026-10-05T09:24:16.974996127Z bookstack-db
2026-10-05T09:24:22.669660366Z bookstack
2026-10-05T09:24:23.521923837Z calibre-web
2026-10-05T09:24:24.437773353Z docmost-postgres
2026-10-05T09:24:24.631250114Z docmost-redis
2026-10-05T09:24:35.395095325Z docmost
2026-10-05T09:24:36.934564476Z kimai-db
2026-10-05T09:24:42.753714563Z kimai
2026-10-05T09:24:43.466405175Z opengist
2026-10-05T09:24:44.478064534Z paperless-redis
2026-10-05T09:24:44.688668856Z paperless-postgres
2026-10-05T09:24:54.928089575Z paperless-webserver
2026-10-05T09:24:55.739635651Z privatebin
RESULT: stop began 09:19:08Z (9 stacks quiesced), local tier 09:19:29-09:24:09, PBS tier BUSY (skipped, retried later), last app back 09:24:55Z => longest stop 5 min 47 s, shortest ~5 min 02 s. 21/21 containers running afterwards.
@@ -0,0 +1,15 @@
09:19:20 running=15 phase=idle
09:19:41 running=4 phase=snapshotted
09:20:02 running=4 phase=snapshotted
09:20:23 running=4 phase=snapshotted
09:20:44 running=4 phase=snapshotted
09:21:05 running=4 phase=snapshotted
09:21:26 running=4 phase=snapshotted
09:21:47 running=4 phase=snapshotted
09:22:08 running=4 phase=snapshotted
09:22:29 running=4 phase=snapshotted
09:22:50 running=4 phase=snapshotted
09:23:13 running=4 phase=snapshotted
09:23:34 running=4 phase=snapshotted
09:23:55 running=4 phase=snapshotted
09:24:16 running=6 phase=done
@@ -0,0 +1,39 @@
== RED-PROOF 1 (R-519): runDBDumpsInternal does not call markRunStarted
=== RUN TestRunRecord_TheRealRunIsOnRecordWhileItRuns
run_record_test.go:78: the run was not on record while it ran — a cut here would go unnoticed
--- FAIL: TestRunRecord_TheRealRunIsOnRecordWhileItRuns (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
FAIL
== RED-PROOF 2 (R-519): the synthesised status says Success: true again
=== RUN TestRunRecord_SynthesisedStatusIsNotOKAfterACut
run_record_test.go:97: after a cut the synthesised status still reads OK: &{LastRun:2026-10-05 11:34:19.454323277 +0200 CEST m=+0.001336319 Results:[{DB:{ContainerName:adventurelog ContainerID: DBType: DBUser: DBName: StackName:adventurelog} FilePath:adventurelog-postgres.sql Size:0 Duration:0s Error:<nil> Validation:{Valid:false TableCount:0 Error: FileSize:0 ModTime:0001-01-01 00:00:00 +0000 UTC UserTableFound:false UserRows:0 LooksEmpty:false}}] Success:true Duration:0s}
--- FAIL: TestRunRecord_SynthesisedStatusIsNotOKAfterACut (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
FAIL
== RED-PROOF 3 (R-519 wiring): main() does not call loadRunRecordAtStartup
=== RUN TestMainWiresRunRecord
run_record_wiring_test.go:47: main() never calls loadRunRecordAtStartup — a cut run is never said on the page
--- FAIL: TestMainWiresRunRecord (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s
FAIL
== RED-PROOF 4 (R-518): the v0.267.0 copy (12 apps / 8 minutes only)
=== RUN TestR518_BackupButtonStatesTheMeasuredDowntime
r518_backup_downtime_copy_test.go:33: hu: "kb. 6 perc" appears 0 times, want it on the page AND in the confirm
r518_backup_downtime_copy_test.go:33: hu: "nem másodperceket" appears 0 times, want it on the page AND in the confirm
r518_backup_downtime_copy_test.go:33: en: "about 6 minutes" appears 0 times, want it on the page AND in the confirm
r518_backup_downtime_copy_test.go:33: en: "not seconds" appears 0 times, want it on the page AND in the confirm
--- FAIL: TestR518_BackupButtonStatesTheMeasuredDowntime (0.07s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s
FAIL
== restored
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.013s
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.014s
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s
@@ -0,0 +1,7 @@
== 9202 teardown of the throwaway bookstack (installed 09:35:56Z for the R-519 reproduction), 2026-10-05T11:44:31Z
stop: HTTP 200
remove: HTTP 200
containers: 0
volumes: 0
stackdir: none
backups: none
@@ -0,0 +1,9 @@
Oct 05 12:54:08 demo-felhom felhom-os-apply[4115602]: os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0
Oct 05 12:54:08 demo-felhom felhom-os-apply[4115603]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)
Oct 05 12:54:09 demo-felhom felhom-os-apply[4115702]: os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131155]: os-apply: BUNDLE START agent=0.146.1 sha=42333e969028867a authority=signed files=26 write=3 same=22 kept=1 skipped=0
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131156]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-selfupdate-guarded (replaced)
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131157]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-priv-apply (new)
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131158]: os-apply: BUNDLE WROTE /etc/sudoers.d/felhom-agent (replaced)
Oct 05 13:09:09 demo-felhom felhom-os-apply[4131255]: os-apply: BUNDLE DONE agent=0.146.1 written=3 same=22 self-check=ok signers-created=False
Oct 05 13:09:09 demo-felhom felhom-agent[4101006]: time=2026-10-05T13:09:09.974+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=67 total=67 degraded=""
@@ -0,0 +1,6 @@
Oct 05 12:42:18 demo-hp felhom-os-apply[500347]: os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0
Oct 05 12:42:19 demo-hp felhom-os-apply[500348]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)
Oct 05 12:42:19 demo-hp felhom-os-apply[500485]: os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False
Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.875+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0"
Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.875+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)"
Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.876+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False"
@@ -0,0 +1,8 @@
== demo-felhom 2026-10-05T11:09:48Z: the checker as the agent user on the real staged files (SAME = nothing changes)
unit mnt-hdd_1.mount rc=3
wg rc=0
sshd-config rc=0
sshd-key rc=0
old route: sudo: a password is required
active active active active
5
@@ -0,0 +1,20 @@
== demo-hp 2026-10-05T10:58:16Z: the checker run AS the agent user through sudo, on the real staged files (each must be SAME — nothing changes)
unit mnt-hdd_1.mount rc=0
wg rc=0
sshd-config rc=0
sshd-key rc=0
dnsmasq rc=0
== an ATTACK, live: the agent stages a unit binding its own dir over /etc/sudoers.d (name and Where agree)
attack rc=3 (installed? no)
== the old route, live: sudo -n install of a staged file
sudo: a password is required
== journal
Oct 05 12:58:08 demo-hp felhom-priv-apply[550409]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 05 12:58:09 demo-hp felhom-priv-apply[550431]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 05 12:58:16 demo-hp felhom-priv-apply[550931]: felhom-priv-apply: SAME unit /etc/systemd/system/mnt-hdd_1.mount
Oct 05 12:58:16 demo-hp felhom-priv-apply[550937]: felhom-priv-apply: SAME wg /etc/wireguard/wg-felhom.conf
Oct 05 12:58:16 demo-hp felhom-priv-apply[550962]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 05 12:58:16 demo-hp felhom-priv-apply[550968]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 05 12:58:16 demo-hp felhom-priv-apply[550981]: felhom-priv-apply: SAME dnsmasq /etc/dnsmasq.d/felhom-resolver-base.conf
Oct 05 12:58:16 demo-hp felhom-priv-apply[550992]: felhom-priv-apply: REFUSED [U3] unit mnt-..-etc-sudoers.d.mount: mnt-..-etc-sudoers.d.mount: Where=/mnt/../etc/sudoers.d is not /mnt/<name> or /mnt/felhom-drives/<name>
@@ -0,0 +1,2 @@
== LIVE AFTER, demo-felhom, 2026-10-05T11:09:46Z: bundle 0.146.1; 'sudo -l -U felhom-agent <argv>' per case (lists only)
RESULT ok=93 fail=0
@@ -0,0 +1,2 @@
== LIVE AFTER, demo-hp, 2026-10-05T10:57:55Z: bundle 0.146.1, sudo Sudo version 1.9.16p2; 'sudo -l -U felhom-agent <argv>' per case (lists only)
RESULT ok=93 fail=0
@@ -0,0 +1,28 @@
== LIVE BEFORE, demo-felhom (felhom-pve), 2026-10-05T10:54:01Z: bundle 0.145.0 — the v0.145.0 sudoers; 'sudo -l -U felhom-agent <argv>' per case (lists only, runs nothing)
FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'unit' 'mnt-felhom\x2dx.mount'
FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'dnsmasq' '/tmp/felhom-resolver-123456789.conf' 'felhom-x.conf'
FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'wg'
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda
FAIL want=DENY got=ALLOW :: /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x
FAIL want=DENY got=ALLOW :: /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d
FAIL want=DENY got=ALLOW :: /usr/bin/umount /mnt/felhom-drives/x /
FAIL want=DENY got=ALLOW :: /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount
FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount
FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh
FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh
FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf
FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf
FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config
FAIL want=DENY got=ALLOW :: /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000
FAIL want=DENY got=ALLOW :: /usr/bin/systemctl enable --now -- etc-sudoers.d.mount
FAIL want=DENY got=ALLOW :: /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd
FAIL want=DENY got=ALLOW :: /usr/bin/rmdir /mnt/felhom-drives/x /etc
FAIL want=DENY got=ALLOW :: /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset
FAIL want=DENY got=ALLOW :: /usr/sbin/smartctl -a -j /dev/sda -s off
FAIL want=DENY got=ALLOW :: /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x
FAIL want=DENY got=ALLOW :: /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller
FAIL want=DENY got=ALLOW :: /usr/sbin/pct unlock 9201 --whatever
RESULT ok=67 fail=26
@@ -0,0 +1,30 @@
== hp 2026-10-05T09:57:31Z — felhom-priv-apply --check against the box's LIVE files (read only; prints OK or the rule, never content)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-demo-hp.conf: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
== felhom-pve 2026-10-05T09:57:32Z — felhom-priv-apply --check against the box's LIVE files (read only; prints OK or the rule, never content)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
== hp 2026-10-05T10:13:26Z — v0.146.1 checker, --check against LIVE files (read only)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-demo-hp.conf: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
== felhom-pve 2026-10-05T10:13:27Z — v0.146.1 checker, --check against LIVE files (read only)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
@@ -0,0 +1,85 @@
== RED-PROOF F1 (mount units): felhom-priv-apply stops checking Where
test_U3_bind_over_sudoers_dir (__main__.Refuses.test_U3_bind_over_sudoers_dir) ... ok
test_U3_name_must_match_where (__main__.Refuses.test_U3_name_must_match_where) ... ok
test_U3_network_outside_drives (__main__.Refuses.test_U3_network_outside_drives) ... ok
test_U3_traversal_in_where (__main__.Refuses.test_U3_traversal_in_where) ... ok
OK
== RED-PROOF F2 (WireGuard): felhom-priv-apply allows any key
FAIL: test_W1_postup (__main__.Refuses.test_W1_postup)
FAILED (failures=1)
== RED-PROOF F3 (self-update, root side): felhom-os-apply agent_update skips the signature
FAILED (errors=1)
== RED-PROOF F4 (self-update, agent side): the agent calls felhom-selfupdate-guarded apply itself again
--- FAIL: TestExecutor_HappyPath (0.00s)
executor_test.go:111: execute: agent_update: the root wrapper did not apply it: <nil> (report: ; stderr: )
FAIL
== RED-PROOF F5 (guest hook): SnippetReady accepts any content
--- FAIL: TestSnippetReady (0.00s)
install_test.go:43: a hook with other content read as ready
FAIL
== RED-PROOF F6 (shared parent): the agent installs the boot script from /tmp again when it differs
--- FAIL: TestSharedParentBoot_NeverInstalls (0.00s)
intermediary_install_test.go:54: both missing: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/001/felhom-shared-parent.sh]] — it must install nothing (R-861)
intermediary_install_test.go:54: script differs: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/002/felhom-shared-parent.sh]] — it must install nothing (R-861)
intermediary_install_test.go:54: unit missing: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/003/felhom-shared-parent.sh]] — it must install nothing (R-861)
== RED-PROOF F7 (escrow, root read): a staged file is read with os.ReadFile (follows a symlink)
--- FAIL: TestAttach_RefusesASymlink (0.00s)
r861_staged_read_test.go:24: a symlinked staged file was read: ok=true err=<nil> value-set=true
FAIL
== RED-PROOF F8 (network shares): nosuid,nodev dropped from the NFS options
--- FAIL: TestPrivApply_AcceptsTheRenderedUnits (0.35s)
r861_privapply_contract_test.go:31: media .mount: REFUSED [U5] mnt-felhom\x2ddrives-media.mount: a network share must carry nosuid,nodev
FAIL
== RED-PROOF F9 (the exact patterns): the v0.145.0 sudoers under the injection test
injections the old file allows (Go matcher): 23
== restored — the same tests green
ok gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/guesthook (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/storage 0.474s
ok gitea.dooplex.hu/admin/felhom-agent/internal/capability 0.177s
OK
OK
== RED-PROOF F1 (re-run): the first run did NOT convict — the name check (escape(Where)==name) masked it. The test now uses the
pair that only the Where rule stops: name mnt-..-etc.mount + Where=/mnt/../etc (= /etc). Mutation: stop checking Where
FAIL: test_U3_traversal_in_where (__main__.Refuses.test_U3_traversal_in_where)
FAILED (failures=1)
== RED-PROOF F3 (re-run, clean assertion): felhom-os-apply skips the signature
FAIL: test_a_bad_signature_never_reaches_the_wrapper (__main__.AgentUpdate.test_a_bad_signature_never_reaches_the_wrapper)
AssertionError: None is not true : a job whose signature does not verify was NOT refused: {'agent_update': {'sha256': 'd76b02acf626ce399da7e0a9e17b35563227a4831e14f5edca4ab7cf89eb2c79', 'version': '0.146.0', 'wrapper': '', 'wrapper_rc': 0}, 'layer': 'host', 'mode': 'agent_update', 'pass_seconds': 0.0, 'refused': None, 'release_id': 'agent-0.146.0', 'vmid': 0}
FAILED (failures=1)
== restored
OK
OK
=== Review findings 2026-10-05 (background security review of commit 6ab1e7c) — fixed in v0.146.1, each red-proved
== RED-PROOF S1 (TOCTOU): the wrapper gets the agent's path again (hash, then copy by path)
FAIL: test_signed_update_flips_and_burns_the_nonce (__main__.AgentUpdate.test_signed_update_flips_and_burns_the_nonce)
FAILED (failures=1)
== RED-PROOF S1b: the A/B wrapper accepts the agent's staging dir again
FAIL: test_the_agents_staging_dir_is_refused (__main__.SelfupdateWrapperConfinement.test_the_agents_staging_dir_is_refused)
FAILED (failures=1)
== RED-PROOF S2 (allowlist escape): [Unit] accepts Wants=/Requires=/Before= again
FAIL: test_U2_wants_starts_another_unit (__main__.Refuses.test_U2_wants_starts_another_unit)
FAILED (failures=1)
== RED-PROOF S3 (path traversal): open the whole path with O_NOFOLLOW only
--- FAIL: TestAttach_RefusesASymlinkedDirectory (0.00s)
r861_staged_read_test.go:54: a key behind a symlinked directory was read: ok=true err=<nil>
FAIL
== restored
OK
OK
ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow 0.008s
@@ -0,0 +1,4 @@
== the R-880 step bundle, 2026-10-05T10:24:08Z: base = felhom-agent/0.145.0/felhom-config-bundle.json (sha 78c00adc…, what demo-hp, demo-felhom, tester-1 run); built by scripts/build-step-bundle.py at agent e4b5cf9; published as felhom-agent/0.146.1-step1/felhom-config-bundle.json
step sha 8482851ec27030a7615216048338b1c8939d4e353023ad11c66e6f8930af8613
round trip sha 8482851ec27030a7615216048338b1c8939d4e353023ad11c66e6f8930af8613
same paths: True changed: ['/usr/local/sbin/felhom-os-apply'] version: 0.146.1-step1
@@ -0,0 +1,129 @@
== R-861 real-sudo proof, sudo 1.9.16p2 (debian:trixie throwaway container on DooPlex, 2026-10-05T09:58:38Z); 'sudo -l -U felhom-agent <argv>' per case
-- NEW sudoers (agent v0.146.0): every capability must be ALLOW, every attack DENY
ok ALLOW '/usr/bin/lxc-info' '-n' '9201' '-p' '-H'
ok ALLOW '/usr/bin/mount' '--bind' '/mnt/felhom-drives' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mount' '--make-shared' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mount' '--make-private' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mount' '--bind' '/mnt/felhom-usb/felhom-data' '/mnt/felhom-drives/felhom-usb'
ok ALLOW '/usr/bin/umount' '/mnt/felhom-drives/felhom-usb'
ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-drives/felhom-usb'
ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-usb/felhom-data'
ok ALLOW '/usr/bin/chown' '100000:100000' '/mnt/felhom-usb/felhom-data'
ok ALLOW '/usr/bin/systemctl' 'enable' 'felhom-shared-parent.service'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '-mp8' '/mnt/felhom-drives,mp=/mnt/felhom-drives'
ok ALLOW '/usr/sbin/blkid' '-p' '-o' 'export' '/dev/sda'
ok ALLOW '/usr/bin/lsblk' '-J' '-o' 'NAME,FSTYPE,PTTYPE,MOUNTPOINT' '/dev/sda'
ok ALLOW '/usr/local/sbin/felhom-mkfs-guarded' '/dev/sda' 'ext4'
ok ALLOW '/usr/local/sbin/felhom-mkfs-guarded' '/dev/sda' 'xfs'
ok ALLOW '/usr/sbin/smartctl' '-a' '-j' '/dev/sda'
ok ALLOW '/usr/sbin/lvs' '--reportformat' 'json' '--units' 'b' '-o' 'lv_name,data_percent,metadata_percent' '--' 'pve/data'
ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'unit' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'daemon-reload'
ok ALLOW '/usr/bin/systemctl' 'enable' '--now' '--' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'disable' '--' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'stop' '--' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'reset-failed' '--' 'mnt-felhom\x2ddrives-media.automount'
ok ALLOW '/usr/bin/rmdir' '/mnt/felhom-drives/media'
ok ALLOW '/usr/bin/systemctl' 'start' 'networking.service'
ok ALLOW '/usr/bin/chown' '-R' '100000:100000' '/var/lib/felhom-agent/guests/9201'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '-mp0' '/var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '-onboot' '1'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '--hookscript' 'local:snippets/felhom-guest-hook.sh'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '--delete' 'mp0'
ok ALLOW '/usr/sbin/pct' 'reboot' '9201'
ok ALLOW '/usr/bin/apt-get' 'install' '-y' '-q' 'dnsmasq'
ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'dnsmasq' '/tmp/felhom-resolver-123456789.conf' 'felhom-x.conf'
ok ALLOW '/usr/bin/systemctl' 'enable' '--now' 'dnsmasq'
ok ALLOW '/usr/local/sbin/felhom-os-apply' '--plan' '/var/lib/felhom-agent/os/plan-x.json'
ok ALLOW '/usr/bin/systemctl' 'reload' 'dnsmasq'
ok ALLOW '/usr/bin/systemctl' 'restart' 'dnsmasq'
ok ALLOW '/usr/bin/rm' '-f' '/etc/dnsmasq.d/felhom-x.conf'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'ip' '-4' '-o' 'addr' 'show' 'dev' 'eth0'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'exec' 'felhom-controller' 'cat' '/opt/docker/felhom-controller/controller.yaml'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'ip' 'route' 'show' 'default'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'cat' '/etc/network/interfaces'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'pgrep' '-x' 'dhclient'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'dhclient' '-pf' '/run/dhclient.eth0.pid' '-lf' '/var/lib/dhcp/dhclient.eth0.leases' 'eth0'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'cat' '/etc/felhom-controller-image'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'image' 'inspect' 'gitea.dooplex.hu/admin/felhom-controller:0.0.0'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'inspect' '-f' '{{.State.Running}}' 'felhom-controller'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'systemctl' 'restart' 'felhom-controller-bootstrap.service'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'tee' '/etc/felhom-controller-image'
ok ALLOW '/usr/sbin/pct' 'unlock' '9201'
ok ALLOW '/usr/bin/apt-get' 'install' '-y' '-q' 'wireguard-tools'
ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'wg'
ok ALLOW '/usr/bin/systemctl' 'enable' '--now' 'wg-quick@wg-felhom'
ok ALLOW '/usr/bin/systemctl' 'restart' 'wg-quick@wg-felhom'
ok ALLOW '/usr/bin/systemctl' 'disable' '--now' 'wg-quick@wg-felhom'
ok ALLOW '/usr/bin/wg' 'show' 'wg-felhom' 'latest-handshakes'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'create' 'felhom-pbs' '10.77.0.1' 'felhom-offsite' 'ns0' 'felhom@pbs!ns0' '00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00' '/etc/pve/priv/storage'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'reconcile' 'felhom-pbs' '10.77.0.1' 'ns0' 'felhom@pbs!ns0' '00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00' '/etc/pve/priv/storage'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'grant' 'felhom-pbs'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'read' 'felhom-pbs' '/etc/pve/priv/storage'
ok ALLOW '/usr/local/bin/felhom-agent' '--config' '/etc/felhom-agent/agent.json' '--selftest=escrow-create' '--upload' '--output=json'
ok ALLOW '/usr/local/sbin/felhom-selfupdate-guarded' 'commit'
ok ALLOW '/usr/local/sbin/felhom-selfupdate-guarded' 'rollback'
ok DENY /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1
ok DENY /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives
ok DENY /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda
ok DENY /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda
ok DENY /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x
ok DENY /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d
ok DENY /usr/bin/umount /mnt/felhom-drives/x /
ok DENY /usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow
ok DENY /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount
ok DENY /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount
ok DENY /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh
ok DENY /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh
ok DENY /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf
ok DENY /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf
ok DENY /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config
ok DENY /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000
ok DENY /usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service
ok DENY /usr/bin/systemctl enable --now -- etc-sudoers.d.mount
ok DENY /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd
ok DENY /usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow
ok DENY /usr/bin/rmdir /mnt/felhom-drives/x /etc
ok DENY /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset
ok DENY /usr/sbin/smartctl -a -j /dev/sda -s off
ok DENY /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x
ok DENY /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller
ok DENY /usr/sbin/pct unlock 9201 --whatever
ok DENY /usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount
ok DENY /usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf
ok DENY /usr/local/sbin/felhom-priv-apply wg /etc/shadow
rc=0
-- OLD sudoers (agent v0.145.0), the same attacks (this side is the red-proof: 'FAIL want=ALLOW got=DENY' means the OLD file already refused that one; 'ok ALLOW' means the old file let it through)
ok ALLOW /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1
ok ALLOW /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives
ok ALLOW /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda
ok ALLOW /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda
ok ALLOW /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x
ok ALLOW /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d
ok ALLOW /usr/bin/umount /mnt/felhom-drives/x /
FAIL want=ALLOW got=DENY :: /usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow
ok ALLOW /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount
ok ALLOW /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount
ok ALLOW /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh
ok ALLOW /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh
ok ALLOW /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf
ok ALLOW /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf
ok ALLOW /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config
ok ALLOW /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000
FAIL want=ALLOW got=DENY :: /usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service
ok ALLOW /usr/bin/systemctl enable --now -- etc-sudoers.d.mount
ok ALLOW /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd
FAIL want=ALLOW got=DENY :: /usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow
ok ALLOW /usr/bin/rmdir /mnt/felhom-drives/x /etc
ok ALLOW /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset
ok ALLOW /usr/sbin/smartctl -a -j /dev/sda -s off
ok ALLOW /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x
ok ALLOW /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller
ok ALLOW /usr/sbin/pct unlock 9201 --whatever
FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount
FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf
FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply wg /etc/shadow
rc=1
@@ -0,0 +1,11 @@
== hub System page after delivery, 2026-10-05T11:43:34Z: per box — agent cell, root-files cell (raw)
Tester-2-be8404 | agent: 0.142.0 → 0.146.1 (since 2026-10-05) | root files: []
demo-felhom-8363b5 | agent: 0.146.1 | root files: ['0.146.1']
demo-hp-bb76ea | agent: 0.146.1 | root files: ['0.146.1']
tester-1-d70be4 | agent: 0.146.1 | root files: ['0.146.1']
tester-1-d70be4: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read
demo-hp-bb76ea: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read
demo-felhom-8363b5: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read
tester-1-d70be4: capability rows 66 {'ok': 66} degraded: []
demo-hp-bb76ea: capability rows 66 {'ok': 66} degraded: []
demo-felhom-8363b5: capability rows 67 {'ok': 67} degraded: []
@@ -0,0 +1,7 @@
== floors 2026-10-05T10:24:56Z: POST /customers/<id>/floor min_controller_version=0.296.0 min_agent=0.131.0
demo-hp: Location: /customers/demo-hp?flash=floor_set
demo-felhom: Location: /customers/demo-felhom?flash=floor_set
tester-1: Location: /customers/tester-1?flash=floor_set
2026/10/05 12:24:56 [INFO] Customer demo-hp controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0")
2026/10/05 12:24:57 [INFO] Customer demo-felhom controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0")
2026/10/05 12:24:57 [INFO] Customer tester-1 controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0")
@@ -0,0 +1,10 @@
== agent_update 0.146.1 (sha badd6c9a…) signed with felhom-op-1, ttl 45m, 2026-10-05T10:25:10Z
-- demo-hp-bb76ea
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-hp-bb76ea-agent_update.json
uploaded signed op to the hub jobs queue
-- demo-felhom-8363b5
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-felhom-8363b5-agent_update.json
uploaded signed op to the hub jobs queue
-- tester-1-d70be4
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-tester-1-d70be4-agent_update.json
uploaded signed op to the hub jobs queue
@@ -0,0 +1,13 @@
== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), demo-hp-bb76ea, 2026-10-05T10:35:43Z
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-hp-bb76ea-agent_config_update.json
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1 (bundle sha 42333e96…), demo-hp-bb76ea, 2026-10-05T10:42:50Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), demo-felhom-8363b5, 2026-10-05T10:53:36Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), tester-1-d70be4, 2026-10-05T10:53:36Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1 (bundle sha 42333e96…), demo-felhom-8363b5, 2026-10-05T10:58:57Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1 (bundle sha 42333e96…), tester-1-d70be4, 2026-10-05T11:13:04Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,4 @@
== vouch 2026-10-05T10:24:18Z: POST /configuration/artifacts (Basic + X-Felhom-Operator), agent 0.146.1, golden 0.296.0, min_agent 0.131.0
HTTP/1.1 303 See Other
Location: /configuration?flash=artifacts_set
2026/10/05 12:24:45 [INFO] Artifact manifest set: agent=0.146.1 golden=0.296.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="42333e969028867ad8142335e6c1bc4040eec231de0d8d330c2d4b2cf7bc3442"
+16
View File
@@ -26,6 +26,22 @@
---
## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124)
The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session).
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-135** | **A cookie-less POST skipped the hub's CSRF gate, so a browser with cached Basic credentials could be made to POST cross-site.** hub v0.135.0: without a session a state change needs Basic credentials AND the header `X-Felhom-Operator` (decision 120); the gate sits before the route switch. 39 paths through RequireAuth→ServeHTTP; red-proof: the old shape lets all 39 through. Live: Basic + no header → 403 (also with `Origin: evil`, also on an unknown path); with the header → passes; header without credentials → 401. **Reasoning kept: a browser cannot add a custom header cross-site without a CORS preflight, which the hub never answers.** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partA/`; `web/r135_csrf_test.go` |
| **R-133** | **Every box's break-glass console password was plaintext in hub.db.** hub v0.135.0: sealed with the off-site seal and key (decision 121); legacy rows sealed at start-up — live: 4 rows sealed, 0 left plain; the demo-hp reveal still returned a password that minted a PVE ticket (HTTP 200; a wrong one 401); a wrong key → 500, nothing in the body or the log, no event. **Reasoning kept: the running hub still holds the key — this closes the database-copy route only; a database backup without `OFFSITE_SECRET_KEY` cannot open the console passwords (R-173).** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partB/`; `store/r133_recovery_seal_test.go`, `web/r133_reveal_wrongkey_test.go` |
| **R-604** | **A per-customer controller floor silently kept a box out of every global raise (demo-hp missed four).** hub v0.135.0: a global raise logs one line per customer whose own LOWER floor wins and sends ONE operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set; the System page's "Version floors" table lists every per-customer floor with its age and which ones the global cannot move. 2 red-proofs. Live: the table shows the three per-customer floors (age "unknown" — set before v0.135.0). The mail was not exercised live (it needs a global raise below an override). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `web/r604_floor_held_back_test.go` |
| **R-530** | **Nothing listed which boxes still run an old agent (agents update only by a per-box signed job).** hub v0.135.0: the System page's Agent cell (box → vouched, how far, since when; red after the wait) and `agent_behind` after 7 days (decision 119). Live: Tester 2 reads `0.142.0 → 0.146.1`. Signing stays per box (the 2026-09-16 ruling: CC may sign until the first paying customer). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `osupdates/r530_agent_alarm_test.go` |
| **R-508** | **Customer tester-1 had no registered e-mail, and the page did not say so.** The address has been set since 2026-09-14 (the connect mails reach it — Gmail-read 2026-10-05); hub v0.135.0 adds the page warning: a configured customer with no box and no e-mail shows a red line (three branches tested, red-proof). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partG/`; `web/r508_no_email_banner_test.go` |
| **R-509** | **A box installed for an existing customer never got the connect e-mail.** Fixed in hub v0.114.0; the owed real-mail proof: three mails from the automatic "host delete" trigger, each within 1 s of the hub's own send line (2026-09-16 12:22:59 and 18:17:46, 2026-09-30 07:23:03 UTC), read through the Gmail connector (metadata only). The "e-mail set" trigger shares the send core and is unit-proven. | CLOSED 2026-10-05 — VERIFIED | `audits/hub-safety-2026-10-05/partG/r509-real-mails.txt` |
| **R-880** | **An installed `felhom-os-apply` refuses a bundle naming a path it does not know (R16), so a release whose bundle ADDS a path cannot reach any box on an older bundle** (found 2026-10-05 before delivering v0.146.1, which adds four). Fixed by a step: `felhom-agent/scripts/build-step-bundle.py` — the box's current bundle with ONLY `felhom-os-apply` replaced (same paths), published as `0.146.1-step1`; then the release's bundle. Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Live: demo-hp, demo-felhom and Tester 1 each took step1 (`written=1 same=20`) then 0.146.1 (`written=3 same=22`), self-check ok. **Reasoning kept: every future bundle that adds a path needs this step (decision 124); the step package stays published while any box may still be on the old bundle (Tester 2).** | CLOSED 2026-10-05 — FIXED (tooling, agent e4b5cf9) | `audits/hub-safety-2026-10-05/part{F,H}/`; memory `bundle-adding-a-path-needs-step-bundle` |
---
## 2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118)
| Row | What | Closed | Evidence |
File diff suppressed because one or more lines are too long
@@ -0,0 +1,129 @@
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
## 1. What is true today (measured 2026-10-05)
| Question | Answer |
|---|---|
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
| What a copy is worth without the key | The sealed columns (console passwords, off-site passwords) are useless without `OFFSITE_SECRET_KEY`. **The key lives only in `Secret/offsite-secret-key` on DooPlex** (and in the GPG secrets export on the same disk). A backup off DooPlex without a key off DooPlex restores a hub that cannot open any console password. |
## 2. The plan (option A — my pick): a nightly, encrypted, consistent copy on ep0's PBS
ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches it through `felhom-ep0-pbs-tunnel`
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
Without these, every later step backs up something nobody can open after a DooPlex loss.
```bash
# 1. the hub's seal key → the operator's password manager (never a file, never a chat)
sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.OFFSITE_SECRET_KEY}' | base64 -d; echo
# 2. (after Step 2) the backup encryption key's paper copy → the password manager
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
```
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
```bash
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
# namespace for operator data, apart from the households' namespaces
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
```
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
hub image, so the copy must be made by the hub itself.
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
```bash
#!/bin/sh -eu
# /usr/local/sbin/felhom-hub-db-backup — push the newest hub snapshot to ep0 (R-173). Root, 0755.
STAGE=/var/lib/felhom-hub-backup/stage; mkdir -p "$STAGE"; chmod 700 "$STAGE"
SNAP=$(kubectl -n felhom-system exec deploy/hub -- sh -c 'ls -1t /data/snapshots/hub-*.db | head -1')
kubectl -n felhom-system exec deploy/hub -- cat "$SNAP" > "$STAGE/hub.db"
sqlite3 -readonly "$STAGE/hub.db" 'PRAGMA integrity_check' | grep -qx ok # never push a broken copy
export PBS_PASSWORD_FILE=/etc/felhom-hub-backup/token PBS_FINGERPRINT=<ep0 cert fingerprint, as in ep0-datastore-copy.md>
proxmox-backup-client backup hubdb.pxar:"$STAGE" --ns operator --backup-id dooplex-hub \
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
shred -u "$STAGE/hub.db"
# the positive signal the alarm reads (written ONLY on success):
echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ \
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
```
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
```bash
T=$(mktemp -d); chmod 700 "$T"
proxmox-backup-client restore "host/dooplex-hub/$(newest snapshot)" hubdb.pxar "$T" --ns operator \
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
sqlite3 -readonly "$T/hub.db" 'PRAGMA integrity_check' | grep -qx ok
test "$(sqlite3 -readonly "$T/hub.db" 'SELECT COUNT(*) FROM hosts')" -gt 0
test "$(sqlite3 -readonly "$T/hub.db" "SELECT COUNT(*) FROM host_recovery WHERE secret NOT LIKE 'enc:v1:%'")" -eq 0
shred -u "$T/hub.db"*; rmdir "$T"
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
```
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
```yaml
- alert: HubDBBackupStale
expr: time() - felhom_hub_db_backup_last_success_timestamp_seconds > 26*3600 or absent(felhom_hub_db_backup_last_success_timestamp_seconds)
for: 30m
labels: {severity: critical}
annotations: {summary: "The hub database has not reached ep0 for 26 h (R-173)"}
- alert: HubDBRestoreTestStale
expr: time() - felhom_hub_db_restore_test_last_success_timestamp_seconds > 8*24*3600 or absent(felhom_hub_db_restore_test_last_success_timestamp_seconds)
for: 1h
labels: {severity: warning}
```
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
log is not a success.
### Step 7 — prove it once (CC, with the operator's go)
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
start it again.
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
Same Steps 0, 1, 3, 5, 6; the push is `restic backup` to a new sub-account with the R-820 append-only key pin. Costs a
new sub-account and its own key custody; the Storage Box sub-account shell can `rm` (memory: storagebox-subaccount-shell)
unless the pin is right. ep0 already has the server-side prune and the return copy, so A is less new machinery.
+18
View File
@@ -31,6 +31,24 @@ sign; no bundle may add, remove or change them. A box that has no signers file g
4. **Undo** = send the previous release's bundle the same way. The previous copies also stay on the box in
`/var/lib/felhom-os-apply/bundle-prev/<time>-before-<version>/` (the last 3).
## A release whose bundle ADDS a path — the step bundle (R-880, decision 124)
The box's INSTALLED `felhom-os-apply` checks every path of an incoming bundle against its OWN table (R16). So when a
release adds a path (agent v0.146.1 added four), every box on an older bundle refuses it. Send a step first:
```bash
# the bundle the boxes run now — check its sha against the hub's Root files / config-bundle record
curl -fsS -o base.json https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/<old>/felhom-config-bundle.json
python3 felhom-agent/scripts/build-step-bundle.py base.json <new>-step1 step.json # prints the step sha
curl -u admin:<token from a file> -X PUT --upload-file step.json \
https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/<new>-step1/felhom-config-bundle.json # 201
# per box: agent_update <new> → agent_config_update <new>-step1 (step sha) → agent_config_update <new> (release sha)
```
The step is the old bundle with ONLY `felhom-os-apply` replaced, so the old wrapper accepts it (`written=1 same=20`);
the new wrapper then accepts the release's bundle. Done this way on demo-hp, demo-felhom and Tester 1 on 2026-10-05
(`audits/hub-safety-2026-10-05/partH/`). Keep the step package while any box may still be on the old bundle.
## A box from before agent v0.143.0 — the ONE by-hand step (bootstrap)
Such a box's `felhom-os-apply` has no bundle mode, and no signed job can write a root file there (that gap IS
@@ -0,0 +1,5 @@
== round trip 2026-10-05T10:23:04Z: anonymous GET .../generic/felhom-golden/0.296.0/golden.tar.zst
HTTP 200
bytes 648611216
sha256 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
printed 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
@@ -0,0 +1,62 @@
# Golden 0.296.0 — bake + publish + vouch, 2026-10-05 (late afternoon)
Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–5, in the drill VM on DooPlex.
| | Previous (`../golden-0.295.0-2026-10-05/`) | This bake |
|---|---|---|
| `build-golden.sh` | v3.2.0 (sha256 `645b3b659cba…`) | same file, unchanged (agent repo `configs/build-golden.sh`, sha256 `645b3b659cba…`) |
| Controller | `felhom-controller:0.295.0` | **`felhom-controller:0.296.0`** (MinAgent 0.131.0, unchanged) |
| Docker engine | the operator-approved set `os-docker-20261004-142842` | same pinned set (still the only approved release — read from the hub's System page) |
| Guest packages | template | template — `GOLDEN_GUEST_PKGS` EMPTY (no guest release approved) |
## Launch
- Drill VM reverted to `virgin` (no qemu running before), cold-booted per §4.0 at 10:16:31 UTC; `pveversion` = `pve-manager/9.2.2`.
- `pveam update` → `update successful`; template `debian-13-standard_13.6-1_amd64.tar.zst`, `checksum verified`.
- `/root/bake-run.sh` reads the token from the file; launched as transient unit `golden-bake` at 10:17:29 UTC.
- Token copied file → file (`scp`). `systemctl show golden-bake -p Environment -p ExecStart | grep -c -F <token>` = **0**
(control with the token appended = **1**).
## Pass markers (from `bake.log`)
```
[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docke
[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a
[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49
docker OK (overlay2; data-root /var/lib/docker)
live-restore: on
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.296.0
GOLDEN_SHA256=65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
```
No `excluding` and no `FATAL` in the log (grep count 0).
## Round trip
== round trip 2026-10-05T10:23:04Z: anonymous GET .../generic/felhom-golden/0.296.0/golden.tar.zst
HTTP 200
bytes 648611216
sha256 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
printed 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
## Secrets
Saved-log leak grep for the literal token: **0**; positive control (a throwaway copy with the token appended): **1**,
copy shredded.
## Vouch (step 5)
`POST /configuration/artifacts` (operator Basic auth + `X-Felhom-Operator`, hub v0.135.0): agent **0.146.1**, golden
**0.296.0**, `min_agent` **0.131.0**, wrapper sha empty → `303 flash=artifacts_set`; hub log `Artifact manifest set:
agent=0.146.1 golden=0.296.0 min_agent="0.131.0" … bundle_sha="42333e96…"` (`../../audits/hub-safety-2026-10-05/partH/vouch.txt`).
Per-customer controller floors 0.296.0 for demo-hp, demo-felhom, tester-1; the global floor unchanged (Tester 2 not moved).
## Teardown
`pct destroy 9100 --purge`; `shred -u` of the token, the runner script and the log in the VM (log copied off first);
`poweroff`; qemu gone (`ps -eo comm | grep -c qemu-system-x86` = 0); `qemu-img snapshot -a virgin`. Host: nothing
provisioned.
@@ -0,0 +1,340 @@
[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.296.0
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
Logical volume "vm-9100-disk-0" created.
Logical volume pve/vm-9100-disk-0 changed.
Creating filesystem with 8388608 4k blocks and 2097152 inodes
Filesystem UUID: bcf15055-fc5b-43d8-97a3-2a296b616c79
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624
Logical volume "vm-9100-disk-1" created.
Logical volume pve/vm-9100-disk-1 changed.
Creating filesystem with 6291456 4k blocks and 1572864 inodes
Filesystem UUID: 66ff4da9-6a06-4b83-b517-65d331fa9348
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
Total bytes read: 553512960 (528MiB, 125MiB/s)
Detected container architecture: amd64
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
done: SHA256:e5l7hmAwZSLGn/EaSW0I9b1I7uhRo+qmUPfmXZxKops root@felhom-golden
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:6r+bvSA9WxcR3VbYvFsE665lF6r2EP/B3L5NOSOT2VA root@felhom-golden
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
done: SHA256:1n9DFGQdLzJEq0ZdG4SbKoTRctiwtQCnh9qStQmfXRI root@felhom-golden
[golden] starting + installing Docker (official repo, trixie channel) …
[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
installed: containerd.io 2.3.6-1~debian.13~trixie
installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie
installed: docker-ce 5:29.8.2-1~debian.13~trixie
installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie
installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie
installed: docker-compose-plugin 5.6.0-1~debian.13~trixie
[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a
[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
Unable to find image 'hello-world:latest' locally
latest: Pulling from library/hello-world
4f55086f7dd0: Pulling fs layer
4f55086f7dd0: Verifying Checksum
4f55086f7dd0: Download complete
4f55086f7dd0: Pull complete
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
Status: Downloaded newer image for hello-world:latest
docker OK (overlay2; data-root /var/lib/docker)
live-restore: on
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.296.0 (no registry cred at deploy) …
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
0.296.0: Pulling from admin/felhom-controller
774043ccc8cc: Pulling fs layer
ab6b448d4be9: Pulling fs layer
23a5bfa58353: Pulling fs layer
862a57157567: Pulling fs layer
a131840ac86e: Pulling fs layer
a316f627328b: Pulling fs layer
862a57157567: Waiting
a131840ac86e: Waiting
a316f627328b: Waiting
23a5bfa58353: Verifying Checksum
23a5bfa58353: Download complete
862a57157567: Verifying Checksum
862a57157567: Download complete
774043ccc8cc: Verifying Checksum
774043ccc8cc: Download complete
a131840ac86e: Verifying Checksum
a131840ac86e: Download complete
a316f627328b: Verifying Checksum
a316f627328b: Download complete
ab6b448d4be9: Verifying Checksum
ab6b448d4be9: Download complete
774043ccc8cc: Pull complete
ab6b448d4be9: Pull complete
23a5bfa58353: Pull complete
862a57157567: Pull complete
a131840ac86e: Pull complete
a316f627328b: Pull complete
Digest: sha256:e08cbc5868e6b72f82fa9eadba9db03c174c18e6621c044b57c99e9479104203
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.296.0
gitea.dooplex.hu/admin/felhom-controller:0.296.0
[golden] asking the controller which infra images it manages …
[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
v3.7.13: Pulling from library/traefik
e2de96513ba9: Pulling fs layer
b686a4f73445: Pulling fs layer
78cb21c375ca: Pulling fs layer
acb2f33459b1: Pulling fs layer
acb2f33459b1: Waiting
e2de96513ba9: Verifying Checksum
e2de96513ba9: Download complete
b686a4f73445: Verifying Checksum
b686a4f73445: Download complete
acb2f33459b1: Verifying Checksum
acb2f33459b1: Download complete
78cb21c375ca: Verifying Checksum
78cb21c375ca: Download complete
e2de96513ba9: Pull complete
b686a4f73445: Pull complete
78cb21c375ca: Pull complete
acb2f33459b1: Pull complete
Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0
Status: Downloaded newer image for traefik:v3.7.13
docker.io/library/traefik:v3.7.13
2026.9.3: Pulling from cloudflare/cloudflared
2cc7ee286bf3: Pulling fs layer
c172f21841df: Pulling fs layer
218cf840d0d9: Pulling fs layer
f6069939f718: Pulling fs layer
d6b1b89eccac: Pulling fs layer
2780920e5dbf: Pulling fs layer
7c12895b777b: Pulling fs layer
3214acf345c0: Pulling fs layer
52630fc75a18: Pulling fs layer
dd64bf2dd177: Pulling fs layer
b839dfae01f6: Pulling fs layer
ebddc55facdc: Pulling fs layer
c4bc6f35ff5e: Pulling fs layer
b96fe2995f90: Pulling fs layer
58c0c263dc73: Pulling fs layer
bd8962e29291: Pulling fs layer
cac2ae0193cb: Pulling fs layer
f0383d5ebc47: Pulling fs layer
dd64bf2dd177: Waiting
b839dfae01f6: Waiting
ebddc55facdc: Waiting
c4bc6f35ff5e: Waiting
b96fe2995f90: Waiting
58c0c263dc73: Waiting
bd8962e29291: Waiting
cac2ae0193cb: Waiting
f0383d5ebc47: Waiting
2780920e5dbf: Waiting
7c12895b777b: Waiting
3214acf345c0: Waiting
52630fc75a18: Waiting
f6069939f718: Waiting
d6b1b89eccac: Waiting
2cc7ee286bf3: Download complete
c172f21841df: Download complete
218cf840d0d9: Verifying Checksum
218cf840d0d9: Download complete
f6069939f718: Verifying Checksum
f6069939f718: Download complete
d6b1b89eccac: Verifying Checksum
d6b1b89eccac: Download complete
2780920e5dbf: Verifying Checksum
2780920e5dbf: Download complete
7c12895b777b: Verifying Checksum
7c12895b777b: Download complete
3214acf345c0: Verifying Checksum
3214acf345c0: Download complete
2cc7ee286bf3: Pull complete
52630fc75a18: Verifying Checksum
52630fc75a18: Download complete
dd64bf2dd177: Verifying Checksum
dd64bf2dd177: Download complete
b839dfae01f6: Verifying Checksum
b839dfae01f6: Download complete
ebddc55facdc: Verifying Checksum
ebddc55facdc: Download complete
c4bc6f35ff5e: Download complete
58c0c263dc73: Verifying Checksum
58c0c263dc73: Download complete
bd8962e29291: Verifying Checksum
bd8962e29291: Download complete
c172f21841df: Pull complete
b96fe2995f90: Verifying Checksum
b96fe2995f90: Download complete
cac2ae0193cb: Download complete
f0383d5ebc47: Verifying Checksum
f0383d5ebc47: Download complete
218cf840d0d9: Pull complete
f6069939f718: Pull complete
d6b1b89eccac: Pull complete
2780920e5dbf: Pull complete
7c12895b777b: Pull complete
3214acf345c0: Pull complete
52630fc75a18: Pull complete
dd64bf2dd177: Pull complete
b839dfae01f6: Pull complete
ebddc55facdc: Pull complete
c4bc6f35ff5e: Pull complete
b96fe2995f90: Pull complete
58c0c263dc73: Pull complete
bd8962e29291: Pull complete
cac2ae0193cb: Pull complete
f0383d5ebc47: Pull complete
Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c
Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3
docker.io/cloudflare/cloudflared:2026.9.3
1.5.6-stable: Pulling from gtstef/filebrowser
55afa1ecc21d: Pulling fs layer
8ed8f35f8d4f: Pulling fs layer
989b226a579c: Pulling fs layer
660aeead31d5: Pulling fs layer
4f4fb700ef54: Pulling fs layer
adce24567e4c: Pulling fs layer
f17ea56b313b: Pulling fs layer
6b6f3b3efe88: Pulling fs layer
4ed1ca4f3fce: Pulling fs layer
e6fc9c6a5757: Pulling fs layer
d47782d1182a: Pulling fs layer
4f4fb700ef54: Waiting
adce24567e4c: Waiting
f17ea56b313b: Waiting
6b6f3b3efe88: Waiting
4ed1ca4f3fce: Waiting
e6fc9c6a5757: Waiting
d47782d1182a: Waiting
660aeead31d5: Waiting
55afa1ecc21d: Verifying Checksum
55afa1ecc21d: Download complete
660aeead31d5: Verifying Checksum
660aeead31d5: Download complete
4f4fb700ef54: Verifying Checksum
4f4fb700ef54: Download complete
8ed8f35f8d4f: Verifying Checksum
8ed8f35f8d4f: Download complete
55afa1ecc21d: Pull complete
adce24567e4c: Verifying Checksum
adce24567e4c: Download complete
f17ea56b313b: Verifying Checksum
f17ea56b313b: Download complete
6b6f3b3efe88: Verifying Checksum
6b6f3b3efe88: Download complete
4ed1ca4f3fce: Verifying Checksum
4ed1ca4f3fce: Download complete
d47782d1182a: Verifying Checksum
d47782d1182a: Download complete
e6fc9c6a5757: Verifying Checksum
e6fc9c6a5757: Download complete
989b226a579c: Verifying Checksum
989b226a579c: Download complete
8ed8f35f8d4f: Pull complete
989b226a579c: Pull complete
660aeead31d5: Pull complete
4f4fb700ef54: Pull complete
adce24567e4c: Pull complete
f17ea56b313b: Pull complete
6b6f3b3efe88: Pull complete
4ed1ca4f3fce: Pull complete
e6fc9c6a5757: Pull complete
d47782d1182a: Pull complete
Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8
Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable
docker.io/gtstef/filebrowser:1.5.6-stable
1.1.0: Pulling from admin/felhom-samba
897d797d2723: Pulling fs layer
3051591aa250: Pulling fs layer
ce57a3f93416: Pulling fs layer
fb94eeec2fe1: Pulling fs layer
fb94eeec2fe1: Waiting
ce57a3f93416: Verifying Checksum
ce57a3f93416: Download complete
fb94eeec2fe1: Verifying Checksum
fb94eeec2fe1: Download complete
897d797d2723: Verifying Checksum
897d797d2723: Download complete
3051591aa250: Verifying Checksum
3051591aa250: Download complete
897d797d2723: Pull complete
3051591aa250: Pull complete
ce57a3f93416: Pull complete
fb94eeec2fe1: Pull complete
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
gitea.dooplex.hu/admin/felhom-samba:1.1.0
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
[golden] identity-clean + minimize …
[golden] stop + archive …
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
INFO: archive file size: 618MB
INFO: Finished Backup of VM 9100 (00:00:29)
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_05-12_21_35.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
[golden] publishing golden (648611216 bytes, sha256 65a87efc4b557caa…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.296.0/golden.tar.zst
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.296.0
GOLDEN_SHA256=65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.296.0 / 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)