hub-safety session: R-135/R-133/R-604/R-530/R-508/R-509/R-880 closed, R-861/R-173/R-518/R-519 narrowed, R-879/R-881 opened (336 → 332); 03 §3.1, 05 §16, golden 0.296.0, the hub-DB off-site plan, STATUS
gates / gates (push) Successful in 32s
gates / gates (push) Successful in 32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -196,10 +196,11 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 |
|
||||
| A second login step for the dashboard (a TOTP code or a passkey) | — | **MISSING** | — | One password, one bcrypt hash (`controller/internal/web/auth.go:37-44`) → R-811 (added 2026-10-03) |
|
||||
| The household can leave Felhom, or outlive it — the box runs without the hub, the household owns its domain, tunnel and off-site account, and can export everything | — | **MISSING** (as a written answer) | the LOST-hub half only: `_recovery-inventory-2026-07-28.md` §D2.4, `07` §8 row 11b | Leaving and hand-over are answered nowhere → R-810 (spike, added 2026-10-03) |
|
||||
| **The host agent cannot reach root without the operator key: exact sudo patterns, no agent-written file installed where root reads it without a content check, the agent binary only by an operator-signed update checked as root** | agent **v0.146.1** (R-861; delivered by a step bundle, R-880) | **PROVEN-LIVE on both demo boxes (2026-10-05) — with three named residuals** | `audits/hub-safety-2026-10-05/partF/` (real sudo: 23 of 29 attacks allowed before, 0 after; 64 capability commands allowed; `sudo -l` 93/93 on demo-hp and demo-felhom after the bundle; capability probe 67/67; a staged unit over `/etc/sudoers.d` refused live) | `03` §3.1: the controller-swap image ref (guest-scoped), the felhom-op SSH key (hub-delivered, unsigned; felhom-op's sudo is scoped), the escrow ceremony relays R → R-861 |
|
||||
| WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` |
|
||||
| The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe |
|
||||
| The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing |
|
||||
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → **R-133** |
|
||||
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. **Sealed at rest since hub v0.135.0 (R-133 CLOSED):** the off-site seal and key; live 2026-10-05: 4 legacy rows sealed at start-up, 0 left plain, and the demo-hp reveal still minted a PVE ticket (HTTP 200; a wrong password 401) — `audits/hub-safety-2026-10-05/partB/`. A database backup now needs `OFFSITE_SECRET_KEY` too → R-173 |
|
||||
|
||||
## F. Notifications & monitoring
|
||||
|
||||
@@ -233,6 +234,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
|
||||
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
@@ -72,6 +72,59 @@ Explicitly does **not**:
|
||||
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
|
||||
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
|
||||
|
||||
### 3.1 The agent's admin commands, group by group — and how each is narrowed (R-861, agent v0.146.1) `[DESIGN — 2026-10-05, CC; operator may reverse]`
|
||||
|
||||
**The question.** Can a compromised agent PROCESS (running as the `felhom-agent` user) become root on its host without
|
||||
the operator's key? Read on 2026-10-04 (R-861) and measured 2026-10-05 with the real sudo 1.9.16 in a throwaway
|
||||
container: **with the v0.145.0 sudoers, yes — 23 of 29 attack command lines were allowed.** Two shapes did it:
|
||||
|
||||
1. **A glob in the arguments.** Sudo's `*` in arguments also matches spaces, so one grant smuggled extra options:
|
||||
`pct set [0-9]* -onboot 1` allowed `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest);
|
||||
`mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` allowed a `..` path onto `/etc/sudoers.d`; `nft add element …
|
||||
*` allowed `; flush ruleset`.
|
||||
2. **A file the agent wrote, installed where root reads it.** A `.mount` unit (bind any directory over `/etc`), a
|
||||
dnsmasq drop-in (`dhcp-script=` runs as root), the wg-quick config (`PostUp=` runs as root), the OOB sshd config
|
||||
(`AuthorizedKeysFile` + `StrictModes no`), the guest pre-start hook (Proxmox runs it as root), the shared-parent boot
|
||||
script, and the agent BINARY itself (the escrow ceremony and the guest hook run it as root; `apply` took a sha the
|
||||
agent passed).
|
||||
|
||||
**The rule after v0.146.1** (the R-861 fix direction: each becomes a root-owned wrapper that checks its own input, or a
|
||||
fixed file, delivered by the signed config bundle):
|
||||
|
||||
- Every varying argument list is a **sudo regular expression** (`^…$`): one value per slot, a fixed character set, no
|
||||
`..`, no extra argument. Literal lines stay literal.
|
||||
- **No file the agent wrote is installed where root reads it.** Either the content is FIXED and comes with the signed
|
||||
bundle (the hook, the shared parent), or a root wrapper checks the CONTENT against the agent's own renderers before
|
||||
installing it (`felhom-priv-apply`), or the operator's signature is checked as root (`felhom-os-apply agent_update`).
|
||||
- The pins: `TestManifestCoveredBySudoers` (every command the agent runs is still allowed), `TestSudoersRefusesTheR861
|
||||
Injections` (the 29 attacks are not), the real-sudo run of both (`audits/hub-safety-2026-10-05/partF/
|
||||
sudo-container-proof.txt`), and `sudo -l -U felhom-agent` on both demo boxes after the bundle.
|
||||
|
||||
| Group | What it is for | How it is narrowed (v0.146.1) | Left open |
|
||||
|---|---|---|---|
|
||||
| `FELHOM_MOUNT` | fs-UUID mount units for enrolled drives | install only via `felhom-priv-apply unit <name>`: `[Unit]` only Description + `After=local-fs-pre.target`, `Where=` `/mnt/<name>` or `/mnt/felhom-drives/<name>` and equal to the unit name, `What=` a UUID or a network source, no `bind`/`suid`/`dev`, no continuation lines; systemctl verbs on `mnt-…\.mount` only | — |
|
||||
| `FELHOM_NETMOUNT` | NAS automount pairs, re-arm, clean-up | same checker; a network share must carry `nosuid,nodev` (the agent now renders them); `rm`/`rmdir`/`reset-failed` one exact name | — |
|
||||
| `FELHOM_DISK` | SMART, thin-pool, PV and pool reads | exact device / LV patterns (no extra options such as `smartctl -s off`, `lvs --config`) | read-only |
|
||||
| `FELHOM_PROVISION` | bootstrap config mount, autostart | exact `mpN` spec (`…/guests/<vmid>/bootstrap,mp=/…[,ro=1]`) — no smuggled `--dev0` | — |
|
||||
| `FELHOM_FORMAT` | data-bearing probe + guarded mkfs | one device path, no space; `mkfs` only through `felhom-mkfs-guarded` (its own root checks) with `ext4`/`xfs` | blkid/lsblk read any `/dev` path (read-only) |
|
||||
| `FELHOM_DNSMASQ` | the LAN split-horizon resolver | drop-ins via `felhom-priv-apply dnsmasq` (only `bind-interfaces`, `no-resolv`, `listen-address`, `server`, `local`, `address`); `rm` one exact name; exact `pct exec` reads | — |
|
||||
| `FELHOM_GUESTHOOK` | the pre-start self-heal hook | the hook is a FIXED bundle file; the agent only checks it (`SnippetReady`) and registers it; exact vmid/slot | — |
|
||||
| `FELHOM_INTERMEDIARY` | the shared drive parent + live drive binds | boot script + unit are FIXED bundle files (the agent only enables the unit); one-segment drive names (no leading dot, no `..`) | — |
|
||||
| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen (pinned-registry) image ref and restart the guest's bootstrap — the household's data, not host root |
|
||||
| `FELHOM_STALELOCK` / `FELHOM_SCRATCH_TEARDOWN` | stale-lock clear; failed restore-test scratch | exact vmid; the scratch band `99000[0-9]` was already exact | — |
|
||||
| `FELHOM_WG` | the off-site tunnel | conf via `felhom-priv-apply wg` (only the keys `renderConf` writes; no `PostUp`/`PreUp`/`DNS`/`Table`; `/32` only) | — |
|
||||
| `FELHOM_SELFUPDATE` | commit / rollback of the A/B flip | **`apply` removed**: the flip runs only inside `felhom-os-apply agent_update`, after the operator signature, host, window and nonce are checked as root and the staged bytes are hashed ONCE and copied to a root-owned dir (`/var/lib/felhom-os-apply/agent-update/`); the wrapper accepts only that dir | — |
|
||||
| `FELHOM_SSHD` | the out-of-band operator sshd | config via `felhom-priv-apply sshd-config` (the ONE template, only the Port varies, never 22); the felhom-op key via `sshd-key` (one plain key, no `command=`/`from=` options) | felhom-op's key itself is hub-delivered, not signed: a compromised agent can install its own key for **felhom-op** — whose sudo is scoped (`felhom-op.sudoers`), not root |
|
||||
| `FELHOM_OOB` | the OOB firewall sets | `add element` takes exactly `{ <ip>[/n] }` or `{ <port> }` — no chained command | — |
|
||||
| `FELHOM_PBSDR` / `FELHOM_BACKUPTARGET` | PBS DR entry; whole-system backup target | unchanged: the arguments stay coarse, and the root wrappers (`felhom-pbs-apply`, `felhom-backup-target-apply`) are the gate (fixed verbs, own validation) | coarse argv into a checking wrapper |
|
||||
| `FELHOM_ESCROW` | the recovery-code ceremony (runs the agent binary as root) | the binary is only ever an operator-signed one (`FELHOM_SELFUPDATE`); as root it pins the PVE secret dir and the WG state dir, refuses a storage id that is a path, and reads its two staged files by walking the path with `openat(O_NOFOLLOW)` (no symlink anywhere) | **by design the agent relays R**, so a compromised agent can still learn this box's PBS key through the ceremony — not root, but the backup key |
|
||||
| `FELHOM_SELFHEAL` / `FELHOM_GUESTNET` / `FELHOM_OSAPPLY` | networking restart; guest DHCP watchdog; OS updates | exact; `felhom-os-apply --plan …` stays the glob line on purpose — the bundle's own self-check reads that exact text, and the wrapper refuses any other plan path (R1) | — |
|
||||
|
||||
**What this does not change.** The operator key (`/etc/felhom/operator-signers`, root-owned, never a bundle path) stays
|
||||
the one trust root; the agent's API token is untouched; a box gets the new rule only through the signed
|
||||
`agent_config_update` (the order: signed `agent_update` first, then the bundle — after the bundle, an agent below
|
||||
0.146.0 cannot update itself on that box).
|
||||
|
||||
## 4. Control model — reconcile + signed destructive ops
|
||||
|
||||
Two channels, split by **reversibility**, not by transport.
|
||||
@@ -655,7 +708,8 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea
|
||||
runs as root at guest start (and `pct reboot` is granted); `FELHOM_INTERMEDIARY` installs a script and a systemd unit
|
||||
that run as root at boot; `FELHOM_ESCROW` runs the agent binary as root, and `FELHOM_SELFUPDATE apply` accepts a sha
|
||||
the agent itself passes. So a compromised agent PROCESS is root on its host; the root-owned trust files (decision 93,
|
||||
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants.
|
||||
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. **Narrowed in agent
|
||||
v0.146.1 — §3.1 lists every group, the rule now, and what stays open.**
|
||||
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
|
||||
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
|
||||
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
|
||||
|
||||
@@ -130,6 +130,17 @@ R-216. Either way the box's reported agent must meet the chosen MinAgent, else t
|
||||
the Hosts page and logged once per change as `managed floor SERVED`. The rules the operator follows:
|
||||
`runbooks/publish-train-rules.md` rule 1.
|
||||
|
||||
**Boxes left behind (hub v0.135.0, R-604 + R-530).** A per-customer floor wins over the global one, so a global
|
||||
raise does not move a box whose OWN floor is lower — and `managed floor SERVED` is logged once per change, so that box
|
||||
was silent (demo-hp missed four raises, 2026-09-21). Now the raise logs one line per such customer and sends ONE
|
||||
operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set
|
||||
(`customer_configs.min_controller_set_at`; "unknown" for one set before v0.135.0), and the System page's "Version
|
||||
floors" table lists the global floor, every per-customer floor with its age, and which ones the global cannot move.
|
||||
Agents are a separate train (they update only by a per-box signed job, R-530's ruling): the System page shows each
|
||||
box's agent against the vouched one ("0.142.0 → 0.145.0 (since …)", amber, red after the wait), and a box behind the
|
||||
vouched agent for 7 days raises `agent_behind` (warning, operator-only; `OS_ALARM_AGENT_BEHIND_AFTER`; the clock starts
|
||||
when the hub first sees the box behind). `[DESIGN — CC 2026-10-05, operator may reverse; `09` decision 119]`
|
||||
|
||||
## 6. Authorization — signed-op queue + editing flow
|
||||
|
||||
Implements Part 4's gate on the hub side. The hub holds **no signing key**.
|
||||
@@ -403,3 +414,35 @@ count appeared was the bind page's passphrase hint, and its English half is now
|
||||
phrase you received from your operator during setup") — because "five words" stops being true for an
|
||||
English household, and was already wrong for one whose passphrase predates this release. The
|
||||
Hungarian „öt szó" is correct and unchanged.
|
||||
|
||||
## 16. The operator surface's own safety [DESIGN — hub v0.135.0, CC 2026-10-05, operator may reverse]
|
||||
|
||||
### 16.1 Form protection (CSRF) on both login paths (R-135)
|
||||
|
||||
The operator logs in two ways: a browser session (`hub_session` cookie + a per-session token on every form) and HTTP
|
||||
Basic for scripts. Until v0.135.0 a state-changing request with NO cookie skipped the token check — on the reasoning
|
||||
that it must be a script. It need not be: a browser caches Basic credentials per origin and resends them on a
|
||||
cross-site form POST (SameSite does not govern the Authorization header). Now a request without a session passes
|
||||
only with Basic credentials AND the header `X-Felhom-Operator` (any value; scripts send `cli`). A page on another site
|
||||
cannot add a custom header without a CORS preflight, which the hub never answers. The gate sits in `ServeHTTP` before
|
||||
the route switch, so it covers every route at once (38 state-changing routes + an unknown path in
|
||||
`r135_csrf_test.go`); `/login` and the public `/bind/<token>` stay exempt (no operator session to ride; the bind token
|
||||
is the capability). The other choice — dropping browser-usable Basic auth entirely — was not taken: CC's headless runs
|
||||
and the runbooks drive the hub with Basic auth, and the header costs them one flag.
|
||||
|
||||
### 16.2 Secrets at rest in `hub.db` (R-133, R-821)
|
||||
|
||||
Two columns are SEALED (AES-256-GCM, `enc:v1:` + nonce, one key: `OFFSITE_SECRET_KEY` from `Secret/offsite-secret-key`,
|
||||
never in the database or git): the off-site sub-account passwords (`one_time_secrets.value`, v0.127.0) and, since
|
||||
v0.135.0, each box's break-glass console password (`host_recovery.secret`) — the same helpers, not a second scheme.
|
||||
Legacy rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, measured live: 4 rows). No key → a save is
|
||||
refused; a wrong key → a reveal is a 500 with nothing in the body or the log. Both retrieval paths (the operator page and
|
||||
the global-key API) open through `GetHostRecoveryCredential`, so the break-glass route still works with the UI down.
|
||||
|
||||
**What a copy of `hub.db` still holds readable** (R-879): each box's hub API key, each household's owner passphrase
|
||||
(`customer_configs.retrieval_password`) and API key, and the PBS-DR token values (`host_pbs_secrets.value`, kept after
|
||||
use). **And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only
|
||||
with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export
|
||||
on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md`
|
||||
(R-173 — awaiting the operator's decision).
|
||||
|
||||
|
||||
@@ -133,6 +133,16 @@ startup a record still marked running becomes a failed, interrupted result („A
|
||||
restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the
|
||||
precedent is kept for what it was written for.
|
||||
|
||||
**A backup RUN cut off by a power cut or a restart is said (R-519, controller v0.296.0, `09` decision 123).** The
|
||||
same shape for the app-data run: `appdata-run.json` beside the restore record, written at both ends of a run. A start
|
||||
that finds it still running turns it into a notice on /backups and /backups/apps („A legutóbbi mentés (…) megszakadt,
|
||||
mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every step OK. Measured BIGNIGHT F2 (2026-09-14):
|
||||
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
|
||||
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
|
||||
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
|
||||
*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the
|
||||
operator (R-519 narrowed).*
|
||||
|
||||
### Lane 2 — the operator: guest and host recovery
|
||||
|
||||
Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator
|
||||
@@ -661,6 +671,11 @@ successes only. After an agent restart the success is read back from the tier's
|
||||
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
|
||||
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
|
||||
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
|
||||
**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the
|
||||
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
|
||||
longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both
|
||||
measurements (≈6 min / 9 apps, ≈8 min / 12 apps), say "minutes, not seconds", and that the off-site copy in the same
|
||||
run makes it longer.
|
||||
|
||||
### 6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, `09` §3 decision 36)
|
||||
|
||||
|
||||
@@ -327,6 +327,8 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||||
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
|
||||
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
|
||||
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
|
||||
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
|
||||
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
|
||||
|
||||
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
|
||||
breaks a `not_running` run.
|
||||
@@ -334,9 +336,9 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
|
||||
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
|
||||
what the box cannot heal — a running container with no connection (wrong token, blocked network).
|
||||
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
|
||||
next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`,
|
||||
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended,
|
||||
operator may reverse* (`11` §8.3).
|
||||
next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`,
|
||||
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`,
|
||||
`OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119).
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -853,6 +853,34 @@ its length, and both fixes cost something the household would notice — operato
|
||||
a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs
|
||||
one call (pinned by a test). agent v0.145.0.
|
||||
|
||||
### 2026-10-05 (afternoon) — decided by CC — operator may reverse (the hub-safety / R-861 brief)
|
||||
|
||||
119. **Boxes left behind (R-604, R-530).** Options for the agent alarm's wait: 3 days (a box off for a long weekend
|
||||
alarms), 7 days (one week, the same wait as the bundle-behind alarm, R-840), 14 days. **Chosen 7 days**
|
||||
(`OS_ALARM_AGENT_BEHIND_AFTER`), counted from when the hub first sees the box behind. A global floor raise names,
|
||||
in one operator mail, every box whose own LOWER floor it cannot move. hub v0.135.0, `05` §5.
|
||||
120. **How the Basic-auth operator path is protected from cross-site POSTs (R-135).** Options: (a) drop browser-usable
|
||||
Basic auth (CC's headless runs and every runbook POST break); (b) require a custom header on a cookie-less
|
||||
state change (a browser cannot add one cross-site without a CORS preflight the hub never answers; scripts add one
|
||||
flag). **Chosen (b)**, header `X-Felhom-Operator`. hub v0.135.0, `05` §16.1.
|
||||
121. **The console password's seal (R-133).** Options: (a) a second key and scheme for `host_recovery`; (b) the off-site
|
||||
seal and key already in force (R-821). **Chosen (b)** — the brief asked for the existing pattern, and one key is
|
||||
one custody question. Consequence named: a database backup needs this key off DooPlex too (R-173). `05` §16.2.
|
||||
122. **How the agent's admin commands are narrowed (R-861).** Options per group: (a) a root wrapper per group with its
|
||||
own argv; (b) exact sudo regex patterns for every varying argument + ONE content checker for every agent-written
|
||||
file root reads + fixed bundle files where the content never varies + the signed update verified as root by the
|
||||
existing `felhom-os-apply`. **Chosen (b)** — fewest new root programs, and the checker is pinned to the agent's
|
||||
own renderers by contract tests. Delivery order: signed `agent_update` first, then the bundle. agent v0.146.1,
|
||||
`03` §3.1.
|
||||
123. **How long a cut-off backup run is said on the page (R-519).** Options: until the next run of any kind (a failed
|
||||
run would clear the warning), until the next run that ends with every step OK, or until dismissed. **Chosen: until
|
||||
a run ends with every step OK.** controller v0.296.0.
|
||||
124. **How a bundle that adds a path reaches a box (R-880).** An installed `felhom-os-apply` refuses any path not in its
|
||||
own table (R16). Options: (a) copy the new wrapper onto each box by hand as root (does not scale, leaves the signed
|
||||
route); (b) a STEP bundle — the box's current bundle with only `felhom-os-apply` replaced, published as
|
||||
`<ver>-step1` — then the release's bundle, both by signed jobs. **Chosen (b)**, `felhom-agent/scripts/build-step-bundle.py`;
|
||||
delivered to demo-hp, demo-felhom and Tester 1 on 2026-10-05. `11` §5.4.2 rule unchanged.
|
||||
|
||||
### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief)
|
||||
|
||||
100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) —
|
||||
|
||||
Reference in New Issue
Block a user