diff --git a/REPORT.md b/REPORT.md index b396376..5dfe8ee 100644 --- a/REPORT.md +++ b/REPORT.md @@ -2,42 +2,32 @@ > **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md). -## RUNBOOK-peti-return — Phase 0a/1 executed, **STOPPED at Gate P1 for a ruling** — 2026-07-13 +## SPIKE — controller-driven escrow ceremony (daemon-context invocation mechanics) — 2026-07-13 -Operational run, read-only throughout (hub API with the bearer + a `hub.db` snapshot queried on -180 and deleted after). Nothing signed, pushed, or changed on any box or in the hub. Full -evidence: `documentation/pilot/RUNBOOK-peti-return-2026-07-13.md` §EXECUTION RECORD. +Docs-only commit; findings at `documentation/audits/SPIKE-controller-escrow-2026-07-13.md`. +No production code written; agent/controller repos untouched. **All five mechanisms under test: GO** +— the production spec (local-API ceremony endpoint + `--output=json` + job/one-shot-R-claim) can be +written on observed behavior. -### Gate P1 FAILS — controller 0.115.0 is NOT converging to the 0.122.0 floor - -- The hub **is serving** the floor (DB floor 0.122.0; agent 0.81.0 ≥ MinAgent 0.81.0; no HELD - lines in live hub logs). The stall is box-side. -- Self-update worked until 07-11 13:38 (0.110→0.112→0.113→0.115), then the **controller→agent - :8443 channel died at 07-11 14:01** ("no route to host 192.168.1.170:8443") and is **still - refusing today** (fresh `agent_channel_unreachable` 16:15 local, hours after host-reports - resumed). The floor auto-update delegates the container swap to the agent over that channel; - a failed attempt is persisted **once-per-floor** (anti-flapping) and never retried. -- **Root cause below it: the cluster split the runbook itself declares a STOP.** The agent's node - ("proxmox", 128 GB RAM) rebooted ~3.5 h ago and reports **0 guests**; the guest's controller - reports 4.2 days uptime on an i5-2500/8 GB — **the felhom guest is not on the node the agent is - enrolled on** (proxmox1/proxmox2 shape; single host row verified). - -### Phase 0a record (the notification-pipeline first-real-customer answer) - -- Claim code: issued gen 1 + **emailed 2026-07-12 16:49:09Z**, NOT yet claimed. -- **FINDING P1-F1: the motioneye 100 % warning NEVER emailed Peti** — all 16 `storage_fill_critical` - notifications went operator-channel only; no `customer_notifications` row exists pre-claim. - Follow-up task: define intended pre-claim customer-notification behavior. -- Collateral: P1-F2 agent local vzdump of 9201 fails every cycle (`not a UPID: "OK"` — cross-node - vmid); P1-F3 `wg-handshake-read` DEGRADED (wg tools missing on his node); P1-F4 the MinAgent - conditional floor is served from last-known agent_version with no freshness check (floor was - served while his agent was 40 h dead, guaranteeing the failed swap). motioneye VG: 1.2 TB at - 99.66 % (4.09 GB free). - -### Ruling needed (all queued in the runbook doc) - -1. Cluster split: migrate guest back / move the agent / accelerate agent-follows-guest. -2. Post-heal convergence path: floor bump (0.125.0 earmarked), manual trigger, or state clear. -3. Optional: hub controller-log pull to confirm the stall variant (CC did not trigger — write op). -4. Phase 0b/0c (message Peti, call slot) — Viktor, unaffected. Phase 4 rotations independent — - can proceed on GO. +- **Target:** the drill VM (qm 300, `192.168.0.152`, take-two end state, agent **0.87.0**); full + probe plan P0–P8 ran, nothing skipped; demo host + live demo escrow untouched (hub row verified). +- **SQ1 (load-bearing): the PTY-driven PBS re-key works with NO controlling terminal** — 3/3 + ceremonies exit 0 under `systemd-run --uid=felhom-agent` + `sudo -n` (the daemon shape), ~2.3 s + each, self-verify included. +- **Refusal matrix 5/5:** the fixed-argv sudoers line refused every altered argv (value change, + extra flag, reorder, `--config` omitted, alternate config path) — direct self-invocation is + sound; no guarded wrapper needed. Caller must pin argv order + `--` spelling byte-exact. +- **R pipe-capture round-trip PROVEN:** R parsed from the captured stdout opened the blob via + `escrow-consume` (fingerprint-gated, recovered key = live key size). Banner parsing is brittle → + `--output=json` is mandatory for production. +- **env_reset clean** (pinned `--config`; WG key auto-captured; hub upload authenticated); + **`--upload` hub-verified** (`host_escrow` row timestamped to the second, 383 B blob, empty + `restic_pw_sha256` — no staged secret existed, as gated in P0). +- **Timings:** 2.28–2.37 s incl. upload → job pattern with 2 s poll / 60 s timeout recommended; + ceremony process peaks ~264 MiB. +- **Correction to the task premise:** capability probes are LIST-mode (`sudo -n -l`), never + executing (manifest.go, verified live) — the escrow grant can get a normal probe entry. +- **Hygiene:** zero R fragments in journal/auth logs/state dir; all probe artifacts removed + (sudoers drop-in gone + post-removal refusal re-proven; scratch shredded; VM left running). +- **Recorded side effect (deliberate, drill scratch):** the drill box's take-two escrow blob is + superseded by a hash-less one; the take-two paper R no longer opens the CURRENT blob. diff --git a/documentation/audits/SPIKE-controller-escrow-2026-07-13.md b/documentation/audits/SPIKE-controller-escrow-2026-07-13.md new file mode 100644 index 0000000..971d8a5 --- /dev/null +++ b/documentation/audits/SPIKE-controller-escrow-2026-07-13.md @@ -0,0 +1,226 @@ +# SPIKE — Controller-driven escrow ceremony: daemon-context invocation mechanics (2026-07-13) + +Empirical validation of every mechanism the planned controller-driven escrow ceremony +(controller → agent local API → `sudo` self-invocation → ceremony → R returned once) depends on, +run BEFORE the production spec is written. **All five mechanisms under test: GO.** No production +code was written; the deployed agent binary was exercised as-is. + +> Hygiene note, stated up front: this spike deliberately violated the ceremony runbook's +> "never pipe the output" rule (RUNBOOK-escrow-ceremony.md §Do NOT) — R was captured on a pipe, +> programmatically parsed, and round-trip-consumed. This happened ONLY on the non-production +> drill VM; every captured R is a throwaway test-box secret and every capture file was shredded +> (P8). No R value, blob bytes, or fingerprint+R pair appears in this document. + +## 1. Target + baselines + +| Item | Recorded | +|---|---| +| Environment | **Drill VM** (qm 300 `drill-day0` on felhom-pve, nested PVE) — running at spike start, LAN `192.168.0.152`, root key-SSH (the DRILL-day0-take2 access path, reused verbatim); left running at spike end | +| Box state at start | The DRILL-day0-take2-2026-07-12 end state: host `demo-vm-felhom-2f4b00`, escrowed, offsite round-trip proven | +| Deployed agent | **felhom-agent 0.87.0** (service active, non-root `felhom-agent` uid 999) | +| `escrow.pbs_storage_id` | `felhom-pbs` (agent.json; no secret values read) | +| PBS key file | `/etc/pve/priv/storage/felhom-pbs.enc` present (255 B, 0600 root:www-data) + `.pw` | +| `age` | `/usr/bin/age` 1.2.1 | +| Hub reachability | `https://hub.felhom.eu/` → 302 from the target | +| Staged restic password | **ABSENT** (`/var/lib/felhom-agent/escrow-stage/` empty — the take-two ceremony wiped it). Drill target → proceed per the P0 gate; consequence: the P5 blob carries an **empty `restic_pw_sha256`** (verified hub-side, §2.5) | +| sudo / ptmx | sudo 1.9.16p2; `/dev/ptmx` 0666 | +| Probes run | **P0–P8 all ran, none skipped** (drill target → full plan incl. P5 `--upload`) | +| Repo baselines | felhom-agent `main` @ `adf7882f7dd6` v0.87.0 (untouched); felhom.eu docs-only commit | + +Demo-host blast radius: nothing on 192.168.0.162 (host level) or the live demo escrow was touched; +`demo-felhom-01`'s hub escrow row still timestamps `2026-07-09 14:16:16` after the spike (§2.5). + +## 2. SQ verdicts + +### 2.1 SQ1 — PTY allocation under NO controlling terminal (the load-bearing result): **WORKS** + +The full ceremony — including the PBS `key change-passphrase` re-key that `pty_linux.go` +drives over a manually allocated `/dev/ptmx` pty — succeeded from a daemon-equivalent context +(systemd transient unit, service uid/gid, no TTY anywhere in the chain), 3/3 runs: + +``` +systemd-run --wait --pipe --collect --uid=felhom-agent --gid=felhom-agent \ + /usr/bin/sudo -n /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json \ + --selftest=escrow-create --offline +``` + +Every run: exit 0; R banner + a plausible EFF-wordlist R on stdout (one dash-joined token, +10 words / 69 chars, `~129 bits R` printed); `blob: 383 bytes … key fingerprint f2:87:…:f7:8e · +posture zero_knowledge`; the `self-verify: the blob unwraps back to the key with R` line present +(self-verify itself exercises a SECOND pty round for the unwrap re-key — both directions work +no-TTY); `identity escrow: 450 bytes … self-verify OK`. The pty path needs no controlling +terminal at all: `Setsid` + `Setctty` on the slave gives the CHILD its controlling terminal +regardless of the parent having none. + +stderr per run: only the two slog INFO lines (`+wg_private_key`, `creating zero-knowledge +recovery-code escrow` with field names, never values) + systemd-run framing. **No secret +material on stderr.** + +### 2.2 sudoers exact-argv refusal matrix (the spike's red-proof): **5/5 REFUSED** + +Temp drop-in `/etc/sudoers.d/zz-felhom-escrow-spike` (0440 root:root, `visudo -cf`-gated), +exactly two fixed-argv lines (recorded verbatim): + +``` +felhom-agent ALL=(root) NOPASSWD: /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json --selftest=escrow-create --offline +felhom-agent ALL=(root) NOPASSWD: /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json --selftest=escrow-create --upload +``` + +Each probe run as the `felhom-agent` user via `sudo -n` (daemon shape): + +| # | Altered argv | Result | +|---|---|---| +| (a) | value altered: `--selftest=escrow-consume --offline` | **REFUSED** — exit 1, `sudo: a password is required`, ceremony never spawned | +| (b) | extra flag appended: `… --offline --paperkey` | **REFUSED** — same | +| (c) | argv reordered: `--selftest… --config … --offline` | **REFUSED** — same | +| (d) | `--config` omitted | **REFUSED** — same | +| (e) | different config path: `--config /tmp/evil.json` | **REFUSED** — same | + +The exact allowed argv was accepted (that is P3/P5 themselves), and after the P8 removal the +same argv is refused again (post-cleanup re-check: `sudo: a password is required`). + +**Consequences confirmed:** sudoers matches the argument vector byte-for-byte — the production +caller must pin argv ORDER exactly as written in the sudoers line. Extra trap recorded: Go's +`flag` package accepts both `-flag` and `--flag` spellings, but sudoers only matches the literal +form in the line — the caller must emit the exact `--`-spelled argv, never normalize. + +### 2.3 R pipe-capture fidelity + round-trip: **PROVEN** + +From one P3 run's captured stdout, entirely on the box (R never left it): +R parsed programmatically (the first single-token line after the banner box's `└` edge); +fingerprint parsed from the `blob:` line; the `--offline` base64 decoded to a root-0600 scratch +blob (383 B). Then, as root directly: + +``` +FELHOM_RECOVERY_CODE='' felhom-agent --config /etc/felhom-agent/agent.json \ + --selftest=escrow-consume --blob --fingerprint f2:87:… --keydest <0600 scratch> +``` + +→ exit 0, `[OK] recovered key installed … (fingerprint-gated, 0600)` — the recovered key is +255 bytes, the live key file's exact size. **The R that crossed a pipe is the R that opens the +blob** — the exact property the controller endpoint relies on. Both scratch files shredded +immediately (keydest right after the consume, blob at P8). + +**Banner-parse brittleness (motivates `--output=json`):** R is identifiable only positionally +(after Unicode box-drawing lines) as "the indented single-token line"; the fingerprint sits +inside a `·`-separated human line; the offline copy is "the long base64-shaped line". Every one +of these breaks on any cosmetic banner change. A machine mode is not optional for production. + +### 2.4 Ceremony under `sudo -n` env_reset: **CLEAN** + +With env_reset stripping the environment (including `FELHOM_AGENT_CONFIG` — which is why the +sudoers line pins `--config` explicitly): config discovery worked from the pinned path; +`escrow.pbs_storage_id` resolved; the live WG key was auto-captured (`identity bundle: ++wg_private_key`, `identity=true` in the banner); the hub upload leg (P5) authenticated and +completed. Nothing in the ceremony depends on inherited environment. + +### 2.5 Full `--upload` path (P5, drill VM only): **WORKS**, hub-verified + +Same systemd-run shape with the `--upload` line: exit 0, 2.35 s wall, +`uploaded the opaque blob(s) to the hub (host record); the hub cannot open them`. +Hub-side proof (read-only sqlite query of `host_escrow` on the k3s node's Longhorn mount): +`demo-vm-felhom-2f4b00 | blob 383 B | restic_pw_sha256 EMPTY | updated_at 2026-07-13 15:06:50Z` +— the timestamp is the P5 run to the second. The empty hash is correct: no staged secret existed +(P0), so nothing was sealed for auto-confirm to match. `demo-felhom-01`'s row unchanged. + +### 2.6 Timings (→ the production job budget) + +| Run | Mode | Exit | Wall (SSH-side) | Service runtime (systemd) | +|---|---|---|---|---| +| P3-1 | `--offline` | 0 | 2.28 s | 2.26 s | +| P3-2 | `--offline` | 0 | 2.30 s | — | +| P3-3 | `--offline` | 0 | 2.37 s | — | +| P5 | `--upload` | 0 | **2.35 s** | 2.33 s | + +min/median/max (offline): 2.28 / 2.30 / 2.37 s; upload adds ≈ nothing on LAN (budget for WAN +upload latency anyway). Memory peak of the ceremony process: **~264 MiB** (systemd accounting) — +the job endpoint spawns a second full agent process; worth remembering on small hosts. + +### 2.7 P6 — staged-dir permission interplay: **NO STRANDING** + +`/var/lib/felhom-agent/escrow-stage/` is 0700 `felhom-agent:felhom-agent`. Root created a 0600 +root:root probe file (`spike-probe`, deliberately NOT the real staged filename) inside it; the +`felhom-agent` user unlinked it cleanly (unlink permission comes from the directory, which the +daemon owns). The production root-ceremony-wipes / daemon-re-stages cycle has no permission trap. + +### 2.8 P7 — R hygiene sweep: **ZERO leaks** + +For each of the 4 captured Rs (3× P3 + P5): the full R and its first two words grepped across +the ENTIRE journal (`journalctl`, incl. `-u felhom-agent`), `auth.log*`, and +`/var/lib/felhom-agent/` → **0 hits total**. sudo logs the argv only (verified lines carry the +fixed flags, no secrets); journald never sees R. + +## 3. Recommendations for the implementation spec + +1. **Production sudoers line** — a single fixed-argv `--upload` variant (the `--offline` line was + probe-only), as a new alias in `configs/felhom-agent.sudoers`, refusal-matrix-validated shape: + + ``` + Cmnd_Alias FELHOM_ESCROW = \ + /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json --selftest=escrow-create --upload + ``` + + The caller must exec exactly this argv (order + `--` spelling pinned; §2.2). `--config` stays + pinned explicitly: env_reset strips `FELHOM_AGENT_CONFIG`, and the pin closes alternate-config + injection (probe (e)). + +2. **`--output=json` machine mode** — grounded in §2.3's brittleness: a SINGLE JSON object on + stdout (`{recovery_code, key_fingerprint, entropy_bits, blob_bytes, identity_blob_bytes, + restic_pw_sealed, uploaded}`), banner suppressed, everything human-facing to stderr. All + human printing already lives in `runSelftestEscrowCreate` (the CLI shell), not in + `escrow.Create` — the variant is a thin switch in main.go, no library change. + +3. **Job + one-shot in-memory R-claim endpoint** (the netstorage job pattern — mandatory anyway, + the agentapi client timeout is 15 s): measured ceremony ≈ 2.4 s incl. upload, so a + **poll interval 2 s, job timeout 60 s** is a ≥25× margin over LAN reality while absorbing WAN + upload latency. R held in memory only, single-flight mutex (one ceremony at a time per host), + one-shot claim (read once → wiped), **TTL ~10 min**; recovery story: "R unclaimed → ceremony + void → a re-run supersedes" — which §2.5 shows is exactly how the blob store behaves (a new + upload replaces the row; the superseded R opens only pre-existing history). + +4. **Capability-manifest caveat — CORRECTED by this spike.** The task premise ("manifest.go + probes EXECUTE their argv") is FALSE at live source: `internal/capability/manifest.go` + probes are **list-mode** (`sudo -n -l `, "never executing" — its own header, L3–4/L46–47), + verified live here (`sudo -n -l` on the escrow line: exit 0, echoes the grant, no ceremony + ran). **The escrow line can therefore get a NORMAL probe entry safely** — recommend adding a + standard `escrow-ceremony` capability row so degradation is visible on the hub like every + other grant, no special representation needed. + +5. **Ruling F1 (2026-07-13), carried:** R transiting the Cloudflare tunnel is an accepted, + documented risk (same trust class as the claim code / login password). The ceremony spec MUST + carry a threat-model paragraph documenting this acceptance explicitly. + +6. Minor spec inputs: exit codes observed/confirmed in source — 2 = usage/config error, 1 = + operational failure, 0 = success; the ceremony process peaks ~264 MiB (§2.6); stderr is + log-clean (§2.1) so the job runner may capture it for diagnostics without an R filter, but + stdout must be treated as secret-bearing until parsed + wiped. + +## 4. Cleanup checklist (P8) + +- [x] `/etc/sudoers.d/zz-felhom-escrow-spike` removed; `visudo -c` on the remaining set: parsed OK +- [x] Post-removal red-check: the previously allowed argv is refused again +- [x] All scratch files shredded (`spike-p3-run{1,2,3}.{out,err}`, `spike-p5-run.{out,err}`, + `spike-p4-blob`; `spike-p4-keydest` shredded immediately after the P4 consume; + P6 `spike-probe` unlinked; P2 temp files removed) +- [x] Drill VM left in its prior power state (running — the take-two end state, as found) +- [x] Demo host / live demo escrow untouched (`demo-felhom-01` hub escrow row still 2026-07-09; + no ceremony was run outside the drill VM) — the demo-path §P8 check is otherwise N/A on + the drill target + +## 5. Observations (out of scope — documented, NOT acted on) + +- **The drill box's escrow blob was superseded by this spike** (deliberately, P5): the take-two + paper R (2026-07-12 22:46) no longer opens the CURRENT hub blob (it still opens the superseded + one for pre-existing history, per the R-supersede rule). The new blob has an empty + `restic_pw_sha256`. The drill guest's controller showed **no reaction within ~6 min** of the + upload (no escrow log lines); whether an `escrowed` controller re-warns on a later ACK carrying + a hash-less blob is version behavior worth one glance at the next drill reset — the box is + re-drill scratch either way. +- The hub-side verification path used here (read-only `sqlite3` query of `host_escrow` via the + node's Longhorn mount) is a useful CC-side check pattern for hub state that the + password-gated UI otherwise hides. +- systemd-run reports the transient unit's resource envelope for free (runtime, CPU, memory + peak) — handy shape for the job runner's own accounting/logging. +- The spike's sudo journal lines confirm sudo logs the FULL argv of both accepted and refused + invocations — with fixed-argv design this is secrets-free by construction and doubles as an + audit trail of ceremony invocations.