From 0826e41b31fe7c93b4811f349ea5e706b6b48793 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 13:47:57 +0200 Subject: [PATCH] =?UTF-8?q?hub-safety=20session:=20R-135/R-133/R-604/R-530?= =?UTF-8?q?/R-508/R-509/R-880=20closed,=20R-861/R-173/R-518/R-519=20narrow?= =?UTF-8?q?ed,=20R-879/R-881=20opened=20(336=20=E2=86=92=20332);=2003=20?= =?UTF-8?q?=C2=A73.1,=2005=20=C2=A716,=20golden=200.296.0,=20the=20hub-DB?= =?UTF-8?q?=20off-site=20plan,=20STATUS?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 12 + REPORT-hub-safety-2026-10-05.md | 131 +++++++ STATUS.md | 60 +++- .../architecture/00-capability-map.md | 5 +- documentation/architecture/03-host-agent.md | 56 ++- .../architecture/05-hub-architecture.md | 43 +++ .../architecture/07-backup-architecture.md | 15 + documentation/architecture/08-alarm-ladder.md | 8 +- .../architecture/09-update-architecture.md | 28 ++ .../hub-safety-2026-10-05/partA/live.txt | 8 + .../partA/route-table.md | 50 +++ .../hub-safety-2026-10-05/partB/live-db.txt | 6 + .../partB/live-reveal.txt | 6 + .../hub-safety-2026-10-05/partC/readings.txt | 42 +++ .../partD/live-system-page.txt | 9 + .../partE/r518-measure.txt | 28 ++ .../partE/r518-watch.log | 15 + .../hub-safety-2026-10-05/partE/red-proof.txt | 39 ++ .../partE/teardown-9202.txt | 7 + .../partF/demo-felhom-steps.txt | 9 + .../partF/demo-hp-step1.txt | 6 + .../partF/live-checker-demo-felhom.txt | 8 + .../partF/live-checker-demo-hp.txt | 20 ++ .../partF/live-sudo-after-demo-felhom.txt | 2 + .../partF/live-sudo-after-demo-hp.txt | 2 + .../partF/live-sudo-before-demo-felhom.txt | 28 ++ .../partF/preflight-live-files.txt | 30 ++ .../hub-safety-2026-10-05/partF/red-proof.txt | 85 +++++ .../partF/step-bundle.txt | 4 + .../partF/sudo-container-proof.txt | 129 +++++++ .../partH/fleet-after.txt | 11 + .../hub-safety-2026-10-05/partH/floors.txt | 7 + .../partH/sign-agent-update.txt | 10 + .../partH/sign-bundle.txt | 13 + .../hub-safety-2026-10-05/partH/vouch.txt | 4 + documentation/backlog/CLOSED-ITEMS.md | 16 + documentation/backlog/OPEN-ITEMS.md | 28 +- .../runbooks/RUNBOOK-hub-db-offsite-backup.md | 129 +++++++ documentation/runbooks/config-bundle.md | 18 + .../02-round-trip.txt | 5 + .../tests/golden-0.296.0-2026-10-05/README.md | 62 ++++ .../tests/golden-0.296.0-2026-10-05/bake.log | 340 ++++++++++++++++++ hub/CHANGELOG.md | 6 +- hub/internal/web/server.go | 2 +- 44 files changed, 1511 insertions(+), 31 deletions(-) create mode 100644 REPORT-hub-safety-2026-10-05.md create mode 100644 documentation/audits/hub-safety-2026-10-05/partA/live.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partA/route-table.md create mode 100644 documentation/audits/hub-safety-2026-10-05/partB/live-db.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partB/live-reveal.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partC/readings.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partD/live-system-page.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partE/r518-measure.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partE/r518-watch.log create mode 100644 documentation/audits/hub-safety-2026-10-05/partE/red-proof.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partE/teardown-9202.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/demo-felhom-steps.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/demo-hp-step1.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-felhom.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-hp.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-felhom.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-hp.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/live-sudo-before-demo-felhom.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/preflight-live-files.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/step-bundle.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partF/sudo-container-proof.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partH/fleet-after.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partH/floors.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partH/sign-agent-update.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partH/sign-bundle.txt create mode 100644 documentation/audits/hub-safety-2026-10-05/partH/vouch.txt create mode 100644 documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md create mode 100644 documentation/tests/golden-0.296.0-2026-10-05/02-round-trip.txt create mode 100644 documentation/tests/golden-0.296.0-2026-10-05/README.md create mode 100644 documentation/tests/golden-0.296.0-2026-10-05/bake.log diff --git a/CONTEXT.md b/CONTEXT.md index 30996970..6b1eba3e 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,18 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller +> v0.296.0, agent v0.146.1 + bundle `42333e96…`, golden 0.296.0 vouched with agent 0.146.1, min_agent 0.131.0).** CC decisions +> 119–124, *operator may reverse*. Hub: R-135 a cookie-less state change needs Basic + `X-Felhom-Operator` (`05` §16.1); R-133 +> console passwords sealed with the off-site seal/key (`05` §16.2; 4 rows sealed live); R-604/R-530 the System page's "Version +> floors" table + Agent cell, `agent_behind` (7 d) and `floor_raise_skipped` (`05` §5, `08` §6.3); R-508 no-e-mail banner. +> Controller: R-519 `run_record.go` (a cut run said until a complete one; live cut refused by the permission check — operator +> asked), R-518 copy with today's measurement (5 min 47 s, demo-hp). Agent: R-861 narrowed (`03` §3.1 — exact sudo regexes, +> `felhom-priv-apply`, fixed hook/parent files, signed update verified as root, escrow paths pinned; three residuals named); +> **R-880: a bundle that adds a path needs a STEP bundle** (`scripts/build-step-bundle.py`, `runbooks/config-bundle.md`). +> R-173 measured (backed up only on DooPlex, by label drift) + `runbooks/RUNBOOK-hub-db-offsite-backup.md`; decision in +> STATUS. Closed R-133, R-135, R-508, R-509, R-530, R-604, R-880; opened R-879, R-881; register 336 → 332. + > **2026-10-05 (afternoon) — a box that is not always on (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0 > vouched with agent 0.145.0, min_agent 0.131.0).** Rulings 109–111 (`09` §3: R-871 option A — a missed night runs once when the > box comes back; the household's banner; Tester 2 read only). CC decisions 112–118, *operator may reverse*. Design `07` diff --git a/REPORT-hub-safety-2026-10-05.md b/REPORT-hub-safety-2026-10-05.md new file mode 100644 index 00000000..e6355671 --- /dev/null +++ b/REPORT-hub-safety-2026-10-05.md @@ -0,0 +1,131 @@ +# REPORT — the hub's own safety, boxes left behind, honest backup wording, two onboarding rows, and the agent's admin permissions (R-135, R-133, R-173/R-232, R-604, R-530, R-518, R-519, R-861, R-508, R-509) — 2026-10-05, late afternoon + +Brief: "an open-items batch — the hub's own safety (CSRF, the console credential at rest, the hub database in +backups), a fleet view that shows boxes left behind, honest backup wording, two stale onboarding rows; and the agent +permission fix (R-861) as its own Part". Evidence: `documentation/audits/hub-safety-2026-10-05/part{A..H}/`, the golden +`documentation/tests/golden-0.296.0-2026-10-05/`. Architecture read before the claims: `05-hub-architecture.md`, +`_hub-review.md`, `04-control-plane-authorization.md`, `03-host-agent.md` §3/§11, `07` §6.1, `08` §6.3, +`runbooks/target-selection.md`, `runbooks/secrets.md`, `runbooks/ep0-datastore-copy.md`, +`audits/RECON-dooplex-backup-2026-08-06.md`. + +Baselines (re-verified at the start): felhom.eu `9bb45eaaa2`, felhom-agent `61345790ed` (v0.145.0), felhom-controller +`477e2548db` (v0.295.0); register 336 rows, highest R-878. + +## 1. The Part table + +| Part | State | Note | +|---|---|---| +| A — CSRF (R-135) | **done** | hub v0.135.0: no session → Basic credentials + `X-Felhom-Operator` (decision 120); every state-changing route in one table (`partA/route-table.md`, 38 routes + an unknown path); red-proof: the old shape lets 39 of 39 through; live: 403 / pass / 401 | +| B — console password at rest (R-133) | **done** | the off-site seal and key reused (decision 121); 4 legacy rows sealed live, 0 left plain; reveal still opens demo-hp's Proxmox; wrong key fails closed; 2 red-proofs. What a DB backup still holds readable → R-879 | +| C — hub DB in backups (R-173, R-232) | **done (read only) — decision with you** | it IS backed up, only on DooPlex, by a label drift; no failure alarm; steps in `runbooks/RUNBOOK-hub-db-offsite-backup.md`; decision in STATUS | +| D — boxes left behind (R-604, R-530) | **done** | System page "Version floors" + Agent cell (live), `agent_behind` 7 d + `floor_raise_skipped` (tests, 3 red-proofs). The mail was not exercised live (needs a global raise) | +| E — honest backup wording (R-518, R-519) | **done, changed** | R-518: the copy was already honest; today's measurement added (5 min 47 s). R-519: dating was already fixed (v0.275.0); the page notice + the synthesised status fixed (v0.296.0, 4 red-proofs). **The live cut on 9202 was refused by the permission check — asked** | +| F — the agent's admin permissions (R-861) | **done, changed** | nine root paths, not four; `03` §3.1 written AFTER the build (not "design first"); agent v0.146.1 (after a review found three holes in v0.146.0); delivered by a two-step bundle (R-880); live on both demo boxes: sudo 93/93, capability 67/67; three residuals named, row stays open narrowed | +| G — onboarding rows (R-508, R-509) | **done** | R-509 closed by three matched real mails; R-508 closed (e-mail set since 09-14 + a new page warning, red-proof) | +| H — release and records | **done, changed** | hub 0.135.0, controller 0.296.0, agent 0.146.1 (+ 0.146.0 never delivered); golden 0.296.0 baked + vouched; floors + signed jobs for demo-hp, demo-felhom, tester-1; docs `00`, `03`, `05`, `07`, `08`, `09`, `11`-runbook | + +## 2. Claims in the brief that turned out wrong + +1. **"The hub's own database is in no backup."** It is in one — only on DooPlex. Longhorn's `backup-daily` / + `backup-weekly` copy `hub-data` every night (last 2026-10-05 02:06 UTC, Completed, 713 MB) to DooPlex's own `sda1`. + R-173's "excluded" is the PVC label (`recurring-job-group.longhorn.io/default: disabled`, set 2026-02-16 with no + reason); the live Longhorn Volume carries `enabled` — a hand-set drift that keeps the backup alive and can be undone + by any sync. Nothing leaves DooPlex, and nothing alarms if it fails (R-232 stands). +2. **"The backup page promises 'a few seconds'."** Not since controller v0.243.0 / v0.267.0: the text already said + "several minutes (about 8 minutes on a 12-app box)". Measured today on demo-hp (9 apps): 5 min 47 s, local tier only. + v0.296.0 adds today's figure and "minutes, not seconds". +3. **"R-519: fix the dating."** The dating was already fixed in controller v0.275.0 (R-696): a restore point carries the + time of its OLDEST part. What was still missing was the sentence on the page, and the page's synthesised "last + database backup … OK" after a restart — both fixed in v0.296.0. +4. **"R-861: four admin-command groups."** It was nine ways to root, not four: besides the four named (guest hook, + intermediary script/unit, escrow, self-update), the mount units, the dnsmasq drop-ins, the WireGuard config and the + OOB sshd config were each installed from agent-written files, and almost every `*` in the arguments matched spaces + (measured with real sudo 1.9.16: 23 of 29 attack lines allowed). +5. **"Deliver the agent fix by the signed bundle."** Not possible in one step: an installed `felhom-os-apply` refuses + a bundle naming a path it does not know (R16), and v0.146.1's bundle adds four. Delivered by a step bundle (R-880). +6. **"Design first" (Part F).** I built first and wrote the `03` §3.1 section after the code, in the same session — the + section records what was built, group by group. + +## 3. Per Part — tests, red-proofs, live proof + +**A.** `hub/internal/web/r135_csrf_test.go` (5 tests). Red-proof `partA/red-proof.txt` (39 of 39 convicted). Live +`partA/live.txt` (hub 0.135.0, ClusterIP): Basic, no header, `Origin: evil` → **403**; unknown path, no header → **403**; +with `X-Felhom-Operator: cli` → **404** (passed the gate); header without credentials → **401**; GET → 200. The skill and the +memory note now carry the header. + +**B.** `store/r133_recovery_seal_test.go` (4), `web/r133_reveal_wrongkey_test.go`, `cmd/hub/r133_wiring_test.go`. Red-proofs +`partB/red-proof.txt` (plaintext save — the first attempt did not compile, re-run with a compiling mutation; the wiring). +Live `partB/live-db.txt`: hub start `console passwords sealed at rest (4 legacy plaintext row(s) sealed now)`; the live DB +copy (scratch, shredded) shows 4 rows `enc:v1:`, 0 not sealed. `partB/live-reveal.txt`: reveal on demo-hp → 200, a +32-char password that minted a PVE ticket (200; a wrong one 401); the timeline event recorded. +**What a hub DB backup now holds:** the console and off-site passwords sealed (useless without `OFFSITE_SECRET_KEY`, which +exists only on DooPlex); still readable: box API keys, owner passphrases + customer API keys, PBS-DR token values (R-879). + +**C.** Readings `partC/readings.txt` (read only). Steps `runbooks/RUNBOOK-hub-db-offsite-backup.md`: keys off the box first; +fix the PVC label in git; a write-only namespace on ep0's PBS; a hub `VACUUM INTO` nightly snapshot (a later hub release); +the encrypted push via the existing tunnel; a weekly restore test (`PRAGMA integrity_check`, row counts, every console +password still sealed); two Prometheus alarms through the existing mail receiver (`absent()` included); a proof run. + +**D.** `osupdates/r530_agent_alarm_test.go` (3), `web/r604_floor_held_back_test.go` (4), `cmd/hub` wiring. 3 red-proofs +(`partD/red-proof.txt`). Live `partD/live-system-page.txt`: global floor 0.292.0, three per-customer floors 0.295.0 (age +"unknown" — set before v0.135.0); Tester 2 `0.142.0 → 0.145.0 (since 2026-10-05)`, the demo boxes "current". + +**E.** `internal/backup/run_record_test.go` (3), `cmd/controller/run_record_wiring_test.go` (2), `TestR518_*`, parity +cases. 4 red-proofs (`partE/red-proof.txt`). Measurement `partE/r518-measure.txt`. The 9202 reproduction: a throwaway +bookstack installed (09:35:56Z) and a complete baseline run (09:37, 35 s); the cut was refused by the permission check; +bookstack removed through the product (`partE/teardown-9202.txt`: no container, volume, folder or backup left). + +**F.** Design `03` §3.1. `configs/test_felhom_priv_apply.py` (32), `AgentUpdate` (8), `SelfupdateWrapperConfinement`, +`StepBundle` (3), Go contract tests (4 packages), `TestSudoersRefusesTheR861Injections`, `TestManifestCoveredBySudoers`. +Red-proofs F1–F9 + S1–S3 (`partF/red-proof.txt`; F1 masked on its first run — strengthened; F3 errored rather than +failed — clean assertion added). Real sudo, container (`partF/sudo-container-proof.txt`): old 23/29 attacks allowed, new +0/29, 64/64 commands allowed. Pre-flight on both boxes' live files: all OK. **Live after the bundle:** `sudo -l` 93/93 on +demo-hp and demo-felhom (`partF/live-sudo-after-*.txt`; before, on demo-felhom: 23 attacks allowed — +`live-sudo-before-demo-felhom.txt`); the checker run as the agent user → SAME on every real file (on demo-felhom the drive unit has no staged copy — an older path wrote it — so that one read `[P1] no staged file`; that box's `/mnt/hdd_1` is the whole-system backup storage, not a household drive, so no bind under `/mnt/felhom-drives` is expected); a staged unit over +`/etc/sudoers.d` refused `[U3]`, nothing installed; the old `install` route → `a password is required`. +**Capability check after the bundle: demo-hp 67/67, demo-felhom 67/67, Tester 1 all ok (hub page), nothing degraded.** +A gap seen on demo-felhom: between the new agent (~10:45 UTC) and the bundle (11:09) its OOB-sshd reconcile logged "install +failed" every minute (the expected gap); 0 errors after the bundle. + +**G.** `partG/r509-real-mails.txt` (three host-delete sends matched to mailbox arrivals within 1 s), `web/r508_no_email_banner_test.go` +(3 branches) + red-proof. + +## 4. Release and delivery + +- **hub v0.135.0** (`3d7a2761` code, `2b30733b` manifest) — built from the pushed commit, ArgoCD Synced/Healthy, image tag + 0.135.0. (A later comment-only change in `server.go` is in this session's docs commit; the image is unchanged by it.) +- **controller v0.296.0** (`ff69074`), MinAgent 0.131.0. Floors 0.296.0 for demo-hp, demo-felhom, tester-1 → both demo + boxes ran 0.296.0 (healthy) within ~30 min; 9202 (scratch) stays 0.295.0. +- **agent v0.146.0** (`6ab1e7c`, released, **never vouched or delivered**) → **v0.146.1** (`fdd8717`) after the review. + Step bundle `0.146.1-step1` (sha `8482851e…`, built from the 0.145.0 bundle, only `felhom-os-apply` replaced). +- **golden 0.296.0** baked and vouched with agent 0.146.1 / min_agent 0.131.0 (`documentation/tests/golden-0.296.0-2026-10-05/`). +- **Per box, signed with felhom-op-1:** `agent_update` 0.146.1 → `agent_config_update` 0.146.1-step1 (`written=1 same=20`, + self-check ok) → `agent_config_update` 0.146.1 (`written=3 same=22`, self-check ok). demo-hp, demo-felhom, Tester 1 all + report agent 0.146.1 and root files 0.146.1 (`partH/fleet-after.txt`). **Tester 2: DOWN all session, nothing sent.** + +## 5. Rows + +Register **336 → 332**. Closed (7): R-133, R-135, R-508, R-509, R-530, R-604, and R-880 (opened and closed today). +Narrowed: R-861 (three residuals), R-173 (measured; waiting on you), R-518 (copy; per-tier quiesce left), R-519 (live cut +left). Opened (2): R-879 (hub.db still holds readable secrets), R-881 (installer uninstall misses `felhom-priv-apply`). +The section counts in `OPEN-ITEMS.md` were recomputed (several were already out of date). + +## 6. Slips of mine, said plainly + +- **Two agent releases** (0.146.0, 0.146.1) against "one per repo". 0.146.0 had three security holes a background review + found after I pushed it; it was never vouched or sent. +- **I did not design Part F first** as the brief asked; the `03` section was written after the build. +- **My first version waiter read the wrong page cell** and reported the boxes as not updated; I re-read the right cell. +- **Two red-proofs did not convict on the first run** (F1 masked, the R-133 plaintext mutation did not compile); both + were fixed and re-run. + +## 7. Teardown, three layers + +- **Machines:** 9202 — the throwaway bookstack removed through the product, nothing left; its controller stays 0.295.0. + The demo boxes keep their real files (the checker reported SAME; the one staged attack file was deleted). + Bake VM: CT 9100 destroyed, token/script/log shredded, qemu stopped, disk back to `virgin`. +- **Hosts:** nothing provisioned. The pre-flight copies of the checker (`/tmp/felhom-priv-apply-check`) and the case + files were removed from both hosts. +- **Hub:** hub 0.135.0 deployed; artifacts vouched (agent 0.146.1, golden 0.296.0); floors 0.296.0 for three customers; + 9 signed jobs (3 × agent_update, 6 × agent_config_update), all consumed. The step package `felhom-agent/0.146.1-step1` + stays published on purpose (Tester 2 will need it). No customer or appliance record created. diff --git a/STATUS.md b/STATUS.md index d4ee7fd6..f6ad432e 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,13 +1,61 @@ # STATUS — what works, what's broken, what's next -**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) stayed offline all day; nothing -was sent to it.** +**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was +sent to it.** -**Updated 2026-10-05 (afternoon, the catch-up session): every box of ours healthy. Built and proven live: a box that was -off at its backup time makes the backups up once when it comes back; the household's banner; the OS update repairs -itself after a power cut (second crash on demo-hp, with your go). Report: `REPORT-catchup-2026-10-05.md`.** +**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form +posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make +itself root. Report: `REPORT-hub-safety-2026-10-05.md`.** -## Today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut +## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights + +**Decisions I took myself (you may reverse each — `09` decisions 119–124):** +- A box behind the approved agent for 7 days raises an alarm to you; a global version raise that cannot move a box + sends you one mail naming it. +- Scripts that post to the hub with the password must now send one extra header; a web page on another site cannot. +- The console passwords are locked with the same key as the off-site passwords (one key to keep safe, not two). +- The agent's rights are narrowed with exact rules and one checking helper, not one helper per command. +- A cut-off backup is shown on the backup page until a backup runs all the way through. +- A new agent whose root files add a file is delivered in two signed steps (the old box would refuse it in one). + +**What works now (proven live):** +- **Form protection:** a password post without the header is refused (403); a browser on another site cannot add it. +- **Console passwords locked:** all 4 were sealed at the hub's start; the demo-hp one still opens its Proxmox (checked). +- **Boxes left behind:** the System page lists the three per-box version floors and shows Tester 2's agent 4 releases + behind. The alarm and the mail are proven by tests only (they need 7 days / a global raise). +- **The agent cannot make itself root any more:** before, the real sudo let 23 of 29 attack commands through; now 0, on + demo-hp and demo-felhom, and every agent feature still passes its check (67 of 67) on all three boxes. +- Agent 0.146.1, controller 0.296.0 and hub 0.135.0 on demo-hp, demo-felhom and Tester 1; new-install image 0.296.0. + +**Found today:** +- **The hub database is backed up — but only inside DooPlex**, and only because a hand-set label says so; nothing tells + anyone if that backup fails. (Your decision below.) +- **A new agent's root files could not reach any box in one step** (an older box refuses files it does not know). Fixed + with a two-step delivery; written down for next time. +- **A security review of my own agent change found three holes** before it went to any box; fixed in a second agent + release (0.146.1). Two agent releases today, against "one per repo" — the first was never sent anywhere. +- **The backup page already said "about 8 minutes"**, not "a few seconds". Measured today on demo-hp (9 apps): about 6 + minutes. Both figures are on the page now. + +**Needs you:** +1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings): + - **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an + alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager. + - **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A. + - **If you decide nothing:** the database stays only on DooPlex. A fire or theft there loses every box's console + password, the escrow records and the customer settings; each box would need re-pairing by hand. Steps: + `documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`. + - **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy + of the database cannot open the console passwords. +2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the + controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays + proven by tests only. +3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own + guest runs; install the operator SSH key for the limited `felhom-op` user; see the box's backup key during the + recovery-code ceremony. If you do nothing, they stay as they are until before the first paying customer. +4. **Tester 2's one-time step** is unchanged (below). + +## Earlier today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut **Decisions I took myself (you may reverse each — `09` decisions 112–118):** - The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 538cc62a..3a287d45 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -196,10 +196,11 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 | | A second login step for the dashboard (a TOTP code or a passkey) | — | **MISSING** | — | One password, one bcrypt hash (`controller/internal/web/auth.go:37-44`) → R-811 (added 2026-10-03) | | The household can leave Felhom, or outlive it — the box runs without the hub, the household owns its domain, tunnel and off-site account, and can export everything | — | **MISSING** (as a written answer) | the LOST-hub half only: `_recovery-inventory-2026-07-28.md` §D2.4, `07` §8 row 11b | Leaving and hand-over are answered nowhere → R-810 (spike, added 2026-10-03) | +| **The host agent cannot reach root without the operator key: exact sudo patterns, no agent-written file installed where root reads it without a content check, the agent binary only by an operator-signed update checked as root** | agent **v0.146.1** (R-861; delivered by a step bundle, R-880) | **PROVEN-LIVE on both demo boxes (2026-10-05) — with three named residuals** | `audits/hub-safety-2026-10-05/partF/` (real sudo: 23 of 29 attacks allowed before, 0 after; 64 capability commands allowed; `sudo -l` 93/93 on demo-hp and demo-felhom after the bundle; capability probe 67/67; a staged unit over `/etc/sudoers.d` refused live) | `03` §3.1: the controller-swap image ref (guest-scoped), the felhom-op SSH key (hub-delivered, unsigned; felhom-op's sudo is scoped), the escrow ceremony relays R → R-861 | | WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` | | The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe | | The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing | -| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → **R-133** | +| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. **Sealed at rest since hub v0.135.0 (R-133 CLOSED):** the off-site seal and key; live 2026-10-05: 4 legacy rows sealed at start-up, 0 left plain, and the demo-hp reveal still minted a PVE ticket (HTTP 200; a wrong password 401) — `audits/hub-safety-2026-10-05/partB/`. A database backup now needs `OFFSITE_SECRET_KEY` too → R-173 | ## F. Notifications & monitoring @@ -233,6 +234,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) | | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | +| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live | +| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | | | Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | | **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. | diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 1f379644..84aa8dff 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -72,6 +72,59 @@ Explicitly does **not**: - **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table. - ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins"). +### 3.1 The agent's admin commands, group by group — and how each is narrowed (R-861, agent v0.146.1) `[DESIGN — 2026-10-05, CC; operator may reverse]` + +**The question.** Can a compromised agent PROCESS (running as the `felhom-agent` user) become root on its host without +the operator's key? Read on 2026-10-04 (R-861) and measured 2026-10-05 with the real sudo 1.9.16 in a throwaway +container: **with the v0.145.0 sudoers, yes — 23 of 29 attack command lines were allowed.** Two shapes did it: + +1. **A glob in the arguments.** Sudo's `*` in arguments also matches spaces, so one grant smuggled extra options: + `pct set [0-9]* -onboot 1` allowed `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest); + `mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` allowed a `..` path onto `/etc/sudoers.d`; `nft add element … + *` allowed `; flush ruleset`. +2. **A file the agent wrote, installed where root reads it.** A `.mount` unit (bind any directory over `/etc`), a + dnsmasq drop-in (`dhcp-script=` runs as root), the wg-quick config (`PostUp=` runs as root), the OOB sshd config + (`AuthorizedKeysFile` + `StrictModes no`), the guest pre-start hook (Proxmox runs it as root), the shared-parent boot + script, and the agent BINARY itself (the escrow ceremony and the guest hook run it as root; `apply` took a sha the + agent passed). + +**The rule after v0.146.1** (the R-861 fix direction: each becomes a root-owned wrapper that checks its own input, or a +fixed file, delivered by the signed config bundle): + +- Every varying argument list is a **sudo regular expression** (`^…$`): one value per slot, a fixed character set, no + `..`, no extra argument. Literal lines stay literal. +- **No file the agent wrote is installed where root reads it.** Either the content is FIXED and comes with the signed + bundle (the hook, the shared parent), or a root wrapper checks the CONTENT against the agent's own renderers before + installing it (`felhom-priv-apply`), or the operator's signature is checked as root (`felhom-os-apply agent_update`). +- The pins: `TestManifestCoveredBySudoers` (every command the agent runs is still allowed), `TestSudoersRefusesTheR861 + Injections` (the 29 attacks are not), the real-sudo run of both (`audits/hub-safety-2026-10-05/partF/ + sudo-container-proof.txt`), and `sudo -l -U felhom-agent` on both demo boxes after the bundle. + +| Group | What it is for | How it is narrowed (v0.146.1) | Left open | +|---|---|---|---| +| `FELHOM_MOUNT` | fs-UUID mount units for enrolled drives | install only via `felhom-priv-apply unit `: `[Unit]` only Description + `After=local-fs-pre.target`, `Where=` `/mnt/` or `/mnt/felhom-drives/` and equal to the unit name, `What=` a UUID or a network source, no `bind`/`suid`/`dev`, no continuation lines; systemctl verbs on `mnt-…\.mount` only | — | +| `FELHOM_NETMOUNT` | NAS automount pairs, re-arm, clean-up | same checker; a network share must carry `nosuid,nodev` (the agent now renders them); `rm`/`rmdir`/`reset-failed` one exact name | — | +| `FELHOM_DISK` | SMART, thin-pool, PV and pool reads | exact device / LV patterns (no extra options such as `smartctl -s off`, `lvs --config`) | read-only | +| `FELHOM_PROVISION` | bootstrap config mount, autostart | exact `mpN` spec (`…/guests//bootstrap,mp=/…[,ro=1]`) — no smuggled `--dev0` | — | +| `FELHOM_FORMAT` | data-bearing probe + guarded mkfs | one device path, no space; `mkfs` only through `felhom-mkfs-guarded` (its own root checks) with `ext4`/`xfs` | blkid/lsblk read any `/dev` path (read-only) | +| `FELHOM_DNSMASQ` | the LAN split-horizon resolver | drop-ins via `felhom-priv-apply dnsmasq` (only `bind-interfaces`, `no-resolv`, `listen-address`, `server`, `local`, `address`); `rm` one exact name; exact `pct exec` reads | — | +| `FELHOM_GUESTHOOK` | the pre-start self-heal hook | the hook is a FIXED bundle file; the agent only checks it (`SnippetReady`) and registers it; exact vmid/slot | — | +| `FELHOM_INTERMEDIARY` | the shared drive parent + live drive binds | boot script + unit are FIXED bundle files (the agent only enables the unit); one-segment drive names (no leading dot, no `..`) | — | +| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen (pinned-registry) image ref and restart the guest's bootstrap — the household's data, not host root | +| `FELHOM_STALELOCK` / `FELHOM_SCRATCH_TEARDOWN` | stale-lock clear; failed restore-test scratch | exact vmid; the scratch band `99000[0-9]` was already exact | — | +| `FELHOM_WG` | the off-site tunnel | conf via `felhom-priv-apply wg` (only the keys `renderConf` writes; no `PostUp`/`PreUp`/`DNS`/`Table`; `/32` only) | — | +| `FELHOM_SELFUPDATE` | commit / rollback of the A/B flip | **`apply` removed**: the flip runs only inside `felhom-os-apply agent_update`, after the operator signature, host, window and nonce are checked as root and the staged bytes are hashed ONCE and copied to a root-owned dir (`/var/lib/felhom-os-apply/agent-update/`); the wrapper accepts only that dir | — | +| `FELHOM_SSHD` | the out-of-band operator sshd | config via `felhom-priv-apply sshd-config` (the ONE template, only the Port varies, never 22); the felhom-op key via `sshd-key` (one plain key, no `command=`/`from=` options) | felhom-op's key itself is hub-delivered, not signed: a compromised agent can install its own key for **felhom-op** — whose sudo is scoped (`felhom-op.sudoers`), not root | +| `FELHOM_OOB` | the OOB firewall sets | `add element` takes exactly `{ [/n] }` or `{ }` — no chained command | — | +| `FELHOM_PBSDR` / `FELHOM_BACKUPTARGET` | PBS DR entry; whole-system backup target | unchanged: the arguments stay coarse, and the root wrappers (`felhom-pbs-apply`, `felhom-backup-target-apply`) are the gate (fixed verbs, own validation) | coarse argv into a checking wrapper | +| `FELHOM_ESCROW` | the recovery-code ceremony (runs the agent binary as root) | the binary is only ever an operator-signed one (`FELHOM_SELFUPDATE`); as root it pins the PVE secret dir and the WG state dir, refuses a storage id that is a path, and reads its two staged files by walking the path with `openat(O_NOFOLLOW)` (no symlink anywhere) | **by design the agent relays R**, so a compromised agent can still learn this box's PBS key through the ceremony — not root, but the backup key | +| `FELHOM_SELFHEAL` / `FELHOM_GUESTNET` / `FELHOM_OSAPPLY` | networking restart; guest DHCP watchdog; OS updates | exact; `felhom-os-apply --plan …` stays the glob line on purpose — the bundle's own self-check reads that exact text, and the wrapper refuses any other plan path (R1) | — | + +**What this does not change.** The operator key (`/etc/felhom/operator-signers`, root-owned, never a bundle path) stays +the one trust root; the agent's API token is untouched; a box gets the new rule only through the signed +`agent_config_update` (the order: signed `agent_update` first, then the bundle — after the bundle, an agent below +0.146.0 cannot update itself on that box). + ## 4. Control model — reconcile + signed destructive ops Two channels, split by **reversibility**, not by transport. @@ -655,7 +708,8 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea runs as root at guest start (and `pct reboot` is granted); `FELHOM_INTERMEDIARY` installs a script and a systemd unit that run as root at boot; `FELHOM_ESCROW` runs the agent binary as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. So a compromised agent PROCESS is root on its host; the root-owned trust files (decision 93, - the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. + the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. **Narrowed in agent + v0.146.1 — §3.1 lists every group, the rule now, and what stays open.** - **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle, so the **agent updates the controller**: snapshot-before-update (free rollback, because the controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback diff --git a/documentation/architecture/05-hub-architecture.md b/documentation/architecture/05-hub-architecture.md index 1b2c9907..586c64ce 100644 --- a/documentation/architecture/05-hub-architecture.md +++ b/documentation/architecture/05-hub-architecture.md @@ -130,6 +130,17 @@ R-216. Either way the box's reported agent must meet the chosen MinAgent, else t the Hosts page and logged once per change as `managed floor SERVED`. The rules the operator follows: `runbooks/publish-train-rules.md` rule 1. +**Boxes left behind (hub v0.135.0, R-604 + R-530).** A per-customer floor wins over the global one, so a global +raise does not move a box whose OWN floor is lower — and `managed floor SERVED` is logged once per change, so that box +was silent (demo-hp missed four raises, 2026-09-21). Now the raise logs one line per such customer and sends ONE +operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set +(`customer_configs.min_controller_set_at`; "unknown" for one set before v0.135.0), and the System page's "Version +floors" table lists the global floor, every per-customer floor with its age, and which ones the global cannot move. +Agents are a separate train (they update only by a per-box signed job, R-530's ruling): the System page shows each +box's agent against the vouched one ("0.142.0 → 0.145.0 (since …)", amber, red after the wait), and a box behind the +vouched agent for 7 days raises `agent_behind` (warning, operator-only; `OS_ALARM_AGENT_BEHIND_AFTER`; the clock starts +when the hub first sees the box behind). `[DESIGN — CC 2026-10-05, operator may reverse; `09` decision 119]` + ## 6. Authorization — signed-op queue + editing flow Implements Part 4's gate on the hub side. The hub holds **no signing key**. @@ -403,3 +414,35 @@ count appeared was the bind page's passphrase hint, and its English half is now phrase you received from your operator during setup") — because "five words" stops being true for an English household, and was already wrong for one whose passphrase predates this release. The Hungarian „öt szó" is correct and unchanged. + +## 16. The operator surface's own safety [DESIGN — hub v0.135.0, CC 2026-10-05, operator may reverse] + +### 16.1 Form protection (CSRF) on both login paths (R-135) + +The operator logs in two ways: a browser session (`hub_session` cookie + a per-session token on every form) and HTTP +Basic for scripts. Until v0.135.0 a state-changing request with NO cookie skipped the token check — on the reasoning +that it must be a script. It need not be: a browser caches Basic credentials per origin and resends them on a +cross-site form POST (SameSite does not govern the Authorization header). Now a request without a session passes +only with Basic credentials AND the header `X-Felhom-Operator` (any value; scripts send `cli`). A page on another site +cannot add a custom header without a CORS preflight, which the hub never answers. The gate sits in `ServeHTTP` before +the route switch, so it covers every route at once (38 state-changing routes + an unknown path in +`r135_csrf_test.go`); `/login` and the public `/bind/` stay exempt (no operator session to ride; the bind token +is the capability). The other choice — dropping browser-usable Basic auth entirely — was not taken: CC's headless runs +and the runbooks drive the hub with Basic auth, and the header costs them one flag. + +### 16.2 Secrets at rest in `hub.db` (R-133, R-821) + +Two columns are SEALED (AES-256-GCM, `enc:v1:` + nonce, one key: `OFFSITE_SECRET_KEY` from `Secret/offsite-secret-key`, +never in the database or git): the off-site sub-account passwords (`one_time_secrets.value`, v0.127.0) and, since +v0.135.0, each box's break-glass console password (`host_recovery.secret`) — the same helpers, not a second scheme. +Legacy rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, measured live: 4 rows). No key → a save is +refused; a wrong key → a reveal is a 500 with nothing in the body or the log. Both retrieval paths (the operator page and +the global-key API) open through `GetHostRecoveryCredential`, so the break-glass route still works with the UI down. + +**What a copy of `hub.db` still holds readable** (R-879): each box's hub API key, each household's owner passphrase +(`customer_configs.retrieval_password`) and API key, and the PBS-DR token values (`host_pbs_secrets.value`, kept after +use). **And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only +with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export +on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md` +(R-173 — awaiting the operator's decision). + diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index da44dfe1..cbadf1cd 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -133,6 +133,16 @@ startup a record still marked running becomes a failed, interrupted result („A restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the precedent is kept for what it was written for. +**A backup RUN cut off by a power cut or a restart is said (R-519, controller v0.296.0, `09` decision 123).** The +same shape for the app-data run: `appdata-run.json` beside the restore record, written at both ends of a run. A start +that finds it still running turns it into a notice on /backups and /backups/apps („A legutóbbi mentés (…) megszakadt, +mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every step OK. Measured BIGNIGHT F2 (2026-09-14): +before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql` +the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was +already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump. +*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the +operator (R-519 narrowed).* + ### Lane 2 — the operator: guest and host recovery Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator @@ -661,6 +671,11 @@ successes only. After an agent restart the success is read back from the tier's **What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped. **Still open:** quiescing per tier, so a slow second tier does not keep every app down. +**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the +local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the +longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both +measurements (≈6 min / 9 apps, ≈8 min / 12 apps), say "minutes, not seconds", and that the off-site copy in the same +run makes it longer. ### 6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, `09` §3 decision 36) diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index f13b7dae..bc9cd0e5 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -327,6 +327,8 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by | `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` | | `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` | | `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` | +| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` | +| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` | - **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report breaks a `not_running` run. @@ -334,9 +336,9 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches what the box cannot heal — a running container with no connection (wrong token, blocked network). - The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the - next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`, - `OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended, - operator may reverse* (`11` §8.3). + next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`, + `OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`, + `OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119). --- diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index d8f2491b..b4a3ecc6 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -853,6 +853,34 @@ its length, and both fixes cost something the household would notice — operato a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs one call (pinned by a test). agent v0.145.0. +### 2026-10-05 (afternoon) — decided by CC — operator may reverse (the hub-safety / R-861 brief) + +119. **Boxes left behind (R-604, R-530).** Options for the agent alarm's wait: 3 days (a box off for a long weekend + alarms), 7 days (one week, the same wait as the bundle-behind alarm, R-840), 14 days. **Chosen 7 days** + (`OS_ALARM_AGENT_BEHIND_AFTER`), counted from when the hub first sees the box behind. A global floor raise names, + in one operator mail, every box whose own LOWER floor it cannot move. hub v0.135.0, `05` §5. +120. **How the Basic-auth operator path is protected from cross-site POSTs (R-135).** Options: (a) drop browser-usable + Basic auth (CC's headless runs and every runbook POST break); (b) require a custom header on a cookie-less + state change (a browser cannot add one cross-site without a CORS preflight the hub never answers; scripts add one + flag). **Chosen (b)**, header `X-Felhom-Operator`. hub v0.135.0, `05` §16.1. +121. **The console password's seal (R-133).** Options: (a) a second key and scheme for `host_recovery`; (b) the off-site + seal and key already in force (R-821). **Chosen (b)** — the brief asked for the existing pattern, and one key is + one custody question. Consequence named: a database backup needs this key off DooPlex too (R-173). `05` §16.2. +122. **How the agent's admin commands are narrowed (R-861).** Options per group: (a) a root wrapper per group with its + own argv; (b) exact sudo regex patterns for every varying argument + ONE content checker for every agent-written + file root reads + fixed bundle files where the content never varies + the signed update verified as root by the + existing `felhom-os-apply`. **Chosen (b)** — fewest new root programs, and the checker is pinned to the agent's + own renderers by contract tests. Delivery order: signed `agent_update` first, then the bundle. agent v0.146.1, + `03` §3.1. +123. **How long a cut-off backup run is said on the page (R-519).** Options: until the next run of any kind (a failed + run would clear the warning), until the next run that ends with every step OK, or until dismissed. **Chosen: until + a run ends with every step OK.** controller v0.296.0. +124. **How a bundle that adds a path reaches a box (R-880).** An installed `felhom-os-apply` refuses any path not in its + own table (R16). Options: (a) copy the new wrapper onto each box by hand as root (does not scale, leaves the signed + route); (b) a STEP bundle — the box's current bundle with only `felhom-os-apply` replaced, published as + `-step1` — then the release's bundle, both by signed jobs. **Chosen (b)**, `felhom-agent/scripts/build-step-bundle.py`; + delivered to demo-hp, demo-felhom and Tester 1 on 2026-10-05. `11` §5.4.2 rule unchanged. + ### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief) 100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) — diff --git a/documentation/audits/hub-safety-2026-10-05/partA/live.txt b/documentation/audits/hub-safety-2026-10-05/partA/live.txt new file mode 100644 index 00000000..61e496ad --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partA/live.txt @@ -0,0 +1,8 @@ +== Part A live, hub 0.135.0, 2026-10-05T09:15:34Z, ClusterIP, Basic auth (password from the credentials file, not printed) +POST /configuration/global-floor, Basic, NO header, Origin evil (empty form): 403 +POST /no-such-route, Basic, NO header: 403 +POST /no-such-route, Basic + X-Felhom-Operator: cli (passes the gate → router 404): 404 +POST /no-such-route, header but NO credentials: 401 +GET /system, Basic, no header: 200 +2026/10/05 11:15:34 [WARN] CSRF rejected: POST /configuration/global-floor from 10.42.0.1:36344 +2026/10/05 11:15:34 [WARN] CSRF rejected: POST /no-such-route from 10.42.0.1:41323 diff --git a/documentation/audits/hub-safety-2026-10-05/partA/route-table.md b/documentation/audits/hub-safety-2026-10-05/partA/route-table.md new file mode 100644 index 00000000..0e794378 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partA/route-table.md @@ -0,0 +1,50 @@ +# Hub state-changing routes and how each is protected (hub v0.135.0, R-135) + +Every route below is reached through `RequireAuth` → `ServeHTTP`; the CSRF gate is the first thing `ServeHTTP` does for any +method other than GET/HEAD/OPTIONS, before the route switch — so the protection is the same for every route, and an +unknown path is refused by the gate before it can 404. + +| Route (representative path) | Protected how | Test | +|---|---|---| +| `POST /configuration` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /apps/demo/reset-telemetry` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /apps/demo/dismiss-issues` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/endpoints` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/endpoints/1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /appliances/1/bind` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /appliances/1/discard` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /hosts/h1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /hosts/h1/reveal-recovery-credential` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /hosts/h1/request-logs` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/block` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/selfbind-link` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/unblock` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/geo/disable` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/create-config` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /customers/c1/request-log-tail` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/new` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configuration/global-floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configuration/artifacts` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configuration/password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/edit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/offsite-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/claim-resend` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/pbsdr-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/offsite-freeze` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/regen-password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /configs/c1/reset` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/remove-unpinned/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/abandon-cancel/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/window-grant/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/windows-enabled` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /offsite/key-audit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /os/ring/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /os/enabled/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /os/approve-now` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /os/approve-docker` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /no-such-route` | the gate, before routing (not a route) | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` | +| `POST /login` | exempt (no session to ride; a wrong password is 401) | — | +| `POST /bind/` | exempt (public self-bind; the e-mailed URL token is the capability, rate-limited) | existing `selfbind_test.go` | +| `GET` routes | not gated by design (a GET must not change state). **Not audited in this session** for a GET that writes — the route switch sends POST-only actions to handlers that check `MethodPost`, but the GET renderers were not read line by line | `TestR135_GetIsNotGated` | diff --git a/documentation/audits/hub-safety-2026-10-05/partB/live-db.txt b/documentation/audits/hub-safety-2026-10-05/partB/live-db.txt new file mode 100644 index 00000000..21f00e17 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partB/live-db.txt @@ -0,0 +1,6 @@ +== Part B live, 2026-10-05T09:16:00Z: live hub.db copied to scratch, only prefix + length selected, copy shredded after +Tester-2-be8404|enc:v1:|87|2026-10-04 16:07:15 +demo-felhom-8363b5|enc:v1:|87|2026-07-18 16:30:41 +demo-hp-bb76ea|enc:v1:|87|2026-07-21 16:24:27 +tester-1-d70be4|enc:v1:|87|2026-10-04 19:40:27 +rows NOT sealed: 0 diff --git a/documentation/audits/hub-safety-2026-10-05/partB/live-reveal.txt b/documentation/audits/hub-safety-2026-10-05/partB/live-reveal.txt new file mode 100644 index 00000000..1725244d --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partB/live-reveal.txt @@ -0,0 +1,6 @@ +== Part B live reveal, 2026-10-05T09:16:15Z: POST /hosts/demo-hp-bb76ea/reveal-recovery-credential (Basic + X-Felhom-Operator), body to a 0600 scratch file, shredded after +reveal HTTP 200 +username root@pam password length 32 set_at 2026-07-21T16:24:27Z +the revealed password logs in to demo-hp's Proxmox API (POST /api2/json/access/ticket, root@pam): HTTP 200 +control, a wrong password: HTTP 401 +2026/10/05 11:16:15 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars) diff --git a/documentation/audits/hub-safety-2026-10-05/partC/readings.txt b/documentation/audits/hub-safety-2026-10-05/partC/readings.txt new file mode 100644 index 00000000..3fc16c06 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partC/readings.txt @@ -0,0 +1,42 @@ +== Part C readings on DooPlex, READ ONLY, 2026-10-05T09:17:24Z +-- where the hub database lives +hub-data pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 1Gi longhorn +-rw-r--r-- 1 root root 374534144 Oct 5 11:14 hub.db +-rw-r--r-- 1 root root 32768 Oct 5 11:16 hub.db-shm +-rw-r--r-- 1 root root 313152 Oct 5 11:16 hub.db-wal + 973.4M 373.4M 584.1M 39% /data +-- the exclusion label: PVC (git, manifests/hub.yaml:47, commit 868e8465 2026-02-16 'updated hub yaml', no reason given) vs the live Longhorn Volume +PVC label: disabled +Volume label: enabled +-- recurring jobs +backup-daily backup 0 4 * * * 1 [default] +backup-weekly backup 0 5 * * 0 1 [default] +-- backups of the hub volume (Longhorn backupstore) +2026-10-04T03:05:01Z Completed 708837376 +2026-10-05T02:06:26Z Completed 713031680 +-- backup target +nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc?nfsOptions=soft,timeo=330,retrans=3 true +-- which disk holds the target +/dev/sda1 +/dev/sdb1 +-- DooPlex's own backup service +Mon 2026-10-05 11:30:00 CEST 12min Mon 2026-10-05 11:15:00 CEST 2min 25s ago backup-freshness.timer backup-freshness.service +Tue 2026-10-06 03:19:15 CEST 16h Mon 2026-10-05 03:15:21 CEST 8h ago dooplex-backup.timer dooplex-backup.service +Result=success +ExecMainStatus=0 +-- what tells anyone when a backup fails +80:export NOTIFY_ON_FAILURE="true" +81:# export NOTIFY_WEBHOOK_URL="https://your-webhook-url" +137: if [ "${NOTIFY_ON_FAILURE}" = "true" ] && [ -n "${NOTIFY_WEBHOOK_URL}" ]; then +140: "${NOTIFY_WEBHOOK_URL}" || true +prometheus rule backup-freshness-alerts.yml MinecraftBackupStale +prometheus rule backup-freshness-alerts.yml BackupFreshnessExporterDead +prometheus rule longhorn-alerts.yml LonghornVolumeSpaceCritical +prometheus rule longhorn-alerts.yml LonghornVolumeSpaceWarning +prometheus rule longhorn-alerts.yml LonghornVolumeDegraded +prometheus rule longhorn-alerts.yml LonghornNodeStoragePressure +(no rule watches a Longhorn BACKUP's success or age, nor dooplex-backup.service; the only backup-freshness rule is MinecraftBackupStale) +-- does anything leave DooPlex for the hub DB? the off-site route that exists today: ep0 PBS reached through felhom-ep0-pbs-tunnel (pull only, ep0 -> DooPlex) +active +/usr/bin/proxmox-backup-client +/usr/bin/sqlite3 diff --git a/documentation/audits/hub-safety-2026-10-05/partD/live-system-page.txt b/documentation/audits/hub-safety-2026-10-05/partD/live-system-page.txt new file mode 100644 index 00000000..e4b650ad --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partD/live-system-page.txt @@ -0,0 +1,9 @@ +== Part D live: GET /system on hub 0.135.0 (Basic auth); extracted, no tokens +Version floors: Version floors Global controller floor: 0.292.0 · vouched agent: 0.145.0 Customer Own floor Set Global floor moves it? demo-felhom Demo Ügyfél 0.295.0 unknown no — its own floor applies (at or above the global) demo-hp Demo HP 0.295.0 unknown no — its own floor applies (at or above the global) tester-1 Tester 1 0.295.0 unknown no — its own floor applies (at or above the global) +Agent cell Tester-2-be8404: [('c-warn', '0.142.0 → 0.145.0 (since 2026-10-05)', '3 minor releases behind — sign an agent_update for this box')] +Agent cell demo-felhom-8363b5: [] +Agent cell demo-hp-bb76ea: [] +Agent cell tester-1-d70be4: [] +0.145.0 +0.145.0 +0.145.0 diff --git a/documentation/audits/hub-safety-2026-10-05/partE/r518-measure.txt b/documentation/audits/hub-safety-2026-10-05/partE/r518-measure.txt new file mode 100644 index 00000000..f2781a09 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partE/r518-measure.txt @@ -0,0 +1,28 @@ +== R-518 measure, demo-hp guest 9201, controller 0.295.0, 2026-10-05: POST /api/guest-backup/trigger (the button's call) at 09:19:05Z +2026/10/05 09:19:07 backup_handlers.go:349: [INFO] [web] manual whole-guest backup triggered (quiesce loop) +2026/10/05 09:19:07 quiesce.go:427: [INFO] [quiesce] manual backup requested — quiescing now +2026/10/05 09:19:08 quiesce.go:517: [INFO] [quiesce] backup due on 2 tier(s) — quiescing 9 stack(s): [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin] +2026/10/05 09:19:29 quiesce.go:566: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 started — polling +2026/10/05 09:23:44 quiesce.go:210: [INFO] [quiesce] a backup cycle is already running — skipping this scheduled check +2026/10/05 09:24:09 quiesce.go:643: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 done — next tier may start (app still quiesced) +2026/10/05 09:24:09 quiesce.go:307: [INFO] [quiesce] tier felhom-pbs is BUSY — the agent refused the backup because a concurrent heavy operation holds it. This is contention, NOT a failure: the tier stays due and retries in 15m0s (contended for 0s) +2026/10/05 09:24:09 quiesce.go:504: [INFO] [quiesce] unquiescing (last tier is busy — deferring to a later cycle): restarting 9 stack(s) +-- container StartedAt after the backup (the apps the quiesce stopped): +2026-10-05T09:24:10.019684173Z adventurelog-postgres +2026-10-05T09:24:10.200085281Z adventurelog-frontend +2026-10-05T09:24:15.705236204Z adventurelog +2026-10-05T09:24:16.331987016Z bentopdf +2026-10-05T09:24:16.974996127Z bookstack-db +2026-10-05T09:24:22.669660366Z bookstack +2026-10-05T09:24:23.521923837Z calibre-web +2026-10-05T09:24:24.437773353Z docmost-postgres +2026-10-05T09:24:24.631250114Z docmost-redis +2026-10-05T09:24:35.395095325Z docmost +2026-10-05T09:24:36.934564476Z kimai-db +2026-10-05T09:24:42.753714563Z kimai +2026-10-05T09:24:43.466405175Z opengist +2026-10-05T09:24:44.478064534Z paperless-redis +2026-10-05T09:24:44.688668856Z paperless-postgres +2026-10-05T09:24:54.928089575Z paperless-webserver +2026-10-05T09:24:55.739635651Z privatebin +RESULT: stop began 09:19:08Z (9 stacks quiesced), local tier 09:19:29-09:24:09, PBS tier BUSY (skipped, retried later), last app back 09:24:55Z => longest stop 5 min 47 s, shortest ~5 min 02 s. 21/21 containers running afterwards. diff --git a/documentation/audits/hub-safety-2026-10-05/partE/r518-watch.log b/documentation/audits/hub-safety-2026-10-05/partE/r518-watch.log new file mode 100644 index 00000000..932370dd --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partE/r518-watch.log @@ -0,0 +1,15 @@ +09:19:20 running=15 phase=idle +09:19:41 running=4 phase=snapshotted +09:20:02 running=4 phase=snapshotted +09:20:23 running=4 phase=snapshotted +09:20:44 running=4 phase=snapshotted +09:21:05 running=4 phase=snapshotted +09:21:26 running=4 phase=snapshotted +09:21:47 running=4 phase=snapshotted +09:22:08 running=4 phase=snapshotted +09:22:29 running=4 phase=snapshotted +09:22:50 running=4 phase=snapshotted +09:23:13 running=4 phase=snapshotted +09:23:34 running=4 phase=snapshotted +09:23:55 running=4 phase=snapshotted +09:24:16 running=6 phase=done diff --git a/documentation/audits/hub-safety-2026-10-05/partE/red-proof.txt b/documentation/audits/hub-safety-2026-10-05/partE/red-proof.txt new file mode 100644 index 00000000..5d0f0d7d --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partE/red-proof.txt @@ -0,0 +1,39 @@ +== RED-PROOF 1 (R-519): runDBDumpsInternal does not call markRunStarted +=== RUN TestRunRecord_TheRealRunIsOnRecordWhileItRuns + run_record_test.go:78: the run was not on record while it ran — a cut here would go unnoticed +--- FAIL: TestRunRecord_TheRealRunIsOnRecordWhileItRuns (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s +FAIL + +== RED-PROOF 2 (R-519): the synthesised status says Success: true again +=== RUN TestRunRecord_SynthesisedStatusIsNotOKAfterACut + run_record_test.go:97: after a cut the synthesised status still reads OK: &{LastRun:2026-10-05 11:34:19.454323277 +0200 CEST m=+0.001336319 Results:[{DB:{ContainerName:adventurelog ContainerID: DBType: DBUser: DBName: StackName:adventurelog} FilePath:adventurelog-postgres.sql Size:0 Duration:0s Error: Validation:{Valid:false TableCount:0 Error: FileSize:0 ModTime:0001-01-01 00:00:00 +0000 UTC UserTableFound:false UserRows:0 LooksEmpty:false}}] Success:true Duration:0s} +--- FAIL: TestRunRecord_SynthesisedStatusIsNotOKAfterACut (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s +FAIL + +== RED-PROOF 3 (R-519 wiring): main() does not call loadRunRecordAtStartup +=== RUN TestMainWiresRunRecord + run_record_wiring_test.go:47: main() never calls loadRunRecordAtStartup — a cut run is never said on the page +--- FAIL: TestMainWiresRunRecord (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s +FAIL + +== RED-PROOF 4 (R-518): the v0.267.0 copy (12 apps / 8 minutes only) +=== RUN TestR518_BackupButtonStatesTheMeasuredDowntime + r518_backup_downtime_copy_test.go:33: hu: "kb. 6 perc" appears 0 times, want it on the page AND in the confirm + r518_backup_downtime_copy_test.go:33: hu: "nem másodperceket" appears 0 times, want it on the page AND in the confirm + r518_backup_downtime_copy_test.go:33: en: "about 6 minutes" appears 0 times, want it on the page AND in the confirm + r518_backup_downtime_copy_test.go:33: en: "not seconds" appears 0 times, want it on the page AND in the confirm +--- FAIL: TestR518_BackupButtonStatesTheMeasuredDowntime (0.07s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s +FAIL + +== restored +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.013s +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.014s +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s diff --git a/documentation/audits/hub-safety-2026-10-05/partE/teardown-9202.txt b/documentation/audits/hub-safety-2026-10-05/partE/teardown-9202.txt new file mode 100644 index 00000000..4444ec16 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partE/teardown-9202.txt @@ -0,0 +1,7 @@ +== 9202 teardown of the throwaway bookstack (installed 09:35:56Z for the R-519 reproduction), 2026-10-05T11:44:31Z +stop: HTTP 200 +remove: HTTP 200 +containers: 0 +volumes: 0 +stackdir: none +backups: none diff --git a/documentation/audits/hub-safety-2026-10-05/partF/demo-felhom-steps.txt b/documentation/audits/hub-safety-2026-10-05/partF/demo-felhom-steps.txt new file mode 100644 index 00000000..522292f8 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/demo-felhom-steps.txt @@ -0,0 +1,9 @@ +Oct 05 12:54:08 demo-felhom felhom-os-apply[4115602]: os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0 +Oct 05 12:54:08 demo-felhom felhom-os-apply[4115603]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced) +Oct 05 12:54:09 demo-felhom felhom-os-apply[4115702]: os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False +Oct 05 13:09:08 demo-felhom felhom-os-apply[4131155]: os-apply: BUNDLE START agent=0.146.1 sha=42333e969028867a authority=signed files=26 write=3 same=22 kept=1 skipped=0 +Oct 05 13:09:08 demo-felhom felhom-os-apply[4131156]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-selfupdate-guarded (replaced) +Oct 05 13:09:08 demo-felhom felhom-os-apply[4131157]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-priv-apply (new) +Oct 05 13:09:08 demo-felhom felhom-os-apply[4131158]: os-apply: BUNDLE WROTE /etc/sudoers.d/felhom-agent (replaced) +Oct 05 13:09:09 demo-felhom felhom-os-apply[4131255]: os-apply: BUNDLE DONE agent=0.146.1 written=3 same=22 self-check=ok signers-created=False +Oct 05 13:09:09 demo-felhom felhom-agent[4101006]: time=2026-10-05T13:09:09.974+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=67 total=67 degraded="" diff --git a/documentation/audits/hub-safety-2026-10-05/partF/demo-hp-step1.txt b/documentation/audits/hub-safety-2026-10-05/partF/demo-hp-step1.txt new file mode 100644 index 00000000..8a044c77 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/demo-hp-step1.txt @@ -0,0 +1,6 @@ +Oct 05 12:42:18 demo-hp felhom-os-apply[500347]: os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0 +Oct 05 12:42:19 demo-hp felhom-os-apply[500348]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced) +Oct 05 12:42:19 demo-hp felhom-os-apply[500485]: os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False +Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.875+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0" +Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.875+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)" +Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.876+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False" diff --git a/documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-felhom.txt b/documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-felhom.txt new file mode 100644 index 00000000..895af674 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-felhom.txt @@ -0,0 +1,8 @@ +== demo-felhom 2026-10-05T11:09:48Z: the checker as the agent user on the real staged files (SAME = nothing changes) +unit mnt-hdd_1.mount rc=3 +wg rc=0 +sshd-config rc=0 +sshd-key rc=0 +old route: sudo: a password is required +active active active active +5 diff --git a/documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-hp.txt b/documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-hp.txt new file mode 100644 index 00000000..f60489ce --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/live-checker-demo-hp.txt @@ -0,0 +1,20 @@ +== demo-hp 2026-10-05T10:58:16Z: the checker run AS the agent user through sudo, on the real staged files (each must be SAME — nothing changes) +unit mnt-hdd_1.mount rc=0 +wg rc=0 +sshd-config rc=0 +sshd-key rc=0 +dnsmasq rc=0 + +== an ATTACK, live: the agent stages a unit binding its own dir over /etc/sudoers.d (name and Where agree) +attack rc=3 (installed? no) +== the old route, live: sudo -n install of a staged file +sudo: a password is required +== journal +Oct 05 12:58:08 demo-hp felhom-priv-apply[550409]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op +Oct 05 12:58:09 demo-hp felhom-priv-apply[550431]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config +Oct 05 12:58:16 demo-hp felhom-priv-apply[550931]: felhom-priv-apply: SAME unit /etc/systemd/system/mnt-hdd_1.mount +Oct 05 12:58:16 demo-hp felhom-priv-apply[550937]: felhom-priv-apply: SAME wg /etc/wireguard/wg-felhom.conf +Oct 05 12:58:16 demo-hp felhom-priv-apply[550962]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config +Oct 05 12:58:16 demo-hp felhom-priv-apply[550968]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op +Oct 05 12:58:16 demo-hp felhom-priv-apply[550981]: felhom-priv-apply: SAME dnsmasq /etc/dnsmasq.d/felhom-resolver-base.conf +Oct 05 12:58:16 demo-hp felhom-priv-apply[550992]: felhom-priv-apply: REFUSED [U3] unit mnt-..-etc-sudoers.d.mount: mnt-..-etc-sudoers.d.mount: Where=/mnt/../etc/sudoers.d is not /mnt/ or /mnt/felhom-drives/ diff --git a/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-felhom.txt b/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-felhom.txt new file mode 100644 index 00000000..10f512a9 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-felhom.txt @@ -0,0 +1,2 @@ +== LIVE AFTER, demo-felhom, 2026-10-05T11:09:46Z: bundle 0.146.1; 'sudo -l -U felhom-agent ' per case (lists only) +RESULT ok=93 fail=0 diff --git a/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-hp.txt b/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-hp.txt new file mode 100644 index 00000000..bbed320c --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-after-demo-hp.txt @@ -0,0 +1,2 @@ +== LIVE AFTER, demo-hp, 2026-10-05T10:57:55Z: bundle 0.146.1, sudo Sudo version 1.9.16p2; 'sudo -l -U felhom-agent ' per case (lists only) +RESULT ok=93 fail=0 diff --git a/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-before-demo-felhom.txt b/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-before-demo-felhom.txt new file mode 100644 index 00000000..a9023afd --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/live-sudo-before-demo-felhom.txt @@ -0,0 +1,28 @@ +== LIVE BEFORE, demo-felhom (felhom-pve), 2026-10-05T10:54:01Z: bundle 0.145.0 — the v0.145.0 sudoers; 'sudo -l -U felhom-agent ' per case (lists only, runs nothing) +FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'unit' 'mnt-felhom\x2dx.mount' +FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'dnsmasq' '/tmp/felhom-resolver-123456789.conf' 'felhom-x.conf' +FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'wg' +FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1 +FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives +FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda +FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda +FAIL want=DENY got=ALLOW :: /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x +FAIL want=DENY got=ALLOW :: /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d +FAIL want=DENY got=ALLOW :: /usr/bin/umount /mnt/felhom-drives/x / +FAIL want=DENY got=ALLOW :: /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount +FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount +FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh +FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh +FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf +FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf +FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config +FAIL want=DENY got=ALLOW :: /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000 +FAIL want=DENY got=ALLOW :: /usr/bin/systemctl enable --now -- etc-sudoers.d.mount +FAIL want=DENY got=ALLOW :: /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd +FAIL want=DENY got=ALLOW :: /usr/bin/rmdir /mnt/felhom-drives/x /etc +FAIL want=DENY got=ALLOW :: /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset +FAIL want=DENY got=ALLOW :: /usr/sbin/smartctl -a -j /dev/sda -s off +FAIL want=DENY got=ALLOW :: /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x +FAIL want=DENY got=ALLOW :: /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller +FAIL want=DENY got=ALLOW :: /usr/sbin/pct unlock 9201 --whatever +RESULT ok=67 fail=26 diff --git a/documentation/audits/hub-safety-2026-10-05/partF/preflight-live-files.txt b/documentation/audits/hub-safety-2026-10-05/partF/preflight-live-files.txt new file mode 100644 index 00000000..95ad3a1b --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/preflight-live-files.txt @@ -0,0 +1,30 @@ +== hp 2026-10-05T09:57:31Z — felhom-priv-apply --check against the box's LIVE files (read only; prints OK or the rule, never content) +unit mnt-hdd_1.mount: OK +dnsmasq felhom-demo-hp.conf: OK +dnsmasq felhom-guest-9201.conf: OK +dnsmasq felhom-resolver-base.conf: OK +wg: OK +sshd-config: OK +sshd-key: OK +== felhom-pve 2026-10-05T09:57:32Z — felhom-priv-apply --check against the box's LIVE files (read only; prints OK or the rule, never content) +unit mnt-hdd_1.mount: OK +dnsmasq felhom-guest-9201.conf: OK +dnsmasq felhom-resolver-base.conf: OK +wg: OK +sshd-config: OK +sshd-key: OK +== hp 2026-10-05T10:13:26Z — v0.146.1 checker, --check against LIVE files (read only) +unit mnt-hdd_1.mount: OK +dnsmasq felhom-demo-hp.conf: OK +dnsmasq felhom-guest-9201.conf: OK +dnsmasq felhom-resolver-base.conf: OK +wg: OK +sshd-config: OK +sshd-key: OK +== felhom-pve 2026-10-05T10:13:27Z — v0.146.1 checker, --check against LIVE files (read only) +unit mnt-hdd_1.mount: OK +dnsmasq felhom-guest-9201.conf: OK +dnsmasq felhom-resolver-base.conf: OK +wg: OK +sshd-config: OK +sshd-key: OK diff --git a/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt b/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt new file mode 100644 index 00000000..d6b1e685 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/red-proof.txt @@ -0,0 +1,85 @@ +== RED-PROOF F1 (mount units): felhom-priv-apply stops checking Where +test_U3_bind_over_sudoers_dir (__main__.Refuses.test_U3_bind_over_sudoers_dir) ... ok +test_U3_name_must_match_where (__main__.Refuses.test_U3_name_must_match_where) ... ok +test_U3_network_outside_drives (__main__.Refuses.test_U3_network_outside_drives) ... ok +test_U3_traversal_in_where (__main__.Refuses.test_U3_traversal_in_where) ... ok +OK + +== RED-PROOF F2 (WireGuard): felhom-priv-apply allows any key +FAIL: test_W1_postup (__main__.Refuses.test_W1_postup) +FAILED (failures=1) + +== RED-PROOF F3 (self-update, root side): felhom-os-apply agent_update skips the signature +FAILED (errors=1) + +== RED-PROOF F4 (self-update, agent side): the agent calls felhom-selfupdate-guarded apply itself again +--- FAIL: TestExecutor_HappyPath (0.00s) + executor_test.go:111: execute: agent_update: the root wrapper did not apply it: (report: ; stderr: ) +FAIL + +== RED-PROOF F5 (guest hook): SnippetReady accepts any content +--- FAIL: TestSnippetReady (0.00s) + install_test.go:43: a hook with other content read as ready +FAIL + +== RED-PROOF F6 (shared parent): the agent installs the boot script from /tmp again when it differs +--- FAIL: TestSharedParentBoot_NeverInstalls (0.00s) + intermediary_install_test.go:54: both missing: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/001/felhom-shared-parent.sh]] — it must install nothing (R-861) + intermediary_install_test.go:54: script differs: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/002/felhom-shared-parent.sh]] — it must install nothing (R-861) + intermediary_install_test.go:54: unit missing: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/003/felhom-shared-parent.sh]] — it must install nothing (R-861) + +== RED-PROOF F7 (escrow, root read): a staged file is read with os.ReadFile (follows a symlink) +--- FAIL: TestAttach_RefusesASymlink (0.00s) + r861_staged_read_test.go:24: a symlinked staged file was read: ok=true err= value-set=true +FAIL + +== RED-PROOF F8 (network shares): nosuid,nodev dropped from the NFS options +--- FAIL: TestPrivApply_AcceptsTheRenderedUnits (0.35s) + r861_privapply_contract_test.go:31: media .mount: REFUSED [U5] mnt-felhom\x2ddrives-media.mount: a network share must carry nosuid,nodev +FAIL + +== RED-PROOF F9 (the exact patterns): the v0.145.0 sudoers under the injection test +injections the old file allows (Go matcher): 23 + +== restored — the same tests green +ok gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate (cached) +ok gitea.dooplex.hu/admin/felhom-agent/internal/guesthook (cached) +ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi (cached) +ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow (cached) +ok gitea.dooplex.hu/admin/felhom-agent/internal/storage 0.474s +ok gitea.dooplex.hu/admin/felhom-agent/internal/capability 0.177s +OK +OK + +== RED-PROOF F1 (re-run): the first run did NOT convict — the name check (escape(Where)==name) masked it. The test now uses the + pair that only the Where rule stops: name mnt-..-etc.mount + Where=/mnt/../etc (= /etc). Mutation: stop checking Where +FAIL: test_U3_traversal_in_where (__main__.Refuses.test_U3_traversal_in_where) +FAILED (failures=1) + +== RED-PROOF F3 (re-run, clean assertion): felhom-os-apply skips the signature +FAIL: test_a_bad_signature_never_reaches_the_wrapper (__main__.AgentUpdate.test_a_bad_signature_never_reaches_the_wrapper) +AssertionError: None is not true : a job whose signature does not verify was NOT refused: {'agent_update': {'sha256': 'd76b02acf626ce399da7e0a9e17b35563227a4831e14f5edca4ab7cf89eb2c79', 'version': '0.146.0', 'wrapper': '', 'wrapper_rc': 0}, 'layer': 'host', 'mode': 'agent_update', 'pass_seconds': 0.0, 'refused': None, 'release_id': 'agent-0.146.0', 'vmid': 0} +FAILED (failures=1) + +== restored +OK +OK + +=== Review findings 2026-10-05 (background security review of commit 6ab1e7c) — fixed in v0.146.1, each red-proved +== RED-PROOF S1 (TOCTOU): the wrapper gets the agent's path again (hash, then copy by path) +FAIL: test_signed_update_flips_and_burns_the_nonce (__main__.AgentUpdate.test_signed_update_flips_and_burns_the_nonce) +FAILED (failures=1) +== RED-PROOF S1b: the A/B wrapper accepts the agent's staging dir again +FAIL: test_the_agents_staging_dir_is_refused (__main__.SelfupdateWrapperConfinement.test_the_agents_staging_dir_is_refused) +FAILED (failures=1) +== RED-PROOF S2 (allowlist escape): [Unit] accepts Wants=/Requires=/Before= again +FAIL: test_U2_wants_starts_another_unit (__main__.Refuses.test_U2_wants_starts_another_unit) +FAILED (failures=1) +== RED-PROOF S3 (path traversal): open the whole path with O_NOFOLLOW only +--- FAIL: TestAttach_RefusesASymlinkedDirectory (0.00s) + r861_staged_read_test.go:54: a key behind a symlinked directory was read: ok=true err= +FAIL +== restored +OK +OK +ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow 0.008s diff --git a/documentation/audits/hub-safety-2026-10-05/partF/step-bundle.txt b/documentation/audits/hub-safety-2026-10-05/partF/step-bundle.txt new file mode 100644 index 00000000..4a3715c1 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/step-bundle.txt @@ -0,0 +1,4 @@ +== the R-880 step bundle, 2026-10-05T10:24:08Z: base = felhom-agent/0.145.0/felhom-config-bundle.json (sha 78c00adc…, what demo-hp, demo-felhom, tester-1 run); built by scripts/build-step-bundle.py at agent e4b5cf9; published as felhom-agent/0.146.1-step1/felhom-config-bundle.json +step sha 8482851ec27030a7615216048338b1c8939d4e353023ad11c66e6f8930af8613 +round trip sha 8482851ec27030a7615216048338b1c8939d4e353023ad11c66e6f8930af8613 +same paths: True changed: ['/usr/local/sbin/felhom-os-apply'] version: 0.146.1-step1 diff --git a/documentation/audits/hub-safety-2026-10-05/partF/sudo-container-proof.txt b/documentation/audits/hub-safety-2026-10-05/partF/sudo-container-proof.txt new file mode 100644 index 00000000..ad89bd73 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partF/sudo-container-proof.txt @@ -0,0 +1,129 @@ +== R-861 real-sudo proof, sudo 1.9.16p2 (debian:trixie throwaway container on DooPlex, 2026-10-05T09:58:38Z); 'sudo -l -U felhom-agent ' per case + +-- NEW sudoers (agent v0.146.0): every capability must be ALLOW, every attack DENY +ok ALLOW '/usr/bin/lxc-info' '-n' '9201' '-p' '-H' +ok ALLOW '/usr/bin/mount' '--bind' '/mnt/felhom-drives' '/mnt/felhom-drives' +ok ALLOW '/usr/bin/mount' '--make-shared' '/mnt/felhom-drives' +ok ALLOW '/usr/bin/mount' '--make-private' '/mnt/felhom-drives' +ok ALLOW '/usr/bin/mount' '--bind' '/mnt/felhom-usb/felhom-data' '/mnt/felhom-drives/felhom-usb' +ok ALLOW '/usr/bin/umount' '/mnt/felhom-drives/felhom-usb' +ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-drives' +ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-drives/felhom-usb' +ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-usb/felhom-data' +ok ALLOW '/usr/bin/chown' '100000:100000' '/mnt/felhom-usb/felhom-data' +ok ALLOW '/usr/bin/systemctl' 'enable' 'felhom-shared-parent.service' +ok ALLOW '/usr/sbin/pct' 'set' '9201' '-mp8' '/mnt/felhom-drives,mp=/mnt/felhom-drives' +ok ALLOW '/usr/sbin/blkid' '-p' '-o' 'export' '/dev/sda' +ok ALLOW '/usr/bin/lsblk' '-J' '-o' 'NAME,FSTYPE,PTTYPE,MOUNTPOINT' '/dev/sda' +ok ALLOW '/usr/local/sbin/felhom-mkfs-guarded' '/dev/sda' 'ext4' +ok ALLOW '/usr/local/sbin/felhom-mkfs-guarded' '/dev/sda' 'xfs' +ok ALLOW '/usr/sbin/smartctl' '-a' '-j' '/dev/sda' +ok ALLOW '/usr/sbin/lvs' '--reportformat' 'json' '--units' 'b' '-o' 'lv_name,data_percent,metadata_percent' '--' 'pve/data' +ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'unit' 'mnt-felhom\x2dx.mount' +ok ALLOW '/usr/bin/systemctl' 'daemon-reload' +ok ALLOW '/usr/bin/systemctl' 'enable' '--now' '--' 'mnt-felhom\x2dx.mount' +ok ALLOW '/usr/bin/systemctl' 'disable' '--' 'mnt-felhom\x2dx.mount' +ok ALLOW '/usr/bin/systemctl' 'stop' '--' 'mnt-felhom\x2dx.mount' +ok ALLOW '/usr/bin/systemctl' 'reset-failed' '--' 'mnt-felhom\x2ddrives-media.automount' +ok ALLOW '/usr/bin/rmdir' '/mnt/felhom-drives/media' +ok ALLOW '/usr/bin/systemctl' 'start' 'networking.service' +ok ALLOW '/usr/bin/chown' '-R' '100000:100000' '/var/lib/felhom-agent/guests/9201' +ok ALLOW '/usr/sbin/pct' 'set' '9201' '-mp0' '/var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1' +ok ALLOW '/usr/sbin/pct' 'set' '9201' '-onboot' '1' +ok ALLOW '/usr/sbin/pct' 'set' '9201' '--hookscript' 'local:snippets/felhom-guest-hook.sh' +ok ALLOW '/usr/sbin/pct' 'set' '9201' '--delete' 'mp0' +ok ALLOW '/usr/sbin/pct' 'reboot' '9201' +ok ALLOW '/usr/bin/apt-get' 'install' '-y' '-q' 'dnsmasq' +ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'dnsmasq' '/tmp/felhom-resolver-123456789.conf' 'felhom-x.conf' +ok ALLOW '/usr/bin/systemctl' 'enable' '--now' 'dnsmasq' +ok ALLOW '/usr/local/sbin/felhom-os-apply' '--plan' '/var/lib/felhom-agent/os/plan-x.json' +ok ALLOW '/usr/bin/systemctl' 'reload' 'dnsmasq' +ok ALLOW '/usr/bin/systemctl' 'restart' 'dnsmasq' +ok ALLOW '/usr/bin/rm' '-f' '/etc/dnsmasq.d/felhom-x.conf' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'ip' '-4' '-o' 'addr' 'show' 'dev' 'eth0' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'exec' 'felhom-controller' 'cat' '/opt/docker/felhom-controller/controller.yaml' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'ip' 'route' 'show' 'default' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'cat' '/etc/network/interfaces' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'pgrep' '-x' 'dhclient' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'dhclient' '-pf' '/run/dhclient.eth0.pid' '-lf' '/var/lib/dhcp/dhclient.eth0.leases' 'eth0' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'cat' '/etc/felhom-controller-image' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'image' 'inspect' 'gitea.dooplex.hu/admin/felhom-controller:0.0.0' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'inspect' '-f' '{{.State.Running}}' 'felhom-controller' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'systemctl' 'restart' 'felhom-controller-bootstrap.service' +ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'tee' '/etc/felhom-controller-image' +ok ALLOW '/usr/sbin/pct' 'unlock' '9201' +ok ALLOW '/usr/bin/apt-get' 'install' '-y' '-q' 'wireguard-tools' +ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'wg' +ok ALLOW '/usr/bin/systemctl' 'enable' '--now' 'wg-quick@wg-felhom' +ok ALLOW '/usr/bin/systemctl' 'restart' 'wg-quick@wg-felhom' +ok ALLOW '/usr/bin/systemctl' 'disable' '--now' 'wg-quick@wg-felhom' +ok ALLOW '/usr/bin/wg' 'show' 'wg-felhom' 'latest-handshakes' +ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'create' 'felhom-pbs' '10.77.0.1' 'felhom-offsite' 'ns0' 'felhom@pbs!ns0' '00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00' '/etc/pve/priv/storage' +ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'reconcile' 'felhom-pbs' '10.77.0.1' 'ns0' 'felhom@pbs!ns0' '00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00' '/etc/pve/priv/storage' +ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'grant' 'felhom-pbs' +ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'read' 'felhom-pbs' '/etc/pve/priv/storage' +ok ALLOW '/usr/local/bin/felhom-agent' '--config' '/etc/felhom-agent/agent.json' '--selftest=escrow-create' '--upload' '--output=json' +ok ALLOW '/usr/local/sbin/felhom-selfupdate-guarded' 'commit' +ok ALLOW '/usr/local/sbin/felhom-selfupdate-guarded' 'rollback' +ok DENY /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1 +ok DENY /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives +ok DENY /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda +ok DENY /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda +ok DENY /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x +ok DENY /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d +ok DENY /usr/bin/umount /mnt/felhom-drives/x / +ok DENY /usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow +ok DENY /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount +ok DENY /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount +ok DENY /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh +ok DENY /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh +ok DENY /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf +ok DENY /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf +ok DENY /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config +ok DENY /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000 +ok DENY /usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service +ok DENY /usr/bin/systemctl enable --now -- etc-sudoers.d.mount +ok DENY /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd +ok DENY /usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow +ok DENY /usr/bin/rmdir /mnt/felhom-drives/x /etc +ok DENY /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset +ok DENY /usr/sbin/smartctl -a -j /dev/sda -s off +ok DENY /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x +ok DENY /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller +ok DENY /usr/sbin/pct unlock 9201 --whatever +ok DENY /usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount +ok DENY /usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf +ok DENY /usr/local/sbin/felhom-priv-apply wg /etc/shadow +rc=0 + +-- OLD sudoers (agent v0.145.0), the same attacks (this side is the red-proof: 'FAIL want=ALLOW got=DENY' means the OLD file already refused that one; 'ok ALLOW' means the old file let it through) +ok ALLOW /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1 +ok ALLOW /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives +ok ALLOW /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda +ok ALLOW /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda +ok ALLOW /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x +ok ALLOW /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d +ok ALLOW /usr/bin/umount /mnt/felhom-drives/x / +FAIL want=ALLOW got=DENY :: /usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow +ok ALLOW /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount +ok ALLOW /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount +ok ALLOW /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh +ok ALLOW /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh +ok ALLOW /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf +ok ALLOW /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf +ok ALLOW /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config +ok ALLOW /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000 +FAIL want=ALLOW got=DENY :: /usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service +ok ALLOW /usr/bin/systemctl enable --now -- etc-sudoers.d.mount +ok ALLOW /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd +FAIL want=ALLOW got=DENY :: /usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow +ok ALLOW /usr/bin/rmdir /mnt/felhom-drives/x /etc +ok ALLOW /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset +ok ALLOW /usr/sbin/smartctl -a -j /dev/sda -s off +ok ALLOW /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x +ok ALLOW /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller +ok ALLOW /usr/sbin/pct unlock 9201 --whatever +FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount +FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf +FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply wg /etc/shadow +rc=1 diff --git a/documentation/audits/hub-safety-2026-10-05/partH/fleet-after.txt b/documentation/audits/hub-safety-2026-10-05/partH/fleet-after.txt new file mode 100644 index 00000000..322b54f0 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partH/fleet-after.txt @@ -0,0 +1,11 @@ +== hub System page after delivery, 2026-10-05T11:43:34Z: per box — agent cell, root-files cell (raw) +Tester-2-be8404 | agent: 0.142.0 → 0.146.1 (since 2026-10-05) | root files: [] +demo-felhom-8363b5 | agent: 0.146.1 | root files: ['0.146.1'] +demo-hp-bb76ea | agent: 0.146.1 | root files: ['0.146.1'] +tester-1-d70be4 | agent: 0.146.1 | root files: ['0.146.1'] +tester-1-d70be4: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read +demo-hp-bb76ea: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read +demo-felhom-8363b5: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read +tester-1-d70be4: capability rows 66 {'ok': 66} degraded: [] +demo-hp-bb76ea: capability rows 66 {'ok': 66} degraded: [] +demo-felhom-8363b5: capability rows 67 {'ok': 67} degraded: [] diff --git a/documentation/audits/hub-safety-2026-10-05/partH/floors.txt b/documentation/audits/hub-safety-2026-10-05/partH/floors.txt new file mode 100644 index 00000000..1954f804 --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partH/floors.txt @@ -0,0 +1,7 @@ +== floors 2026-10-05T10:24:56Z: POST /customers//floor min_controller_version=0.296.0 min_agent=0.131.0 +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/05 12:24:56 [INFO] Customer demo-hp controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0") +2026/10/05 12:24:57 [INFO] Customer demo-felhom controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0") +2026/10/05 12:24:57 [INFO] Customer tester-1 controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0") diff --git a/documentation/audits/hub-safety-2026-10-05/partH/sign-agent-update.txt b/documentation/audits/hub-safety-2026-10-05/partH/sign-agent-update.txt new file mode 100644 index 00000000..16d3780c --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partH/sign-agent-update.txt @@ -0,0 +1,10 @@ +== agent_update 0.146.1 (sha badd6c9a…) signed with felhom-op-1, ttl 45m, 2026-10-05T10:25:10Z +-- demo-hp-bb76ea +wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-hp-bb76ea-agent_update.json +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-felhom-8363b5-agent_update.json +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-tester-1-d70be4-agent_update.json +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/hub-safety-2026-10-05/partH/sign-bundle.txt b/documentation/audits/hub-safety-2026-10-05/partH/sign-bundle.txt new file mode 100644 index 00000000..8ff0b86e --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partH/sign-bundle.txt @@ -0,0 +1,13 @@ +== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), demo-hp-bb76ea, 2026-10-05T10:35:43Z +wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-hp-bb76ea-agent_config_update.json +uploaded signed op to the hub jobs queue +== agent_config_update 0.146.1 (bundle sha 42333e96…), demo-hp-bb76ea, 2026-10-05T10:42:50Z +uploaded signed op to the hub jobs queue +== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), demo-felhom-8363b5, 2026-10-05T10:53:36Z +uploaded signed op to the hub jobs queue +== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), tester-1-d70be4, 2026-10-05T10:53:36Z +uploaded signed op to the hub jobs queue +== agent_config_update 0.146.1 (bundle sha 42333e96…), demo-felhom-8363b5, 2026-10-05T10:58:57Z +uploaded signed op to the hub jobs queue +== agent_config_update 0.146.1 (bundle sha 42333e96…), tester-1-d70be4, 2026-10-05T11:13:04Z +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/hub-safety-2026-10-05/partH/vouch.txt b/documentation/audits/hub-safety-2026-10-05/partH/vouch.txt new file mode 100644 index 00000000..e4ac9b6b --- /dev/null +++ b/documentation/audits/hub-safety-2026-10-05/partH/vouch.txt @@ -0,0 +1,4 @@ +== vouch 2026-10-05T10:24:18Z: POST /configuration/artifacts (Basic + X-Felhom-Operator), agent 0.146.1, golden 0.296.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +2026/10/05 12:24:45 [INFO] Artifact manifest set: agent=0.146.1 golden=0.296.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="42333e969028867ad8142335e6c1bc4040eec231de0d8d330c2d4b2cf7bc3442" diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 42a4d021..19fe9dd1 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,22 @@ --- +## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124) + +The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session). + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-135** | **A cookie-less POST skipped the hub's CSRF gate, so a browser with cached Basic credentials could be made to POST cross-site.** hub v0.135.0: without a session a state change needs Basic credentials AND the header `X-Felhom-Operator` (decision 120); the gate sits before the route switch. 39 paths through RequireAuth→ServeHTTP; red-proof: the old shape lets all 39 through. Live: Basic + no header → 403 (also with `Origin: evil`, also on an unknown path); with the header → passes; header without credentials → 401. **Reasoning kept: a browser cannot add a custom header cross-site without a CORS preflight, which the hub never answers.** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partA/`; `web/r135_csrf_test.go` | +| **R-133** | **Every box's break-glass console password was plaintext in hub.db.** hub v0.135.0: sealed with the off-site seal and key (decision 121); legacy rows sealed at start-up — live: 4 rows sealed, 0 left plain; the demo-hp reveal still returned a password that minted a PVE ticket (HTTP 200; a wrong one 401); a wrong key → 500, nothing in the body or the log, no event. **Reasoning kept: the running hub still holds the key — this closes the database-copy route only; a database backup without `OFFSITE_SECRET_KEY` cannot open the console passwords (R-173).** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partB/`; `store/r133_recovery_seal_test.go`, `web/r133_reveal_wrongkey_test.go` | +| **R-604** | **A per-customer controller floor silently kept a box out of every global raise (demo-hp missed four).** hub v0.135.0: a global raise logs one line per customer whose own LOWER floor wins and sends ONE operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set; the System page's "Version floors" table lists every per-customer floor with its age and which ones the global cannot move. 2 red-proofs. Live: the table shows the three per-customer floors (age "unknown" — set before v0.135.0). The mail was not exercised live (it needs a global raise below an override). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `web/r604_floor_held_back_test.go` | +| **R-530** | **Nothing listed which boxes still run an old agent (agents update only by a per-box signed job).** hub v0.135.0: the System page's Agent cell (box → vouched, how far, since when; red after the wait) and `agent_behind` after 7 days (decision 119). Live: Tester 2 reads `0.142.0 → 0.146.1`. Signing stays per box (the 2026-09-16 ruling: CC may sign until the first paying customer). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `osupdates/r530_agent_alarm_test.go` | +| **R-508** | **Customer tester-1 had no registered e-mail, and the page did not say so.** The address has been set since 2026-09-14 (the connect mails reach it — Gmail-read 2026-10-05); hub v0.135.0 adds the page warning: a configured customer with no box and no e-mail shows a red line (three branches tested, red-proof). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partG/`; `web/r508_no_email_banner_test.go` | +| **R-509** | **A box installed for an existing customer never got the connect e-mail.** Fixed in hub v0.114.0; the owed real-mail proof: three mails from the automatic "host delete" trigger, each within 1 s of the hub's own send line (2026-09-16 12:22:59 and 18:17:46, 2026-09-30 07:23:03 UTC), read through the Gmail connector (metadata only). The "e-mail set" trigger shares the send core and is unit-proven. | CLOSED 2026-10-05 — VERIFIED | `audits/hub-safety-2026-10-05/partG/r509-real-mails.txt` | +| **R-880** | **An installed `felhom-os-apply` refuses a bundle naming a path it does not know (R16), so a release whose bundle ADDS a path cannot reach any box on an older bundle** (found 2026-10-05 before delivering v0.146.1, which adds four). Fixed by a step: `felhom-agent/scripts/build-step-bundle.py` — the box's current bundle with ONLY `felhom-os-apply` replaced (same paths), published as `0.146.1-step1`; then the release's bundle. Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Live: demo-hp, demo-felhom and Tester 1 each took step1 (`written=1 same=20`) then 0.146.1 (`written=3 same=22`), self-check ok. **Reasoning kept: every future bundle that adds a path needs this step (decision 124); the step package stays published while any box may still be on the old bundle (Tester 2).** | CLOSED 2026-10-05 — FIXED (tooling, agent e4b5cf9) | `audits/hub-safety-2026-10-05/part{F,H}/`; memory `bundle-adding-a-path-needs-step-bundle` | + +--- + ## 2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118) | Row | What | Closed | Evidence | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index dc070c3e..36935088 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -103,12 +103,10 @@ match what the reader sees is how an instrument stops being believed (R-421). It stopping line that lies. -## Install & onboarding — 18 rows (P2 2, P3 11, P4 5) +## Install & onboarding — 17 rows (P3 11, P4 6) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-508** | Install & onboarding | P2 | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** | — | — | CC + operator | -| **R-509** | Install & onboarding | P2 | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). **SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL:** triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as `selfbind_link_sent` and shown on the Setup tab. Unit-proven with a red-proof (`TestSelfBind_EmailSetOnWaitingCustomerSendsLink`). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 `deprovision` + Cloudflare), which is fenced without the operator's word. **Closes on:** one real mail from either trigger. | **VERIFY** (2026-10-03 triage: Fix shipped: hub sends the self-bind link on e-mail set/change for a waiting customer and on host delete — felhom.eu CHANGELOG.md:289-297 (hub v0.114.0, R-509)) — **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** **Re-ranked 2026-10-03: P1→P2: the fix shipped in hub v0.114.0 with a unit red-proof; only one real-mail proof is owed before customers.** | — | — | CC + operator | | **R-130** | Install & onboarding | P3 | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC | | **R-179** | Install & onboarding | P3 | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | | **R-180** | Install & onboarding | P3 | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | @@ -125,8 +123,9 @@ stopping line that lies. | **R-503** | Install & onboarding | P4 | **[P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian".** Offered to the operator 2026-09-14 and **not chosen**: the 2026-07-31 ruling (a person chooses the disk) stands. Recorded so a reversal starts from measurements, not from the offer. **What must be measured first:** (1) whether the Proxmox auto-installer's HTTP answer mode can serve a per-machine answer from posted system info without network being a precondition a volunteer can miss; (2) whether USB transport is reliably visible in sysfs (`/sys/block/*/device` path, `removable`) where udev properties were measured blind (SPIKE-universal-iso-1 §3.2); (3) whether any refusal can be shown in Hungarian without modifying the Proxmox installer squashfs. **Reverses two rulings if built — needs an operator word.** | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (ruling), CC (spike)** **Re-ranked 2026-10-03: P3→P4: an unchosen idea that would reverse two rulings.** | — | — | operator | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | | **R-725** | Install & onboarding | P4 | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). **FIXED 2026-09-30:** the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). **NARROWED — remaining:** the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). | **NARROWED — three small copy items; owner: CC** **Re-ranked 2026-10-03: P3→P4: three small copy slips left; nothing blocks the household.** | — | — | CC | +| **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | -## Apps & catalog — 34 rows (P3 15, P4 20) +## Apps & catalog — 36 rows (P3 15, P4 21) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -191,7 +190,7 @@ stopping line that lies. | **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 52 rows (P2 8, P3 23, P4 21) +## Backup & restore — 54 rows (P2 9, P3 23, P4 22) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -200,8 +199,8 @@ stopping line that lies. | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | -| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller)** | — | — | CC | -| **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | — | — | CC | +| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** | — | — | CC | +| **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **NARROWED 2026-10-05 — FIXED controller v0.296.0, unit-proven (3 red-proofs): a run cut off by a stop is said on /backups and /backups/apps until a run ends with every step OK; the synthesised "last database backup" no longer reads OK over a cut run; the restore point's time was already its oldest part (v0.275.0). LEFT: the live cut on 9202 — the permission check refused restarting the controller mid-backup; the operator is asked (STATUS).** | — | the operator's go for one controller restart mid-backup on 9202 | CC | | **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC | | **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | @@ -267,14 +266,12 @@ stopping line that lies. | **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks//deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC | | **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | -## Security & access — 32 rows (P2 4, P3 25, P4 3) +## Security & access — 31 rows (P2 2, P3 26, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-133** | Security & access | P2 | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | -| **R-135** | Security & access | P2 | **`validateCSRF` returns TRUE when there is no session cookie** (`hub/internal/web/server.go:678-683`) — measured live: `POST` with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 | READY (S) — **security** | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and `SameSite` does not govern the `Authorization` header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 | CC | | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | -| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **READY — design + operator go; owner: CC** | — | the operator ranks it against the first paying customer | CC | +| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | @@ -303,13 +300,12 @@ stopping line that lies. | **R-134** | Security & access | P4 | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | +| **R-879** | Security & access | P3 | **A copy of hub.db still holds secrets readable without any key:** each box's hub API key (`hosts.api_key`), each household's owner passphrase and API key (`customer_configs.retrieval_password`, `api_key`), and the PBS-DR token values (`host_pbs_secrets.value`, kept after use). Found 2026-10-05 while answering "what does a hub database backup now contain" (R-133 sealed the console passwords; the off-site passwords were sealed by R-821). A stolen database copy lets an attacker report as any box and read every owner passphrase. `05` §16.2 | READY | — | Seal or hash each with the same seal (the box keys and the passphrases are compared, so a hash may fit; the PBS token is served once, so it could be deleted after use) | CC | -## Box system & updates — 18 rows (P2 3, P3 13, P4 2) +## Box system & updates — 19 rows (P2 1, P3 15, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator | -| **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC | | **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator | | **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | @@ -365,7 +361,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-173** | Hub & operator | P2 | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **READY (M) — NEW 2026-08-02** | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | +| **R-173** | Hub & operator | P2 | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **NARROWED 2026-10-05 — WAITING-ON-OPERATOR.** Measured (read only): the database IS backed up nightly — only on DooPlex: Longhorn `backup-daily`/`backup-weekly` (last 2026-10-05 02:06 UTC, Completed, 713 MB) to DooPlex's own `sda1`, because the live Volume carries `default: enabled` although the PVC (git, 2026-02-16, no reason) says `disabled` — a hand-set drift any sync may undo. Nothing leaves DooPlex; nothing alarms on failure (R-232). Since hub v0.135.0 the console passwords are sealed, so a copy is worth little without `OFFSITE_SECRET_KEY`, which also exists only on DooPlex. The off-site plan (keys off the box, the label fixed, a consistent `VACUUM INTO` snapshot pushed encrypted to ep0's PBS, a weekly restore test, Prometheus alarms): `runbooks/RUNBOOK-hub-db-offsite-backup.md`; the decision is in STATUS. `audits/hub-safety-2026-10-05/partC/` | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | | **R-30** | Hub & operator | P3 | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk.** | — | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | CC | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | @@ -401,7 +397,7 @@ stopping line that lies. | **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 86 rows (P3 4, P4 82) +## Process & tooling — 88 rows (P3 4, P4 84) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| diff --git a/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md b/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md new file mode 100644 index 00000000..9d03cfb0 --- /dev/null +++ b/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md @@ -0,0 +1,129 @@ +# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go + +> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes +> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on: +> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only). + +## 1. What is true today (measured 2026-10-05) + +| Question | Answer | +|---|---| +| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. | +| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. | +| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. | +| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). | +| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. | +| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. | +| What a copy is worth without the key | The sealed columns (console passwords, off-site passwords) are useless without `OFFSITE_SECRET_KEY`. **The key lives only in `Secret/offsite-secret-key` on DooPlex** (and in the GPG secrets export on the same disk). A backup off DooPlex without a key off DooPlex restores a hub that cannot open any console password. | + +## 2. The plan (option A — my pick): a nightly, encrypted, consistent copy on ep0's PBS + +ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches it through `felhom-ep0-pbs-tunnel` +(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places, +one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees. + +### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) + +Without these, every later step backs up something nobody can open after a DooPlex loss. + +```bash +# 1. the hub's seal key → the operator's password manager (never a file, never a chat) +sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.OFFSITE_SECRET_KEY}' | base64 -d; echo +# 2. (after Step 2) the backup encryption key's paper copy → the password manager +sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text +``` + +### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) + +`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check: +`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`. +This keeps today's on-DooPlex copy alive; it is not the off-site copy. + +### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) + +```bash +proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)" +proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file +# namespace for operator data, apart from the households' namespaces +proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite' +proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push' +# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune) +proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \ + --schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8 +``` + +### Step 3 — a consistent snapshot of the live database (CC, a hub release) + +`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub +gets a nightly `VACUUM INTO '/data/snapshots/hub-.db'` (keeps 2, logs size and duration) — one SQLite +statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the +hub image, so the copy must be made by the hub itself. + +### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves) + +```bash +#!/bin/sh -eu +# /usr/local/sbin/felhom-hub-db-backup — push the newest hub snapshot to ep0 (R-173). Root, 0755. +STAGE=/var/lib/felhom-hub-backup/stage; mkdir -p "$STAGE"; chmod 700 "$STAGE" +SNAP=$(kubectl -n felhom-system exec deploy/hub -- sh -c 'ls -1t /data/snapshots/hub-*.db | head -1') +kubectl -n felhom-system exec deploy/hub -- cat "$SNAP" > "$STAGE/hub.db" +sqlite3 -readonly "$STAGE/hub.db" 'PRAGMA integrity_check' | grep -qx ok # never push a broken copy +export PBS_PASSWORD_FILE=/etc/felhom-hub-backup/token PBS_FINGERPRINT= +proxmox-backup-client backup hubdb.pxar:"$STAGE" --ns operator --backup-id dooplex-hub \ + --keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' +shred -u "$STAGE/hub.db" +# the positive signal the alarm reads (written ONLY on success): +echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ \ + && mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom +``` + +### Step 5 — the restore test (weekly, Sun 04:30, same unit family) + +```bash +T=$(mktemp -d); chmod 700 "$T" +proxmox-backup-client restore "host/dooplex-hub/$(newest snapshot)" hubdb.pxar "$T" --ns operator \ + --keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' +sqlite3 -readonly "$T/hub.db" 'PRAGMA integrity_check' | grep -qx ok +test "$(sqlite3 -readonly "$T/hub.db" 'SELECT COUNT(*) FROM hosts')" -gt 0 +test "$(sqlite3 -readonly "$T/hub.db" "SELECT COUNT(*) FROM host_recovery WHERE secret NOT LIKE 'enc:v1:%'")" -eq 0 +shred -u "$T/hub.db"*; rmdir "$T" +echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv +``` + +The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner). + +### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader) + +```yaml +- alert: HubDBBackupStale + expr: time() - felhom_hub_db_backup_last_success_timestamp_seconds > 26*3600 or absent(felhom_hub_db_backup_last_success_timestamp_seconds) + for: 30m + labels: {severity: critical} + annotations: {summary: "The hub database has not reached ep0 for 26 h (R-173)"} +- alert: HubDBRestoreTestStale + expr: time() - felhom_hub_db_restore_test_last_success_timestamp_seconds > 8*24*3600 or absent(felhom_hub_db_restore_test_last_success_timestamp_seconds) + for: 1h + labels: {severity: warning} +``` + +Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty +log is not a success. + +### Step 7 — prove it once (CC, with the operator's go) + +Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore +test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then +start it again. + +## 3. Bringing the hub back from this copy (the procedure the plan exists for) + +1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy). +2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into + the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line + `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches. + +## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account + +Same Steps 0, 1, 3, 5, 6; the push is `restic backup` to a new sub-account with the R-820 append-only key pin. Costs a +new sub-account and its own key custody; the Storage Box sub-account shell can `rm` (memory: storagebox-subaccount-shell) +unless the pin is right. ep0 already has the server-side prune and the return copy, so A is less new machinery. diff --git a/documentation/runbooks/config-bundle.md b/documentation/runbooks/config-bundle.md index 3b65d3bf..e243ab41 100644 --- a/documentation/runbooks/config-bundle.md +++ b/documentation/runbooks/config-bundle.md @@ -31,6 +31,24 @@ sign; no bundle may add, remove or change them. A box that has no signers file g 4. **Undo** = send the previous release's bundle the same way. The previous copies also stay on the box in `/var/lib/felhom-os-apply/bundle-prev/