hub-safety session: R-135/R-133/R-604/R-530/R-508/R-509/R-880 closed, R-861/R-173/R-518/R-519 narrowed, R-879/R-881 opened (336 → 332); 03 §3.1, 05 §16, golden 0.296.0, the hub-DB off-site plan, STATUS
gates / gates (push) Successful in 32s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 13:47:57 +02:00
parent 2b30733b0d
commit 0826e41b31
44 changed files with 1511 additions and 31 deletions
+12
View File
@@ -16,6 +16,18 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller
> v0.296.0, agent v0.146.1 + bundle `42333e96…`, golden 0.296.0 vouched with agent 0.146.1, min_agent 0.131.0).** CC decisions
> 119–124, *operator may reverse*. Hub: R-135 a cookie-less state change needs Basic + `X-Felhom-Operator` (`05` §16.1); R-133
> console passwords sealed with the off-site seal/key (`05` §16.2; 4 rows sealed live); R-604/R-530 the System page's "Version
> floors" table + Agent cell, `agent_behind` (7 d) and `floor_raise_skipped` (`05` §5, `08` §6.3); R-508 no-e-mail banner.
> Controller: R-519 `run_record.go` (a cut run said until a complete one; live cut refused by the permission check — operator
> asked), R-518 copy with today's measurement (5 min 47 s, demo-hp). Agent: R-861 narrowed (`03` §3.1 — exact sudo regexes,
> `felhom-priv-apply`, fixed hook/parent files, signed update verified as root, escrow paths pinned; three residuals named);
> **R-880: a bundle that adds a path needs a STEP bundle** (`scripts/build-step-bundle.py`, `runbooks/config-bundle.md`).
> R-173 measured (backed up only on DooPlex, by label drift) + `runbooks/RUNBOOK-hub-db-offsite-backup.md`; decision in
> STATUS. Closed R-133, R-135, R-508, R-509, R-530, R-604, R-880; opened R-879, R-881; register 336 → 332.
> **2026-10-05 (afternoon) — a box that is not always on (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0
> vouched with agent 0.145.0, min_agent 0.131.0).** Rulings 109–111 (`09` §3: R-871 option A — a missed night runs once when the
> box comes back; the household's banner; Tester 2 read only). CC decisions 112–118, *operator may reverse*. Design `07`
+131
View File
@@ -0,0 +1,131 @@
# REPORT — the hub's own safety, boxes left behind, honest backup wording, two onboarding rows, and the agent's admin permissions (R-135, R-133, R-173/R-232, R-604, R-530, R-518, R-519, R-861, R-508, R-509) — 2026-10-05, late afternoon
Brief: "an open-items batch — the hub's own safety (CSRF, the console credential at rest, the hub database in
backups), a fleet view that shows boxes left behind, honest backup wording, two stale onboarding rows; and the agent
permission fix (R-861) as its own Part". Evidence: `documentation/audits/hub-safety-2026-10-05/part{A..H}/`, the golden
`documentation/tests/golden-0.296.0-2026-10-05/`. Architecture read before the claims: `05-hub-architecture.md`,
`_hub-review.md`, `04-control-plane-authorization.md`, `03-host-agent.md` §3/§11, `07` §6.1, `08` §6.3,
`runbooks/target-selection.md`, `runbooks/secrets.md`, `runbooks/ep0-datastore-copy.md`,
`audits/RECON-dooplex-backup-2026-08-06.md`.
Baselines (re-verified at the start): felhom.eu `9bb45eaaa2`, felhom-agent `61345790ed` (v0.145.0), felhom-controller
`477e2548db` (v0.295.0); register 336 rows, highest R-878.
## 1. The Part table
| Part | State | Note |
|---|---|---|
| A — CSRF (R-135) | **done** | hub v0.135.0: no session → Basic credentials + `X-Felhom-Operator` (decision 120); every state-changing route in one table (`partA/route-table.md`, 38 routes + an unknown path); red-proof: the old shape lets 39 of 39 through; live: 403 / pass / 401 |
| B — console password at rest (R-133) | **done** | the off-site seal and key reused (decision 121); 4 legacy rows sealed live, 0 left plain; reveal still opens demo-hp's Proxmox; wrong key fails closed; 2 red-proofs. What a DB backup still holds readable → R-879 |
| C — hub DB in backups (R-173, R-232) | **done (read only) — decision with you** | it IS backed up, only on DooPlex, by a label drift; no failure alarm; steps in `runbooks/RUNBOOK-hub-db-offsite-backup.md`; decision in STATUS |
| D — boxes left behind (R-604, R-530) | **done** | System page "Version floors" + Agent cell (live), `agent_behind` 7 d + `floor_raise_skipped` (tests, 3 red-proofs). The mail was not exercised live (needs a global raise) |
| E — honest backup wording (R-518, R-519) | **done, changed** | R-518: the copy was already honest; today's measurement added (5 min 47 s). R-519: dating was already fixed (v0.275.0); the page notice + the synthesised status fixed (v0.296.0, 4 red-proofs). **The live cut on 9202 was refused by the permission check — asked** |
| F — the agent's admin permissions (R-861) | **done, changed** | nine root paths, not four; `03` §3.1 written AFTER the build (not "design first"); agent v0.146.1 (after a review found three holes in v0.146.0); delivered by a two-step bundle (R-880); live on both demo boxes: sudo 93/93, capability 67/67; three residuals named, row stays open narrowed |
| G — onboarding rows (R-508, R-509) | **done** | R-509 closed by three matched real mails; R-508 closed (e-mail set since 09-14 + a new page warning, red-proof) |
| H — release and records | **done, changed** | hub 0.135.0, controller 0.296.0, agent 0.146.1 (+ 0.146.0 never delivered); golden 0.296.0 baked + vouched; floors + signed jobs for demo-hp, demo-felhom, tester-1; docs `00`, `03`, `05`, `07`, `08`, `09`, `11`-runbook |
## 2. Claims in the brief that turned out wrong
1. **"The hub's own database is in no backup."** It is in one — only on DooPlex. Longhorn's `backup-daily` /
`backup-weekly` copy `hub-data` every night (last 2026-10-05 02:06 UTC, Completed, 713 MB) to DooPlex's own `sda1`.
R-173's "excluded" is the PVC label (`recurring-job-group.longhorn.io/default: disabled`, set 2026-02-16 with no
reason); the live Longhorn Volume carries `enabled` — a hand-set drift that keeps the backup alive and can be undone
by any sync. Nothing leaves DooPlex, and nothing alarms if it fails (R-232 stands).
2. **"The backup page promises 'a few seconds'."** Not since controller v0.243.0 / v0.267.0: the text already said
"several minutes (about 8 minutes on a 12-app box)". Measured today on demo-hp (9 apps): 5 min 47 s, local tier only.
v0.296.0 adds today's figure and "minutes, not seconds".
3. **"R-519: fix the dating."** The dating was already fixed in controller v0.275.0 (R-696): a restore point carries the
time of its OLDEST part. What was still missing was the sentence on the page, and the page's synthesised "last
database backup … OK" after a restart — both fixed in v0.296.0.
4. **"R-861: four admin-command groups."** It was nine ways to root, not four: besides the four named (guest hook,
intermediary script/unit, escrow, self-update), the mount units, the dnsmasq drop-ins, the WireGuard config and the
OOB sshd config were each installed from agent-written files, and almost every `*` in the arguments matched spaces
(measured with real sudo 1.9.16: 23 of 29 attack lines allowed).
5. **"Deliver the agent fix by the signed bundle."** Not possible in one step: an installed `felhom-os-apply` refuses
a bundle naming a path it does not know (R16), and v0.146.1's bundle adds four. Delivered by a step bundle (R-880).
6. **"Design first" (Part F).** I built first and wrote the `03` §3.1 section after the code, in the same session — the
section records what was built, group by group.
## 3. Per Part — tests, red-proofs, live proof
**A.** `hub/internal/web/r135_csrf_test.go` (5 tests). Red-proof `partA/red-proof.txt` (39 of 39 convicted). Live
`partA/live.txt` (hub 0.135.0, ClusterIP): Basic, no header, `Origin: evil` → **403**; unknown path, no header → **403**;
with `X-Felhom-Operator: cli` → **404** (passed the gate); header without credentials → **401**; GET → 200. The skill and the
memory note now carry the header.
**B.** `store/r133_recovery_seal_test.go` (4), `web/r133_reveal_wrongkey_test.go`, `cmd/hub/r133_wiring_test.go`. Red-proofs
`partB/red-proof.txt` (plaintext save — the first attempt did not compile, re-run with a compiling mutation; the wiring).
Live `partB/live-db.txt`: hub start `console passwords sealed at rest (4 legacy plaintext row(s) sealed now)`; the live DB
copy (scratch, shredded) shows 4 rows `enc:v1:`, 0 not sealed. `partB/live-reveal.txt`: reveal on demo-hp → 200, a
32-char password that minted a PVE ticket (200; a wrong one 401); the timeline event recorded.
**What a hub DB backup now holds:** the console and off-site passwords sealed (useless without `OFFSITE_SECRET_KEY`, which
exists only on DooPlex); still readable: box API keys, owner passphrases + customer API keys, PBS-DR token values (R-879).
**C.** Readings `partC/readings.txt` (read only). Steps `runbooks/RUNBOOK-hub-db-offsite-backup.md`: keys off the box first;
fix the PVC label in git; a write-only namespace on ep0's PBS; a hub `VACUUM INTO` nightly snapshot (a later hub release);
the encrypted push via the existing tunnel; a weekly restore test (`PRAGMA integrity_check`, row counts, every console
password still sealed); two Prometheus alarms through the existing mail receiver (`absent()` included); a proof run.
**D.** `osupdates/r530_agent_alarm_test.go` (3), `web/r604_floor_held_back_test.go` (4), `cmd/hub` wiring. 3 red-proofs
(`partD/red-proof.txt`). Live `partD/live-system-page.txt`: global floor 0.292.0, three per-customer floors 0.295.0 (age
"unknown" — set before v0.135.0); Tester 2 `0.142.0 → 0.145.0 (since 2026-10-05)`, the demo boxes "current".
**E.** `internal/backup/run_record_test.go` (3), `cmd/controller/run_record_wiring_test.go` (2), `TestR518_*`, parity
cases. 4 red-proofs (`partE/red-proof.txt`). Measurement `partE/r518-measure.txt`. The 9202 reproduction: a throwaway
bookstack installed (09:35:56Z) and a complete baseline run (09:37, 35 s); the cut was refused by the permission check;
bookstack removed through the product (`partE/teardown-9202.txt`: no container, volume, folder or backup left).
**F.** Design `03` §3.1. `configs/test_felhom_priv_apply.py` (32), `AgentUpdate` (8), `SelfupdateWrapperConfinement`,
`StepBundle` (3), Go contract tests (4 packages), `TestSudoersRefusesTheR861Injections`, `TestManifestCoveredBySudoers`.
Red-proofs F1–F9 + S1–S3 (`partF/red-proof.txt`; F1 masked on its first run — strengthened; F3 errored rather than
failed — clean assertion added). Real sudo, container (`partF/sudo-container-proof.txt`): old 23/29 attacks allowed, new
0/29, 64/64 commands allowed. Pre-flight on both boxes' live files: all OK. **Live after the bundle:** `sudo -l` 93/93 on
demo-hp and demo-felhom (`partF/live-sudo-after-*.txt`; before, on demo-felhom: 23 attacks allowed —
`live-sudo-before-demo-felhom.txt`); the checker run as the agent user → SAME on every real file (on demo-felhom the drive unit has no staged copy — an older path wrote it — so that one read `[P1] no staged file`; that box's `/mnt/hdd_1` is the whole-system backup storage, not a household drive, so no bind under `/mnt/felhom-drives` is expected); a staged unit over
`/etc/sudoers.d` refused `[U3]`, nothing installed; the old `install` route → `a password is required`.
**Capability check after the bundle: demo-hp 67/67, demo-felhom 67/67, Tester 1 all ok (hub page), nothing degraded.**
A gap seen on demo-felhom: between the new agent (~10:45 UTC) and the bundle (11:09) its OOB-sshd reconcile logged "install
failed" every minute (the expected gap); 0 errors after the bundle.
**G.** `partG/r509-real-mails.txt` (three host-delete sends matched to mailbox arrivals within 1 s), `web/r508_no_email_banner_test.go`
(3 branches) + red-proof.
## 4. Release and delivery
- **hub v0.135.0** (`3d7a2761` code, `2b30733b` manifest) — built from the pushed commit, ArgoCD Synced/Healthy, image tag
0.135.0. (A later comment-only change in `server.go` is in this session's docs commit; the image is unchanged by it.)
- **controller v0.296.0** (`ff69074`), MinAgent 0.131.0. Floors 0.296.0 for demo-hp, demo-felhom, tester-1 → both demo
boxes ran 0.296.0 (healthy) within ~30 min; 9202 (scratch) stays 0.295.0.
- **agent v0.146.0** (`6ab1e7c`, released, **never vouched or delivered**) → **v0.146.1** (`fdd8717`) after the review.
Step bundle `0.146.1-step1` (sha `8482851e…`, built from the 0.145.0 bundle, only `felhom-os-apply` replaced).
- **golden 0.296.0** baked and vouched with agent 0.146.1 / min_agent 0.131.0 (`documentation/tests/golden-0.296.0-2026-10-05/`).
- **Per box, signed with felhom-op-1:** `agent_update` 0.146.1 → `agent_config_update` 0.146.1-step1 (`written=1 same=20`,
self-check ok) → `agent_config_update` 0.146.1 (`written=3 same=22`, self-check ok). demo-hp, demo-felhom, Tester 1 all
report agent 0.146.1 and root files 0.146.1 (`partH/fleet-after.txt`). **Tester 2: DOWN all session, nothing sent.**
## 5. Rows
Register **336 → 332**. Closed (7): R-133, R-135, R-508, R-509, R-530, R-604, and R-880 (opened and closed today).
Narrowed: R-861 (three residuals), R-173 (measured; waiting on you), R-518 (copy; per-tier quiesce left), R-519 (live cut
left). Opened (2): R-879 (hub.db still holds readable secrets), R-881 (installer uninstall misses `felhom-priv-apply`).
The section counts in `OPEN-ITEMS.md` were recomputed (several were already out of date).
## 6. Slips of mine, said plainly
- **Two agent releases** (0.146.0, 0.146.1) against "one per repo". 0.146.0 had three security holes a background review
found after I pushed it; it was never vouched or sent.
- **I did not design Part F first** as the brief asked; the `03` section was written after the build.
- **My first version waiter read the wrong page cell** and reported the boxes as not updated; I re-read the right cell.
- **Two red-proofs did not convict on the first run** (F1 masked, the R-133 plaintext mutation did not compile); both
were fixed and re-run.
## 7. Teardown, three layers
- **Machines:** 9202 — the throwaway bookstack removed through the product, nothing left; its controller stays 0.295.0.
The demo boxes keep their real files (the checker reported SAME; the one staged attack file was deleted).
Bake VM: CT 9100 destroyed, token/script/log shredded, qemu stopped, disk back to `virgin`.
- **Hosts:** nothing provisioned. The pre-flight copies of the checker (`/tmp/felhom-priv-apply-check`) and the case
files were removed from both hosts.
- **Hub:** hub 0.135.0 deployed; artifacts vouched (agent 0.146.1, golden 0.296.0); floors 0.296.0 for three customers;
9 signed jobs (3 × agent_update, 6 × agent_config_update), all consumed. The step package `felhom-agent/0.146.1-step1`
stays published on purpose (Tester 2 will need it). No customer or appliance record created.
+54 -6
View File
@@ -1,13 +1,61 @@
# STATUS — what works, what's broken, what's next
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) stayed offline all day; nothing
was sent to it.**
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (afternoon, the catch-up session): every box of ours healthy. Built and proven live: a box that was
off at its backup time makes the backups up once when it comes back; the household's banner; the OS update repairs
itself after a power cut (second crash on demo-hp, with your go). Report: `REPORT-catchup-2026-10-05.md`.**
**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form
posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make
itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
## Today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
**Decisions I took myself (you may reverse each — `09` decisions 119–124):**
- A box behind the approved agent for 7 days raises an alarm to you; a global version raise that cannot move a box
sends you one mail naming it.
- Scripts that post to the hub with the password must now send one extra header; a web page on another site cannot.
- The console passwords are locked with the same key as the off-site passwords (one key to keep safe, not two).
- The agent's rights are narrowed with exact rules and one checking helper, not one helper per command.
- A cut-off backup is shown on the backup page until a backup runs all the way through.
- A new agent whose root files add a file is delivered in two signed steps (the old box would refuse it in one).
**What works now (proven live):**
- **Form protection:** a password post without the header is refused (403); a browser on another site cannot add it.
- **Console passwords locked:** all 4 were sealed at the hub's start; the demo-hp one still opens its Proxmox (checked).
- **Boxes left behind:** the System page lists the three per-box version floors and shows Tester 2's agent 4 releases
behind. The alarm and the mail are proven by tests only (they need 7 days / a global raise).
- **The agent cannot make itself root any more:** before, the real sudo let 23 of 29 attack commands through; now 0, on
demo-hp and demo-felhom, and every agent feature still passes its check (67 of 67) on all three boxes.
- Agent 0.146.1, controller 0.296.0 and hub 0.135.0 on demo-hp, demo-felhom and Tester 1; new-install image 0.296.0.
**Found today:**
- **The hub database is backed up — but only inside DooPlex**, and only because a hand-set label says so; nothing tells
anyone if that backup fails. (Your decision below.)
- **A new agent's root files could not reach any box in one step** (an older box refuses files it does not know). Fixed
with a two-step delivery; written down for next time.
- **A security review of my own agent change found three holes** before it went to any box; fixed in a second agent
release (0.146.1). Two agent releases today, against "one per repo" — the first was never sent anywhere.
- **The backup page already said "about 8 minutes"**, not "a few seconds". Measured today on demo-hp (9 apps): about 6
minutes. Both figures are on the page now.
**Needs you:**
1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
- **If you decide nothing:** the database stays only on DooPlex. A fire or theft there loses every box's console
password, the escrow records and the customer settings; each box would need re-pairing by hand. Steps:
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
of the database cannot open the console passwords.
2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
proven by tests only.
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own
guest runs; install the operator SSH key for the limited `felhom-op` user; see the box's backup key during the
recovery-code ceremony. If you do nothing, they stay as they are until before the first paying customer.
4. **Tester 2's one-time step** is unchanged (below).
## Earlier today (2026-10-05, afternoon): a box that is not always on; the self-repair after a power cut
**Decisions I took myself (you may reverse each — `09` decisions 112–118):**
- The make-up run starts 15 minutes after the box comes back; a backup due within 30 minutes is left to its normal time.
@@ -196,10 +196,11 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Multiple household users / per-person accounts | — | **MISSING** | — | Single dashboard password; acceptable for alpha → R-15 |
| A second login step for the dashboard (a TOTP code or a passkey) | — | **MISSING** | — | One password, one bcrypt hash (`controller/internal/web/auth.go:37-44`) → R-811 (added 2026-10-03) |
| The household can leave Felhom, or outlive it — the box runs without the hub, the household owns its domain, tunnel and off-site account, and can export everything | — | **MISSING** (as a written answer) | the LOST-hub half only: `_recovery-inventory-2026-07-28.md` §D2.4, `07` §8 row 11b | Leaving and hand-over are answered nowhere → R-810 (spike, added 2026-10-03) |
| **The host agent cannot reach root without the operator key: exact sudo patterns, no agent-written file installed where root reads it without a content check, the agent binary only by an operator-signed update checked as root** | agent **v0.146.1** (R-861; delivered by a step bundle, R-880) | **PROVEN-LIVE on both demo boxes (2026-10-05) — with three named residuals** | `audits/hub-safety-2026-10-05/partF/` (real sudo: 23 of 29 attacks allowed before, 0 after; 64 capability commands allowed; `sudo -l` 93/93 on demo-hp and demo-felhom after the bundle; capability probe 67/67; a staged unit over `/etc/sudoers.d` refused live) | `03` §3.1: the controller-swap image ref (guest-scoped), the felhom-op SSH key (hub-delivered, unsigned; felhom-op's sudo is scoped), the escrow ceremony relays R → R-861 |
| WireGuard base infra always-on; OOB operator access (felhom-sshd, /32 peer) | agent v0.72, hub v0.35 | **IMPLEMENTED** | `SPIKE-oob-wg-operator-peer-2026-07-05`, `SPIKE-felhom-sshd-2026-07-05` | Mutual-repair desired-state arc not built → R-13. **CHECKED 2026-08-08 (R-260) and this row was NOT claiming something untrue** — it claims the capability is implemented, never that it is monitored, so no correction was owed. What WAS untrue is narrower and sat one layer down: **the hub's own OOB health check could not see whether the operator's key was installed.** `HostOOBRow` mirrored five of the agent's eight OOB fields, so `operator_key_configured` — emitted every heartbeat since agent v0.72.0, i.e. from this row's own vintage — was discarded by `encoding/json` on arrival, and `oobDegraded` returned `ok` for a box with felhom-sshd active, reachable, a valid config, a configured peer and **no operator key at all**. `operator_peer_configured`, which it did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Fixed hub v0.99.0; the missing key now degrades and the alert NAMES it; a stanza too old to carry the field is reported distinctly and is never a silent ok. Pinned end-to-end from raw report JSON by `TestHostOOB_MissingOperatorKey_EndToEnd` and `TestHostOOB_NoKeyField_IsNotSilentlyOK_EndToEnd` |
| The operator can see WHERE a managed host is — its LAN address and its WireGuard address, on the host page | agent **v0.119.0**, hub **v0.85.0** | **PROVEN-LIVE** (2026-07-31) | `audits/host-addresses-visible-2026-07-31.md` | Before this the LAN IP was **not reportable at all** — `HostMetrics` carried no address of any kind — and the WG IP existed only in `/offsite`'s peer table keyed by pubkey (peer→host, never host→peer). New wire field `addresses[]`, one row per (interface, address); `IsGlobalUnicast()` is the whole filter, chosen by MEASURING both demo boxes, and it needs no veth/fwbr denylist because that plumbing carries no IP. Rendered live on both 0.119.0 hosts matching their `ip addr` ground truth exactly. **Two honesty properties carry the risk and are both red-proofed:** WireGuard shows the hub ALLOCATION and whether the box CONFIRMS holding it (allocation alone cannot tell a live tunnel from a peer never applied), and an agent below 0.119.0 renders **UNKNOWN, never "no addresses"** — proven live on `drill-r50-0a4f9a` (0.113.0). **Not covered:** a two-LAN-bridge box and a real WG drift, neither of which exists to observe |
| The operator can see whether a managed host's **guests still have working networking** — and **how often the watchdog had to repair them** | agent **v0.92.0** (emitter, 2026-07-21), hub **v0.104.0** (reader, 2026-08-13) | **IMPLEMENTED** | `backlog/OPEN-ITEMS.md` R-319; `hub/internal/web/hosts_guestnet_test.go` (7 tests, fixtures copied verbatim from `demo-felhom-8363b5`'s live `host_reports` row) | The agent emitted `guest_net` on every heartbeat for **twenty-three days** while the string occurred **nowhere** in `felhom.eu/hub/` — stored as raw text in `report_json`, read by nothing (R-260/R-264, the first of that census's readers to be built). **The fact that carries the risk is `heals_last_hour`, not `state`:** a guest the watchdog keeps repairing is healthy at every instant anyone looks, so rendering the state alone would give it a green tick — the failed-disk-drawn-as-a-healthy-empty-disk shape. `heal_succeeded` is decoded beside it, because six FAILED repairs is a guest that is down while six successful ones is a nuisance. **Unknown is never drawn as healthy:** three absences, three sentences (agent < 0.92.0; a capable agent that sent nothing; a guest whose own state the watchdog did not assert), and a malformed stanza degrades to unknown without a 500. **Three red-proofs, each mutation asserted applied by grep before its run**, including the one that matters — removing the unknown branches and watching a silent machine render as healthy. **Positive control that it is WIRED and not merely written: the wire-contract gate's checked-tag count rose 182 → 190** as the eight `guest_net` allowlist entries were deleted (an allowlisted tag is skipped, so leaving them would have meant these fields were never checked) | **IMPLEMENTED, not PROVEN-LIVE, and the distinction is the honest half.** Every scenario is proven against the real wire in tests, and the healthy case renders correctly for the live fleet — but **no machine has ever been observed with a climbing repair count on this card**, because neither demo box has needed a repair since the watchdog shipped. The signal this card exists for has therefore never been seen firing on hardware. It moves to PROVEN-LIVE the first time a real repair count is watched appearing. **No alarm was added, deliberately** (R-319): the incident behind this was about nobody being able to SEE the condition, and a new email on a fleet of two demo machines is untested noise — revisit when a third machine exists or when a count is seen climbing |
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. The vaulted secret is plaintext at rest → **R-133** |
| Break-glass management-plane recovery | agent v0.71, hub v0.84 | **IMPLEMENTED** | `runbooks/break-glass.md` | hub v0.84.0 adds an **operator-SESSION** retrieval path (host page → Console access → Reveal; `POST /hosts/{id}/reveal-recovery-credential`, CSRF-gated, writes a customer-visible `recovery_credential_revealed` event) beside the pre-existing **global-key** one (`GET /api/v1/admin/hosts/{id}/recovery-credential`), which is untouched and stays the route for when the hub UI itself is down. **The credential half is now PROVEN (2026-07-31):** the vaulted `demo-hp-bb76ea` password was verified against the box's own `/etc/shadow` hash AND minted a real PVE ticket — `POST /api2/json/access/ticket` → **HTTP 200, `root@pam`, 367-char ticket**, the exact API the login form submits to. Still IMPLEMENTED rather than PROVEN-LIVE overall, because the path has not been exercised on a REAL lockout (SSH was available throughout). Discovered during that check: the card's Copy button shipped `disabled` until a Reveal and silently no-opped, leaving ANOTHER host's password in the clipboard — fixed in hub v0.86.0. **Sealed at rest since hub v0.135.0 (R-133 CLOSED):** the off-site seal and key; live 2026-10-05: 4 legacy rows sealed at start-up, 0 left plain, and the demo-hp reveal still minted a PVE ticket (HTTP 200; a wrong password 401) — `audits/hub-safety-2026-10-05/partB/`. A database backup now needs `OFFSITE_SECRET_KEY` too → R-173 |
## F. Notifications & monitoring
@@ -233,6 +234,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
+55 -1
View File
@@ -72,6 +72,59 @@ Explicitly does **not**:
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent does not manage it; it READS its health (R-841, agent v0.141.0). Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
### 3.1 The agent's admin commands, group by group — and how each is narrowed (R-861, agent v0.146.1) `[DESIGN — 2026-10-05, CC; operator may reverse]`
**The question.** Can a compromised agent PROCESS (running as the `felhom-agent` user) become root on its host without
the operator's key? Read on 2026-10-04 (R-861) and measured 2026-10-05 with the real sudo 1.9.16 in a throwaway
container: **with the v0.145.0 sudoers, yes — 23 of 29 attack command lines were allowed.** Two shapes did it:
1. **A glob in the arguments.** Sudo's `*` in arguments also matches spaces, so one grant smuggled extra options:
`pct set [0-9]* -onboot 1` allowed `pct set 100 --dev0 /dev/sda -onboot 1` (a raw host disk for the guest);
`mount --bind /mnt/*/felhom-data /mnt/felhom-drives/*` allowed a `..` path onto `/etc/sudoers.d`; `nft add element …
*` allowed `; flush ruleset`.
2. **A file the agent wrote, installed where root reads it.** A `.mount` unit (bind any directory over `/etc`), a
dnsmasq drop-in (`dhcp-script=` runs as root), the wg-quick config (`PostUp=` runs as root), the OOB sshd config
(`AuthorizedKeysFile` + `StrictModes no`), the guest pre-start hook (Proxmox runs it as root), the shared-parent boot
script, and the agent BINARY itself (the escrow ceremony and the guest hook run it as root; `apply` took a sha the
agent passed).
**The rule after v0.146.1** (the R-861 fix direction: each becomes a root-owned wrapper that checks its own input, or a
fixed file, delivered by the signed config bundle):
- Every varying argument list is a **sudo regular expression** (`^…$`): one value per slot, a fixed character set, no
`..`, no extra argument. Literal lines stay literal.
- **No file the agent wrote is installed where root reads it.** Either the content is FIXED and comes with the signed
bundle (the hook, the shared parent), or a root wrapper checks the CONTENT against the agent's own renderers before
installing it (`felhom-priv-apply`), or the operator's signature is checked as root (`felhom-os-apply agent_update`).
- The pins: `TestManifestCoveredBySudoers` (every command the agent runs is still allowed), `TestSudoersRefusesTheR861
Injections` (the 29 attacks are not), the real-sudo run of both (`audits/hub-safety-2026-10-05/partF/
sudo-container-proof.txt`), and `sudo -l -U felhom-agent` on both demo boxes after the bundle.
| Group | What it is for | How it is narrowed (v0.146.1) | Left open |
|---|---|---|---|
| `FELHOM_MOUNT` | fs-UUID mount units for enrolled drives | install only via `felhom-priv-apply unit <name>`: `[Unit]` only Description + `After=local-fs-pre.target`, `Where=` `/mnt/<name>` or `/mnt/felhom-drives/<name>` and equal to the unit name, `What=` a UUID or a network source, no `bind`/`suid`/`dev`, no continuation lines; systemctl verbs on `mnt-…\.mount` only | — |
| `FELHOM_NETMOUNT` | NAS automount pairs, re-arm, clean-up | same checker; a network share must carry `nosuid,nodev` (the agent now renders them); `rm`/`rmdir`/`reset-failed` one exact name | — |
| `FELHOM_DISK` | SMART, thin-pool, PV and pool reads | exact device / LV patterns (no extra options such as `smartctl -s off`, `lvs --config`) | read-only |
| `FELHOM_PROVISION` | bootstrap config mount, autostart | exact `mpN` spec (`…/guests/<vmid>/bootstrap,mp=/…[,ro=1]`) — no smuggled `--dev0` | — |
| `FELHOM_FORMAT` | data-bearing probe + guarded mkfs | one device path, no space; `mkfs` only through `felhom-mkfs-guarded` (its own root checks) with `ext4`/`xfs` | blkid/lsblk read any `/dev` path (read-only) |
| `FELHOM_DNSMASQ` | the LAN split-horizon resolver | drop-ins via `felhom-priv-apply dnsmasq` (only `bind-interfaces`, `no-resolv`, `listen-address`, `server`, `local`, `address`); `rm` one exact name; exact `pct exec` reads | — |
| `FELHOM_GUESTHOOK` | the pre-start self-heal hook | the hook is a FIXED bundle file; the agent only checks it (`SnippetReady`) and registers it; exact vmid/slot | — |
| `FELHOM_INTERMEDIARY` | the shared drive parent + live drive binds | boot script + unit are FIXED bundle files (the agent only enables the unit); one-segment drive names (no leading dot, no `..`) | — |
| `FELHOM_CONTROLLERSWAP` | the managed controller update | exact vmid; image ref pinned to `gitea.dooplex.hu/admin/felhom-controller:X.Y.Z` for the image check; the inspect template stays free text | **guest-scoped by design**: a compromised agent can still `tee` a chosen (pinned-registry) image ref and restart the guest's bootstrap — the household's data, not host root |
| `FELHOM_STALELOCK` / `FELHOM_SCRATCH_TEARDOWN` | stale-lock clear; failed restore-test scratch | exact vmid; the scratch band `99000[0-9]` was already exact | — |
| `FELHOM_WG` | the off-site tunnel | conf via `felhom-priv-apply wg` (only the keys `renderConf` writes; no `PostUp`/`PreUp`/`DNS`/`Table`; `/32` only) | — |
| `FELHOM_SELFUPDATE` | commit / rollback of the A/B flip | **`apply` removed**: the flip runs only inside `felhom-os-apply agent_update`, after the operator signature, host, window and nonce are checked as root and the staged bytes are hashed ONCE and copied to a root-owned dir (`/var/lib/felhom-os-apply/agent-update/`); the wrapper accepts only that dir | — |
| `FELHOM_SSHD` | the out-of-band operator sshd | config via `felhom-priv-apply sshd-config` (the ONE template, only the Port varies, never 22); the felhom-op key via `sshd-key` (one plain key, no `command=`/`from=` options) | felhom-op's key itself is hub-delivered, not signed: a compromised agent can install its own key for **felhom-op** — whose sudo is scoped (`felhom-op.sudoers`), not root |
| `FELHOM_OOB` | the OOB firewall sets | `add element` takes exactly `{ <ip>[/n] }` or `{ <port> }` — no chained command | — |
| `FELHOM_PBSDR` / `FELHOM_BACKUPTARGET` | PBS DR entry; whole-system backup target | unchanged: the arguments stay coarse, and the root wrappers (`felhom-pbs-apply`, `felhom-backup-target-apply`) are the gate (fixed verbs, own validation) | coarse argv into a checking wrapper |
| `FELHOM_ESCROW` | the recovery-code ceremony (runs the agent binary as root) | the binary is only ever an operator-signed one (`FELHOM_SELFUPDATE`); as root it pins the PVE secret dir and the WG state dir, refuses a storage id that is a path, and reads its two staged files by walking the path with `openat(O_NOFOLLOW)` (no symlink anywhere) | **by design the agent relays R**, so a compromised agent can still learn this box's PBS key through the ceremony — not root, but the backup key |
| `FELHOM_SELFHEAL` / `FELHOM_GUESTNET` / `FELHOM_OSAPPLY` | networking restart; guest DHCP watchdog; OS updates | exact; `felhom-os-apply --plan …` stays the glob line on purpose — the bundle's own self-check reads that exact text, and the wrapper refuses any other plan path (R1) | — |
**What this does not change.** The operator key (`/etc/felhom/operator-signers`, root-owned, never a bundle path) stays
the one trust root; the agent's API token is untouched; a box gets the new rule only through the signed
`agent_config_update` (the order: signed `agent_update` first, then the bundle — after the bundle, an agent below
0.146.0 cannot update itself on that box).
## 4. Control model — reconcile + signed destructive ops
Two channels, split by **reversibility**, not by transport.
@@ -655,7 +708,8 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea
runs as root at guest start (and `pct reboot` is granted); `FELHOM_INTERMEDIARY` installs a script and a systemd unit
that run as root at boot; `FELHOM_ESCROW` runs the agent binary as root, and `FELHOM_SELFUPDATE apply` accepts a sha
the agent itself passes. So a compromised agent PROCESS is root on its host; the root-owned trust files (decision 93,
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants.
the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. **Narrowed in agent
v0.146.1 — §3.1 lists every group, the rule now, and what stays open.**
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
@@ -130,6 +130,17 @@ R-216. Either way the box's reported agent must meet the chosen MinAgent, else t
the Hosts page and logged once per change as `managed floor SERVED`. The rules the operator follows:
`runbooks/publish-train-rules.md` rule 1.
**Boxes left behind (hub v0.135.0, R-604 + R-530).** A per-customer floor wins over the global one, so a global
raise does not move a box whose OWN floor is lower — and `managed floor SERVED` is logged once per change, so that box
was silent (demo-hp missed four raises, 2026-09-21). Now the raise logs one line per such customer and sends ONE
operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set
(`customer_configs.min_controller_set_at`; "unknown" for one set before v0.135.0), and the System page's "Version
floors" table lists the global floor, every per-customer floor with its age, and which ones the global cannot move.
Agents are a separate train (they update only by a per-box signed job, R-530's ruling): the System page shows each
box's agent against the vouched one ("0.142.0 → 0.145.0 (since …)", amber, red after the wait), and a box behind the
vouched agent for 7 days raises `agent_behind` (warning, operator-only; `OS_ALARM_AGENT_BEHIND_AFTER`; the clock starts
when the hub first sees the box behind). `[DESIGN — CC 2026-10-05, operator may reverse; `09` decision 119]`
## 6. Authorization — signed-op queue + editing flow
Implements Part 4's gate on the hub side. The hub holds **no signing key**.
@@ -403,3 +414,35 @@ count appeared was the bind page's passphrase hint, and its English half is now
phrase you received from your operator during setup") — because "five words" stops being true for an
English household, and was already wrong for one whose passphrase predates this release. The
Hungarian „öt szó" is correct and unchanged.
## 16. The operator surface's own safety [DESIGN — hub v0.135.0, CC 2026-10-05, operator may reverse]
### 16.1 Form protection (CSRF) on both login paths (R-135)
The operator logs in two ways: a browser session (`hub_session` cookie + a per-session token on every form) and HTTP
Basic for scripts. Until v0.135.0 a state-changing request with NO cookie skipped the token check — on the reasoning
that it must be a script. It need not be: a browser caches Basic credentials per origin and resends them on a
cross-site form POST (SameSite does not govern the Authorization header). Now a request without a session passes
only with Basic credentials AND the header `X-Felhom-Operator` (any value; scripts send `cli`). A page on another site
cannot add a custom header without a CORS preflight, which the hub never answers. The gate sits in `ServeHTTP` before
the route switch, so it covers every route at once (38 state-changing routes + an unknown path in
`r135_csrf_test.go`); `/login` and the public `/bind/<token>` stay exempt (no operator session to ride; the bind token
is the capability). The other choice — dropping browser-usable Basic auth entirely — was not taken: CC's headless runs
and the runbooks drive the hub with Basic auth, and the header costs them one flag.
### 16.2 Secrets at rest in `hub.db` (R-133, R-821)
Two columns are SEALED (AES-256-GCM, `enc:v1:` + nonce, one key: `OFFSITE_SECRET_KEY` from `Secret/offsite-secret-key`,
never in the database or git): the off-site sub-account passwords (`one_time_secrets.value`, v0.127.0) and, since
v0.135.0, each box's break-glass console password (`host_recovery.secret`) — the same helpers, not a second scheme.
Legacy rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, measured live: 4 rows). No key → a save is
refused; a wrong key → a reveal is a 500 with nothing in the body or the log. Both retrieval paths (the operator page and
the global-key API) open through `GetHostRecoveryCredential`, so the break-glass route still works with the UI down.
**What a copy of `hub.db` still holds readable** (R-879): each box's hub API key, each household's owner passphrase
(`customer_configs.retrieval_password`) and API key, and the PBS-DR token values (`host_pbs_secrets.value`, kept after
use). **And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only
with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export
on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md`
(R-173 — awaiting the operator's decision).
@@ -133,6 +133,16 @@ startup a record still marked running becomes a failed, interrupted result („A
restore, and raised once as `restore_interrupted`. **Notification cooldowns stay in memory** — the
precedent is kept for what it was written for.
**A backup RUN cut off by a power cut or a restart is said (R-519, controller v0.296.0, `09` decision 123).** The
same shape for the app-data run: `appdata-run.json` beside the restore record, written at both ends of a run. A start
that finds it still running turns it into a notice on /backups and /backups/apps („A legutóbbi mentés (…) megszakadt,
mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every step OK. Measured BIGNIGHT F2 (2026-09-14):
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the
operator (R-519 narrowed).*
### Lane 2 — the operator: guest and host recovery
Rebuilding an LXC guest, rebuilding a host, and re-establishing a box's identity are **operator
@@ -661,6 +671,11 @@ successes only. After an agent restart the success is read back from the tier's
**What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before
anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped.
**Still open:** quiescing per tier, so a slow second tier does not keep every app down.
**Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the
local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the
longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both
measurements (≈6 min / 9 apps, ≈8 min / 12 apps), say "minutes, not seconds", and that the off-site copy in the same
run makes it longer.
### 6.5 Kept data — what a removed app leaves on the drive (controller v0.274.0, `09` §3 decision 36)
@@ -327,6 +327,8 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
| `agent_behind` | warning | the box has run an agent OLDER than the vouched one for **7 days** (from when the hub first saw it behind; an unreadable version never counts; nothing vouched → nothing behind) — agents update only by a per-box signed job (R-530), so this is the "nobody signed for this box" alarm (hub v0.135.0) | the box reports the vouched agent (or newer) | `osupdates/r530_agent_alarm_test.go` |
| `floor_raise_skipped` | warning | a GLOBAL controller floor was raised and one or more boxes keep their own LOWER per-customer floor, so the raise does not move them — ONE mail naming them all (R-604, hub v0.135.0) | — (one per raise) | `web/r604_floor_held_back_test.go` |
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
breaks a `not_running` run.
@@ -334,9 +336,9 @@ one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
protected-container check recreates it within 5 minutes, and the host reports every 15. So `tunnel_down` catches
what the box cannot heal — a running container with no connection (wrong token, blocked network).
- The OS alarms are checked **hourly**, re-sent at most **once a week** while true, and forgotten when false, so the
next occurrence is announced again. The four numbers are configuration (`OS_ALARM_STALE_AFTER`,
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`) — *decided by CC unattended,
operator may reverse* (`11` §8.3).
next occurrence is announced again. The numbers are configuration (`OS_ALARM_STALE_AFTER`,
`OS_ALARM_REBOOT_AFTER`, `OS_ALARM_RING0_STALL_AFTER`, `OS_ALARM_NOT_COVERED_AFTER`, `OS_ALARM_BUNDLE_BEHIND_AFTER`,
`OS_ALARM_AGENT_BEHIND_AFTER`) — *decided by CC unattended, operator may reverse* (`11` §8.3; `09` decision 119).
---
@@ -853,6 +853,34 @@ its length, and both fixes cost something the household would notice — operato
a belt repair + retry once when apt itself says "dpkg was interrupted". **Chosen (b)**: a clean pass still costs
one call (pinned by a test). agent v0.145.0.
### 2026-10-05 (afternoon) — decided by CC — operator may reverse (the hub-safety / R-861 brief)
119. **Boxes left behind (R-604, R-530).** Options for the agent alarm's wait: 3 days (a box off for a long weekend
alarms), 7 days (one week, the same wait as the bundle-behind alarm, R-840), 14 days. **Chosen 7 days**
(`OS_ALARM_AGENT_BEHIND_AFTER`), counted from when the hub first sees the box behind. A global floor raise names,
in one operator mail, every box whose own LOWER floor it cannot move. hub v0.135.0, `05` §5.
120. **How the Basic-auth operator path is protected from cross-site POSTs (R-135).** Options: (a) drop browser-usable
Basic auth (CC's headless runs and every runbook POST break); (b) require a custom header on a cookie-less
state change (a browser cannot add one cross-site without a CORS preflight the hub never answers; scripts add one
flag). **Chosen (b)**, header `X-Felhom-Operator`. hub v0.135.0, `05` §16.1.
121. **The console password's seal (R-133).** Options: (a) a second key and scheme for `host_recovery`; (b) the off-site
seal and key already in force (R-821). **Chosen (b)** — the brief asked for the existing pattern, and one key is
one custody question. Consequence named: a database backup needs this key off DooPlex too (R-173). `05` §16.2.
122. **How the agent's admin commands are narrowed (R-861).** Options per group: (a) a root wrapper per group with its
own argv; (b) exact sudo regex patterns for every varying argument + ONE content checker for every agent-written
file root reads + fixed bundle files where the content never varies + the signed update verified as root by the
existing `felhom-os-apply`. **Chosen (b)** — fewest new root programs, and the checker is pinned to the agent's
own renderers by contract tests. Delivery order: signed `agent_update` first, then the bundle. agent v0.146.1,
`03` §3.1.
123. **How long a cut-off backup run is said on the page (R-519).** Options: until the next run of any kind (a failed
run would clear the warning), until the next run that ends with every step OK, or until dismissed. **Chosen: until
a run ends with every step OK.** controller v0.296.0.
124. **How a bundle that adds a path reaches a box (R-880).** An installed `felhom-os-apply` refuses any path not in its
own table (R16). Options: (a) copy the new wrapper onto each box by hand as root (does not scale, leaves the signed
route); (b) a STEP bundle — the box's current bundle with only `felhom-os-apply` replaced, published as
`<ver>-step1` — then the release's bundle, both by signed jobs. **Chosen (b)**, `felhom-agent/scripts/build-step-bundle.py`;
delivered to demo-hp, demo-felhom and Tester 1 on 2026-10-05. `11` §5.4.2 rule unchanged.
### 2026-10-05 (06:49) — four operator rulings (recorded before the work; the night-fixes brief)
100. **Tester 1's Cloudflare tokens, shown in the 2026-10-04 night session's output, are NOT rotated** (option B) —
@@ -0,0 +1,8 @@
== Part A live, hub 0.135.0, 2026-10-05T09:15:34Z, ClusterIP, Basic auth (password from the credentials file, not printed)
POST /configuration/global-floor, Basic, NO header, Origin evil (empty form): 403
POST /no-such-route, Basic, NO header: 403
POST /no-such-route, Basic + X-Felhom-Operator: cli (passes the gate → router 404): 404
POST /no-such-route, header but NO credentials: 401
GET /system, Basic, no header: 200
2026/10/05 11:15:34 [WARN] CSRF rejected: POST /configuration/global-floor from 10.42.0.1:36344
2026/10/05 11:15:34 [WARN] CSRF rejected: POST /no-such-route from 10.42.0.1:41323
@@ -0,0 +1,50 @@
# Hub state-changing routes and how each is protected (hub v0.135.0, R-135)
Every route below is reached through `RequireAuth` → `ServeHTTP`; the CSRF gate is the first thing `ServeHTTP` does for any
method other than GET/HEAD/OPTIONS, before the route switch — so the protection is the same for every route, and an
unknown path is refused by the gate before it can 404.
| Route (representative path) | Protected how | Test |
|---|---|---|
| `POST /configuration` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /apps/demo/reset-telemetry` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /apps/demo/dismiss-issues` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/endpoints` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/endpoints/1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /appliances/1/bind` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /appliances/1/discard` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /hosts/h1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /hosts/h1/reveal-recovery-credential` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /hosts/h1/request-logs` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/block` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/selfbind-link` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/unblock` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/geo/disable` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/create-config` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /customers/c1/request-log-tail` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/new` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configuration/global-floor` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configuration/artifacts` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configuration/password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/delete` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/edit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/offsite-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/claim-resend` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/pbsdr-reissue` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/offsite-freeze` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/regen-password` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /configs/c1/reset` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/remove-unpinned/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/abandon-cancel/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/window-grant/c1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/windows-enabled` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /offsite/key-audit` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/ring/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/enabled/h1` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/approve-now` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /os/approve-docker` | session cookie + token, OR Basic auth + `X-Felhom-Operator` | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /no-such-route` | the gate, before routing (not a route) | `TestR135_BasicAuthWithoutHeaderIsRefused`, `TestR135_SessionWithoutTokenIsRefused`, `TestR135_TheTwoAllowedShapesPassTheGate` |
| `POST /login` | exempt (no session to ride; a wrong password is 401) | — |
| `POST /bind/<token>` | exempt (public self-bind; the e-mailed URL token is the capability, rate-limited) | existing `selfbind_test.go` |
| `GET` routes | not gated by design (a GET must not change state). **Not audited in this session** for a GET that writes — the route switch sends POST-only actions to handlers that check `MethodPost`, but the GET renderers were not read line by line | `TestR135_GetIsNotGated` |
@@ -0,0 +1,6 @@
== Part B live, 2026-10-05T09:16:00Z: live hub.db copied to scratch, only prefix + length selected, copy shredded after
Tester-2-be8404|enc:v1:|87|2026-10-04 16:07:15
demo-felhom-8363b5|enc:v1:|87|2026-07-18 16:30:41
demo-hp-bb76ea|enc:v1:|87|2026-07-21 16:24:27
tester-1-d70be4|enc:v1:|87|2026-10-04 19:40:27
rows NOT sealed: 0
@@ -0,0 +1,6 @@
== Part B live reveal, 2026-10-05T09:16:15Z: POST /hosts/demo-hp-bb76ea/reveal-recovery-credential (Basic + X-Felhom-Operator), body to a 0600 scratch file, shredded after
reveal HTTP 200
username root@pam password length 32 set_at 2026-07-21T16:24:27Z
the revealed password logs in to demo-hp's Proxmox API (POST /api2/json/access/ticket, root@pam): HTTP 200
control, a wrong password: HTTP 401
2026/10/05 11:16:15 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars)
@@ -0,0 +1,42 @@
== Part C readings on DooPlex, READ ONLY, 2026-10-05T09:17:24Z
-- where the hub database lives
hub-data pvc-486c9809-4672-4b56-b70e-0bf01d0c3628 1Gi longhorn
-rw-r--r-- 1 root root 374534144 Oct 5 11:14 hub.db
-rw-r--r-- 1 root root 32768 Oct 5 11:16 hub.db-shm
-rw-r--r-- 1 root root 313152 Oct 5 11:16 hub.db-wal
973.4M 373.4M 584.1M 39% /data
-- the exclusion label: PVC (git, manifests/hub.yaml:47, commit 868e8465 2026-02-16 'updated hub yaml', no reason given) vs the live Longhorn Volume
PVC label: disabled
Volume label: enabled
-- recurring jobs
backup-daily backup 0 4 * * * 1 [default]
backup-weekly backup 0 5 * * 0 1 [default]
-- backups of the hub volume (Longhorn backupstore)
2026-10-04T03:05:01Z Completed 708837376
2026-10-05T02:06:26Z Completed 713031680
-- backup target
nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc?nfsOptions=soft,timeo=330,retrans=3 true
-- which disk holds the target
/dev/sda1
/dev/sdb1
-- DooPlex's own backup service
Mon 2026-10-05 11:30:00 CEST 12min Mon 2026-10-05 11:15:00 CEST 2min 25s ago backup-freshness.timer backup-freshness.service
Tue 2026-10-06 03:19:15 CEST 16h Mon 2026-10-05 03:15:21 CEST 8h ago dooplex-backup.timer dooplex-backup.service
Result=success
ExecMainStatus=0
-- what tells anyone when a backup fails
80:export NOTIFY_ON_FAILURE="true"
81:# export NOTIFY_WEBHOOK_URL="https://your-webhook-url"
137: if [ "${NOTIFY_ON_FAILURE}" = "true" ] && [ -n "${NOTIFY_WEBHOOK_URL}" ]; then
140: "${NOTIFY_WEBHOOK_URL}" || true
prometheus rule backup-freshness-alerts.yml MinecraftBackupStale
prometheus rule backup-freshness-alerts.yml BackupFreshnessExporterDead
prometheus rule longhorn-alerts.yml LonghornVolumeSpaceCritical
prometheus rule longhorn-alerts.yml LonghornVolumeSpaceWarning
prometheus rule longhorn-alerts.yml LonghornVolumeDegraded
prometheus rule longhorn-alerts.yml LonghornNodeStoragePressure
(no rule watches a Longhorn BACKUP's success or age, nor dooplex-backup.service; the only backup-freshness rule is MinecraftBackupStale)
-- does anything leave DooPlex for the hub DB? the off-site route that exists today: ep0 PBS reached through felhom-ep0-pbs-tunnel (pull only, ep0 -> DooPlex)
active
/usr/bin/proxmox-backup-client
/usr/bin/sqlite3
@@ -0,0 +1,9 @@
== Part D live: GET /system on hub 0.135.0 (Basic auth); extracted, no tokens
Version floors: Version floors Global controller floor: 0.292.0 · vouched agent: 0.145.0 Customer Own floor Set Global floor moves it? demo-felhom Demo Ügyfél 0.295.0 unknown no — its own floor applies (at or above the global) demo-hp Demo HP 0.295.0 unknown no — its own floor applies (at or above the global) tester-1 Tester 1 0.295.0 unknown no — its own floor applies (at or above the global)
Agent cell Tester-2-be8404: [('c-warn', '0.142.0 → 0.145.0 (since 2026-10-05)', '3 minor releases behind — sign an agent_update for this box')]
Agent cell demo-felhom-8363b5: []
Agent cell demo-hp-bb76ea: []
Agent cell tester-1-d70be4: []
<td title="current (vouched 0.145.0)">0.145.0</td>
<td title="current (vouched 0.145.0)">0.145.0</td>
<td title="current (vouched 0.145.0)">0.145.0</td>
@@ -0,0 +1,28 @@
== R-518 measure, demo-hp guest 9201, controller 0.295.0, 2026-10-05: POST /api/guest-backup/trigger (the button's call) at 09:19:05Z
2026/10/05 09:19:07 backup_handlers.go:349: [INFO] [web] manual whole-guest backup triggered (quiesce loop)
2026/10/05 09:19:07 quiesce.go:427: [INFO] [quiesce] manual backup requested — quiescing now
2026/10/05 09:19:08 quiesce.go:517: [INFO] [quiesce] backup due on 2 tier(s) — quiescing 9 stack(s): [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin]
2026/10/05 09:19:29 quiesce.go:566: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 started — polling
2026/10/05 09:23:44 quiesce.go:210: [INFO] [quiesce] a backup cycle is already running — skipping this scheduled check
2026/10/05 09:24:09 quiesce.go:643: [INFO] [quiesce] tier local: backup job backup-9201-1791191969558324187 done — next tier may start (app still quiesced)
2026/10/05 09:24:09 quiesce.go:307: [INFO] [quiesce] tier felhom-pbs is BUSY — the agent refused the backup because a concurrent heavy operation holds it. This is contention, NOT a failure: the tier stays due and retries in 15m0s (contended for 0s)
2026/10/05 09:24:09 quiesce.go:504: [INFO] [quiesce] unquiescing (last tier is busy — deferring to a later cycle): restarting 9 stack(s)
-- container StartedAt after the backup (the apps the quiesce stopped):
2026-10-05T09:24:10.019684173Z adventurelog-postgres
2026-10-05T09:24:10.200085281Z adventurelog-frontend
2026-10-05T09:24:15.705236204Z adventurelog
2026-10-05T09:24:16.331987016Z bentopdf
2026-10-05T09:24:16.974996127Z bookstack-db
2026-10-05T09:24:22.669660366Z bookstack
2026-10-05T09:24:23.521923837Z calibre-web
2026-10-05T09:24:24.437773353Z docmost-postgres
2026-10-05T09:24:24.631250114Z docmost-redis
2026-10-05T09:24:35.395095325Z docmost
2026-10-05T09:24:36.934564476Z kimai-db
2026-10-05T09:24:42.753714563Z kimai
2026-10-05T09:24:43.466405175Z opengist
2026-10-05T09:24:44.478064534Z paperless-redis
2026-10-05T09:24:44.688668856Z paperless-postgres
2026-10-05T09:24:54.928089575Z paperless-webserver
2026-10-05T09:24:55.739635651Z privatebin
RESULT: stop began 09:19:08Z (9 stacks quiesced), local tier 09:19:29-09:24:09, PBS tier BUSY (skipped, retried later), last app back 09:24:55Z => longest stop 5 min 47 s, shortest ~5 min 02 s. 21/21 containers running afterwards.
@@ -0,0 +1,15 @@
09:19:20 running=15 phase=idle
09:19:41 running=4 phase=snapshotted
09:20:02 running=4 phase=snapshotted
09:20:23 running=4 phase=snapshotted
09:20:44 running=4 phase=snapshotted
09:21:05 running=4 phase=snapshotted
09:21:26 running=4 phase=snapshotted
09:21:47 running=4 phase=snapshotted
09:22:08 running=4 phase=snapshotted
09:22:29 running=4 phase=snapshotted
09:22:50 running=4 phase=snapshotted
09:23:13 running=4 phase=snapshotted
09:23:34 running=4 phase=snapshotted
09:23:55 running=4 phase=snapshotted
09:24:16 running=6 phase=done
@@ -0,0 +1,39 @@
== RED-PROOF 1 (R-519): runDBDumpsInternal does not call markRunStarted
=== RUN TestRunRecord_TheRealRunIsOnRecordWhileItRuns
run_record_test.go:78: the run was not on record while it ran — a cut here would go unnoticed
--- FAIL: TestRunRecord_TheRealRunIsOnRecordWhileItRuns (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
FAIL
== RED-PROOF 2 (R-519): the synthesised status says Success: true again
=== RUN TestRunRecord_SynthesisedStatusIsNotOKAfterACut
run_record_test.go:97: after a cut the synthesised status still reads OK: &{LastRun:2026-10-05 11:34:19.454323277 +0200 CEST m=+0.001336319 Results:[{DB:{ContainerName:adventurelog ContainerID: DBType: DBUser: DBName: StackName:adventurelog} FilePath:adventurelog-postgres.sql Size:0 Duration:0s Error:<nil> Validation:{Valid:false TableCount:0 Error: FileSize:0 ModTime:0001-01-01 00:00:00 +0000 UTC UserTableFound:false UserRows:0 LooksEmpty:false}}] Success:true Duration:0s}
--- FAIL: TestRunRecord_SynthesisedStatusIsNotOKAfterACut (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s
FAIL
== RED-PROOF 3 (R-519 wiring): main() does not call loadRunRecordAtStartup
=== RUN TestMainWiresRunRecord
run_record_wiring_test.go:47: main() never calls loadRunRecordAtStartup — a cut run is never said on the page
--- FAIL: TestMainWiresRunRecord (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s
FAIL
== RED-PROOF 4 (R-518): the v0.267.0 copy (12 apps / 8 minutes only)
=== RUN TestR518_BackupButtonStatesTheMeasuredDowntime
r518_backup_downtime_copy_test.go:33: hu: "kb. 6 perc" appears 0 times, want it on the page AND in the confirm
r518_backup_downtime_copy_test.go:33: hu: "nem másodperceket" appears 0 times, want it on the page AND in the confirm
r518_backup_downtime_copy_test.go:33: en: "about 6 minutes" appears 0 times, want it on the page AND in the confirm
r518_backup_downtime_copy_test.go:33: en: "not seconds" appears 0 times, want it on the page AND in the confirm
--- FAIL: TestR518_BackupButtonStatesTheMeasuredDowntime (0.07s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s
FAIL
== restored
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.013s
ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.014s
ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.080s
@@ -0,0 +1,7 @@
== 9202 teardown of the throwaway bookstack (installed 09:35:56Z for the R-519 reproduction), 2026-10-05T11:44:31Z
stop: HTTP 200
remove: HTTP 200
containers: 0
volumes: 0
stackdir: none
backups: none
@@ -0,0 +1,9 @@
Oct 05 12:54:08 demo-felhom felhom-os-apply[4115602]: os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0
Oct 05 12:54:08 demo-felhom felhom-os-apply[4115603]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)
Oct 05 12:54:09 demo-felhom felhom-os-apply[4115702]: os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131155]: os-apply: BUNDLE START agent=0.146.1 sha=42333e969028867a authority=signed files=26 write=3 same=22 kept=1 skipped=0
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131156]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-selfupdate-guarded (replaced)
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131157]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-priv-apply (new)
Oct 05 13:09:08 demo-felhom felhom-os-apply[4131158]: os-apply: BUNDLE WROTE /etc/sudoers.d/felhom-agent (replaced)
Oct 05 13:09:09 demo-felhom felhom-os-apply[4131255]: os-apply: BUNDLE DONE agent=0.146.1 written=3 same=22 self-check=ok signers-created=False
Oct 05 13:09:09 demo-felhom felhom-agent[4101006]: time=2026-10-05T13:09:09.974+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=67 total=67 degraded=""
@@ -0,0 +1,6 @@
Oct 05 12:42:18 demo-hp felhom-os-apply[500347]: os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0
Oct 05 12:42:19 demo-hp felhom-os-apply[500348]: os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)
Oct 05 12:42:19 demo-hp felhom-os-apply[500485]: os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False
Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.875+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE START agent=0.146.1-step1 sha=8482851ec27030a7 authority=signed files=22 write=1 same=20 kept=1 skipped=0"
Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.875+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)"
Oct 05 12:42:19 demo-hp felhom-agent[453394]: time=2026-10-05T12:42:19.876+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.146.1-step1 written=1 same=20 self-check=ok signers-created=False"
@@ -0,0 +1,8 @@
== demo-felhom 2026-10-05T11:09:48Z: the checker as the agent user on the real staged files (SAME = nothing changes)
unit mnt-hdd_1.mount rc=3
wg rc=0
sshd-config rc=0
sshd-key rc=0
old route: sudo: a password is required
active active active active
5
@@ -0,0 +1,20 @@
== demo-hp 2026-10-05T10:58:16Z: the checker run AS the agent user through sudo, on the real staged files (each must be SAME — nothing changes)
unit mnt-hdd_1.mount rc=0
wg rc=0
sshd-config rc=0
sshd-key rc=0
dnsmasq rc=0
== an ATTACK, live: the agent stages a unit binding its own dir over /etc/sudoers.d (name and Where agree)
attack rc=3 (installed? no)
== the old route, live: sudo -n install of a staged file
sudo: a password is required
== journal
Oct 05 12:58:08 demo-hp felhom-priv-apply[550409]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 05 12:58:09 demo-hp felhom-priv-apply[550431]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 05 12:58:16 demo-hp felhom-priv-apply[550931]: felhom-priv-apply: SAME unit /etc/systemd/system/mnt-hdd_1.mount
Oct 05 12:58:16 demo-hp felhom-priv-apply[550937]: felhom-priv-apply: SAME wg /etc/wireguard/wg-felhom.conf
Oct 05 12:58:16 demo-hp felhom-priv-apply[550962]: felhom-priv-apply: SAME sshd-config /etc/felhom-sshd/sshd_config
Oct 05 12:58:16 demo-hp felhom-priv-apply[550968]: felhom-priv-apply: SAME sshd-key /etc/felhom-sshd/authorized_keys/felhom-op
Oct 05 12:58:16 demo-hp felhom-priv-apply[550981]: felhom-priv-apply: SAME dnsmasq /etc/dnsmasq.d/felhom-resolver-base.conf
Oct 05 12:58:16 demo-hp felhom-priv-apply[550992]: felhom-priv-apply: REFUSED [U3] unit mnt-..-etc-sudoers.d.mount: mnt-..-etc-sudoers.d.mount: Where=/mnt/../etc/sudoers.d is not /mnt/<name> or /mnt/felhom-drives/<name>
@@ -0,0 +1,2 @@
== LIVE AFTER, demo-felhom, 2026-10-05T11:09:46Z: bundle 0.146.1; 'sudo -l -U felhom-agent <argv>' per case (lists only)
RESULT ok=93 fail=0
@@ -0,0 +1,2 @@
== LIVE AFTER, demo-hp, 2026-10-05T10:57:55Z: bundle 0.146.1, sudo Sudo version 1.9.16p2; 'sudo -l -U felhom-agent <argv>' per case (lists only)
RESULT ok=93 fail=0
@@ -0,0 +1,28 @@
== LIVE BEFORE, demo-felhom (felhom-pve), 2026-10-05T10:54:01Z: bundle 0.145.0 — the v0.145.0 sudoers; 'sudo -l -U felhom-agent <argv>' per case (lists only, runs nothing)
FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'unit' 'mnt-felhom\x2dx.mount'
FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'dnsmasq' '/tmp/felhom-resolver-123456789.conf' 'felhom-x.conf'
FAIL want=ALLOW got=DENY :: '/usr/local/sbin/felhom-priv-apply' 'wg'
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda
FAIL want=DENY got=ALLOW :: /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda
FAIL want=DENY got=ALLOW :: /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x
FAIL want=DENY got=ALLOW :: /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d
FAIL want=DENY got=ALLOW :: /usr/bin/umount /mnt/felhom-drives/x /
FAIL want=DENY got=ALLOW :: /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount
FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount
FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh
FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh
FAIL want=DENY got=ALLOW :: /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf
FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf
FAIL want=DENY got=ALLOW :: /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config
FAIL want=DENY got=ALLOW :: /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000
FAIL want=DENY got=ALLOW :: /usr/bin/systemctl enable --now -- etc-sudoers.d.mount
FAIL want=DENY got=ALLOW :: /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd
FAIL want=DENY got=ALLOW :: /usr/bin/rmdir /mnt/felhom-drives/x /etc
FAIL want=DENY got=ALLOW :: /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset
FAIL want=DENY got=ALLOW :: /usr/sbin/smartctl -a -j /dev/sda -s off
FAIL want=DENY got=ALLOW :: /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x
FAIL want=DENY got=ALLOW :: /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller
FAIL want=DENY got=ALLOW :: /usr/sbin/pct unlock 9201 --whatever
RESULT ok=67 fail=26
@@ -0,0 +1,30 @@
== hp 2026-10-05T09:57:31Z — felhom-priv-apply --check against the box's LIVE files (read only; prints OK or the rule, never content)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-demo-hp.conf: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
== felhom-pve 2026-10-05T09:57:32Z — felhom-priv-apply --check against the box's LIVE files (read only; prints OK or the rule, never content)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
== hp 2026-10-05T10:13:26Z — v0.146.1 checker, --check against LIVE files (read only)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-demo-hp.conf: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
== felhom-pve 2026-10-05T10:13:27Z — v0.146.1 checker, --check against LIVE files (read only)
unit mnt-hdd_1.mount: OK
dnsmasq felhom-guest-9201.conf: OK
dnsmasq felhom-resolver-base.conf: OK
wg: OK
sshd-config: OK
sshd-key: OK
@@ -0,0 +1,85 @@
== RED-PROOF F1 (mount units): felhom-priv-apply stops checking Where
test_U3_bind_over_sudoers_dir (__main__.Refuses.test_U3_bind_over_sudoers_dir) ... ok
test_U3_name_must_match_where (__main__.Refuses.test_U3_name_must_match_where) ... ok
test_U3_network_outside_drives (__main__.Refuses.test_U3_network_outside_drives) ... ok
test_U3_traversal_in_where (__main__.Refuses.test_U3_traversal_in_where) ... ok
OK
== RED-PROOF F2 (WireGuard): felhom-priv-apply allows any key
FAIL: test_W1_postup (__main__.Refuses.test_W1_postup)
FAILED (failures=1)
== RED-PROOF F3 (self-update, root side): felhom-os-apply agent_update skips the signature
FAILED (errors=1)
== RED-PROOF F4 (self-update, agent side): the agent calls felhom-selfupdate-guarded apply itself again
--- FAIL: TestExecutor_HappyPath (0.00s)
executor_test.go:111: execute: agent_update: the root wrapper did not apply it: <nil> (report: ; stderr: )
FAIL
== RED-PROOF F5 (guest hook): SnippetReady accepts any content
--- FAIL: TestSnippetReady (0.00s)
install_test.go:43: a hook with other content read as ready
FAIL
== RED-PROOF F6 (shared parent): the agent installs the boot script from /tmp again when it differs
--- FAIL: TestSharedParentBoot_NeverInstalls (0.00s)
intermediary_install_test.go:54: both missing: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/001/felhom-shared-parent.sh]] — it must install nothing (R-861)
intermediary_install_test.go:54: script differs: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/002/felhom-shared-parent.sh]] — it must install nothing (R-861)
intermediary_install_test.go:54: unit missing: the agent ran [[install -m 0755 -- /tmp/felhom-shared-parent-x.sh /tmp/TestSharedParentBoot_NeverInstalls2355530825/003/felhom-shared-parent.sh]] — it must install nothing (R-861)
== RED-PROOF F7 (escrow, root read): a staged file is read with os.ReadFile (follows a symlink)
--- FAIL: TestAttach_RefusesASymlink (0.00s)
r861_staged_read_test.go:24: a symlinked staged file was read: ok=true err=<nil> value-set=true
FAIL
== RED-PROOF F8 (network shares): nosuid,nodev dropped from the NFS options
--- FAIL: TestPrivApply_AcceptsTheRenderedUnits (0.35s)
r861_privapply_contract_test.go:31: media .mount: REFUSED [U5] mnt-felhom\x2ddrives-media.mount: a network share must carry nosuid,nodev
FAIL
== RED-PROOF F9 (the exact patterns): the v0.145.0 sudoers under the injection test
injections the old file allows (Go matcher): 23
== restored — the same tests green
ok gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/guesthook (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow (cached)
ok gitea.dooplex.hu/admin/felhom-agent/internal/storage 0.474s
ok gitea.dooplex.hu/admin/felhom-agent/internal/capability 0.177s
OK
OK
== RED-PROOF F1 (re-run): the first run did NOT convict — the name check (escape(Where)==name) masked it. The test now uses the
pair that only the Where rule stops: name mnt-..-etc.mount + Where=/mnt/../etc (= /etc). Mutation: stop checking Where
FAIL: test_U3_traversal_in_where (__main__.Refuses.test_U3_traversal_in_where)
FAILED (failures=1)
== RED-PROOF F3 (re-run, clean assertion): felhom-os-apply skips the signature
FAIL: test_a_bad_signature_never_reaches_the_wrapper (__main__.AgentUpdate.test_a_bad_signature_never_reaches_the_wrapper)
AssertionError: None is not true : a job whose signature does not verify was NOT refused: {'agent_update': {'sha256': 'd76b02acf626ce399da7e0a9e17b35563227a4831e14f5edca4ab7cf89eb2c79', 'version': '0.146.0', 'wrapper': '', 'wrapper_rc': 0}, 'layer': 'host', 'mode': 'agent_update', 'pass_seconds': 0.0, 'refused': None, 'release_id': 'agent-0.146.0', 'vmid': 0}
FAILED (failures=1)
== restored
OK
OK
=== Review findings 2026-10-05 (background security review of commit 6ab1e7c) — fixed in v0.146.1, each red-proved
== RED-PROOF S1 (TOCTOU): the wrapper gets the agent's path again (hash, then copy by path)
FAIL: test_signed_update_flips_and_burns_the_nonce (__main__.AgentUpdate.test_signed_update_flips_and_burns_the_nonce)
FAILED (failures=1)
== RED-PROOF S1b: the A/B wrapper accepts the agent's staging dir again
FAIL: test_the_agents_staging_dir_is_refused (__main__.SelfupdateWrapperConfinement.test_the_agents_staging_dir_is_refused)
FAILED (failures=1)
== RED-PROOF S2 (allowlist escape): [Unit] accepts Wants=/Requires=/Before= again
FAIL: test_U2_wants_starts_another_unit (__main__.Refuses.test_U2_wants_starts_another_unit)
FAILED (failures=1)
== RED-PROOF S3 (path traversal): open the whole path with O_NOFOLLOW only
--- FAIL: TestAttach_RefusesASymlinkedDirectory (0.00s)
r861_staged_read_test.go:54: a key behind a symlinked directory was read: ok=true err=<nil>
FAIL
== restored
OK
OK
ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow 0.008s
@@ -0,0 +1,4 @@
== the R-880 step bundle, 2026-10-05T10:24:08Z: base = felhom-agent/0.145.0/felhom-config-bundle.json (sha 78c00adc…, what demo-hp, demo-felhom, tester-1 run); built by scripts/build-step-bundle.py at agent e4b5cf9; published as felhom-agent/0.146.1-step1/felhom-config-bundle.json
step sha 8482851ec27030a7615216048338b1c8939d4e353023ad11c66e6f8930af8613
round trip sha 8482851ec27030a7615216048338b1c8939d4e353023ad11c66e6f8930af8613
same paths: True changed: ['/usr/local/sbin/felhom-os-apply'] version: 0.146.1-step1
@@ -0,0 +1,129 @@
== R-861 real-sudo proof, sudo 1.9.16p2 (debian:trixie throwaway container on DooPlex, 2026-10-05T09:58:38Z); 'sudo -l -U felhom-agent <argv>' per case
-- NEW sudoers (agent v0.146.0): every capability must be ALLOW, every attack DENY
ok ALLOW '/usr/bin/lxc-info' '-n' '9201' '-p' '-H'
ok ALLOW '/usr/bin/mount' '--bind' '/mnt/felhom-drives' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mount' '--make-shared' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mount' '--make-private' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mount' '--bind' '/mnt/felhom-usb/felhom-data' '/mnt/felhom-drives/felhom-usb'
ok ALLOW '/usr/bin/umount' '/mnt/felhom-drives/felhom-usb'
ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-drives'
ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-drives/felhom-usb'
ok ALLOW '/usr/bin/mkdir' '-p' '/mnt/felhom-usb/felhom-data'
ok ALLOW '/usr/bin/chown' '100000:100000' '/mnt/felhom-usb/felhom-data'
ok ALLOW '/usr/bin/systemctl' 'enable' 'felhom-shared-parent.service'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '-mp8' '/mnt/felhom-drives,mp=/mnt/felhom-drives'
ok ALLOW '/usr/sbin/blkid' '-p' '-o' 'export' '/dev/sda'
ok ALLOW '/usr/bin/lsblk' '-J' '-o' 'NAME,FSTYPE,PTTYPE,MOUNTPOINT' '/dev/sda'
ok ALLOW '/usr/local/sbin/felhom-mkfs-guarded' '/dev/sda' 'ext4'
ok ALLOW '/usr/local/sbin/felhom-mkfs-guarded' '/dev/sda' 'xfs'
ok ALLOW '/usr/sbin/smartctl' '-a' '-j' '/dev/sda'
ok ALLOW '/usr/sbin/lvs' '--reportformat' 'json' '--units' 'b' '-o' 'lv_name,data_percent,metadata_percent' '--' 'pve/data'
ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'unit' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'daemon-reload'
ok ALLOW '/usr/bin/systemctl' 'enable' '--now' '--' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'disable' '--' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'stop' '--' 'mnt-felhom\x2dx.mount'
ok ALLOW '/usr/bin/systemctl' 'reset-failed' '--' 'mnt-felhom\x2ddrives-media.automount'
ok ALLOW '/usr/bin/rmdir' '/mnt/felhom-drives/media'
ok ALLOW '/usr/bin/systemctl' 'start' 'networking.service'
ok ALLOW '/usr/bin/chown' '-R' '100000:100000' '/var/lib/felhom-agent/guests/9201'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '-mp0' '/var/lib/felhom-agent/guests/9201/bootstrap,mp=/etc/felhom-bootstrap,ro=1'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '-onboot' '1'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '--hookscript' 'local:snippets/felhom-guest-hook.sh'
ok ALLOW '/usr/sbin/pct' 'set' '9201' '--delete' 'mp0'
ok ALLOW '/usr/sbin/pct' 'reboot' '9201'
ok ALLOW '/usr/bin/apt-get' 'install' '-y' '-q' 'dnsmasq'
ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'dnsmasq' '/tmp/felhom-resolver-123456789.conf' 'felhom-x.conf'
ok ALLOW '/usr/bin/systemctl' 'enable' '--now' 'dnsmasq'
ok ALLOW '/usr/local/sbin/felhom-os-apply' '--plan' '/var/lib/felhom-agent/os/plan-x.json'
ok ALLOW '/usr/bin/systemctl' 'reload' 'dnsmasq'
ok ALLOW '/usr/bin/systemctl' 'restart' 'dnsmasq'
ok ALLOW '/usr/bin/rm' '-f' '/etc/dnsmasq.d/felhom-x.conf'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'ip' '-4' '-o' 'addr' 'show' 'dev' 'eth0'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'exec' 'felhom-controller' 'cat' '/opt/docker/felhom-controller/controller.yaml'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'ip' 'route' 'show' 'default'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'cat' '/etc/network/interfaces'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'pgrep' '-x' 'dhclient'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'dhclient' '-pf' '/run/dhclient.eth0.pid' '-lf' '/var/lib/dhcp/dhclient.eth0.leases' 'eth0'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'cat' '/etc/felhom-controller-image'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'image' 'inspect' 'gitea.dooplex.hu/admin/felhom-controller:0.0.0'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'docker' 'inspect' '-f' '{{.State.Running}}' 'felhom-controller'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'systemctl' 'restart' 'felhom-controller-bootstrap.service'
ok ALLOW '/usr/sbin/pct' 'exec' '9201' '--' 'tee' '/etc/felhom-controller-image'
ok ALLOW '/usr/sbin/pct' 'unlock' '9201'
ok ALLOW '/usr/bin/apt-get' 'install' '-y' '-q' 'wireguard-tools'
ok ALLOW '/usr/local/sbin/felhom-priv-apply' 'wg'
ok ALLOW '/usr/bin/systemctl' 'enable' '--now' 'wg-quick@wg-felhom'
ok ALLOW '/usr/bin/systemctl' 'restart' 'wg-quick@wg-felhom'
ok ALLOW '/usr/bin/systemctl' 'disable' '--now' 'wg-quick@wg-felhom'
ok ALLOW '/usr/bin/wg' 'show' 'wg-felhom' 'latest-handshakes'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'create' 'felhom-pbs' '10.77.0.1' 'felhom-offsite' 'ns0' 'felhom@pbs!ns0' '00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00' '/etc/pve/priv/storage'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'reconcile' 'felhom-pbs' '10.77.0.1' 'ns0' 'felhom@pbs!ns0' '00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00:00' '/etc/pve/priv/storage'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'grant' 'felhom-pbs'
ok ALLOW '/usr/local/sbin/felhom-pbs-apply' 'read' 'felhom-pbs' '/etc/pve/priv/storage'
ok ALLOW '/usr/local/bin/felhom-agent' '--config' '/etc/felhom-agent/agent.json' '--selftest=escrow-create' '--upload' '--output=json'
ok ALLOW '/usr/local/sbin/felhom-selfupdate-guarded' 'commit'
ok ALLOW '/usr/local/sbin/felhom-selfupdate-guarded' 'rollback'
ok DENY /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1
ok DENY /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives
ok DENY /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda
ok DENY /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda
ok DENY /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x
ok DENY /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d
ok DENY /usr/bin/umount /mnt/felhom-drives/x /
ok DENY /usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow
ok DENY /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount
ok DENY /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount
ok DENY /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh
ok DENY /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh
ok DENY /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf
ok DENY /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf
ok DENY /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config
ok DENY /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000
ok DENY /usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service
ok DENY /usr/bin/systemctl enable --now -- etc-sudoers.d.mount
ok DENY /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd
ok DENY /usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow
ok DENY /usr/bin/rmdir /mnt/felhom-drives/x /etc
ok DENY /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset
ok DENY /usr/sbin/smartctl -a -j /dev/sda -s off
ok DENY /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x
ok DENY /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller
ok DENY /usr/sbin/pct unlock 9201 --whatever
ok DENY /usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount
ok DENY /usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf
ok DENY /usr/local/sbin/felhom-priv-apply wg /etc/shadow
rc=0
-- OLD sudoers (agent v0.145.0), the same attacks (this side is the red-proof: 'FAIL want=ALLOW got=DENY' means the OLD file already refused that one; 'ok ALLOW' means the old file let it through)
ok ALLOW /usr/sbin/pct set 9201 --dev0 /dev/sda -onboot 1
ok ALLOW /usr/sbin/pct set 9201 --dev0 /dev/sda -mp8 /mnt/felhom-drives
ok ALLOW /usr/sbin/pct set 9201 --delete mp0 --dev0 /dev/sda
ok ALLOW /usr/sbin/pct set 9201 -mp0 /var/lib/felhom-agent/guests/9201/bootstrap,mp=/x --dev0 /dev/sda
ok ALLOW /usr/bin/mount --bind /mnt/../var/lib/felhom-agent/x/felhom-data /mnt/felhom-drives/x
ok ALLOW /usr/bin/mount --bind /mnt/a/felhom-data /mnt/felhom-drives/../../etc/sudoers.d
ok ALLOW /usr/bin/umount /mnt/felhom-drives/x /
FAIL want=ALLOW got=DENY :: /usr/bin/chown 100000:100000 /mnt/a/felhom-data /etc/shadow
ok ALLOW /usr/bin/mkdir -p /mnt/felhom-drives/x /etc/systemd/system/evil.mount
ok ALLOW /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/units/x.mount /etc/systemd/system/etc-sudoers.d.mount
ok ALLOW /usr/bin/install -m 0755 -- /tmp/felhom-guest-hook-1.sh /var/lib/vz/snippets/felhom-guest-hook.sh
ok ALLOW /usr/bin/install -m 0755 -- /tmp/felhom-shared-parent-1.sh /usr/local/sbin/felhom-shared-parent.sh
ok ALLOW /usr/bin/install -m 0644 /tmp/felhom-resolver-1.conf /etc/dnsmasq.d/felhom-x.conf
ok ALLOW /usr/bin/install -o root -g root -m 0600 -- /var/lib/felhom-agent/wg/wg-felhom.conf /etc/wireguard/wg-felhom.conf
ok ALLOW /usr/bin/install -o root -g root -m 0644 -- /var/lib/felhom-agent/felhom-sshd/sshd_config /etc/felhom-sshd/sshd_config
ok ALLOW /usr/local/sbin/felhom-selfupdate-guarded apply /var/lib/felhom-agent/selfupdate/felhom-agent-9.9.9 0000000000000000000000000000000000000000000000000000000000000000
FAIL want=ALLOW got=DENY :: /usr/bin/systemctl enable --now -- mnt-hdd_1.mount evil.service
ok ALLOW /usr/bin/systemctl enable --now -- etc-sudoers.d.mount
ok ALLOW /usr/bin/rm -f /etc/systemd/system/mnt-felhomx /etc/passwd
FAIL want=ALLOW got=DENY :: /usr/bin/rm -f /etc/dnsmasq.d/felhom-x.conf /etc/shadow
ok ALLOW /usr/bin/rmdir /mnt/felhom-drives/x /etc
ok ALLOW /usr/sbin/nft add element inet felhom_oob operator_ips { 10.77.0.250 } ';' flush ruleset
ok ALLOW /usr/sbin/smartctl -a -j /dev/sda -s off
ok ALLOW /usr/sbin/lvs --reportformat json --units b -o lv_name,data_percent,metadata_percent -- pve/data --config x
ok ALLOW /usr/sbin/pct exec 9201 --keep-env -- docker inspect -f x felhom-controller
ok ALLOW /usr/sbin/pct unlock 9201 --whatever
FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply unit ../../etc/x.mount
FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply dnsmasq /etc/shadow felhom-x.conf
FAIL want=ALLOW got=DENY :: /usr/local/sbin/felhom-priv-apply wg /etc/shadow
rc=1
@@ -0,0 +1,11 @@
== hub System page after delivery, 2026-10-05T11:43:34Z: per box — agent cell, root-files cell (raw)
Tester-2-be8404 | agent: 0.142.0 → 0.146.1 (since 2026-10-05) | root files: []
demo-felhom-8363b5 | agent: 0.146.1 | root files: ['0.146.1']
demo-hp-bb76ea | agent: 0.146.1 | root files: ['0.146.1']
tester-1-d70be4 | agent: 0.146.1 | root files: ['0.146.1']
tester-1-d70be4: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read
demo-hp-bb76ea: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read
demo-felhom-8363b5: Capabilities Capability Status Feature / reason controllerswap-image-inspect critical ok controller-swap / managed auto-update controllerswap-inspect critical ok controller-swap / managed auto-update controllerswap-read
tester-1-d70be4: capability rows 66 {'ok': 66} degraded: []
demo-hp-bb76ea: capability rows 66 {'ok': 66} degraded: []
demo-felhom-8363b5: capability rows 67 {'ok': 67} degraded: []
@@ -0,0 +1,7 @@
== floors 2026-10-05T10:24:56Z: POST /customers/<id>/floor min_controller_version=0.296.0 min_agent=0.131.0
demo-hp: Location: /customers/demo-hp?flash=floor_set
demo-felhom: Location: /customers/demo-felhom?flash=floor_set
tester-1: Location: /customers/tester-1?flash=floor_set
2026/10/05 12:24:56 [INFO] Customer demo-hp controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0")
2026/10/05 12:24:57 [INFO] Customer demo-felhom controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0")
2026/10/05 12:24:57 [INFO] Customer tester-1 controller-version floor override set to "0.296.0" (declared MinAgent "0.131.0")
@@ -0,0 +1,10 @@
== agent_update 0.146.1 (sha badd6c9a…) signed with felhom-op-1, ttl 45m, 2026-10-05T10:25:10Z
-- demo-hp-bb76ea
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-hp-bb76ea-agent_update.json
uploaded signed op to the hub jobs queue
-- demo-felhom-8363b5
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-felhom-8363b5-agent_update.json
uploaded signed op to the hub jobs queue
-- tester-1-d70be4
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-tester-1-d70be4-agent_update.json
uploaded signed op to the hub jobs queue
@@ -0,0 +1,13 @@
== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), demo-hp-bb76ea, 2026-10-05T10:35:43Z
wrote envelope to /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/env-demo-hp-bb76ea-agent_config_update.json
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1 (bundle sha 42333e96…), demo-hp-bb76ea, 2026-10-05T10:42:50Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), demo-felhom-8363b5, 2026-10-05T10:53:36Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1-step1 (bundle sha 8482851e…), tester-1-d70be4, 2026-10-05T10:53:36Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1 (bundle sha 42333e96…), demo-felhom-8363b5, 2026-10-05T10:58:57Z
uploaded signed op to the hub jobs queue
== agent_config_update 0.146.1 (bundle sha 42333e96…), tester-1-d70be4, 2026-10-05T11:13:04Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,4 @@
== vouch 2026-10-05T10:24:18Z: POST /configuration/artifacts (Basic + X-Felhom-Operator), agent 0.146.1, golden 0.296.0, min_agent 0.131.0
HTTP/1.1 303 See Other
Location: /configuration?flash=artifacts_set
2026/10/05 12:24:45 [INFO] Artifact manifest set: agent=0.146.1 golden=0.296.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="42333e969028867ad8142335e6c1bc4040eec231de0d8d330c2d4b2cf7bc3442"
+16
View File
@@ -26,6 +26,22 @@
---
## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124)
The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session).
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-135** | **A cookie-less POST skipped the hub's CSRF gate, so a browser with cached Basic credentials could be made to POST cross-site.** hub v0.135.0: without a session a state change needs Basic credentials AND the header `X-Felhom-Operator` (decision 120); the gate sits before the route switch. 39 paths through RequireAuth→ServeHTTP; red-proof: the old shape lets all 39 through. Live: Basic + no header → 403 (also with `Origin: evil`, also on an unknown path); with the header → passes; header without credentials → 401. **Reasoning kept: a browser cannot add a custom header cross-site without a CORS preflight, which the hub never answers.** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partA/`; `web/r135_csrf_test.go` |
| **R-133** | **Every box's break-glass console password was plaintext in hub.db.** hub v0.135.0: sealed with the off-site seal and key (decision 121); legacy rows sealed at start-up — live: 4 rows sealed, 0 left plain; the demo-hp reveal still returned a password that minted a PVE ticket (HTTP 200; a wrong one 401); a wrong key → 500, nothing in the body or the log, no event. **Reasoning kept: the running hub still holds the key — this closes the database-copy route only; a database backup without `OFFSITE_SECRET_KEY` cannot open the console passwords (R-173).** | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partB/`; `store/r133_recovery_seal_test.go`, `web/r133_reveal_wrongkey_test.go` |
| **R-604** | **A per-customer controller floor silently kept a box out of every global raise (demo-hp missed four).** hub v0.135.0: a global raise logs one line per customer whose own LOWER floor wins and sends ONE operator mail naming them (`floor_raise_skipped`); a per-customer floor records when it was set; the System page's "Version floors" table lists every per-customer floor with its age and which ones the global cannot move. 2 red-proofs. Live: the table shows the three per-customer floors (age "unknown" — set before v0.135.0). The mail was not exercised live (it needs a global raise below an override). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `web/r604_floor_held_back_test.go` |
| **R-530** | **Nothing listed which boxes still run an old agent (agents update only by a per-box signed job).** hub v0.135.0: the System page's Agent cell (box → vouched, how far, since when; red after the wait) and `agent_behind` after 7 days (decision 119). Live: Tester 2 reads `0.142.0 → 0.146.1`. Signing stays per box (the 2026-09-16 ruling: CC may sign until the first paying customer). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partD/`; `osupdates/r530_agent_alarm_test.go` |
| **R-508** | **Customer tester-1 had no registered e-mail, and the page did not say so.** The address has been set since 2026-09-14 (the connect mails reach it — Gmail-read 2026-10-05); hub v0.135.0 adds the page warning: a configured customer with no box and no e-mail shows a red line (three branches tested, red-proof). | CLOSED 2026-10-05 — FIXED hub v0.135.0 | `audits/hub-safety-2026-10-05/partG/`; `web/r508_no_email_banner_test.go` |
| **R-509** | **A box installed for an existing customer never got the connect e-mail.** Fixed in hub v0.114.0; the owed real-mail proof: three mails from the automatic "host delete" trigger, each within 1 s of the hub's own send line (2026-09-16 12:22:59 and 18:17:46, 2026-09-30 07:23:03 UTC), read through the Gmail connector (metadata only). The "e-mail set" trigger shares the send core and is unit-proven. | CLOSED 2026-10-05 — VERIFIED | `audits/hub-safety-2026-10-05/partG/r509-real-mails.txt` |
| **R-880** | **An installed `felhom-os-apply` refuses a bundle naming a path it does not know (R16), so a release whose bundle ADDS a path cannot reach any box on an older bundle** (found 2026-10-05 before delivering v0.146.1, which adds four). Fixed by a step: `felhom-agent/scripts/build-step-bundle.py` — the box's current bundle with ONLY `felhom-os-apply` replaced (same paths), published as `0.146.1-step1`; then the release's bundle. Tests `StepBundle` (the R16 refusal reproduced; the step accepted; exactly one file changed). Live: demo-hp, demo-felhom and Tester 1 each took step1 (`written=1 same=20`) then 0.146.1 (`written=3 same=22`), self-check ok. **Reasoning kept: every future bundle that adds a path needs this step (decision 124); the step package stays published while any box may still be on the old bundle (Tester 2).** | CLOSED 2026-10-05 — FIXED (tooling, agent e4b5cf9) | `audits/hub-safety-2026-10-05/part{F,H}/`; memory `bundle-adding-a-path-needs-step-bundle` |
---
## 2026-10-05 (afternoon) — a box that is not always on: the catch-up, the banner, the alarms; the OS update repairs itself after a power cut (controller v0.295.0, agent v0.145.0, hub v0.134.0, golden 0.295.0; rulings 109–111, CC decisions 112–118)
| Row | What | Closed | Evidence |
File diff suppressed because one or more lines are too long
@@ -0,0 +1,129 @@
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
## 1. What is true today (measured 2026-10-05)
| Question | Answer |
|---|---|
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
| What a copy is worth without the key | The sealed columns (console passwords, off-site passwords) are useless without `OFFSITE_SECRET_KEY`. **The key lives only in `Secret/offsite-secret-key` on DooPlex** (and in the GPG secrets export on the same disk). A backup off DooPlex without a key off DooPlex restores a hub that cannot open any console password. |
## 2. The plan (option A — my pick): a nightly, encrypted, consistent copy on ep0's PBS
ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches it through `felhom-ep0-pbs-tunnel`
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
Without these, every later step backs up something nobody can open after a DooPlex loss.
```bash
# 1. the hub's seal key → the operator's password manager (never a file, never a chat)
sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.OFFSITE_SECRET_KEY}' | base64 -d; echo
# 2. (after Step 2) the backup encryption key's paper copy → the password manager
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
```
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
```bash
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
# namespace for operator data, apart from the households' namespaces
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
```
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
hub image, so the copy must be made by the hub itself.
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
```bash
#!/bin/sh -eu
# /usr/local/sbin/felhom-hub-db-backup — push the newest hub snapshot to ep0 (R-173). Root, 0755.
STAGE=/var/lib/felhom-hub-backup/stage; mkdir -p "$STAGE"; chmod 700 "$STAGE"
SNAP=$(kubectl -n felhom-system exec deploy/hub -- sh -c 'ls -1t /data/snapshots/hub-*.db | head -1')
kubectl -n felhom-system exec deploy/hub -- cat "$SNAP" > "$STAGE/hub.db"
sqlite3 -readonly "$STAGE/hub.db" 'PRAGMA integrity_check' | grep -qx ok # never push a broken copy
export PBS_PASSWORD_FILE=/etc/felhom-hub-backup/token PBS_FINGERPRINT=<ep0 cert fingerprint, as in ep0-datastore-copy.md>
proxmox-backup-client backup hubdb.pxar:"$STAGE" --ns operator --backup-id dooplex-hub \
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
shred -u "$STAGE/hub.db"
# the positive signal the alarm reads (written ONLY on success):
echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ \
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
```
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
```bash
T=$(mktemp -d); chmod 700 "$T"
proxmox-backup-client restore "host/dooplex-hub/$(newest snapshot)" hubdb.pxar "$T" --ns operator \
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
sqlite3 -readonly "$T/hub.db" 'PRAGMA integrity_check' | grep -qx ok
test "$(sqlite3 -readonly "$T/hub.db" 'SELECT COUNT(*) FROM hosts')" -gt 0
test "$(sqlite3 -readonly "$T/hub.db" "SELECT COUNT(*) FROM host_recovery WHERE secret NOT LIKE 'enc:v1:%'")" -eq 0
shred -u "$T/hub.db"*; rmdir "$T"
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
```
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
```yaml
- alert: HubDBBackupStale
expr: time() - felhom_hub_db_backup_last_success_timestamp_seconds > 26*3600 or absent(felhom_hub_db_backup_last_success_timestamp_seconds)
for: 30m
labels: {severity: critical}
annotations: {summary: "The hub database has not reached ep0 for 26 h (R-173)"}
- alert: HubDBRestoreTestStale
expr: time() - felhom_hub_db_restore_test_last_success_timestamp_seconds > 8*24*3600 or absent(felhom_hub_db_restore_test_last_success_timestamp_seconds)
for: 1h
labels: {severity: warning}
```
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
log is not a success.
### Step 7 — prove it once (CC, with the operator's go)
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
start it again.
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
Same Steps 0, 1, 3, 5, 6; the push is `restic backup` to a new sub-account with the R-820 append-only key pin. Costs a
new sub-account and its own key custody; the Storage Box sub-account shell can `rm` (memory: storagebox-subaccount-shell)
unless the pin is right. ep0 already has the server-side prune and the return copy, so A is less new machinery.
+18
View File
@@ -31,6 +31,24 @@ sign; no bundle may add, remove or change them. A box that has no signers file g
4. **Undo** = send the previous release's bundle the same way. The previous copies also stay on the box in
`/var/lib/felhom-os-apply/bundle-prev/<time>-before-<version>/` (the last 3).
## A release whose bundle ADDS a path — the step bundle (R-880, decision 124)
The box's INSTALLED `felhom-os-apply` checks every path of an incoming bundle against its OWN table (R16). So when a
release adds a path (agent v0.146.1 added four), every box on an older bundle refuses it. Send a step first:
```bash
# the bundle the boxes run now — check its sha against the hub's Root files / config-bundle record
curl -fsS -o base.json https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/<old>/felhom-config-bundle.json
python3 felhom-agent/scripts/build-step-bundle.py base.json <new>-step1 step.json # prints the step sha
curl -u admin:<token from a file> -X PUT --upload-file step.json \
https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/<new>-step1/felhom-config-bundle.json # 201
# per box: agent_update <new> → agent_config_update <new>-step1 (step sha) → agent_config_update <new> (release sha)
```
The step is the old bundle with ONLY `felhom-os-apply` replaced, so the old wrapper accepts it (`written=1 same=20`);
the new wrapper then accepts the release's bundle. Done this way on demo-hp, demo-felhom and Tester 1 on 2026-10-05
(`audits/hub-safety-2026-10-05/partH/`). Keep the step package while any box may still be on the old bundle.
## A box from before agent v0.143.0 — the ONE by-hand step (bootstrap)
Such a box's `felhom-os-apply` has no bundle mode, and no signed job can write a root file there (that gap IS
@@ -0,0 +1,5 @@
== round trip 2026-10-05T10:23:04Z: anonymous GET .../generic/felhom-golden/0.296.0/golden.tar.zst
HTTP 200
bytes 648611216
sha256 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
printed 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
@@ -0,0 +1,62 @@
# Golden 0.296.0 — bake + publish + vouch, 2026-10-05 (late afternoon)
Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–5, in the drill VM on DooPlex.
| | Previous (`../golden-0.295.0-2026-10-05/`) | This bake |
|---|---|---|
| `build-golden.sh` | v3.2.0 (sha256 `645b3b659cba…`) | same file, unchanged (agent repo `configs/build-golden.sh`, sha256 `645b3b659cba…`) |
| Controller | `felhom-controller:0.295.0` | **`felhom-controller:0.296.0`** (MinAgent 0.131.0, unchanged) |
| Docker engine | the operator-approved set `os-docker-20261004-142842` | same pinned set (still the only approved release — read from the hub's System page) |
| Guest packages | template | template — `GOLDEN_GUEST_PKGS` EMPTY (no guest release approved) |
## Launch
- Drill VM reverted to `virgin` (no qemu running before), cold-booted per §4.0 at 10:16:31 UTC; `pveversion` = `pve-manager/9.2.2`.
- `pveam update` → `update successful`; template `debian-13-standard_13.6-1_amd64.tar.zst`, `checksum verified`.
- `/root/bake-run.sh` reads the token from the file; launched as transient unit `golden-bake` at 10:17:29 UTC.
- Token copied file → file (`scp`). `systemctl show golden-bake -p Environment -p ExecStart | grep -c -F <token>` = **0**
(control with the token appended = **1**).
## Pass markers (from `bake.log`)
```
[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docke
[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a
[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49
docker OK (overlay2; data-root /var/lib/docker)
live-restore: on
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.296.0
GOLDEN_SHA256=65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
```
No `excluding` and no `FATAL` in the log (grep count 0).
## Round trip
== round trip 2026-10-05T10:23:04Z: anonymous GET .../generic/felhom-golden/0.296.0/golden.tar.zst
HTTP 200
bytes 648611216
sha256 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
printed 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
## Secrets
Saved-log leak grep for the literal token: **0**; positive control (a throwaway copy with the token appended): **1**,
copy shredded.
## Vouch (step 5)
`POST /configuration/artifacts` (operator Basic auth + `X-Felhom-Operator`, hub v0.135.0): agent **0.146.1**, golden
**0.296.0**, `min_agent` **0.131.0**, wrapper sha empty → `303 flash=artifacts_set`; hub log `Artifact manifest set:
agent=0.146.1 golden=0.296.0 min_agent="0.131.0" … bundle_sha="42333e96…"` (`../../audits/hub-safety-2026-10-05/partH/vouch.txt`).
Per-customer controller floors 0.296.0 for demo-hp, demo-felhom, tester-1; the global floor unchanged (Tester 2 not moved).
## Teardown
`pct destroy 9100 --purge`; `shred -u` of the token, the runner script and the log in the VM (log copied off first);
`poweroff`; qemu gone (`ps -eo comm | grep -c qemu-system-x86` = 0); `qemu-img snapshot -a virgin`. Host: nothing
provisioned.
@@ -0,0 +1,340 @@
[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.296.0
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
Logical volume "vm-9100-disk-0" created.
Logical volume pve/vm-9100-disk-0 changed.
Creating filesystem with 8388608 4k blocks and 2097152 inodes
Filesystem UUID: bcf15055-fc5b-43d8-97a3-2a296b616c79
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624
Logical volume "vm-9100-disk-1" created.
Logical volume pve/vm-9100-disk-1 changed.
Creating filesystem with 6291456 4k blocks and 1572864 inodes
Filesystem UUID: 66ff4da9-6a06-4b83-b517-65d331fa9348
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
Total bytes read: 553512960 (528MiB, 125MiB/s)
Detected container architecture: amd64
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
done: SHA256:e5l7hmAwZSLGn/EaSW0I9b1I7uhRo+qmUPfmXZxKops root@felhom-golden
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:6r+bvSA9WxcR3VbYvFsE665lF6r2EP/B3L5NOSOT2VA root@felhom-golden
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
done: SHA256:1n9DFGQdLzJEq0ZdG4SbKoTRctiwtQCnh9qStQmfXRI root@felhom-golden
[golden] starting + installing Docker (official repo, trixie channel) …
[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
installed: containerd.io 2.3.6-1~debian.13~trixie
installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie
installed: docker-ce 5:29.8.2-1~debian.13~trixie
installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie
installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie
installed: docker-compose-plugin 5.6.0-1~debian.13~trixie
[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a
[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
Unable to find image 'hello-world:latest' locally
latest: Pulling from library/hello-world
4f55086f7dd0: Pulling fs layer
4f55086f7dd0: Verifying Checksum
4f55086f7dd0: Download complete
4f55086f7dd0: Pull complete
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
Status: Downloaded newer image for hello-world:latest
docker OK (overlay2; data-root /var/lib/docker)
live-restore: on
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.296.0 (no registry cred at deploy) …
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
0.296.0: Pulling from admin/felhom-controller
774043ccc8cc: Pulling fs layer
ab6b448d4be9: Pulling fs layer
23a5bfa58353: Pulling fs layer
862a57157567: Pulling fs layer
a131840ac86e: Pulling fs layer
a316f627328b: Pulling fs layer
862a57157567: Waiting
a131840ac86e: Waiting
a316f627328b: Waiting
23a5bfa58353: Verifying Checksum
23a5bfa58353: Download complete
862a57157567: Verifying Checksum
862a57157567: Download complete
774043ccc8cc: Verifying Checksum
774043ccc8cc: Download complete
a131840ac86e: Verifying Checksum
a131840ac86e: Download complete
a316f627328b: Verifying Checksum
a316f627328b: Download complete
ab6b448d4be9: Verifying Checksum
ab6b448d4be9: Download complete
774043ccc8cc: Pull complete
ab6b448d4be9: Pull complete
23a5bfa58353: Pull complete
862a57157567: Pull complete
a131840ac86e: Pull complete
a316f627328b: Pull complete
Digest: sha256:e08cbc5868e6b72f82fa9eadba9db03c174c18e6621c044b57c99e9479104203
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.296.0
gitea.dooplex.hu/admin/felhom-controller:0.296.0
[golden] asking the controller which infra images it manages …
[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
v3.7.13: Pulling from library/traefik
e2de96513ba9: Pulling fs layer
b686a4f73445: Pulling fs layer
78cb21c375ca: Pulling fs layer
acb2f33459b1: Pulling fs layer
acb2f33459b1: Waiting
e2de96513ba9: Verifying Checksum
e2de96513ba9: Download complete
b686a4f73445: Verifying Checksum
b686a4f73445: Download complete
acb2f33459b1: Verifying Checksum
acb2f33459b1: Download complete
78cb21c375ca: Verifying Checksum
78cb21c375ca: Download complete
e2de96513ba9: Pull complete
b686a4f73445: Pull complete
78cb21c375ca: Pull complete
acb2f33459b1: Pull complete
Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0
Status: Downloaded newer image for traefik:v3.7.13
docker.io/library/traefik:v3.7.13
2026.9.3: Pulling from cloudflare/cloudflared
2cc7ee286bf3: Pulling fs layer
c172f21841df: Pulling fs layer
218cf840d0d9: Pulling fs layer
f6069939f718: Pulling fs layer
d6b1b89eccac: Pulling fs layer
2780920e5dbf: Pulling fs layer
7c12895b777b: Pulling fs layer
3214acf345c0: Pulling fs layer
52630fc75a18: Pulling fs layer
dd64bf2dd177: Pulling fs layer
b839dfae01f6: Pulling fs layer
ebddc55facdc: Pulling fs layer
c4bc6f35ff5e: Pulling fs layer
b96fe2995f90: Pulling fs layer
58c0c263dc73: Pulling fs layer
bd8962e29291: Pulling fs layer
cac2ae0193cb: Pulling fs layer
f0383d5ebc47: Pulling fs layer
dd64bf2dd177: Waiting
b839dfae01f6: Waiting
ebddc55facdc: Waiting
c4bc6f35ff5e: Waiting
b96fe2995f90: Waiting
58c0c263dc73: Waiting
bd8962e29291: Waiting
cac2ae0193cb: Waiting
f0383d5ebc47: Waiting
2780920e5dbf: Waiting
7c12895b777b: Waiting
3214acf345c0: Waiting
52630fc75a18: Waiting
f6069939f718: Waiting
d6b1b89eccac: Waiting
2cc7ee286bf3: Download complete
c172f21841df: Download complete
218cf840d0d9: Verifying Checksum
218cf840d0d9: Download complete
f6069939f718: Verifying Checksum
f6069939f718: Download complete
d6b1b89eccac: Verifying Checksum
d6b1b89eccac: Download complete
2780920e5dbf: Verifying Checksum
2780920e5dbf: Download complete
7c12895b777b: Verifying Checksum
7c12895b777b: Download complete
3214acf345c0: Verifying Checksum
3214acf345c0: Download complete
2cc7ee286bf3: Pull complete
52630fc75a18: Verifying Checksum
52630fc75a18: Download complete
dd64bf2dd177: Verifying Checksum
dd64bf2dd177: Download complete
b839dfae01f6: Verifying Checksum
b839dfae01f6: Download complete
ebddc55facdc: Verifying Checksum
ebddc55facdc: Download complete
c4bc6f35ff5e: Download complete
58c0c263dc73: Verifying Checksum
58c0c263dc73: Download complete
bd8962e29291: Verifying Checksum
bd8962e29291: Download complete
c172f21841df: Pull complete
b96fe2995f90: Verifying Checksum
b96fe2995f90: Download complete
cac2ae0193cb: Download complete
f0383d5ebc47: Verifying Checksum
f0383d5ebc47: Download complete
218cf840d0d9: Pull complete
f6069939f718: Pull complete
d6b1b89eccac: Pull complete
2780920e5dbf: Pull complete
7c12895b777b: Pull complete
3214acf345c0: Pull complete
52630fc75a18: Pull complete
dd64bf2dd177: Pull complete
b839dfae01f6: Pull complete
ebddc55facdc: Pull complete
c4bc6f35ff5e: Pull complete
b96fe2995f90: Pull complete
58c0c263dc73: Pull complete
bd8962e29291: Pull complete
cac2ae0193cb: Pull complete
f0383d5ebc47: Pull complete
Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c
Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3
docker.io/cloudflare/cloudflared:2026.9.3
1.5.6-stable: Pulling from gtstef/filebrowser
55afa1ecc21d: Pulling fs layer
8ed8f35f8d4f: Pulling fs layer
989b226a579c: Pulling fs layer
660aeead31d5: Pulling fs layer
4f4fb700ef54: Pulling fs layer
adce24567e4c: Pulling fs layer
f17ea56b313b: Pulling fs layer
6b6f3b3efe88: Pulling fs layer
4ed1ca4f3fce: Pulling fs layer
e6fc9c6a5757: Pulling fs layer
d47782d1182a: Pulling fs layer
4f4fb700ef54: Waiting
adce24567e4c: Waiting
f17ea56b313b: Waiting
6b6f3b3efe88: Waiting
4ed1ca4f3fce: Waiting
e6fc9c6a5757: Waiting
d47782d1182a: Waiting
660aeead31d5: Waiting
55afa1ecc21d: Verifying Checksum
55afa1ecc21d: Download complete
660aeead31d5: Verifying Checksum
660aeead31d5: Download complete
4f4fb700ef54: Verifying Checksum
4f4fb700ef54: Download complete
8ed8f35f8d4f: Verifying Checksum
8ed8f35f8d4f: Download complete
55afa1ecc21d: Pull complete
adce24567e4c: Verifying Checksum
adce24567e4c: Download complete
f17ea56b313b: Verifying Checksum
f17ea56b313b: Download complete
6b6f3b3efe88: Verifying Checksum
6b6f3b3efe88: Download complete
4ed1ca4f3fce: Verifying Checksum
4ed1ca4f3fce: Download complete
d47782d1182a: Verifying Checksum
d47782d1182a: Download complete
e6fc9c6a5757: Verifying Checksum
e6fc9c6a5757: Download complete
989b226a579c: Verifying Checksum
989b226a579c: Download complete
8ed8f35f8d4f: Pull complete
989b226a579c: Pull complete
660aeead31d5: Pull complete
4f4fb700ef54: Pull complete
adce24567e4c: Pull complete
f17ea56b313b: Pull complete
6b6f3b3efe88: Pull complete
4ed1ca4f3fce: Pull complete
e6fc9c6a5757: Pull complete
d47782d1182a: Pull complete
Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8
Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable
docker.io/gtstef/filebrowser:1.5.6-stable
1.1.0: Pulling from admin/felhom-samba
897d797d2723: Pulling fs layer
3051591aa250: Pulling fs layer
ce57a3f93416: Pulling fs layer
fb94eeec2fe1: Pulling fs layer
fb94eeec2fe1: Waiting
ce57a3f93416: Verifying Checksum
ce57a3f93416: Download complete
fb94eeec2fe1: Verifying Checksum
fb94eeec2fe1: Download complete
897d797d2723: Verifying Checksum
897d797d2723: Download complete
3051591aa250: Verifying Checksum
3051591aa250: Download complete
897d797d2723: Pull complete
3051591aa250: Pull complete
ce57a3f93416: Pull complete
fb94eeec2fe1: Pull complete
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
gitea.dooplex.hu/admin/felhom-samba:1.1.0
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
[golden] identity-clean + minimize …
[golden] stop + archive …
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
INFO: archive file size: 618MB
INFO: Finished Backup of VM 9100 (00:00:29)
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_05-12_21_35.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
[golden] publishing golden (648611216 bytes, sha256 65a87efc4b557caa…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.296.0/golden.tar.zst
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.296.0
GOLDEN_SHA256=65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.296.0 / 65a87efc4b557caaa0199f09fd398144707c8ffc7eb10e02c30d247e76517c1f
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
+3 -3
View File
@@ -3,19 +3,19 @@
**Operator action on deploy: none.** Scripts that POST to the hub with the operator password must now send
`X-Felhom-Operator: cli` (the build-deploy skill and the memory note say so).
- **R-135 (`05` §8.1).** A state-changing request without a session cookie used to pass the CSRF gate unconditionally
- **R-135 (`05` §16.1).** A state-changing request without a session cookie used to pass the CSRF gate unconditionally
(measured: a Basic-auth POST with no cookie reached the handler). Now it passes only with Basic credentials AND the
`X-Felhom-Operator` header — a page on another site cannot add a custom header (no CORS preflight is answered), so a
browser with cached Basic credentials can no longer be made to POST. The session path is unchanged (cookie + token).
`web/r135_csrf_test.go` drives all 38 state-changing routes plus an unknown path (39 paths) through RequireAuth → ServeHTTP;
red-proof: the old `return true` lets all 39 through (none answers 403).
- **R-133 (`05` §8.2).** `host_recovery.secret` (each box's break-glass `root@pam` password) is sealed with the SAME
- **R-133 (`05` §16.2).** `host_recovery.secret` (each box's break-glass `root@pam` password) is sealed with the SAME
AES-256-GCM seal and key as the off-site passwords (`OFFSITE_SECRET_KEY`, `store/offsite_seal.go` — reused, not a second
scheme). Existing rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, beside the off-site one). No key →
the save is refused; a wrong key → the reveal is a 500 with nothing in the body or the log, and no "revealed" event.
Both retrieval paths (the operator page and the global-key API) open it through the same store call.
`store/r133_recovery_seal_test.go`, `web/r133_reveal_wrongkey_test.go`, `cmd/hub/r133_wiring_test.go`; 2 red-proofs.
- **R-604 (`05` §7).** A global floor raise now logs one line per customer whose OWN floor is lower (it wins, so the
- **R-604 (`05` §5).** A global floor raise now logs one line per customer whose OWN floor is lower (it wins, so the
raise does not move that box) and sends ONE operator mail naming them (`floor_raise_skipped`, warning, operator-only).
A per-customer floor now records when it was set (`customer_configs.min_controller_set_at`); the System page has a
"Version floors" table: the global floor, every per-customer floor with its age, and which ones the global cannot
+1 -1
View File
@@ -882,7 +882,7 @@ func (s *Server) handleLogin(w http.ResponseWriter, r *http.Request) {
// does NOT prove the request is programmatic. A page on another site can make the browser POST a
// form with the operator's cached Basic credentials; it cannot add a custom header (that needs a
// CORS preflight, which the hub never answers). So the header is the proof the old check assumed.
// Decided by CC — operator may reverse (`05` §8.1). Pinned by r135_csrf_test.go.
// Decided by CC — operator may reverse (`05` §16.1). Pinned by r135_csrf_test.go.
const OperatorCLIHeader = "X-Felhom-Operator"
// validateCSRF checks a state-changing request (R-135). Two ways pass, nothing else: