diff --git a/CONTEXT.md b/CONTEXT.md index 80ad9d2b..68df1214 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,15 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-04 (night) — R-840 BUILT (agent v0.143.0, hub v0.133.0, installer 1.31.0), R-859 + R-860 FIXED (hub v0.133.0, +> controller v0.293.0), golden 0.293.0 vouched, drill-r50 removed.** The config bundle (`11` §5.4.2): one table of 22 +> root paths in `felhom-os-apply`; signed `agent_config_update`; the trust root never a bundle path; installer installs the +> same bundle; a pre-0.143.0 box needs `scripts/felhom-bundle-bootstrap.sh` once (demo boxes done; Tester 2 = R-862, +> operator). Test approvals (`11` §5.3.1): marked, cancelled at a start without the override (4 cancelled 18:20 UTC; the +> operator's Docker approval kept — CC decision). Docker socket self-heal: controller exits after 60 s of refusals. +> Floors: demo customers 0.293.0, global 0.292.0 (Tester 2 not moved). Tester 2: signed agent_update to 0.143.0 queued. +> R-861 opened (the agent sudoers is root-equivalent). Report: `REPORT-r840-config-bundle-2026-10-04.md`. + > **2026-10-04 (~18:49) — rulings 96–99 (the R-840 / Tester 2 brief), recorded before the work.** `09` §3: **96** build > R-840 now (option A, a signed config bundle; replaces 82); **97** Tester 2 may receive only the signed agent update, > the bundle by the route and the one-time live-restore reload (~19:05: the operator will install the bootstrap file by diff --git a/REPORT-r840-config-bundle-2026-10-04.md b/REPORT-r840-config-bundle-2026-10-04.md new file mode 100644 index 00000000..83588b8f --- /dev/null +++ b/REPORT-r840-config-bundle-2026-10-04.md @@ -0,0 +1,125 @@ +# REPORT — the config bundle (R-840), test approvals (R-859), the Docker-socket self-heal (R-860), Tester 2, drill-r50 + +2026-10-04, evening–night. Brief: "a signed route that brings an installed box's root-owned parts up to date (R-840)…". +Architecture read first: `11-os-updates.md` (§5.3, §5.4, §5.8, §5.9, §8), `03-host-agent.md` (§3, §4, §11), +`04-control-plane-authorization.md` (§3), `07-backup-architecture.md` (§6.4). Evidence: `documentation/audits/r840-config-bundle-2026-10-04/`. + +## The Part table + +| Part | Result | Why / what changed | +|---|---|---| +| A — what Tester 2 has | **done (read through the hub; no route to the box)** | The brief's timestamp reasoning was wrong (UTC vs local). Tester 2 has installer 1.30.0, agent 0.142.0, the crash guard armed, `live-restore` on, Docker 29.8.2, the re-made golden. It lacks only the R-858 wrapper fix. | +| A — the `felhom-pbs` warning | **done: not a fault** | By design (R-723): a new box's first-hour tier skip is recorded, never mailed; the tier came 7 min later and backed up at 16:31 UTC. No row. | +| B — the bundle route | **done, live on both demo boxes** | Agent 0.143.0 (bundle mode in `felhom-os-apply`, signed `agent_config_update`), hub 0.133.0, installer 1.31.0. Changed: the installer is NOT the self-update wrapper (it cannot do it); a one-time bootstrap is needed on older boxes. | +| C — Tester 2 | **act 1 sent; act 2 waits for the operator; act 3 not needed** | Act 1 (signed agent update to 0.143.0): queued 18:35 UTC — see "Tester 2" below. Act 2 needs the one by-hand bootstrap (R-862, the operator's tunnel). Act 3: `live-restore` is already on. | +| D — test approvals end | **done, live** | Hub 0.133.0 cancelled the 4 test approvals of 2026-10-04 at start (18:20 UTC). Changed: the operator's own Docker button approval was not backfilled (CC decision). | +| E — self-heal after any Docker restart | **done, live (9202 + demo-hp)** | Controller 0.293.0. Changed: only a `docker.socket` restart breaks it; a dockerd crash or `systemctl restart docker` does not. | +| F — new installs start with the approved fixes | **built; tonight's golden carries none** | `build-golden.sh` 3.2.0 `GOLDEN_GUEST_PKGS`. No guest release is in force tonight (the test one was cancelled), so the golden keeps the template (49 Debian updates pending for the next real approval). | +| G — drill-r50 | **done** | Deleted through the product (operator's yes: it touched ep0). On ep0: its PBS namespace did not exist; its WireGuard peer left (5 → 4 peers pushed). | +| H — release, golden, records | **done** | Agent 0.143.0, controller 0.293.0, hub 0.133.0, installer 1.31.0, golden 0.293.0 baked + vouched. Floor 0.293.0 for the two demo customers only (Tester 2 untouched). | + +## Claims in the brief that turned out wrong + +1. **The timestamp reasoning about Tester 2 (§1).** Tester 2 was bound at **16:06 UTC = 18:06 local**; the releases were + compared in local time. At 18:06 local, installer 1.30.0 (17:36), agent 0.142.0 (vouched ~17:49) and the re-made golden + (17:48) were all current. Measured (hub records): agent **0.142.0**, crash guard **armed** (`kernel.panic` 10), `live-restore` + **on**, Docker **29.8.2**; the wrapper reports a facts mode (so it is not "old"). Only agent 0.142.1's R-858 wrapper fix is + missing — and the root files of 0.142.0 equal 0.143.0's in every file but `felhom-os-apply`. +2. **"The felhom-pbs warning is a fault."** It is the designed first-hour behaviour (R-723): recorded, not mailed. +3. **"The self-update wrapper can install a bundle safely."** It cannot: it is a fixed sh script that swaps one binary. The + bundle is installed by `felhom-os-apply` (already reachable through the agent's sudoers line). And **no** existing root + helper can install the first bundle on an older box — one by-hand bootstrap per older box (done on both demo boxes). +4. **"A restart of an existing container picks up the new socket."** True — measured twice (controller exit → restart policy; + `docker restart traefik`): both came back on the current inode, same container ids. +5. **"dockerd restarts on its own … the box shows DOWN until a person acts."** Only when the socket FILE is re-created + (`docker.socket` restart, i.e. a docker-ce upgrade or by hand). A dockerd crash or `systemctl restart docker` keeps it (measured). +6. **Act 3 for Tester 2 (`live-restore` reload)** is not needed: it is already on (from the golden). + +## Part A — Tester 2, read back (hub records; CC has no route to the box) + +| Item | Brief expected | Measured | Source | +|---|---|---|---| +| bound / enrolled | 16:06 (read as local) | 16:06:27 UTC bind, 16:07:04 UTC host | `appliance_registrations`, `hosts` | +| installer | 1.29.0 | **1.30.0** (the crash guard is installed — only 1.30.0 does that) | host report `system.facts.host.crash_guard` | +| agent | 0.141.1 | **0.142.0** | `/hosts` | +| PBS wrapper sha | — | `104db0a4…` = every release's | host report `wrapper_sha256` | +| `felhom-os-apply` | old (no facts, no Docker) | 0.142.0's (it answers the facts mode) | facts present | +| sudoers | — | not readable from the hub; 0.142.0's file equals 0.143.0's (tag diff); capability probe all ok | `/hosts/Tester-2-be8404` | +| crash guard / `kernel.panic` | absent / 0 | **armed / 10** | facts | +| operator-signers | absent | not readable from the hub; installer 1.30.0 writes it | inference | +| `live-restore` / Docker | off / unpinned | **on / 29.8.2** | facts | +| golden | first 0.292.0 bake | the re-bake (live-restore + 29.8.2) | facts | +| OS releases installed | TEST-approved | guest `os-guest-20261004-123933` (49 pkgs, 16:24) and host `os-host-20261004-124133` (106, 16:25), both approved "auto" 1.5 h after first seen = the TEST wait | `os_reports`, `os_releases` | + +## Part B — the bundle + +Shape, checks, trust rule, bootstrap: `11` §5.4.2; runbook `runbooks/config-bundle.md`. Red-proofs: 22 of 22 wrapper rules +(`partB/redproof.txt`), hub 6 of 6 (`partB/hub-bundle-redproof.txt`), Go executor tests. Live: + +| Step | Box | Result | +|---|---|---| +| read the box's files vs the bundle | both | 21 of 22 identical; only `felhom-os-apply` differed (`b1`) | +| bootstrap, wrong sha | demo-hp | `STOP`, nothing changed (`b2`) | +| bootstrap | both | the new `felhom-os-apply`, self-check ok (`b2`, `b3`) | +| job A, wrong sha | demo-hp | refused by the agent, nothing changed (`b4`) | +| job B, the 0.143.0 bundle | both | written 0, same 21, kept 1, self-check ok, probe 71/71 (`b4`) | +| a bundle with one deliberate change | demo-hp | written 1 (that file), self-check ok (`b6`) | +| its undo (the 0.143.0 bundle again) | demo-hp | written 1, the release file back (sha `b71d8698…`), probe 71/71 (`b6`) | +| replay of job B | demo-hp | `REJECTED … replay (nonce already seen)` (`b6`) | +| the installer's new path | demo-felhom | written 0, record `installer`, services active (`b5`) | + +## Part D — test approvals + +Marked, cancelled at start, amber, operator event, ring-1 bump; red-proofs 6 of 6 (`partD/d-redproof.txt`). Live at the +18:20 UTC start: `os-guest-20261004-123933`, `os-20261004-091417`, `os-host-20261004-124133`, `os-host-20261004-124034` +cancelled; one mail (three held by the cooldown); the System page lists them; the Docker release stays (`partD/d2`). The +2026-10-04 fact and why the risk was small: `runbooks/os-updates-test-waits.md`. + +## Part E — the measured heal + +| Box | Break | Controller back | traefik back | Container ids | +|---|---|---|---|---| +| 9202 (controller 0.291.0, before the fix) | `systemctl restart docker.socket` | never by itself (blind 2+ min, health "healthy") | never | same | +| 9202 (0.293.0) | same | +73 s | +104 s | same (`e7`) | +| demo-hp 9201 (0.293.0) | same | +88 s | +120 s | 21 of 21 same (`e8`) | + +`systemctl restart docker` and `kill -9 dockerd`: socket inode unchanged, nothing to heal (`e1`, `e2`). Red-proof 9 of 9 (`e6`). + +## Part F — the first-night count + +Golden 0.293.0 (`documentation/tests/golden-0.293.0-2026-10-04/`): no guest release in force → template versions kept; +**49** Debian updates pending in the baked guest, which the next REAL approval (tomorrow, after 24 h + a night) will bring. +A ring-1 box from this golden installs **0** on its first night until then. The host: the agent's first leg already runs +right after the first whole-guest backup (Tester 2: 17 min after enrolment, 106 host packages in 44 s) — an installer pass +would cost ~45 s and run before any backup exists; not built. + +## Part G — drill-r50 + +What the product delete touched (read first, `partG/g1`): hub rows (host, guest, reports 185, telemetry 185, the recovery +credential, the DR recipe, the claim, an appliance registration, the WireGuard peer 10.77.0.4), ep0's PBS tenancy +(`tenantsync.Deprovision`, namespace `drill-r50`), ep0's WireGuard peer list. No Storage Box (no off-site tier), no mail +(no address). Asked the operator (ep0); yes. Result: `deprovision ok (existed=false)` — nothing destroyed on ep0; the peer +left ep0 at the next push (5 → 4); `/hosts` without drill-r50; the delete preview 404s (`partG/`). + +## Rows + +Before **333**, after **334**. Closed: **R-840** (built), **R-859** (opened and closed), **R-860** (opened and closed). +Opened: **R-861** (P2 — the agent's sudoers is root-equivalent; read, not exploited), **R-862** (P3 — Tester 2's bootstrap, +operator). R-857 not fixed: avoided by baking under a new controller version. + +## Decisions taken by CC (operator may reverse) + +1. The one-time backfill marks only the AUTOMATIC test approvals; your Docker button approval of 2026-10-04 stays in force. +2. The bundle alarm fires after 7 days behind (`OS_ALARM_BUNDLE_BEHIND_AFTER`). +3. The bundle lives inside `felhom-os-apply` (no new sudoers line), not in a new tool. +4. The controller fixes Part E (it reaches every box by the floor), not the agent (that would need a sudoers change). +5. The demo customers' floor is 0.293.0; the global floor stays 0.292.0 so Tester 2's controller does not move this session. + +## Teardown + +- Machines: 9202 left on controller 0.293.0, working (all apps up). demo-hp and demo-felhom: agent 0.143.0, bundle 0.143.0, + controller 0.293.0; the test line removed again on demo-hp. Drill VM off and at `virgin`. +- Hosts: the bootstrap script, the installer harness and the test script removed from `/root` on both demo hosts. The + bundles' previous copies stay in `/var/lib/felhom-os-apply/bundle-prev/` by design. +- Hub: the registry's test version `0.143.0-r840test` deleted (204, then 404); drill-r50 removed; the scratchpad copy of + the hub DB deleted at the end. diff --git a/STATUS.md b/STATUS.md index c0658dd3..f1199f55 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,10 +1,44 @@ # STATUS — what works, what's broken, what's next -**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).** +**Ready for the first real tester (Tester-2): yes. Tester 2 is installed and running (since 2026-10-04 18:06).** -**Updated 2026-10-04 (late evening): the hub has a System page with every box's versions and the update buttons; Docker -updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.1 and -controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.1.** +**Updated 2026-10-04 (night): installed boxes can now get new root files by a signed job; test approvals end with the +test; the controller repairs itself after a Docker socket restart. Both demo boxes run host agent 0.143.0 and controller +0.293.0. Hub 0.133.0. Installer 1.31.0. New installs get golden 0.293.0 with agent 0.143.0.** + +## Today (2026-10-04, night): root files for installed boxes, test approvals, Docker self-repair + +**Decisions I took myself (you may reverse each):** +- The hub cancelled only the AUTOMATIC test approvals of today (guest and host fixes). Your own "Approve Docker set" + press stays in force. From now on every approval made during a test wait gets the test mark, button or not. +- A box whose root files are behind the approved ones for 7 days sends you a mail. +- The demo boxes' controller floor is 0.293.0; the fleet floor stays 0.292.0, so Tester 2's controller did not move. + +**What I did:** +- **Root files for installed boxes (your ruling):** a box's root-owned files (permissions list, helper scripts, crash + guard) now travel as one signed package. The box checks your signature, every file and itself; on any problem it + puts the old files back. Proven on both demo boxes: a wrong package refused, a package with one changed line + installed and undone, a copied old job refused. New installs use the same package. +- **One catch:** a box installed before tonight cannot take the FIRST package by itself; it needs one small step by + hand. I did it on both demo boxes. **Tester 2 needs it from you** (below). +- **Tester 2 was not as old as the brief thought:** it was installed at 18:06 local, after the crash guard and the new + image. It already has the crash guard, live-restore and Docker 29.8.2. It lacks only tonight's Docker-update fix. I + sent it the signed agent update. +- **Test approvals end with the test:** the hub cancelled today's 4 test approvals at its restart (you got one mail). + Tester 2 keeps what it installed; no further box installs them. The same fixes get a real approval after 24 h + a night. +- **Docker self-repair:** if Docker's socket is re-created, the controller now restarts itself and the web router within + about 2 minutes. Proven on the scratch box and on demo-hp, no app restarted. +- **drill-r50 is gone** from the hub. On ep0 nothing was destroyed (it had no backups there); its tunnel entry left. +- The "felhom-pbs skipped" line on Tester 2 is normal for a new box's first hour (no mail was sent). +- New-install image 0.293.0 baked and approved. + +**Needs you (nothing breaks if you wait):** +- **Tester 2's one-time step** (5 minutes): connect your tunnel, `ssh -p 8822 felhom-op@10.77.0.5`, reveal the root + password in the hub (Hosts → Tester-2 → Console access; this writes one line on Tester 2's timeline), `su -`, then + run the three commands in `documentation/runbooks/config-bundle.md` ("Tester 2"). Tell me when done; I send the package. + If you wait: Tester 2 keeps working; it just cannot get new root files until then. +- **Found, not fixed:** the agent's permission list is wider than "minimal": a broken-into agent could become root on + its own box. Worth fixing before the first paying customer. ## Today (2026-10-04, ~18:30): you found a bug — the Docker update blinded the box's controller diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 5cf7a441..245de38a 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) | | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | -| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.142.0, hub v0.132.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes need the wrapper, the trust files and the guard by hand (R-840); the kernel lane (R-836); facts reach the hub late after a boot (R-853) | +| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853) | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | | **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. | diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 24a66359..1f379644 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -644,8 +644,18 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea /var/lib/felhom-agent/os/plan-*.json`). The agent has no `apt` grant of its own for this; every rule (no removal, no downgrade, no new or unlisted package, Debian origin only, the box's own customer guest only) is in the wrapper, red-proved per rule. The leg runs after a successful primary whole-guest backup, under the heavy-op gate. - **[FACT] The signed agent update does NOT carry the wrapper or the sudoers line** — only the installer installs - them (R-840). + ~~**[FACT] The signed agent update does NOT carry the wrapper or the sudoers line** — only the installer installs + them (R-840).~~ **[FACT, 2026-10-04, agent v0.143.0] The config bundle does:** a signed `agent_config_update` brings + every root-owned file (sudoers, wrappers, units) as one checked unit; `felhom-os-apply` verifies it itself, keeps the + previous copies and undoes on a failed self-check; the trust root (`/etc/felhom/operator-signers`, `os-trust.json`) is + never a bundle path. The installer installs the same bundle. A box from before 0.143.0 needs one by-hand bootstrap. + Design: `11` §5.4.2; runbook `runbooks/config-bundle.md`. +- **[FACT, 2026-10-04 — R-861] "Root-minimized" overstates it.** Read from `configs/felhom-agent.sudoers`: the agent user + can already reach root without the operator key — `FELHOM_GUESTHOOK` installs a hookscript from `/tmp` that Proxmox + runs as root at guest start (and `pct reboot` is granted); `FELHOM_INTERMEDIARY` installs a script and a systemd unit + that run as root at boot; `FELHOM_ESCROW` runs the agent binary as root, and `FELHOM_SELFUPDATE apply` accepts a sha + the agent itself passes. So a compromised agent PROCESS is root on its host; the root-owned trust files (decision 93, + the bundle's R17) are defence in depth, not a boundary, until R-861 narrows these grants. - **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle, so the **agent updates the controller**: snapshot-before-update (free rollback, because the controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback diff --git a/documentation/architecture/04-control-plane-authorization.md b/documentation/architecture/04-control-plane-authorization.md index c353da2b..cd39ec76 100644 --- a/documentation/architecture/04-control-plane-authorization.md +++ b/documentation/architecture/04-control-plane-authorization.md @@ -113,6 +113,16 @@ Proven twice: `demo-hp-bb76ea` 2026-09-15, `demo-felhom-8363b5` 2026-09-16. **Re sale** — at that point the key belongs behind the operator (or a hardware key, §7), and a fleet rollout step still has to be designed (R-530). +**Who may sign a config bundle (`agent_config_update`, R-840, decision 96, 2026-10-04).** The same operational key +(`felhom-op-1`) and the same custody: CC may sign it under the ruling above (per box, until the first paying +customer). It is destructive-class (the agent's gate requires the operational key) AND the box's root wrapper verifies +the signature itself against the ROOT-OWNED `/etc/felhom/operator-signers` — not the agent's config. The recovery key +does not sign bundles. **A bundle can never change who may sign:** the signers file and `os-trust.json` are not bundle +paths (R17); on a box with no signers file, only a job signed by the installer's pinned `felhom-op-1` is accepted, and +the file is created with exactly that key. **Rotating a signer** (adding `felhom-op-2`, removing `felhom-op-1`) is a +separate act, not built: it needs its own op, signed by the key being replaced (or the recovery key), with its own +design (`11` §5.4.2). + ## 4. Rotation & compromise recovery The agents pin the operator public keys. The danger: rotation must **not** flow as plain hub config, diff --git a/documentation/architecture/11-os-updates.md b/documentation/architecture/11-os-updates.md index 642aea85..755b4d3c 100644 --- a/documentation/architecture/11-os-updates.md +++ b/documentation/architecture/11-os-updates.md @@ -211,6 +211,17 @@ downloadable — but `curl`, `libcurl*` and `libssh2` were "not covered" in the ALREADY ran the newer version and so installed nothing. `[PROPOSAL]` step 1 reports the full installed `package=version` set after the run, and approval covers every version ring 0 runs healthy. +### 5.3.1 Test approvals end with the test — BUILT 2026-10-04 (hub v0.133.0, R-859) `[FACT]` + +An approval made while a TEST override (`OS_APPROVE_AFTER`, `OS_APPROVE_NIGHTS`, `OS_DOCKER_APPROVE_NIGHTS`) is active +carries a `test` mark (amber on the System page). At every hub start WITHOUT an override, every test approval that no +real approval has superseded is cancelled: never served again, ring-1 boxes bumped, one operator event +`os_release_cancelled` each; what boxes installed stays; the ruled wait approves the same set again as a real release. +A one-time backfill marked the AUTOMATIC approvals made under 24 h after first seen; the 2026-10-04 guest and host test +approvals (which Tester 2 installed on its first night) were cancelled at 18:20 UTC. The operator's Docker button +approval of that day stays in force (decided by CC unattended — operator may reverse). Runbook: +`runbooks/os-updates-test-waits.md`. Evidence: `audits/r840-config-bundle-2026-10-04/partD/`. + ### 5.4 Who runs it, and with what permission - **The agent runs every OS update**, for the host and for the guest (`pct exec`). The controller does @@ -289,6 +300,46 @@ os-apply: FAILED rc= step= — dpkg state: docker= +BLIND traefik_inode=18453 host=41134 ++16s controller->docker= +BLIND traefik_inode=18453 host=41134 ++21s controller->docker= +BLIND traefik_inode=18453 host=41134 ++26s controller->docker= +BLIND traefik_inode=18453 host=41134 ++31s controller->docker= +BLIND traefik_inode=18453 host=41134 ++37s controller->docker= +BLIND traefik_inode=18453 host=41134 ++42s controller->docker= +BLIND traefik_inode=18453 host=41134 ++47s controller->docker= +BLIND traefik_inode=18453 host=41134 ++52s controller->docker= +BLIND traefik_inode=18453 host=41134 ++57s controller->docker= +BLIND traefik_inode=18453 host=41134 ++62s controller->docker= +BLIND traefik_inode=18453 host=41134 ++68s controller->docker= +BLIND traefik_inode=18453 host=41134 ++73s controller->docker= +BLIND traefik_inode=18453 host=41134 ++78s controller->docker= +BLIND traefik_inode=18453 host=41134 ++83s controller->docker= +BLIND traefik_inode=18453 host=41134 ++88s controller->docker=29.8.2 traefik_inode=18453 host=41134 ++94s controller->docker=29.8.2 traefik_inode=18453 host=41134 ++99s controller->docker=29.8.2 traefik_inode=18453 host=41134 ++104s controller->docker=29.8.2 traefik_inode=18453 host=41134 ++109s controller->docker=29.8.2 traefik_inode=18453 host=41134 ++114s controller->docker=29.8.2 traefik_inode=18453 host=41134 ++120s controller->docker=29.8.2 traefik_inode=41134 host=41134 +HEALED +== ids before vs after (empty = none changed) +containers after: 21 +2026/10/04 18:23:43 sockheal.go:118: [WARN] [sockheal] Docker refuses the socket (dial unix /var/run/docker.sock: connect: connection refused) — exiting after 1m0s of refusals so Docker restarts this controller on the current socket (R-860) +2026/10/04 18:24:58 sockheal.go:123: [ERROR] [sockheal] Docker has refused the socket for 1m15s (dial unix /var/run/docker.sock: connect: connection refused) — the socket file was re-created and this container holds the old one; EXITING (code 75) so Docker's restart policy brings it back on the current socket (R-860) +2026/10/04 18:25:29 sockheal.go:155: [WARN] [sockheal] traefik holds an old docker socket (inode 18453, current 41134) — restarting it (R-860) +2026/10/04 18:25:30 sockheal.go:160: [INFO] [sockheal] traefik restarted onto the current docker socket diff --git a/documentation/audits/r840-config-bundle-2026-10-04/partG/g5-wgsync-after.txt b/documentation/audits/r840-config-bundle-2026-10-04/partG/g5-wgsync-after.txt new file mode 100644 index 00000000..877f76d0 --- /dev/null +++ b/documentation/audits/r840-config-bundle-2026-10-04/partG/g5-wgsync-after.txt @@ -0,0 +1,4 @@ +2026/10/04 20:25:03 [INFO] wgsync: pushed 4 peers to 167.233.158.164:22 +2026/10/04 20:30:03 [INFO] wgsync: pushed 4 peers to 167.233.158.164:22 +2026/10/04 20:35:03 [INFO] wgsync: pushed 4 peers to 167.233.158.164:22 +wg_peers BEFORE the delete (hub DB copy 16:55 UTC): 5 ['10.77.0.2', '10.77.0.250', '10.77.0.3', '10.77.0.4', '10.77.0.5'] diff --git a/documentation/audits/r840-config-bundle-2026-10-04/partH/h1-vouch.txt b/documentation/audits/r840-config-bundle-2026-10-04/partH/h1-vouch.txt new file mode 100644 index 00000000..dc236531 --- /dev/null +++ b/documentation/audits/r840-config-bundle-2026-10-04/partH/h1-vouch.txt @@ -0,0 +1,2 @@ +Location: /configuration?flash=artifacts_set +2026/10/04 20:21:14 [INFO] Artifact manifest set: agent=0.143.0 golden=0.293.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="8d7273cf5313ef62b867cb6f831c631923a436452d6f90b8ff7f0771170396ba" diff --git a/documentation/audits/r840-config-bundle-2026-10-04/partH/h2-install-manifest.json b/documentation/audits/r840-config-bundle-2026-10-04/partH/h2-install-manifest.json new file mode 100644 index 00000000..cc389b24 --- /dev/null +++ b/documentation/audits/r840-config-bundle-2026-10-04/partH/h2-install-manifest.json @@ -0,0 +1,14 @@ +{ + "agent": { + "version": "0.143.0", + "sha256": "41c0d3060013dfda795262147454248149bee0888c0935170fdb85de6e7a35da" + }, + "golden": { + "version": "0.293.0", + "sha256": "e7966872abeb38db320e7b26bd9baf6527e87270ed2faf3f0e53763bfc4764b7" + }, + "bundle": { + "version": "0.143.0", + "sha256": "8d7273cf5313ef62b867cb6f831c631923a436452d6f90b8ff7f0771170396ba" + } +} diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index f910cc6f..e5ba16d7 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,14 @@ --- +## 2026-10-04 (night) — the config bundle, test approvals, the Docker-socket self-heal (agent v0.143.0, hub v0.133.0, controller v0.293.0, installer 1.31.0; rulings 96–99) + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-840** | **A new root wrapper or sudoers line could not reach an installed box — no product route.** Built (decision 96): the config bundle — every root-owned file the installer writes, one reproducible file beside the binary; a signed `agent_config_update` the box's `felhom-os-apply` verifies itself (signature against the root-owned signers file, sha, the 22-path table, every content check before the first write, self-check, undo); the trust root is never a bundle path; the installer installs the same bundle. Live: both demo boxes (bootstrap, then 0 written of 22, probe 71/71), a deliberate change and its undo on demo-hp, a wrong sha and a replay refused, the installer path on demo-felhom. **Reasoning kept: the route cannot reach a root side that predates it — the first bundle on such a box is one by-hand step, and that must be said, not hidden.** Tester 2's step → R-862. | CLOSED 2026-10-04 — BUILT | `audits/r840-config-bundle-2026-10-04/partB/` | +| **R-859** | **An OS approval made under a TEST wait stayed in force after the test, and a real ring-1 box installed it** (Tester 2, 2026-10-04: 49 guest + 106 host packages). Hub v0.133.0: every approval under an override is marked; a start without the override cancels every unsuperseded test approval (no ring-1 plan, ring-1 boxes bumped, one operator event each); amber on the System page; one-time backfill of the automatic early approvals. Live: 4 cancelled at the 18:20 UTC start; the operator's Docker approval stays (CC decision, operator may reverse). | CLOSED 2026-10-04 — FIXED | `audits/r840-config-bundle-2026-10-04/partD/`; `runbooks/os-updates-test-waits.md` | +| **R-860** | **A re-created Docker socket file left the controller and traefik blind until a person acted** (generalises R-858). Measured: only a `docker.socket` restart re-creates the file; a dockerd crash or `systemctl restart docker` does not. Controller v0.293.0 (`internal/sockheal`): 60 s of refusals → the controller exits and Docker restarts it on the current socket; then it restarts any other socket user on an older inode (traefik). Live: 9202 healed in 104 s, demo-hp 9201 in 120 s, every container id unchanged. **Reasoning kept: count only refusals, never timeouts, and only after Docker answered once — or a slow daemon restarts the controller, and a broken socket loops it.** | CLOSED 2026-10-04 — FIXED | `audits/r840-config-bundle-2026-10-04/partE/` | + ## 2026-10-04 (~18:30) — R-858 incident (agent v0.142.1, ruling 95) | Row | What | Closed | Evidence | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 37b2257e..bbd7189b 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -267,13 +267,14 @@ stopping line that lies. | **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks//deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC | | **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | -## Security & access — 30 rows (P2 3, P3 24, P4 3) +## Security & access — 31 rows (P2 4, P3 24, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-133** | Security & access | P2 | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | | **R-135** | Security & access | P2 | **`validateCSRF` returns TRUE when there is no session cookie** (`hub/internal/web/server.go:678-683`) — measured live: `POST` with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 | READY (S) — **security** | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and `SameSite` does not govern the `Authorization` header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 | CC | | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | +| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **READY — design + operator go; owner: CC** | — | the operator ranks it against the first paying customer | CC | | **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | @@ -302,13 +303,14 @@ stopping line that lies. | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | -## Box system & updates — 20 rows (P2 4, P3 13, P4 3) +## Box system & updates — 20 rows (P2 3, P3 14, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator | | **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC | | **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator | +| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | @@ -326,7 +328,6 @@ stopping line that lies. | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | | **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | -| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC.** 2026-10-04 (evening): the demo boxes got the v0.142.0 wrapper, the crash guard and the two root-owned trust files BY HAND again (`audits/os-docker-crash-2026-10-04/partD/`). **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC | ## Monitoring & notifications — 25 rows (P2 2, P3 16, P4 7) diff --git a/documentation/runbooks/config-bundle.md b/documentation/runbooks/config-bundle.md new file mode 100644 index 00000000..3b65d3bf --- /dev/null +++ b/documentation/runbooks/config-bundle.md @@ -0,0 +1,69 @@ +# RUNBOOK — the config bundle: a box's root-owned files by a signed job (R-840) + +Design: `architecture/11-os-updates.md` §5.4.2. Decision 96 (`09` §3). Built 2026-10-04: agent v0.143.0, hub v0.133.0, +installer 1.31.0. Evidence: `audits/r840-config-bundle-2026-10-04/`. + +## What it is + +Every root-owned file the installer writes for the agent (both sudoers files, the five wrappers, the crash guard and +its units, the agent and rollback units, the start-limit drop-in, the mgmt watchdog, the OOB belt's files) travels as +ONE file, `felhom-config-bundle.json`, published beside the agent binary (`felhom-agent//`). A new box gets it +from the installer; an installed box gets it by a signed `agent_config_update`. The box's own root-owned +`felhom-os-apply` checks the signature, the sha, every path and every file before it writes anything, keeps the +previous copies, self-checks, and puts everything back on a failure. + +**The trust root is never in a bundle.** `/etc/felhom/operator-signers` and `/etc/felhom/os-trust.json` decide who may +sign; no bundle may add, remove or change them. A box that has no signers file gets exactly the installer's pinned key +(`felhom-op-1`), and only through a job that key signed. Signer rotation is a separate act (`04` §3). + +## Send a bundle to a box (CC or the operator, on DooPlex) + +1. The bundle's sha: the release output (`release-agent.sh` prints `bundle :`), or the hub's Configuration page after + vouching (the hub resolves it from the registry by exact name). +2. Sign and queue (key `felhom-op-1`; the hub key is Secret `felhom-system/report-api`, read into a file, never printed): + ```bash + felhom-opsign -op agent_config_update -host -key-id felhom-op-1 -key \ + -agent-version -bundle-sha256 -ttl 45m -upload http://:8080 -hub-key "$(cat )" + ``` +3. The agent takes it on its next job poll (measured 4–15 min). Positive observables in the box's journal: + `os-apply: BUNDLE DONE agent= written= … self-check=ok`, then `capability probe after the config bundle ok=71 + total=71`. The hub's System page "Root files" column shows the version. +4. **Undo** = send the previous release's bundle the same way. The previous copies also stay on the box in + `/var/lib/felhom-os-apply/bundle-prev/