OS updates guest fast lane: records — 11 §8.1 BUILT, 03 cloudflared corrected, 07 §6.1 OS leg, 00 PARTIAL, monthly runbook infra pins, golden 0.291.0 record + vouch, register 331 -> 333 (R-837/838/726/843 closed; R-840/841/842/844/845 opened), STATUS, report
gates / gates (push) Successful in 31s
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03). Design: `architecture/11-os-updates.md` (NOT RATIFIED, 2026-10-04). **Spike done 2026-10-04** (`audits/os-updates-spike-2026-10-04/`): measured, nothing built — still MISSING |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.140.0, hub v0.130.0 | **PARTIAL — the GUEST's Debian fast lane is PROVEN-LIVE (2026-10-04); the host, Docker and the kernel are MISSING** | `audits/os-guest-lane-2026-10-04/` — ring 0 (both demo boxes) installed 53 Debian fixes each, healthy; the hub approved a 272-package release (TEST wait 2 min, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the 3 approved versions; a stopped app → `health_failed` + operator mail. Design `architecture/11-os-updates.md` §8.1 | **No automatic undo** (a customer guest cannot be snapshotted, R-837 → R-842); existing boxes need the wrapper + sudoers by hand (R-840); host / Docker / kernel lanes not built (R-812, R-835, R-836); `felhom-host-install.sh:2133-2136` still runs no host upgrades |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
|
||||
|
||||
@@ -47,7 +47,7 @@ Owns:
|
||||
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
|
||||
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
|
||||
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
|
||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub.
|
||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841.
|
||||
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
|
||||
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
|
||||
7. **Local API** — the per-guest authorization gate the controller calls.
|
||||
@@ -64,7 +64,7 @@ Explicitly does **not**:
|
||||
|
||||
- **Native Go binary, systemd service** on the host: boot-start, `Restart=always`, systemd watchdog (kill+restart on hang), journald logging, resource limits.
|
||||
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
|
||||
- **`cloudflared` is a separate systemd service**, not embedded in the agent. This is what makes the data path survive control-plane death by construction. The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.
|
||||
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent neither manages it nor (see R-841) sees its health. Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
|
||||
|
||||
## 4. Control model — reconcile + signed destructive ops
|
||||
|
||||
@@ -143,7 +143,7 @@ notification) is the control.
|
||||
|
||||
**Box-initiated poll.** The hub never connects inbound. Each poll cycle exchanges:
|
||||
|
||||
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health, agent + controller versions, audit-log tail.
|
||||
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (a host-unit probe that always reads `inactive` — R-841), agent + controller versions, audit-log tail.
|
||||
- **Down:** the current desired state, any pending signed one-shot jobs, and config (poll interval, update window, policy changes).
|
||||
|
||||
**Dead-man's-switch (essential, not optional).** In a box-initiated model the heartbeat
|
||||
@@ -633,6 +633,13 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea
|
||||
never — the binary is what flips. Report field `selfupdate_pending` surfaces a runs-but-never-commits
|
||||
binary. v1 scope-outs: no hub-floor auto-update, no failed-update auto-retry (the operator re-signs),
|
||||
no pending-timeout auto-rollback.
|
||||
- **OS updates, guest fast lane (agent v0.140.0, `11-os-updates.md` §8.1).** A second root-owned wrapper,
|
||||
`/usr/local/sbin/felhom-os-apply`, behind ONE sudoers entry (`FELHOM_OSAPPLY`: `felhom-os-apply --plan
|
||||
/var/lib/felhom-agent/os/plan-*.json`). The agent has no `apt` grant of its own for this; every rule (no removal,
|
||||
no downgrade, no new or unlisted package, Debian origin only, the box's own customer guest only) is in the wrapper,
|
||||
red-proved per rule. The leg runs after a successful primary whole-guest backup, under the heavy-op gate.
|
||||
**[FACT] The signed agent update does NOT carry the wrapper or the sudoers line** — only the installer installs
|
||||
them (R-840).
|
||||
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
|
||||
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
|
||||
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
|
||||
|
||||
@@ -362,6 +362,12 @@ never moved by an update or an undo, and the unit restore accepts it.
|
||||
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
|
||||
|
||||
|
||||
**[FACT, 2026-10-04 — agent v0.140.0, `11-os-updates.md` §8.1] The OS leg closes the night.** After the
|
||||
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
|
||||
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a
|
||||
backup or a restore-test; at most once per 20 h. The backup minutes old is the guest's undo (no snapshot is possible,
|
||||
R-837). A failed or missed backup → no OS leg that night.
|
||||
|
||||
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
|
||||
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
|
||||
(`cmd/controller/main.go:604-607`). Both boxes run W = `02:30`.
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
> | | |
|
||||
> |---|---|
|
||||
> | **Status** | **NOT RATIFIED — a PROPOSAL with one operator ruling, corrected by the 2026-10-04 spike (§7.1, corrections C1–C12 below).** Ratification is Viktor's review, not an editor's. |
|
||||
> | **Status** | **NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1).** Ratification is Viktor's review, not an editor's. |
|
||||
> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. |
|
||||
> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 |
|
||||
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
|
||||
@@ -403,7 +403,8 @@ not covered: 79 Proxmox + 1 Tailscale on the host (slow lane / not ours), 6 Dock
|
||||
Each step returns to the operator for go or no-go.
|
||||
|
||||
1. **Spike** (measure Q1–Q10; no product code).
|
||||
2. **Guest Debian, fast lane.** Lowest risk: a snapshot undo exists.
|
||||
2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub
|
||||
v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured).
|
||||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
|
||||
4. **Fleet view and alarms** (§5.7).
|
||||
5. **Slow lane: Docker engine.**
|
||||
@@ -412,6 +413,47 @@ Each step returns to the operator for go or no-go.
|
||||
|
||||
---
|
||||
|
||||
### 8.1 Step 2 as BUILT (2026-10-04) `[FACT]`
|
||||
|
||||
Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast lane, decisions 78–80 (`09` §3).
|
||||
|
||||
- **The undo (R-837, measured first): none automatic.** PVE refuses ANY snapshot of a customer guest —
|
||||
`PVE/AbstractConfig.pm:755-757` skips non-snapshot mounts only for a snapshot named `vzdump`, and every customer
|
||||
guest carries the host-path binds mp8/mp9. As the agent's token and as root: `snapshot feature is not available`.
|
||||
The token's role HAS `VM.Snapshot` / `VM.Snapshot.Rollback`. So §5.6's first row does not exist: a failed health
|
||||
check stops, reports `health_failed`, and the hub mails the operator; the whole-guest backup taken minutes earlier
|
||||
is the undo, by hand. The choice of a real undo is in STATUS (R-842).
|
||||
- **The wrapper** `felhom-os-apply` (agent repo `configs/`), Python 3 stdlib, per §5.4.1 with refusals R1–R13 (R12:
|
||||
the host layer; R13: dpkg still broken after the repair). **Changed from the draft** *(decided by CC unattended —
|
||||
operator may reverse)*: one sudoers entry (`--plan <file>`); the plan's `mode` field (`inventory` / `apply` /
|
||||
`health`) replaces a separate `--repair-only` (the repair runs first on every apply); Python, because a JSON plan
|
||||
cannot be parsed safely in sh. The box's own customer guest is "the guest that binds `/mnt/felhom-drives`" (R10).
|
||||
- **The leg** (agent `internal/osupdate`): after a SUCCESSFUL primary whole-guest backup, inside the backup's
|
||||
goroutine before the host-wide heavy-op gate is released (so never beside a backup or a restore-test, C10), 90 s
|
||||
after the backup, at most once per 20 h. *Decided by CC unattended — operator may reverse:* the 90 s settle and the
|
||||
20 h gap; the controller's own self-update (04:30) is not detected — the 5-minute health wait absorbs a restart.
|
||||
- **The health rule** (`HealthVerdict`, pinned): docker answers, the guest resolves `deb.debian.org`, the controller's
|
||||
health check is `healthy`, and every container running at the START of the leg runs again (healthy if it was). The
|
||||
baseline merges the inventory's reading with the apply's own — found live: an app stopped between them escaped the
|
||||
first rule. Wait 5 min, poll 15 s.
|
||||
- **Rings and the switch** (hub, per box): ring 0 = demo-hp + demo-felhom, everything else ring 1; switch ON by
|
||||
default; OFF → the box reports, installs nothing. A box with no `os_update` block (older hub) = ring 1, ON, no
|
||||
release.
|
||||
- **The approval rule (the ruled "1–2 day wait")**: every Debian / Debian-Security package=version that ALL ring-0
|
||||
boxes having it agree on; approved when, since that set was first seen, **24 h** passed with every ring-0 report
|
||||
healthy and every ring-0 box completed **1** post-backup night run (`OS_APPROVE_AFTER`, `OS_APPROVE_NIGHTS`; an
|
||||
override is logged as a TEST configuration). The approval time is the snapshot.debian.org timestamp (decision 79).
|
||||
"Approve now" is an operator event. *The 24 h / 1 night numbers: decided by CC unattended within the ruled 1–2 days.*
|
||||
- **The household's line**: hub event `os_update_applied` (info: on the household's hub timeline, never mailed;
|
||||
hu/en in the bundle). There is no surface on the box itself (R-844).
|
||||
- **Measured live**: ring 0 — 53 Debian packages on each demo box (18.7–31.7 s inside the wrapper; 174–226 s for the
|
||||
whole leg incl. the inventory), healthy, 6 Docker updates "not covered"; approval — with a 2-minute TEST wait, a
|
||||
272-package release approved automatically, then the ruled values restored; ring 1 — demo-felhom installed exactly
|
||||
the 3 approved versions it lacked (269 already current) and left the newer Docker packages alone; a deliberately
|
||||
stopped app → `health_failed` after 5 min, the operator mailed, the household line recorded.
|
||||
- **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the
|
||||
installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.
|
||||
|
||||
## 9. Where the rest lives
|
||||
|
||||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||||
|
||||
Reference in New Issue
Block a user