docs: rulings 2026-10-04 ~12:20 (decisions 81-83); R-842 closed (A); R-840 ruling recorded
gates / gates (push) Successful in 28s
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -16,6 +16,10 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **Rulings 2026-10-04 (~12:20) — recorded before the work (host fast lane brief).** `09` §3 decisions **81** (R-842 A:
|
||||
> the undo is the whole-guest backup by hand; R-842 closed), **82** (R-840 not built now — "There are no older boxes";
|
||||
> row kept open with the reviewer's note) and **83** (next: R-841, the host fast lane, OS-update improvements).
|
||||
|
||||
> **2026-10-04 (afternoon) — OS updates, guest fast lane BUILT (`11` §8.1).** Agent v0.140.0: `configs/felhom-os-apply`
|
||||
> (R1–R13, repair first, snapshot.debian.org fallback; FELHOM_OSAPPLY sudoers), `internal/osupdate` (the leg after a
|
||||
> successful primary backup, under the heavy-op gate; health baseline = start of the leg), `--selftest=os-update`.
|
||||
|
||||
@@ -765,6 +765,19 @@ its length, and both fixes cost something the household would notice — operato
|
||||
80. **Next build: the guest fast lane** (`11` §8 step 2), with R-837 (the snapshot undo) measured first and R-838 (the
|
||||
infrastructure images) in the same session. *Operator ruling 2026-10-04 ~10:17.*
|
||||
|
||||
### 2026-10-04 (~12:20) — three operator rulings (recorded before the work)
|
||||
|
||||
81. **The undo for a failed OS update is the whole-guest backup, restored by hand (R-842, option A).** Nothing new is
|
||||
built. **Rejected:** a home-made LVM-thin snapshot behind Proxmox's back (unmeasured; thin-pool risk).
|
||||
*Operator ruling 2026-10-04 ~12:20.*
|
||||
82. **R-840 (a route for wrappers/sudoers to installed boxes) is not built now** — *"There are no older boxes."* No
|
||||
installed box exists outside the two demo boxes, and they take wrapper changes by hand. The row stays open, not
|
||||
re-ranked. *Reviewer's note, recorded as written:* the ruling is right for today, but every box installed from now on
|
||||
becomes an "older box" the first time the wrapper or the sudoers line changes again; R-840 must be solved before the
|
||||
first box that cannot be reached by hand. *Operator ruling 2026-10-04 ~12:20.*
|
||||
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
|
||||
*Operator ruling 2026-10-04 ~12:20.*
|
||||
|
||||
### 2026-09-30 (day) — operator notes, recorded before the work
|
||||
|
||||
- **The day brief runs by day.** Every backup and automatic-update test is started by hand — the night chain's debug
|
||||
|
||||
@@ -26,6 +26,12 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-04 (~12:20) — operator ruling
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-842** | **The guest OS update had no automatic undo (no snapshot of a guest with host binds is possible).** Ruled option A (`09` §3 decision 81): the whole-guest backup, minutes old, restored by hand, is the undo. Nothing new built; a home-made LVM-thin snapshot rejected (unmeasured, thin-pool risk). | CLOSED 2026-10-04 — RULED (decision 81) | `audits/os-guest-lane-2026-10-04/partA/README.md` |
|
||||
|
||||
## 2026-10-04 (day) — OS updates, guest fast lane (agent v0.140.0, hub v0.130.0, controller v0.291.0)
|
||||
|
||||
> Evidence: `audits/os-guest-lane-2026-10-04/`.
|
||||
|
||||
@@ -302,7 +302,7 @@ stopping line that lies.
|
||||
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
|
||||
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
|
||||
|
||||
## Box system & updates — 22 rows (P2 4, P3 15, P4 3)
|
||||
## Box system & updates — 21 rows (P2 4, P3 14, P4 3)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -326,9 +326,8 @@ stopping line that lies.
|
||||
| **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC |
|
||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **READY — measure before the kernel slow lane; owner: CC + operator (reboots)** | — | — | CC |
|
||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** | — | — | CC |
|
||||
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
|
||||
| **R-841** | Box system & updates | P3 | **The agent's `cloudflared` health probe reads a host systemd unit that does not exist — every box reports its tunnel `inactive`.** FOUND 2026-10-04: `felhom-agent/internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST; cloudflared is a container in the GUEST (`11` C8), so demo-hp answers `inactive` / `Unit cloudflared.service could not be found`, and the hub stores that in `cloudflared_status` for every box. The field is equally consistent with "tunnel down" and "never checked" (R-96 rule 3). Fix direction: read the guest's container state (the agent already may `pct exec * -- docker inspect -f *`), or drop the field; and say so in `03` (corrected 2026-10-04). | **READY — owner: CC** | — | — | CC |
|
||||
| **R-842** | Box system & updates | P3 | **The guest OS update has no automatic undo: a customer guest cannot be snapshotted.** MEASURED 2026-10-04 (R-837, closed): PVE refuses any snapshot not named `vzdump` when a guest has host-path binds (mp8, mp9) — as the agent's token (which holds `VM.Snapshot` / `VM.Snapshot.Rollback`) and as root. Today a failed health check stops, reports `health_failed` and mails the operator; the whole-guest backup taken minutes earlier is the undo, by hand. **Options (STATUS):** A — keep it so (no new mechanism; the backup is minutes old; a restore costs ~1–3 min down plus app data written since); B — a root-wrapper LVM-thin snapshot of rootfs + mp0 behind PVE's back, rolled back with the guest stopped (a new mechanism nobody has measured; thin-pool risk, R-672's shape). `audits/os-guest-lane-2026-10-04/partA/README.md` | **WAITING-ON-OPERATOR 2026-10-04 — CC's pick: A** | — | — | operator |
|
||||
|
||||
## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user