R-858 filed (P2): a Docker engine step leaves the controller and traefik on the old docker socket; demo-felhom repaired by restarting the two containers; incident evidence
gates / gates (push) Successful in 28s
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,20 @@
|
||||
== 2026-10-04T15:56:54Z demo-felhom guest 9201
|
||||
1506210 srw-rw---- 1 root docker 0 Oct 4 14:13 /var/run/docker.sock
|
||||
143 srw-rw---- 0 root 991 0 Aug 10 07:25 /var/run/docker.sock
|
||||
cloudflared mounts-socket=no
|
||||
felhom-controller mounts-socket=socket
|
||||
filebrowser mounts-socket=no
|
||||
traefik mounts-socket=socket
|
||||
opengist mounts-socket=no
|
||||
2026-10-04T14:13:33+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:33.518477721Z" level=info msg="Restoring containers: start."
|
||||
2026-10-04T14:13:33+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:33.566951360Z" level=info msg="Deleting nftables IPv4 rules" error="running nft: /dev/stdin:1:17-30: Error: Could not process rule: No such file or directory\ndelete table ip docker-bridges\n ^^^^^^^^^^^^^^\n exit status 1"
|
||||
2026-10-04T14:13:33+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:33.575940033Z" level=info msg="Deleting nftables IPv6 rules" error="running nft: /dev/stdin:1:18-31: Error: Could not process rule: No such file or directory\ndelete table ip6 docker-bridges\n ^^^^^^^^^^^^^^\n exit status 1"
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.060369627Z" level=info msg="there are running containers, updated network configuration will not take affect"
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.060647190Z" level=info msg="Loading containers: done."
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.074298789Z" level=info msg="Docker daemon" commit=8af9fe3 containerd-snapshotter=false storage-driver=overlay2 version=29.8.2
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.074346504Z" level=info msg="Initializing buildkit"
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.074487454Z" level=warning msg="failed check for fsverity support" error="enable fsverity failed: operation not supported" path=/var/lib/docker/buildkit/content
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.219975776Z" level=info msg="Completed buildkit initialization"
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.223588774Z" level=info msg="Daemon has completed initialization"
|
||||
2026-10-04T14:13:34+00:00 demo-felhom dockerd[3505106]: time="2026-10-04T14:13:34.223626300Z" level=info msg="API listen on /run/docker.sock"
|
||||
2026-10-04T14:13:34+00:00 demo-felhom systemd[1]: Started docker.service - Docker Application Container Engine.
|
||||
@@ -0,0 +1,40 @@
|
||||
2026/10/04 11:57:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 11:57:03 [INFO] [scheduler] Job agent-channel-health completed (took 837ms)
|
||||
2026/10/04 11:58:02 [INFO] [scheduler] Running job: system-health
|
||||
2026/10/04 11:58:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 11:58:02 [INFO] [monitor] Health check: status=ok
|
||||
2026/10/04 11:58:02 [INFO] [scheduler] Job system-health completed (took 194ms)
|
||||
2026/10/04 11:58:03 [INFO] [scheduler] Job agent-channel-health completed (took 958ms)
|
||||
2026/10/04 11:58:42 [INFO] Health probes: 1 ok (of 1 probed)
|
||||
2026/10/04 11:59:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 11:59:03 [INFO] [scheduler] Job agent-channel-health completed (took 838ms)
|
||||
2026/10/04 12:00:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:00:03 [INFO] [scheduler] Job agent-channel-health completed (took 833ms)
|
||||
2026/10/04 12:01:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:01:03 [INFO] [scheduler] Job agent-channel-health completed (took 825ms)
|
||||
2026/10/04 12:02:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:02:03 [INFO] [scheduler] Job agent-channel-health completed (took 828ms)
|
||||
2026/10/04 12:03:02 [INFO] [scheduler] Running job: system-health
|
||||
2026/10/04 12:03:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:03:02 [INFO] [monitor] Health check: status=ok
|
||||
2026/10/04 12:03:02 [INFO] [scheduler] Job system-health completed (took 239ms)
|
||||
2026/10/04 12:03:02 [INFO] [monitor] Health check: status=ok
|
||||
2026/10/04 12:03:03 [INFO] [scheduler] Job agent-channel-health completed (took 1.085s)
|
||||
2026/10/04 12:03:42 [INFO] Health probes: 1 ok (of 1 probed)
|
||||
2026/10/04 12:04:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:04:03 [INFO] [scheduler] Job agent-channel-health completed (took 831ms)
|
||||
2026/10/04 12:05:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:05:03 [INFO] [scheduler] Job agent-channel-health completed (took 838ms)
|
||||
2026/10/04 12:06:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:06:03 [INFO] [scheduler] Job agent-channel-health completed (took 860ms)
|
||||
2026/10/04 12:07:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:07:03 [INFO] [scheduler] Job agent-channel-health completed (took 846ms)
|
||||
2026/10/04 12:08:02 [INFO] [scheduler] Running job: system-health
|
||||
2026/10/04 12:08:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:08:02 [INFO] [monitor] Health check: status=ok
|
||||
2026/10/04 12:08:02 [INFO] [scheduler] Job system-health completed (took 188ms)
|
||||
2026/10/04 12:08:03 [INFO] [scheduler] Job agent-channel-health completed (took 968ms)
|
||||
2026/10/04 12:08:52 [INFO] Health probes: 1 ok (of 1 probed)
|
||||
2026/10/04 12:09:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
2026/10/04 12:09:03 [INFO] [scheduler] Job agent-channel-health completed (took 836ms)
|
||||
2026/10/04 12:10:02 [INFO] [scheduler] Running job: agent-channel-health
|
||||
@@ -0,0 +1,14 @@
|
||||
repair 2026-10-04T15:57:13Z: docker restart traefik felhom-controller (containers only; the engine untouched)
|
||||
traefik
|
||||
felhom-controller
|
||||
controller health=healthy
|
||||
1506210 srw-rw---- 1 root 991 0 Oct 4 14:13 /var/run/docker.sock
|
||||
1506210 srw-rw---- 1 root 991 0 Oct 4 14:13 /var/run/docker.sock
|
||||
cloudflared | Up 5 hours (healthy)
|
||||
felhom-controller | Up 9 seconds (healthy)
|
||||
filebrowser | Up 7 hours (healthy)
|
||||
traefik | Up 9 seconds
|
||||
opengist | Up 10 hours (healthy)
|
||||
ActiveEnterTimestamp=Sun 2026-10-04 14:13:34 UTC
|
||||
2026/10/04 15:57:16 [INFO] [monitor] Health check: status=ok
|
||||
2026/10/04 15:57:21 [INFO] [monitor] Health check: status=ok
|
||||
@@ -302,7 +302,7 @@ stopping line that lies.
|
||||
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
|
||||
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
|
||||
|
||||
## Box system & updates — 20 rows (P2 4, P3 13, P4 3)
|
||||
## Box system & updates — 21 rows (P2 5, P3 13, P4 3)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -325,6 +325,7 @@ stopping line that lies.
|
||||
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
|
||||
| **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-858** | Box system & updates | P2 | **A Docker engine step leaves the controller and traefik BLIND: they keep the OLD Docker socket.** FOUND 2026-10-04 by the operator (demo-felhom DOWN on the hub since 14:18 UTC): the package update restarted dockerd at 14:13, which recreates `/run/docker.sock`; `live-restore` (decision 87) kept every container running — including the two that bind-mount the socket FILE (`felhom-controller`, `traefik`), so they kept the deleted inode (controller saw inode 143 from 10 Aug, the guest inode 1506210). The controller then reported "docker not reachable" and every protected container "not running" (health FAIL); traefik stops seeing new routes. demo-hp had the same fault and only healed because the crash test rebooted it. **The Docker health rule (`DockerHealthVerdict`) PASSED** — the controller's own health check stayed "healthy" and the ids were the same — so the step reported healthy. Repaired on demo-felhom by restarting the two containers (15:57 UTC, engine untouched). No Docker step is pending on either box now. Fix direction in STATUS. `audits/os-docker-crash-2026-10-04/partE-incident/` | **READY — operator decision (fix choice); owner: CC** | — | — | CC |
|
||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC.** 2026-10-04 (evening): the demo boxes got the v0.142.0 wrapper, the crash guard and the two root-owned trust files BY HAND again (`audits/os-docker-crash-2026-10-04/partD/`). **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user