R-858 closed (agent v0.142.1, ruling 95): the Docker step restarts the socket users; incident evidence; STATUS; report addendum
gates / gates (push) Successful in 34s
gates / gates (push) Successful in 34s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -16,6 +16,12 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-04 (~18:00) — R-858 incident + ruling 95.** The operator saw demo-felhom DOWN: the 14:13 UTC Docker step
|
||||
> left `felhom-controller` and `traefik` on the OLD socket (live-restore kept them running; their bind-mounted socket
|
||||
> file was recreated). Repaired by restarting the two (15:57 UTC). `09` §3 decision **95**: the wrapper restarts the
|
||||
> socket users after a step + the health rule checks the controller reaches Docker (agent v0.142.1). Demo ring-0 Docker
|
||||
> marks OFF until released.
|
||||
|
||||
> **2026-10-04 (evening) — the System page, the Docker engine slow lane and the crash guard BUILT.** Agent v0.142.0,
|
||||
> hub v0.132.0, installer 1.30.0, `build-golden.sh` 3.1.0. **Decided by CC unattended — operator may reverse:** `09` §3
|
||||
> decisions **90** (crash signal = a clean-shutdown marker; pstore saved nothing), **91** (`panic_on_oops` stays 0; an
|
||||
|
||||
@@ -84,3 +84,16 @@ Checked by `head_sha` over every page of the Gitea `jobs` endpoint: felhom.eu
|
||||
to `21986d0` (re-vouch records) **success** (incl. hub `175ecfc` 1265→, installer, manifests, docs); felhom-agent —
|
||||
`b1746c2` (1273, 1274), `f24dce5` (1275), `42af3ab` (1281) **success**. This report's own commit: checked after the push
|
||||
(see the session's final message).
|
||||
|
||||
## Addendum (~18:30) — R-858, found by the operator
|
||||
|
||||
The N100 showed DOWN from 14:18 UTC. Cause: v0.142.0's Docker step restarted dockerd (14:13); live-restore kept the
|
||||
containers running, but `felhom-controller` and `traefik` bind-mount the socket FILE and kept the deleted inode, so the
|
||||
controller could not reach Docker. My Docker health rule passed it (the controller's own check said healthy) — the rule
|
||||
checked the mechanism, not the consequence. Repaired by restarting the two containers (15:57 UTC). Ruling 95: agent
|
||||
**v0.142.1** (`4950030`, sha256 `003f882a…62bd`) restarts only the socket users after a step and fails health when the
|
||||
controller cannot reach Docker. Proven live on demo-hp before the release (signed undo, then forward): `applied, healthy`,
|
||||
guest / controller / traefik on the same socket inode both times. Ring-0 marks were OFF during the fix, back ON after.
|
||||
Both boxes on 0.142.1; vouched for new installs (golden unchanged). Red-proofs 4/4. Register: R-858 opened and closed
|
||||
(still **333** open). Evidence `partE-incident/`. Not touched: Tester 1 shows DOWN for 4 days on the dashboard — a
|
||||
fenced tester box, outside this brief.
|
||||
|
||||
@@ -3,8 +3,17 @@
|
||||
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
|
||||
|
||||
**Updated 2026-10-04 (late evening): the hub has a System page with every box's versions and the update buttons; Docker
|
||||
updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.0 and
|
||||
controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.0.**
|
||||
updates are built; a crashed box restarts by itself, at most twice an hour. Both demo boxes run host agent 0.142.1 and
|
||||
controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292.0 (re-made: live-restore on, Docker 29.8.2) with agent 0.142.1.**
|
||||
|
||||
## Today (2026-10-04, ~18:30): you found a bug — the Docker update blinded the box's controller
|
||||
|
||||
- **What happened:** the Docker update on the N100 restarted Docker itself. The apps kept running (as designed), but
|
||||
the controller and the web router kept a connection to the OLD Docker, so the controller could not see anything:
|
||||
the hub showed the N100 DOWN from 14:18 to 15:57. demo-hp had the same fault; the crash test happened to heal it.
|
||||
- **Fixed (your choice):** after a Docker update the box now restarts just those two (about 10 seconds, apps untouched),
|
||||
and the update's health check now asks "can the controller really reach Docker?". Released as host agent 0.142.1,
|
||||
proven twice on demo-hp, on both demo boxes now, and approved for new installs.
|
||||
|
||||
## Today (2026-10-04, late evening): the System page, Docker updates, the crash restart
|
||||
|
||||
@@ -34,7 +43,7 @@ controller 0.292.0. Hub 0.132.0. Installer 1.30.0. New installs get golden 0.292
|
||||
an "app stopped" mail besides "restarted after a crash" — whether to calm that is a later choice for you.
|
||||
- **New-install image re-made** with live-restore on and the approved Docker version, and approved in the hub with host
|
||||
agent 0.142.0. New boxes also get the crash guard (installer 1.30.0).
|
||||
- **Rows:** 5 closed, 1 opened-and-closed the same day, 4 opened. The list went from 334 to 333.
|
||||
- **Rows:** 6 closed, 1 opened-and-closed the same day, 5 opened. The list went from 334 to 333.
|
||||
|
||||
**Needs you later (nothing breaks if you wait):**
|
||||
- Installed boxes still get new root files only by hand (the long-standing gap; there are no other boxes today).
|
||||
|
||||
@@ -778,6 +778,15 @@ its length, and both fixes cost something the household would notice — operato
|
||||
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
|
||||
*Operator ruling 2026-10-04 ~12:20.*
|
||||
|
||||
### 2026-10-04 (~18:00) — operator ruling after the R-858 incident
|
||||
|
||||
95. **A Docker engine step restarts the containers that mount the Docker socket** (R-858): after every step that
|
||||
installed something, the wrapper restarts ONLY those containers (today `felhom-controller` and `traefik`, ~10 s;
|
||||
apps untouched, the engine untouched), and the health rule checks that the controller reaches Docker from inside its
|
||||
container. **Rejected:** mounting a socket folder instead of the file (cleaner, but it changes Docker's config, the
|
||||
controller templates and the golden on every box). Until the fix is released, the demo boxes' root-owned ring-0
|
||||
Docker mark is OFF (no unsigned step). *Operator ruling 2026-10-04 ~18:00.*
|
||||
|
||||
### 2026-10-04 (evening) — decided by CC unattended — operator may reverse (System page / Docker / crash-restart brief)
|
||||
|
||||
90. **How does a box tell a crash boot from a clean one?** Options: (a) `pstore` — measured on demo-hp: `efi_pstore` is on,
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
felhom-pve: {"host_id":"demo-felhom-8363b5","ring0_slow_lane":false}
|
||||
felhom-pve: 644 root
|
||||
demo-hp: {"host_id":"demo-hp-bb76ea","ring0_slow_lane":false}
|
||||
demo-hp: 644 root
|
||||
@@ -0,0 +1,2 @@
|
||||
demo-hp wrapper (R-858 fix, pre-release) bdf60f5c79a84db7
|
||||
demo-felhom wrapper v0.142.1 bdf60f5c79a84db7
|
||||
@@ -0,0 +1,8 @@
|
||||
time=2026-10-04T18:02:51.747+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=2f5f79f8a5c22923 op=os_docker_step key_id=felhom-op-1 nonce=f0715c2a4d18668ca2de50258f4ea07c
|
||||
os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)
|
||||
time=2026-10-04T18:04:05.435+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)"
|
||||
time=2026-10-04T18:04:05.466+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=2f5f79f8a5c22923 op=os_docker_step
|
||||
controller sees engine 29.7.2
|
||||
guest socket: 12159 controller: 12159 traefik: 12159
|
||||
felhom-controller Up 17 minutes (healthy)
|
||||
traefik Up 17 minutes
|
||||
@@ -0,0 +1,5 @@
|
||||
[CAUGHT] socket users restarted after a step: DockerLane.test_step_restarts_only_the_socket_users rc=1
|
||||
[CAUGHT] only socket users, not every container: DockerLane.test_step_restarts_only_the_socket_users rc=1
|
||||
[CAUGHT] controller reach check in health: DockerLane.test_health_says_when_the_controller_cannot_reach_docker rc=1
|
||||
[CAUGHT] health rule fails a blind controller (Go)
|
||||
after restore: ok gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.730s
|
||||
@@ -0,0 +1,8 @@
|
||||
time=2026-10-04T18:04:05.435+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)"
|
||||
time=2026-10-04T18:32:54.303+02:00 level=INFO msg="osupdate: START" run=20261004T163252Z layer=docker vmid=9201 ring=0 trigger=signed enabled=true release=r858-proof-forward
|
||||
time=2026-10-04T18:34:03.314+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=r858-proof-forward layer=docker:9201 lane=slow mode=apply select=listed packages=6 authority=signed"
|
||||
time=2026-10-04T18:34:03.314+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: SOCKET-USERS restarted=felhom-controller,traefik rc=0 (R-858: they held the old docker socket)"
|
||||
controller sees engine 29.8.2
|
||||
socket guest=18453 controller=18453 traefik=18453
|
||||
2026/10/04 18:04:05 [INFO] osupdates: demo-hp-bb76ea reported docker run 20261004T160251Z: ring=0 mode=apply outcome=applied healthy=true upgraded=6 pending=6 not-covered=0 restart-needed=1 reboot-needed=false wrapper=71.6s
|
||||
2026/10/04 18:34:03 [INFO] osupdates: demo-hp-bb76ea reported docker run 20261004T163252Z: ring=0 mode=apply outcome=applied healthy=true upgraded=6 pending=0 not-covered=0 restart-needed=1 reboot-needed=false wrapper=68.9s
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
felhom-pve: {"host_id":"demo-felhom-8363b5","ring0_slow_lane":true}
|
||||
demo-hp: {"host_id":"demo-hp-bb76ea","ring0_slow_lane":true}
|
||||
@@ -26,6 +26,12 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-04 (~18:30) — R-858 incident (agent v0.142.1, ruling 95)
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-858** | **A Docker engine step left `felhom-controller` and `traefik` on the OLD docker socket** (found by the operator: demo-felhom DOWN 14:18–15:57 UTC). live-restore kept them running with the deleted, bind-mounted socket file; their own health stayed "healthy", so the step passed. Repaired by restarting the two; agent v0.142.1 restarts ONLY the socket-mounting containers after a step and fails the health rule when the controller cannot reach Docker. Proven live twice on demo-hp (signed undo and forward): both restarted, guest / controller / traefik on the same socket inode, `applied, healthy`. **Reasoning kept: a container that bind-mounts a socket FILE does not follow a restarted daemon — live-restore makes this worse, not better. Check the consequence (the controller reaches Docker), not the mechanism (the container still runs).** | CLOSED 2026-10-04 — FIXED | `audits/os-docker-crash-2026-10-04/partE-incident/` |
|
||||
|
||||
## 2026-10-04 (evening) — System page, Docker slow lane, crash guard (agent v0.142.0, hub v0.132.0, installer 1.30.0)
|
||||
|
||||
> Evidence: `audits/os-docker-crash-2026-10-04/`.
|
||||
|
||||
@@ -302,7 +302,7 @@ stopping line that lies.
|
||||
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
|
||||
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
|
||||
|
||||
## Box system & updates — 21 rows (P2 5, P3 13, P4 3)
|
||||
## Box system & updates — 20 rows (P2 4, P3 13, P4 3)
|
||||
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
@@ -325,7 +325,6 @@ stopping line that lies.
|
||||
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
|
||||
| **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC |
|
||||
| **R-858** | Box system & updates | P2 | **A Docker engine step leaves the controller and traefik BLIND: they keep the OLD Docker socket.** FOUND 2026-10-04 by the operator (demo-felhom DOWN on the hub since 14:18 UTC): the package update restarted dockerd at 14:13, which recreates `/run/docker.sock`; `live-restore` (decision 87) kept every container running — including the two that bind-mount the socket FILE (`felhom-controller`, `traefik`), so they kept the deleted inode (controller saw inode 143 from 10 Aug, the guest inode 1506210). The controller then reported "docker not reachable" and every protected container "not running" (health FAIL); traefik stops seeing new routes. demo-hp had the same fault and only healed because the crash test rebooted it. **The Docker health rule (`DockerHealthVerdict`) PASSED** — the controller's own health check stayed "healthy" and the ids were the same — so the step reported healthy. Repaired on demo-felhom by restarting the two containers (15:57 UTC, engine untouched). No Docker step is pending on either box now. Fix direction in STATUS. `audits/os-docker-crash-2026-10-04/partE-incident/` | **READY — operator decision (fix choice); owner: CC** | — | — | CC |
|
||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC.** 2026-10-04 (evening): the demo boxes got the v0.142.0 wrapper, the crash guard and the two root-owned trust files BY HAND again (`audits/os-docker-crash-2026-10-04/partD/`). **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user