docs: 11 §5.7 System page, §5.8 Docker slow lane BUILT, §5.9 crash restart; 00/03/07/08; decisions 90-94 (CC unattended); runbooks docker-undo + crash-guard; register R-852 R-835 R-848 R-849 R-851 R-854 closed, R-853 R-855 R-856 opened, R-812 R-840 narrowed (334 -> 332); live evidence
gates / gates (push) Successful in 33s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-04 17:36:10 +02:00
parent 245886ee92
commit 0c55336fba
38 changed files with 1594 additions and 16 deletions
+10
View File
@@ -16,6 +16,16 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-04 (evening) — the System page, the Docker engine slow lane and the crash guard BUILT.** Agent v0.142.0,
> hub v0.132.0, installer 1.30.0, `build-golden.sh` 3.1.0. **Decided by CC unattended — operator may reverse:** `09` §3
> decisions **90** (crash signal = a clean-shutdown marker; pstore saved nothing), **91** (`panic_on_oops` stays 0; an
> oops is reported), **92** (the 3rd unclean stop in 60 min leaves the box off; 10 s / 24 h), **93** (the wrapper
> verifies Docker authority against root-owned files), **94** (page colours = alarm thresholds). Live: live-restore on by
> reload with the same ids (9202 6/6, demo-hp 24/24, demo-felhom 5/5); Docker 29.7 → 29.8.2 on both demo boxes, ids kept;
> the operator's button approved `os-docker-20261004-142842` (TEST 0-night wait, reverted); a signed undo to 29.7.2 on
> demo-hp and back; demo-felhom as ring 1 by a signed job; a replay refused; crash 1–2 restarted in 54/53 s, the guard
> tripped, crash 3 stayed off, the hub mailed the trip, re-armed. `REPORT-os-docker-crash-2026-10-04.md`.
> **Rulings 2026-10-04 (~15:17) — recorded before the work (System page + Docker lane + crash restart brief).** `09` §3
> decisions **87** (Docker `live-restore` ON fleet-wide, never off by a plain restart), **88** (a crashed host restarts by
> itself, with a limit — R-851) and **89** (the operator sees the boxes' OS / Proxmox versions in the hub — R-852).
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.141.1, hub v0.131.1 | **PARTIAL — the GUEST and HOST Debian fast lanes are PROVEN-LIVE (2026-10-04), with the fleet view and four operator alarms; Docker and the kernel are MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/` — ring 0 on both demo hosts (demo-felhom 108 Debian host packages, healthy, 0 Proxmox-origin); a 605-package host release approved (TEST wait, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the one host version it lacked; the by-hand host undo proved (`runbooks/os-updates-host-undo.md`); `tunnel_down` / `tunnel_recovered` live. Design `architecture/11-os-updates.md` §8.1–§8.3 | **No automatic undo** (guest: last night's backup, decision 81; host: the by-hand runbook); existing boxes need the wrapper + sudoers by hand (R-840); Docker and kernel lanes not built (R-812, R-835, R-836); a host panic is not restarted (R-851) |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.142.0, hub v0.132.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes need the wrapper, the trust files and the guard by hand (R-840); the kernel lane (R-836); facts reach the hub late after a boot (R-853) |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
+7 -1
View File
@@ -47,7 +47,13 @@ Owns:
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line.
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841. **[FACT, FIXED agent v0.141.0 / controller v0.292.0, R-841 CLOSED]** The agent now reads the guest's `cloudflared` container through the existing `pct exec [0-9]* -- docker inspect -f *` sudoers line: state, exit code and the Docker health status of a check the controller adds (`cloudflared tunnel --metrics localhost:20241 ready` → cloudflared's own `/ready`, 200 only with a connection). Three states: `running` (healthy), `not_running` (stopped, absent, or running but NOT connected), `unknown` (could not ask, or the check is still starting) — `unknown` never alarms. No new sudoers line. **[FACT, agent v0.142.0, R-852]** The host report gains `system`: the Proxmox version and kernel from the
Proxmox API, plus the OS wrapper's read-only `facts` (host Debian, next-boot kernel, held packages, taint, the crash
guard; guest Debian, Docker engine, containerd, live-restore), at most every 10 min. **Still no new sudoers line** — the
facts, `live-restore-on` and the Docker layer all ride `FELHOM_OSAPPLY`. New ROOT-owned files instead (installer 1.30.0;
by hand on installed boxes, R-840): `/etc/felhom/os-trust.json`, `/etc/felhom/operator-signers` (the wrapper verifies
Docker authority against them, never against the agent-writable config — `09` decision 93), and the crash guard
(`/usr/local/sbin/felhom-crash-guard`, its units, `/etc/felhom/crash-guard.conf`).
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
7. **Local API** — the per-guest authorization gate the controller calls.
@@ -369,7 +369,10 @@ backup or a restore-test; at most once per 20 h. The backup minutes old is the g
R-837). A failed or missed backup → no OS leg that night. **[FACT, 2026-10-04 — agent v0.141.1, `11` §8.2]** On an
appliance, the HOST step follows under the same gate: after a healthy guest step only (a failed or unhealthy guest
step skips it), Debian-origin fixes only, never a kernel, boot or firmware package, never a reboot. Measured: both
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s.
steps with nothing to install, 23–32 s; a 108-package host pass, 70 s. **[FACT, 2026-10-04 — agent v0.142.0, `11` §5.8]** On a RING-0 box a third step
follows, still under the gate: the guest's Docker engine set (live-restore turned on first, once, by reload; every
container must keep its id). Ring 1 never takes it in the night leg. Measured: the three steps with nothing to install,
38–51 s; a six-package Docker step, ~50 s.
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
@@ -312,7 +312,8 @@ message, not a wider cooldown.
## 6.3 Box alarms outside the app ladder: the tunnel and OS updates [DESIGN, hub v0.131.0, 2026-10-04]
These are **operator-only** (the household can act on none of them; `operatorOnlyEvents`, pinned by
These are **operator-only** (the household can act on none of them — except `host_restarted_after_crash`, the household's
one info line beside `host_crash_restart`; `operatorOnlyEvents`, pinned by
`TestOSUpdateEvents_OperatorOnlyExceptApplied`). They go through the same dispatcher and severity contract (§6.1):
`info` is recorded and never mailed; `warning` and `error` are mailed.
@@ -323,6 +324,9 @@ These are **operator-only** (the household can act on none of them; `operatorOnl
| `os_reboot_needed` | warning | the host has needed a reboot for **14 days** (from the FIRST scanned report that said so) | a scanned pass that finds nothing (agent ≥ 0.141.1 scans the host every pass) | `TestAlarm_RebootNeeded`, `TestRebootNeeded_ClearedByAScannedPass` |
| `os_ring0_stalled` | error | ring 0 approved nothing for **7 days** in a layer while it has pending FAST-lane updates (a pending kernel does not count) | a new release | `TestAlarm_Ring0Stalled` |
| `os_not_covered` | warning | a ring-1 box has had fast-lane packages no approved release names for **14 days** | the packages are covered or gone | `TestAlarm_NotCovered` |
| `host_crash_restart` | warning | the box's crash guard reports a NEW unclean boot (a crash, a power cut or a hard reset; hub v0.132.0) | — (one per boot) | `api/crash_test.go` |
| `host_crash_guard_tripped` | error | the guard tripped: the next crash leaves the box OFF | the re-arm → `host_crash_guard_rearmed` (info) | `api/crash_test.go` |
| `host_kernel_oops` | warning | a kernel oops this boot (taint D) — the box keeps running | — (once per boot) | `api/crash_test.go` |
- **`unknown` never alarms** (R-96 rule 3): a probe that could not ask is neither up nor down. An `unknown` report
breaks a `not_running` run.
@@ -778,6 +778,28 @@ its length, and both fixes cost something the household would notice — operato
83. **Next: the tunnel status (R-841), the host fast lane (`11` §8 step 3), and other OS-update improvements.**
*Operator ruling 2026-10-04 ~12:20.*
### 2026-10-04 (evening) — decided by CC unattended — operator may reverse (System page / Docker / crash-restart brief)
90. **How does a box tell a crash boot from a clean one?** Options: (a) `pstore` — measured on demo-hp: `efi_pstore` is on,
yet a real panic saved NOTHING; (b) the previous boot's journal ends without shutdown lines — readable only after
the journal is up, and slow; (c) a marker written by an ExecStop at every orderly shutdown. **Chosen (c):** cheap,
early, measured to work; cost: a power cut and a hard reset count as a crash too (they cannot be told apart on these
boxes). `11` §5.9.
91. **`kernel.panic_on_oops`?** Options: set it (an oops becomes a restart, which may loop on a bad driver) or leave it
0 and REPORT an oops (taint bit D) to the operator. **Chosen: leave 0, report** — a box with an oops usually keeps
serving the household; the hub mails `host_kernel_oops`. `11` §5.9.
92. **The guard numbers.** The brief said both "at 3 it sets panic=0 so the next crash leaves the box off" (= the 4th)
and "if it crashes 3 times within one hour, it stays off" (= the 3rd, the operator's own page). **Chosen: the
operator's words** — the 3rd unclean stop within 60 minutes leaves the box off (the guard trips at the 2nd crash
boot); `kernel.panic` 10 s; re-arm after 24 h of normal running. All four are `/etc/felhom/crash-guard.conf`.
93. **Who may authorize a Docker step on the box itself?** Options: the agent's word (its config is agent-writable);
the wrapper verifying the signed job against ROOT-owned files. **Chosen: the wrapper verifies** — `ssh-keygen -Y
verify` against `/etc/felhom/operator-signers`, bound to `/etc/felhom/os-trust.json` `host_id`, a root-owned nonce
record; an unsigned ring-0 step only with that file's `ring0_slow_lane: true`, set by hand on the demo boxes only.
Cost: two more root files per box (the installer writes them; R-840 for installed boxes). `11` §5.8.
94. **The System page colours** use the alarm thresholds themselves (red = an `08` alarm would fire; amber = worth a
look; `unknown` = amber with its reason). No second set of numbers. `11` §5.7.
### 2026-10-04 (~15:17) — three operator rulings (recorded before the work)
87. **Docker `live-restore` is ON for every box** (`11` §5.8, option A). It is turned on once — by the golden and by a
+41 -2
View File
@@ -332,11 +332,31 @@ must never overlap a backup, a restore-test or a self-update.~~
- **Operator:** a hub event for each update and each failure; a fleet view showing each box's OS release,
how far behind it is, and whether it needs a reboot. A box that is more than N days behind the newest
approved release raises an alarm (`08`).
- **[FACT, hub v0.132.0] The System page** (`/system`, R-852, decision 89): one row per box — ring and switch with
buttons, the tunnel, host Proxmox / running and next-boot kernel / Debian / release / pending / not covered / held /
reboot needed / `kernel.panic` / oops / crash restarts / the guard, guest Debian / release / pending / restart needed,
Docker engine / containerd / live-restore / release, the last leg — and above it the releases, what ring 0 runs, "Approve
now" and "Approve Docker set". Colours are the alarm thresholds (decision 94). Hosts shows Proxmox / kernel too.
- **Household:** one line on the timeline in both languages, informal voice: what was updated and
whether the box restarted. Telling households in advance that the box may restart at night is a
**promise to users**. That is the operator's decision when the slow lane is built.
### 5.8 The Docker engine slow lane — DESIGN (2026-10-04, nothing built) `[PROPOSAL]`
### 5.8 The Docker engine slow lane — BUILT 2026-10-04 (agent v0.142.0, hub v0.132.0) `[FACT]`
**As built** (evidence `audits/os-docker-crash-2026-10-04/partB/`; decisions 87, 93):
- **live-restore ON**: the golden bakes it (`build-golden.sh` 3.1.0, fail-closed assertion); an installed box gets it once
by the wrapper's `live-restore-on` (merge into daemon.json + `systemctl reload docker`). Measured: the same container ids
after (9202 6/6 by hand — R10 refuses a scratch guest by design —, demo-hp 24/24, demo-felhom 5/5).
- **The step** runs in the night leg after a healthy guest and host step, **ring 0 only**, `select pending-docker`,
allowed by the wrapper only with the root-owned `ring0_slow_lane` mark. **Ring 1 and every undo** only through a signed
`os_docker_step` the wrapper re-verifies itself (decision 93). Measured: 29.7.x → 29.8.2 on both demo boxes, every id
kept; a signed undo to 29.7.2 on demo-hp (41 s, ids kept) and back; demo-felhom as ring 1 by a signed job; a replayed
job refused by the agent (nonce).
- **Approval**: the hub never approves a Docker set automatically; the System page's button works after every ring-0 box
ran the set in 2 healthy night Docker steps (`OS_DOCKER_APPROVE_NIGHTS` TEST override, logged). An approval nudges no
box. Undo: `runbooks/os-updates-docker-undo.md`.
**The design as written before the build:**
Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `live-restore` on, fleet-wide.**
@@ -360,6 +380,25 @@ Built from C5 (measured) and R-835. **The one decision it needs is in STATUS: `l
- **Not covered here:** the golden's own engine (baked weekly; a new golden carries the approved set), and BYO hosts
(the guest is ours on both, so the lane applies there too).
### 5.9 A crashed host restarts, with a limit — BUILT 2026-10-04 (agent v0.142.0, installer 1.30.0) `[FACT]`
Decision 88 (R-851): *"Yes, but maybe not indefinitely."* Evidence `audits/os-docker-crash-2026-10-04/partC/`.
- **Measured first (demo-hp, the operator's word before each crash):** `kernel.panic = 10` + `echo c >
/proc/sysrq-trigger` → the box restarted by itself in 54 s, on the same kernel (the saved default), the agent up 18 s
after the boot. `kernel.panic` set by `sysctl -w` is gone after the restart (back to 0) — it must be set at every boot.
- **The crash signal** (decision 90): `efi_pstore` is on, yet a real panic saved NOTHING; the journal and `last` show
only "no shutdown". The guard uses a **clean-stop marker** (the unit's ExecStop at every orderly shutdown); a boot
without it followed a crash, a power cut or a hard reset — counted alike.
- **The guard** (`felhom-crash-guard`, early boot unit + hourly re-arm timer): armed → `kernel.panic = 10`; the 2nd
unclean boot within 60 minutes TRIPS it (`kernel.panic = 0`), so **the 3rd crash within the hour leaves the box off**
(decision 92 — the operator's words); re-arms after 24 h of normal running or `felhom-crash-guard rearm`. A crash before
the unit runs (very early boot) leaves the box off: the safe side. `panic_on_oops` stays 0; an oops is reported
(decision 91).
- **Telling people:** each unclean boot = an operator mail (`host_crash_restart`) + the household's line; the trip = an
alarm (`host_crash_guard_tripped`); re-arm and oops announced once (`08` §6.3). The System page shows the guard.
- **Measured live:** crash 1 → back in 54 s; crash 2 → back in 53 s, guard tripped; crash 3 → **stayed off** until the
operator switched it on; the hub mailed the trip (15 min after the boot — R-853); re-armed by hand.
---
## 6. Risks and edge cases
@@ -432,7 +471,7 @@ Each step returns to the operator for go or no-go.
3. **Host Debian, fast lane** (no kernel, no Proxmox packages). **BUILT 2026-10-04** — agent v0.141.1, hub v0.131.1;
§8.2.
4. **Fleet view and alarms** (§5.7). **BUILT 2026-10-04** — hub v0.131.0/v0.131.1; §8.3.
5. **Slow lane: Docker engine.**
5. **Slow lane: Docker engine.** **BUILT 2026-10-04** — agent v0.142.0, hub v0.132.0; §5.8.
6. **Slow lane: host kernel and Proxmox packages, with the reboot.**
7. **Later:** the Proxmox major upgrade (PVE 9 → 10), drilled on ring 0 first.
@@ -0,0 +1,143 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Hosts — Felhom Hub</title>
<link rel="stylesheet" href="/style.css?v=0.132.0">
</head>
<body>
<svg xmlns="http://www.w3.org/2000/svg" style="display:none" aria-hidden="true">
<symbol id="i-triangle-alert" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m21.73 18-8-14a2 2 0 0 0-3.48 0l-8 14A2 2 0 0 0 4 21h16a2 2 0 0 0 1.73-3" /> <path d="M12 9v4" /> <path d="M12 17h.01" /></symbol>
<symbol id="i-check" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M20 6 9 17l-5-5" /></symbol>
<symbol id="i-server" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="20" height="8" x="2" y="2" rx="2" ry="2" /> <rect width="20" height="8" x="2" y="14" rx="2" ry="2" /> <line x1="6" x2="6.01" y1="6" y2="6" /> <line x1="6" x2="6.01" y1="18" y2="18" /></symbol>
<symbol id="i-settings" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M9.671 4.136a2.34 2.34 0 0 1 4.659 0 2.34 2.34 0 0 0 3.319 1.915 2.34 2.34 0 0 1 2.33 4.033 2.34 2.34 0 0 0 0 3.831 2.34 2.34 0 0 1-2.33 4.033 2.34 2.34 0 0 0-3.319 1.915 2.34 2.34 0 0 1-4.659 0 2.34 2.34 0 0 0-3.32-1.915 2.34 2.34 0 0 1-2.33-4.033 2.34 2.34 0 0 0 0-3.831A2.34 2.34 0 0 1 6.35 6.051a2.34 2.34 0 0 0 3.319-1.915" /> <circle cx="12" cy="12" r="3" /></symbol>
<symbol id="i-x" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M18 6 6 18" /> <path d="m6 6 12 12" /></symbol>
<symbol id="i-info" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 16v-4" /> <path d="M12 8h.01" /></symbol>
<symbol id="i-hard-drive" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M10 16h.01" /> <path d="M2.212 11.577a2 2 0 0 0-.212.896V18a2 2 0 0 0 2 2h16a2 2 0 0 0 2-2v-5.527a2 2 0 0 0-.212-.896L18.55 5.11A2 2 0 0 0 16.76 4H7.24a2 2 0 0 0-1.79 1.11z" /> <path d="M21.946 12.013H2.054" /> <path d="M6 16h.01" /></symbol>
<symbol id="i-cpu" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 20v2" /> <path d="M12 2v2" /> <path d="M17 20v2" /> <path d="M17 2v2" /> <path d="M2 12h2" /> <path d="M2 17h2" /> <path d="M2 7h2" /> <path d="M20 12h2" /> <path d="M20 17h2" /> <path d="M20 7h2" /> <path d="M7 20v2" /> <path d="M7 2v2" /> <rect x="4" y="4" width="16" height="16" rx="2" /> <rect x="8" y="8" width="8" height="8" rx="1" /></symbol>
<symbol id="i-clock" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 6v6l4 2" /></symbol>
<symbol id="i-boxes" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M2.97 12.92A2 2 0 0 0 2 14.63v3.24a2 2 0 0 0 .97 1.71l3 1.8a2 2 0 0 0 2.06 0L12 19v-5.5l-5-3-4.03 2.42Z" /> <path d="m7 16.5-4.74-2.85" /> <path d="m7 16.5 5-3" /> <path d="M7 16.5v5.17" /> <path d="M12 13.5V19l3.97 2.38a2 2 0 0 0 2.06 0l3-1.8a2 2 0 0 0 .97-1.71v-3.24a2 2 0 0 0-.97-1.71L17 10.5l-5 3Z" /> <path d="m17 16.5-5-3" /> <path d="m17 16.5 4.74-2.85" /> <path d="M17 16.5v5.17" /> <path d="M7.97 4.42A2 2 0 0 0 7 6.13v4.37l5 3 5-3V6.13a2 2 0 0 0-.97-1.71l-3-1.8a2 2 0 0 0-2.06 0l-3 1.8Z" /> <path d="M12 8 7.26 5.15" /> <path d="m12 8 4.74-2.85" /> <path d="M12 13.5V8" /></symbol>
<symbol id="i-users" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M16 21v-2a4 4 0 0 0-4-4H6a4 4 0 0 0-4 4v2" /> <path d="M16 3.128a4 4 0 0 1 0 7.744" /> <path d="M22 21v-2a4 4 0 0 0-3-3.87" /> <circle cx="9" cy="7" r="4" /></symbol>
</svg>
<script>
function felhomConfirm(el,question,onYes){
if(!el||el.dataset.fcOpen)return;
el.dataset.fcOpen='1';
var wrap=document.createElement('span');wrap.className='inline-confirm';
var q=document.createElement('span');q.className='inline-confirm-q';q.textContent=question;
var yes=document.createElement('button');yes.type='button';yes.className='btn btn-sm btn-danger';yes.textContent='Igen';
var no=document.createElement('button');no.type='button';no.className='btn btn-sm btn-outline';no.textContent='Mégse';
wrap.appendChild(q);wrap.appendChild(yes);wrap.appendChild(no);
el.style.display='none';el.parentNode.insertBefore(wrap,el.nextSibling);
function close(){wrap.remove();el.style.display='';delete el.dataset.fcOpen;}
no.addEventListener('click',close);
yes.addEventListener('click',function(){close();onYes();});
}
document.addEventListener('click',function(e){
var btn=e.target.closest?e.target.closest('[data-confirm]'):null;
if(!btn)return;
e.preventDefault();
felhomConfirm(btn,btn.getAttribute('data-confirm'),function(){
var form=btn.closest('form');
if(form){if(form.requestSubmit)form.requestSubmit(btn);else form.submit();}
});
});
</script>
<div class="container">
<header>
<h1>Felhom <span>Hub</span></h1>
<nav class="nav-links">
<a href="/" class="nav-link">Dashboard</a>
<a href="/configs" class="nav-link">Customers</a>
<a href="/apps" class="nav-link">Apps</a>
<a href="/hosts" class="nav-link active">Hosts</a>
<a href="/system" class="nav-link">System</a>
<a href="/offsite" class="nav-link">Offsite</a>
<a href="/configuration" class="nav-link">Configuration</a>
</nav>
</header>
<h2 style="margin-bottom: 1rem;">Hosts</h2>
<section class="card" style="padding: 0; overflow: hidden;">
<table class="data-table">
<thead>
<tr>
<th>Host</th>
<th>Customer</th>
<th>Agent</th>
<th>Proxmox / kernel</th>
<th>Status</th>
<th>Guests</th>
<th>CPU</th>
<th>Mem</th>
<th>Disk</th>
<th>Cloudflared</th>
<th>Worst Storage</th>
</tr>
</thead>
<tbody>
<tr onclick="window.location='/hosts/demo-felhom-8363b5'" style="cursor: pointer;">
<td><a href="/hosts/demo-felhom-8363b5">demo-felhom-8363b5</a></td>
<td>Demo Ügyfél</td>
<td><code>0.142.0</code></td>
<td style="font-size: 0.85em;">9.2.2<br><span class="text-muted">7.0.2-6-pve</span></td>
<td><span class="status-badge status-badge-ok">ONLINE</span></td>
<td>1/2</td>
<td>0%</td>
<td>17%</td>
<td>29%</td>
<td>running (connected)</td>
<td>29% <span class="text-muted">local</span></td>
</tr>
<tr onclick="window.location='/hosts/demo-hp-bb76ea'" style="cursor: pointer;">
<td><a href="/hosts/demo-hp-bb76ea">demo-hp-bb76ea</a></td>
<td>Demo HP</td>
<td><code>0.142.0</code></td>
<td style="font-size: 0.85em;">9.2.2<br><span class="text-muted">7.0.14-20-pve</span></td>
<td><span class="status-badge status-badge-ok">ONLINE</span></td>
<td>1/2</td>
<td>16%</td>
<td>19%</td>
<td>51%</td>
<td>running (connected)</td>
<td>63% <span class="text-muted">local-lvm</span></td>
</tr>
<tr onclick="window.location='/hosts/drill-r50-0a4f9a'" style="cursor: pointer;">
<td><a href="/hosts/drill-r50-0a4f9a">drill-r50-0a4f9a</a></td>
<td>drill-r50</td>
<td><code>0.129.0</code> <span class="status-badge status-warn" title="held: agent 0.129.0 &lt; MinAgent 0.131.0">floor held</span></td>
<td style="font-size: 0.85em;">unknown<br><span class="text-muted">unknown</span></td>
<td><span class="status-badge status-badge-down">DOWN</span></td>
<td>1/1</td>
<td>20%</td>
<td>20%</td>
<td>10%</td>
<td>inactive</td>
<td>10% <span class="text-muted">local</span></td>
</tr>
</tbody>
</table>
</section>
<footer style="margin-top: 2rem; color: var(--text-muted); font-size: 0.8rem; text-align: center;">
Felhom Hub <span style="font-family: var(--font-mono)">0.132.0</span>
</footer>
</div>
</body>
</html>
@@ -0,0 +1,235 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>System — Felhom Hub</title>
<link rel="stylesheet" href="/style.css?v=0.132.0">
<style>
.sys td, .sys th { white-space: nowrap; font-size: 0.82em; vertical-align: top; }
.sys .grp { border-left: 2px solid var(--border, #444); }
.c-warn { color: var(--warn); font-weight: 600; }
.c-bad { color: var(--danger, #e5534b); font-weight: 700; }
.sys form { display: inline; }
.rel-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(16rem, 1fr)); gap: 0.75rem; }
</style>
</head>
<body>
<svg xmlns="http://www.w3.org/2000/svg" style="display:none" aria-hidden="true">
<symbol id="i-triangle-alert" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m21.73 18-8-14a2 2 0 0 0-3.48 0l-8 14A2 2 0 0 0 4 21h16a2 2 0 0 0 1.73-3" /> <path d="M12 9v4" /> <path d="M12 17h.01" /></symbol>
<symbol id="i-check" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M20 6 9 17l-5-5" /></symbol>
<symbol id="i-server" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="20" height="8" x="2" y="2" rx="2" ry="2" /> <rect width="20" height="8" x="2" y="14" rx="2" ry="2" /> <line x1="6" x2="6.01" y1="6" y2="6" /> <line x1="6" x2="6.01" y1="18" y2="18" /></symbol>
<symbol id="i-settings" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M9.671 4.136a2.34 2.34 0 0 1 4.659 0 2.34 2.34 0 0 0 3.319 1.915 2.34 2.34 0 0 1 2.33 4.033 2.34 2.34 0 0 0 0 3.831 2.34 2.34 0 0 1-2.33 4.033 2.34 2.34 0 0 0-3.319 1.915 2.34 2.34 0 0 1-4.659 0 2.34 2.34 0 0 0-3.32-1.915 2.34 2.34 0 0 1-2.33-4.033 2.34 2.34 0 0 0 0-3.831A2.34 2.34 0 0 1 6.35 6.051a2.34 2.34 0 0 0 3.319-1.915" /> <circle cx="12" cy="12" r="3" /></symbol>
<symbol id="i-x" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M18 6 6 18" /> <path d="m6 6 12 12" /></symbol>
<symbol id="i-info" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 16v-4" /> <path d="M12 8h.01" /></symbol>
<symbol id="i-hard-drive" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M10 16h.01" /> <path d="M2.212 11.577a2 2 0 0 0-.212.896V18a2 2 0 0 0 2 2h16a2 2 0 0 0 2-2v-5.527a2 2 0 0 0-.212-.896L18.55 5.11A2 2 0 0 0 16.76 4H7.24a2 2 0 0 0-1.79 1.11z" /> <path d="M21.946 12.013H2.054" /> <path d="M6 16h.01" /></symbol>
<symbol id="i-cpu" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 20v2" /> <path d="M12 2v2" /> <path d="M17 20v2" /> <path d="M17 2v2" /> <path d="M2 12h2" /> <path d="M2 17h2" /> <path d="M2 7h2" /> <path d="M20 12h2" /> <path d="M20 17h2" /> <path d="M20 7h2" /> <path d="M7 20v2" /> <path d="M7 2v2" /> <rect x="4" y="4" width="16" height="16" rx="2" /> <rect x="8" y="8" width="8" height="8" rx="1" /></symbol>
<symbol id="i-clock" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 6v6l4 2" /></symbol>
<symbol id="i-boxes" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M2.97 12.92A2 2 0 0 0 2 14.63v3.24a2 2 0 0 0 .97 1.71l3 1.8a2 2 0 0 0 2.06 0L12 19v-5.5l-5-3-4.03 2.42Z" /> <path d="m7 16.5-4.74-2.85" /> <path d="m7 16.5 5-3" /> <path d="M7 16.5v5.17" /> <path d="M12 13.5V19l3.97 2.38a2 2 0 0 0 2.06 0l3-1.8a2 2 0 0 0 .97-1.71v-3.24a2 2 0 0 0-.97-1.71L17 10.5l-5 3Z" /> <path d="m17 16.5-5-3" /> <path d="m17 16.5 4.74-2.85" /> <path d="M17 16.5v5.17" /> <path d="M7.97 4.42A2 2 0 0 0 7 6.13v4.37l5 3 5-3V6.13a2 2 0 0 0-.97-1.71l-3-1.8a2 2 0 0 0-2.06 0l-3 1.8Z" /> <path d="M12 8 7.26 5.15" /> <path d="m12 8 4.74-2.85" /> <path d="M12 13.5V8" /></symbol>
<symbol id="i-users" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M16 21v-2a4 4 0 0 0-4-4H6a4 4 0 0 0-4 4v2" /> <path d="M16 3.128a4 4 0 0 1 0 7.744" /> <path d="M22 21v-2a4 4 0 0 0-3-3.87" /> <circle cx="9" cy="7" r="4" /></symbol>
</svg>
<script>
function felhomConfirm(el,question,onYes){
if(!el||el.dataset.fcOpen)return;
el.dataset.fcOpen='1';
var wrap=document.createElement('span');wrap.className='inline-confirm';
var q=document.createElement('span');q.className='inline-confirm-q';q.textContent=question;
var yes=document.createElement('button');yes.type='button';yes.className='btn btn-sm btn-danger';yes.textContent='Igen';
var no=document.createElement('button');no.type='button';no.className='btn btn-sm btn-outline';no.textContent='Mégse';
wrap.appendChild(q);wrap.appendChild(yes);wrap.appendChild(no);
el.style.display='none';el.parentNode.insertBefore(wrap,el.nextSibling);
function close(){wrap.remove();el.style.display='';delete el.dataset.fcOpen;}
no.addEventListener('click',close);
yes.addEventListener('click',function(){close();onYes();});
}
document.addEventListener('click',function(e){
var btn=e.target.closest?e.target.closest('[data-confirm]'):null;
if(!btn)return;
e.preventDefault();
felhomConfirm(btn,btn.getAttribute('data-confirm'),function(){
var form=btn.closest('form');
if(form){if(form.requestSubmit)form.requestSubmit(btn);else form.submit();}
});
});
</script>
<div class="container">
<header>
<h1>Felhom <span>Hub</span></h1>
<nav class="nav-links">
<a href="/" class="nav-link">Dashboard</a>
<a href="/configs" class="nav-link">Customers</a>
<a href="/apps" class="nav-link">Apps</a>
<a href="/hosts" class="nav-link">Hosts</a>
<a href="/system" class="nav-link active">System</a>
<a href="/offsite" class="nav-link">Offsite</a>
<a href="/configuration" class="nav-link">Configuration</a>
</nav>
</header>
<h2 style="margin-bottom: 1rem;">System — versions and OS updates</h2>
<section class="card" style="margin-bottom: 1.5rem;">
<h3 style="margin-top: 0;">Approved releases</h3>
<div class="rel-grid">
<div><strong>guest</strong>: <code>os-guest-20261004-123933</code><br>
<span class="text-muted">272 packages · 2026-10-04 12:39 UTC · by auto</span></div>
<div><strong>host</strong>: <code>os-host-20261004-124133</code><br>
<span class="text-muted">607 packages · 2026-10-04 12:41 UTC · by auto</span></div>
</div>
<h3>What ring 0 runs now</h3>
<div class="rel-grid">
<div><strong>guest</strong>:
0 packages, first seen 2026-10-04 14:14 UTC<br>
<span class="text-muted">healthy for 10m0s of 24h0m0s</span>
</div>
<div><strong>host</strong>:
607 packages, first seen 2026-10-04 12:24 UTC<br>
<span class="text-muted">approved as os-host-20261004-124133</span>
</div>
<div><strong>docker</strong>:
<span class="text-muted">—</span><br>
<span class="text-muted">a ring-0 box has not reported a Docker step</span>
</div>
</div>
<form method="POST" action="/os/approve-now" style="margin-top: 0.8rem;">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<button type="submit" class="btn btn-sm btn-danger" data-confirm="Approve the guest and host sets ring 0 runs NOW, without the 24 h + 1 night wait? Every ring-1 box installs them at its next night run.">Approve now (guest + host)</button>
<span class="text-muted" style="font-size: 0.85em;">An urgent fix only — normally the hub approves after 24 h and one night.</span>
</form>
</section>
<section class="card" style="padding: 0; overflow-x: auto;">
<table class="data-table sys">
<thead>
<tr>
<th>Box</th><th>Ring / updates</th><th>Tunnel</th>
<th class="grp">Proxmox</th><th>Kernel (running)</th><th>Kernel (next boot)</th><th>Debian</th><th>Felhom release</th><th>Pending</th><th>Not covered</th><th>Held</th><th>Reboot needed</th><th>kernel.panic</th><th>Oops</th><th>Crash restarts 24 h</th><th>Crash guard</th>
<th class="grp">Guest Debian</th><th>Felhom release</th><th>Pending</th><th>Restart needed</th>
<th class="grp">Docker</th><th>containerd</th><th>live-restore</th><th>Docker release</th>
<th class="grp">Last OS leg</th>
</tr>
<tr class="text-muted"><th></th><th></th><th></th><th class="grp" colspan="13">host</th><th class="grp" colspan="4">guest</th><th class="grp" colspan="4">Docker engine</th><th class="grp"></th></tr>
</thead>
<tbody>
<tr>
<td><a href="/hosts/demo-felhom-8363b5">demo-felhom-8363b5</a><br><span class="text-muted">Demo Ügyfél</span>
</td>
<td>
ring 0
<form method="POST" action="/os/ring/demo-felhom-8363b5">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="1"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make demo-felhom-8363b5 a normal (ring 1) box? It then installs only approved releases.">→ normal</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/demo-felhom-8363b5">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for demo-felhom-8363b5? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td>running</td>
<td class="grp ">9.2.2</td>
<td>7.0.2-6-pve</td><td>7.0.2-6-pve</td><td>13.7</td>
<td>ring0-20261004T141237Z</td><td>80</td><td>0</td>
<td>none</td><td class="c-warn">since 2026-10-04</td><td>10 s</td><td>no</td>
<td>0</td><td>armed</td>
<td class="grp ">13.7</td>
<td>ring0-20261004T141237Z</td><td>0</td><td>10</td>
<td class="grp ">29.8.2</td>
<td>2.3.6-1~debian.13~trixie</td><td>on</td><td>—</td>
<td class="grp " title="last successful leg: 10 min ago">10 min ago · applied · 34 s</td>
</tr>
<tr>
<td><a href="/hosts/demo-hp-bb76ea">demo-hp-bb76ea</a><br><span class="text-muted">Demo HP</span>
</td>
<td>
ring 0
<form method="POST" action="/os/ring/demo-hp-bb76ea">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="1"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make demo-hp-bb76ea a normal (ring 1) box? It then installs only approved releases.">→ normal</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/demo-hp-bb76ea">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for demo-hp-bb76ea? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td>running</td>
<td class="grp ">9.2.2</td>
<td>7.0.14-20-pve</td><td>7.0.14-20-pve</td><td>13.7</td>
<td>ring0-20261004T140947Z</td><td>78</td><td>0</td>
<td>none</td><td>no</td><td>10 s</td><td>no</td>
<td>0</td><td>armed</td>
<td class="grp ">13.7</td>
<td>ring0-20261004T140947Z</td><td>0</td><td>1</td>
<td class="grp ">29.8.2</td>
<td>2.3.6-1~debian.13~trixie</td><td>on</td><td>—</td>
<td class="grp " title="last successful leg: 13 min ago">13 min ago · applied · 52 s</td>
</tr>
<tr>
<td><a href="/hosts/drill-r50-0a4f9a">drill-r50-0a4f9a</a><br><span class="text-muted">drill-r50</span>
<br><span class="c-warn" style="font-weight: normal;">no versions reported (agent older than v0.142.0)</span></td>
<td>
ring 1
<form method="POST" action="/os/ring/drill-r50-0a4f9a">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make drill-r50-0a4f9a a DEMO (ring 0) box? It then installs every new fix first and takes unsigned Docker steps if its root-owned ring-0 mark allows.">→ demo</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/drill-r50-0a4f9a">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for drill-r50-0a4f9a? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td class="c-bad">inactive</td>
<td class="grp c-warn">unknown</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td>
<td>—</td><td>0</td><td>0</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>no</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>no</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn">not installed</td>
<td class="grp c-warn">unknown</td>
<td>—</td><td>0</td><td>0</td>
<td class="grp c-warn">unknown</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>—</td>
<td class="grp " title="last successful leg: never">never · — · 0 s</td>
</tr>
</tbody>
</table>
</section>
<p class="text-muted" style="font-size: 0.85em;">Amber: worth a look. Red: an operator alarm fires (`08` §6.3). "unknown": the box could not read the value — never a guess.</p>
<footer style="margin-top: 2rem; color: var(--text-muted); font-size: 0.8rem; text-align: center;">
Felhom Hub <span style="font-family: var(--font-mono)">0.132.0</span>
</footer>
</div>
</body>
</html>
@@ -0,0 +1,2 @@
HTTP/1.1 303 See Other
Location: /system?flash=approved+os-docker-20261004-142842
@@ -0,0 +1,240 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>System — Felhom Hub</title>
<link rel="stylesheet" href="/style.css?v=0.132.0">
<style>
.sys td, .sys th { white-space: nowrap; font-size: 0.82em; vertical-align: top; }
.sys .grp { border-left: 2px solid var(--border, #444); }
.c-warn { color: var(--warn); font-weight: 600; }
.c-bad { color: var(--danger, #e5534b); font-weight: 700; }
.sys form { display: inline; }
.rel-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(16rem, 1fr)); gap: 0.75rem; }
</style>
</head>
<body>
<svg xmlns="http://www.w3.org/2000/svg" style="display:none" aria-hidden="true">
<symbol id="i-triangle-alert" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m21.73 18-8-14a2 2 0 0 0-3.48 0l-8 14A2 2 0 0 0 4 21h16a2 2 0 0 0 1.73-3" /> <path d="M12 9v4" /> <path d="M12 17h.01" /></symbol>
<symbol id="i-check" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M20 6 9 17l-5-5" /></symbol>
<symbol id="i-server" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="20" height="8" x="2" y="2" rx="2" ry="2" /> <rect width="20" height="8" x="2" y="14" rx="2" ry="2" /> <line x1="6" x2="6.01" y1="6" y2="6" /> <line x1="6" x2="6.01" y1="18" y2="18" /></symbol>
<symbol id="i-settings" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M9.671 4.136a2.34 2.34 0 0 1 4.659 0 2.34 2.34 0 0 0 3.319 1.915 2.34 2.34 0 0 1 2.33 4.033 2.34 2.34 0 0 0 0 3.831 2.34 2.34 0 0 1-2.33 4.033 2.34 2.34 0 0 0-3.319 1.915 2.34 2.34 0 0 1-4.659 0 2.34 2.34 0 0 0-3.32-1.915 2.34 2.34 0 0 1-2.33-4.033 2.34 2.34 0 0 0 0-3.831A2.34 2.34 0 0 1 6.35 6.051a2.34 2.34 0 0 0 3.319-1.915" /> <circle cx="12" cy="12" r="3" /></symbol>
<symbol id="i-x" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M18 6 6 18" /> <path d="m6 6 12 12" /></symbol>
<symbol id="i-info" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 16v-4" /> <path d="M12 8h.01" /></symbol>
<symbol id="i-hard-drive" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M10 16h.01" /> <path d="M2.212 11.577a2 2 0 0 0-.212.896V18a2 2 0 0 0 2 2h16a2 2 0 0 0 2-2v-5.527a2 2 0 0 0-.212-.896L18.55 5.11A2 2 0 0 0 16.76 4H7.24a2 2 0 0 0-1.79 1.11z" /> <path d="M21.946 12.013H2.054" /> <path d="M6 16h.01" /></symbol>
<symbol id="i-cpu" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 20v2" /> <path d="M12 2v2" /> <path d="M17 20v2" /> <path d="M17 2v2" /> <path d="M2 12h2" /> <path d="M2 17h2" /> <path d="M2 7h2" /> <path d="M20 12h2" /> <path d="M20 17h2" /> <path d="M20 7h2" /> <path d="M7 20v2" /> <path d="M7 2v2" /> <rect x="4" y="4" width="16" height="16" rx="2" /> <rect x="8" y="8" width="8" height="8" rx="1" /></symbol>
<symbol id="i-clock" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 6v6l4 2" /></symbol>
<symbol id="i-boxes" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M2.97 12.92A2 2 0 0 0 2 14.63v3.24a2 2 0 0 0 .97 1.71l3 1.8a2 2 0 0 0 2.06 0L12 19v-5.5l-5-3-4.03 2.42Z" /> <path d="m7 16.5-4.74-2.85" /> <path d="m7 16.5 5-3" /> <path d="M7 16.5v5.17" /> <path d="M12 13.5V19l3.97 2.38a2 2 0 0 0 2.06 0l3-1.8a2 2 0 0 0 .97-1.71v-3.24a2 2 0 0 0-.97-1.71L17 10.5l-5 3Z" /> <path d="m17 16.5-5-3" /> <path d="m17 16.5 4.74-2.85" /> <path d="M17 16.5v5.17" /> <path d="M7.97 4.42A2 2 0 0 0 7 6.13v4.37l5 3 5-3V6.13a2 2 0 0 0-.97-1.71l-3-1.8a2 2 0 0 0-2.06 0l-3 1.8Z" /> <path d="M12 8 7.26 5.15" /> <path d="m12 8 4.74-2.85" /> <path d="M12 13.5V8" /></symbol>
<symbol id="i-users" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M16 21v-2a4 4 0 0 0-4-4H6a4 4 0 0 0-4 4v2" /> <path d="M16 3.128a4 4 0 0 1 0 7.744" /> <path d="M22 21v-2a4 4 0 0 0-3-3.87" /> <circle cx="9" cy="7" r="4" /></symbol>
</svg>
<script>
function felhomConfirm(el,question,onYes){
if(!el||el.dataset.fcOpen)return;
el.dataset.fcOpen='1';
var wrap=document.createElement('span');wrap.className='inline-confirm';
var q=document.createElement('span');q.className='inline-confirm-q';q.textContent=question;
var yes=document.createElement('button');yes.type='button';yes.className='btn btn-sm btn-danger';yes.textContent='Igen';
var no=document.createElement('button');no.type='button';no.className='btn btn-sm btn-outline';no.textContent='Mégse';
wrap.appendChild(q);wrap.appendChild(yes);wrap.appendChild(no);
el.style.display='none';el.parentNode.insertBefore(wrap,el.nextSibling);
function close(){wrap.remove();el.style.display='';delete el.dataset.fcOpen;}
no.addEventListener('click',close);
yes.addEventListener('click',function(){close();onYes();});
}
document.addEventListener('click',function(e){
var btn=e.target.closest?e.target.closest('[data-confirm]'):null;
if(!btn)return;
e.preventDefault();
felhomConfirm(btn,btn.getAttribute('data-confirm'),function(){
var form=btn.closest('form');
if(form){if(form.requestSubmit)form.requestSubmit(btn);else form.submit();}
});
});
</script>
<div class="container">
<header>
<h1>Felhom <span>Hub</span></h1>
<nav class="nav-links">
<a href="/" class="nav-link">Dashboard</a>
<a href="/configs" class="nav-link">Customers</a>
<a href="/apps" class="nav-link">Apps</a>
<a href="/hosts" class="nav-link">Hosts</a>
<a href="/system" class="nav-link active">System</a>
<a href="/offsite" class="nav-link">Offsite</a>
<a href="/configuration" class="nav-link">Configuration</a>
</nav>
</header>
<h2 style="margin-bottom: 1rem;">System — versions and OS updates</h2>
<section class="card" style="margin-bottom: 1.5rem;">
<h3 style="margin-top: 0;">Approved releases</h3>
<div class="rel-grid">
<div><strong>guest</strong>: <code>os-guest-20261004-123933</code><br>
<span class="text-muted">272 packages · 2026-10-04 12:39 UTC · by auto</span></div>
<div><strong>host</strong>: <code>os-host-20261004-124133</code><br>
<span class="text-muted">607 packages · 2026-10-04 12:41 UTC · by auto</span></div>
</div>
<h3>What ring 0 runs now</h3>
<div class="rel-grid">
<div><strong>guest</strong>:
272 packages, first seen 2026-10-04 11:07 UTC<br>
<span class="text-muted">approved as os-guest-20261004-123933</span>
</div>
<div><strong>host</strong>:
607 packages, first seen 2026-10-04 12:24 UTC<br>
<span class="text-muted">approved as os-host-20261004-124133</span>
</div>
<div><strong>docker</strong>:
6 packages, first seen 2026-10-04 14:28 UTC<br>
<form method="POST" action="/os/approve-docker" style="margin-top: 0.3rem;">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<button type="submit" class="btn btn-sm" data-confirm="Approve this Docker engine set? Ring-1 boxes take it only through a signed operator job.">Approve Docker set</button>
</form>
</div>
</div>
<form method="POST" action="/os/approve-now" style="margin-top: 0.8rem;">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<button type="submit" class="btn btn-sm btn-danger" data-confirm="Approve the guest and host sets ring 0 runs NOW, without the 24 h + 1 night wait? Every ring-1 box installs them at its next night run.">Approve now (guest + host)</button>
<span class="text-muted" style="font-size: 0.85em;">An urgent fix only — normally the hub approves after 24 h and one night.</span>
</form>
</section>
<section class="card" style="padding: 0; overflow-x: auto;">
<table class="data-table sys">
<thead>
<tr>
<th>Box</th><th>Ring / updates</th><th>Tunnel</th>
<th class="grp">Proxmox</th><th>Kernel (running)</th><th>Kernel (next boot)</th><th>Debian</th><th>Felhom release</th><th>Pending</th><th>Not covered</th><th>Held</th><th>Reboot needed</th><th>kernel.panic</th><th>Oops</th><th>Crash restarts 24 h</th><th>Crash guard</th>
<th class="grp">Guest Debian</th><th>Felhom release</th><th>Pending</th><th>Restart needed</th>
<th class="grp">Docker</th><th>containerd</th><th>live-restore</th><th>Docker release</th>
<th class="grp">Last OS leg</th>
</tr>
<tr class="text-muted"><th></th><th></th><th></th><th class="grp" colspan="13">host</th><th class="grp" colspan="4">guest</th><th class="grp" colspan="4">Docker engine</th><th class="grp"></th></tr>
</thead>
<tbody>
<tr>
<td><a href="/hosts/demo-felhom-8363b5">demo-felhom-8363b5</a><br><span class="text-muted">Demo Ügyfél</span>
</td>
<td>
ring 0
<form method="POST" action="/os/ring/demo-felhom-8363b5">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="1"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make demo-felhom-8363b5 a normal (ring 1) box? It then installs only approved releases.">→ normal</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/demo-felhom-8363b5">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for demo-felhom-8363b5? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td>running</td>
<td class="grp ">9.2.2</td>
<td>7.0.2-6-pve</td><td>7.0.2-6-pve</td><td>13.7</td>
<td>ring0-20261004T142539Z</td><td>80</td><td>0</td>
<td>none</td><td class="c-warn">since 2026-10-04</td><td>10 s</td><td>no</td>
<td>0</td><td>armed</td>
<td class="grp ">13.7</td>
<td>ring0-20261004T142539Z</td><td>0</td><td>10</td>
<td class="grp ">29.8.2</td>
<td>2.3.6-1~debian.13~trixie</td><td>on</td><td>ring0-20261004T142539Z</td>
<td class="grp " title="last successful leg: 2 min ago">2 min ago · nothing · 13 s</td>
</tr>
<tr>
<td><a href="/hosts/demo-hp-bb76ea">demo-hp-bb76ea</a><br><span class="text-muted">Demo HP</span>
</td>
<td>
ring 0
<form method="POST" action="/os/ring/demo-hp-bb76ea">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="1"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make demo-hp-bb76ea a normal (ring 1) box? It then installs only approved releases.">→ normal</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/demo-hp-bb76ea">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for demo-hp-bb76ea? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td>running</td>
<td class="grp ">9.2.2</td>
<td>7.0.14-20-pve</td><td>7.0.14-20-pve</td><td>13.7</td>
<td>ring0-20261004T142618Z</td><td>78</td><td>0</td>
<td>none</td><td>no</td><td>10 s</td><td>no</td>
<td>0</td><td>armed</td>
<td class="grp ">13.7</td>
<td>ring0-20261004T142618Z</td><td>0</td><td>1</td>
<td class="grp ">29.8.2</td>
<td>2.3.6-1~debian.13~trixie</td><td>on</td><td>ring0-20261004T142618Z</td>
<td class="grp " title="last successful leg: 2 min ago">1 min ago · nothing · 17 s</td>
</tr>
<tr>
<td><a href="/hosts/drill-r50-0a4f9a">drill-r50-0a4f9a</a><br><span class="text-muted">drill-r50</span>
<br><span class="c-warn" style="font-weight: normal;">no versions reported (agent older than v0.142.0)</span></td>
<td>
ring 1
<form method="POST" action="/os/ring/drill-r50-0a4f9a">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make drill-r50-0a4f9a a DEMO (ring 0) box? It then installs every new fix first and takes unsigned Docker steps if its root-owned ring-0 mark allows.">→ demo</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/drill-r50-0a4f9a">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for drill-r50-0a4f9a? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td class="c-bad">inactive</td>
<td class="grp c-warn">unknown</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td>
<td>—</td><td>0</td><td>0</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>no</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>no</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn">not installed</td>
<td class="grp c-warn">unknown</td>
<td>—</td><td>0</td><td>0</td>
<td class="grp c-warn">unknown</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>—</td>
<td class="grp " title="last successful leg: never">never · — · 0 s</td>
</tr>
</tbody>
</table>
</section>
<p class="text-muted" style="font-size: 0.85em;">Amber: worth a look. Red: an operator alarm fires (`08` §6.3). "unknown": the box could not read the value — never a guess.</p>
<footer style="margin-top: 2rem; color: var(--text-muted); font-size: 0.8rem; text-align: center;">
Felhom Hub <span style="font-family: var(--font-mono)">0.132.0</span>
</footer>
</div>
</body>
</html>
@@ -0,0 +1,163 @@
=== felhom-agent 0.142.0 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
time=2026-10-04T16:26:18.795+02:00 level=INFO msg="osupdate: START" run=20261004T142618Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T142618Z
time=2026-10-04T16:26:34.572+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T142618Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T16:26:34.572+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T16:26:34.572+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:26:34.572+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:26:34.574+02:00 level=INFO msg="osupdate: DONE" run=20261004T142618Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=1 reboot_needed=false wrapper_seconds=15.7
time=2026-10-04T16:26:34.583+02:00 level=INFO msg="osupdate: START" run=20261004T142618Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T142618Z
time=2026-10-04T16:26:49.369+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T142618Z layer=host lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T16:26:49.369+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T16:26:49.369+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:26:49.369+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:26:50.399+02:00 level=INFO msg="osupdate: DONE" run=20261004T142618Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=78 not_covered=78 restart_needed=0 reboot_needed=false wrapper_seconds=14.7
time=2026-10-04T16:26:52.343+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: LIVE-RESTORE already on"
time=2026-10-04T16:26:52.343+02:00 level=INFO msg="osupdate: live-restore" vmid=9201 result="{\"result\": \"already on\"}"
time=2026-10-04T16:26:52.343+02:00 level=INFO msg="osupdate: START" run=20261004T142618Z layer=docker vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T142618Z
time=2026-10-04T16:27:09.714+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T142618Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0"
time=2026-10-04T16:27:09.715+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T16:27:09.715+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:27:09.715+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:27:09.716+02:00 level=INFO msg="osupdate: DONE" run=20261004T142618Z layer=docker vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=1 reboot_needed=false wrapper_seconds=17.3
--- os-update report (guest) ---
{
"authority": "",
"docker_engine": "",
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": null,
"outcome": "nothing",
"pending": 0,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T142618Z",
"restart_needed": [
"containerd-shim"
],
"ring": 0,
"run_id": "20261004T142618Z",
"upgraded": [],
"wrapper_seconds": 15.7
}
--- os-update report (host) ---
{
"authority": "",
"docker_engine": "",
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"proxmox-secure-boot-support",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"amd64-microcode",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"pve-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed"
],
"outcome": "nothing",
"pending": 78,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T142618Z",
"restart_needed": [],
"ring": 0,
"run_id": "20261004T142618Z",
"upgraded": [],
"wrapper_seconds": 14.7
}
--- os-update report (docker) ---
{
"authority": "ring0",
"docker_engine": "29.8.2",
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": null,
"outcome": "nothing",
"pending": 0,
"reboot_needed": false,
"refused": null,
"release_id": "ring0-20261004T142618Z",
"restart_needed": [
"containerd-shim"
],
"ring": 0,
"run_id": "20261004T142618Z",
"upgraded": [],
"wrapper_seconds": 17.3
}
pass took 50.9s
WALL=51.005070963
@@ -0,0 +1,218 @@
=== felhom-agent 0.142.0 selftest=os-update vmid=9201 ring=0 enabled=true guest-release=false host-release=false appliance=true ===
time=2026-10-04T16:25:39.774+02:00 level=INFO msg="osupdate: START" run=20261004T142539Z layer=guest vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T142539Z
time=2026-10-04T16:25:52.120+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T142539Z layer=guest:9201 lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T16:25:52.120+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T16:25:52.120+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:25:52.120+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:25:52.120+02:00 level=INFO msg="osupdate: DONE" run=20261004T142539Z layer=guest vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=10 reboot_needed=true wrapper_seconds=12.3
time=2026-10-04T16:25:52.131+02:00 level=INFO msg="osupdate: START" run=20261004T142539Z layer=host vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T142539Z
time=2026-10-04T16:26:02.540+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T142539Z layer=host lane=fast mode=apply select=pending-fast packages=0"
time=2026-10-04T16:26:02.540+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T16:26:02.540+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:26:02.540+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:26:03.376+02:00 level=INFO msg="osupdate: DONE" run=20261004T142539Z layer=host vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=10.4
time=2026-10-04T16:26:04.999+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: LIVE-RESTORE already on"
time=2026-10-04T16:26:04.999+02:00 level=INFO msg="osupdate: live-restore" vmid=9201 result="{\"result\": \"already on\"}"
time=2026-10-04T16:26:04.999+02:00 level=INFO msg="osupdate: START" run=20261004T142539Z layer=docker vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T142539Z
time=2026-10-04T16:26:18.174+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T142539Z layer=docker:9201 lane=slow mode=apply select=pending-docker packages=0 authority=ring0"
time=2026-10-04T16:26:18.174+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
time=2026-10-04T16:26:18.174+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:26:18.174+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:26:18.174+02:00 level=INFO msg="osupdate: DONE" run=20261004T142539Z layer=docker vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=10 reboot_needed=true wrapper_seconds=13.1
--- os-update report (guest) ---
{
"authority": "",
"docker_engine": "",
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": null,
"outcome": "nothing",
"pending": 0,
"reboot_needed": true,
"refused": null,
"release_id": "ring0-20261004T142539Z",
"restart_needed": [
"agetty",
"containerd-shim",
"cron",
"dbus-daemon",
"dhclient",
"sshd",
"systemd",
"systemd-journal",
"systemd-logind",
"systemd-network"
],
"ring": 0,
"run_id": "20261004T142539Z",
"upgraded": [],
"wrapper_seconds": 12.3
}
--- os-update report (host) ---
{
"authority": "",
"docker_engine": "",
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": [
"frr",
"shim-signed-common",
"shim-unsigned",
"shim-helpers-amd64-signed",
"shim-signed",
"libradosstriper1",
"librgw2",
"ceph-common",
"librbd1",
"librados2",
"python3-cephfs",
"libcephfs2",
"python3-rgw",
"python3-rados",
"python3-ceph-argparse",
"python3-ceph-common",
"python3-rbd",
"ceph-fuse",
"chrony",
"libcorosync-common4",
"libcfg7",
"libcmap4",
"libcpg4",
"libknet1t64",
"libnozzle1t64",
"libquorum5",
"libvotequorum8",
"corosync",
"frr-pythontools",
"libjs-extjs",
"libnvpair3linux",
"libproxmox-acme-plugins",
"libproxmox-backup-qemu0",
"pve-qemu-kvm",
"libpve-notify-perl",
"libpve-cluster-api-perl",
"libpve-cluster-perl",
"pve-cluster",
"libpve-access-control",
"libpve-apiclient-perl",
"librados2-perl",
"proxmox-backup-client",
"proxmox-backup-file-restore",
"pve-manager",
"libproxmox-acme-perl",
"libpve-common-perl",
"libpve-guest-common-perl",
"qemu-server",
"libpve-storage-perl",
"pve-edk2-firmware-legacy",
"pve-edk2-firmware-ovmf",
"libpve-network-api-perl",
"libpve-network-perl",
"proxmox-firewall-data",
"pve-firewall",
"pve-container",
"pve-ha-manager",
"novnc-pve",
"proxmox-enterprise-support-keyring",
"proxmox-mini-journalreader",
"proxmox-widget-toolkit",
"pve-docs",
"pve-i18n",
"pve-xtermjs",
"pve-yew-mobile-i18n",
"pve-yew-mobile-gui",
"libuutil3linux",
"libzfs7linux",
"libzpool7linux",
"proxmox-first-boot",
"pve-firmware",
"proxmox-kernel-7.0.14-20-pve-signed",
"proxmox-kernel-7.0",
"proxmox-kernel-helper",
"pve-edk2-firmware-aarch64",
"pve-edk2-firmware",
"zfs-initramfs",
"zfsutils-linux",
"zfs-zed",
"tailscale"
],
"outcome": "nothing",
"pending": 80,
"reboot_needed": true,
"refused": null,
"release_id": "ring0-20261004T142539Z",
"restart_needed": [
"agetty",
"blkmapd",
"chronyd",
"cron",
"dbus-daemon",
"dmeventd",
"ksmtuned",
"lxc-monitord",
"lxc-start",
"lxcfs",
"pmxcfs",
"proxmox-firewal",
"pve-firewall",
"pve-ha-crm",
"pve-ha-lrm",
"pve-lxc-syscall",
"pvedaemon",
"pvedaemon worke",
"pvefw-logger",
"pveproxy",
"pveproxy worker",
"pvescheduler",
"pvestatd",
"qmeventd",
"rpcbind",
"rrdcached",
"smartd",
"spiceproxy",
"spiceproxy work",
"sshd",
"systemd-logind",
"systemd-udevd",
"watchdog-mux",
"zed"
],
"ring": 0,
"run_id": "20261004T142539Z",
"upgraded": [],
"wrapper_seconds": 10.4
}
--- os-update report (docker) ---
{
"authority": "ring0",
"docker_engine": "29.8.2",
"health_reason": "",
"healthy": true,
"mode": "apply",
"not_covered": null,
"outcome": "nothing",
"pending": 0,
"reboot_needed": true,
"refused": null,
"release_id": "ring0-20261004T142539Z",
"restart_needed": [
"agetty",
"containerd-shim",
"cron",
"dbus-daemon",
"dhclient",
"sshd",
"systemd",
"systemd-journal",
"systemd-logind",
"systemd-network"
],
"ring": 0,
"run_id": "20261004T142539Z",
"upgraded": [],
"wrapper_seconds": 13.1
}
pass took 38.4s
WALL=38.491768907
@@ -0,0 +1 @@
nonce=cb4a902da3007b8f4c3a440d3f2d9a60
@@ -0,0 +1 @@
{"release_id":"os-docker-20261004-142842","undo":false,"packages":[{"name":"containerd.io","version":"2.3.6-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-buildx-plugin","version":"0.37.1-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-ce","version":"5:29.8.2-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-ce-cli","version":"5:29.8.2-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-ce-rootless-extras","version":"5:29.8.2-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-compose-plugin","version":"5.6.0-1~debian.13~trixie","origin":"Docker CE"}]}
@@ -0,0 +1,2 @@
HTTP/1.1 303 See Other
Location: /system?flash=done
@@ -0,0 +1,11 @@
=== felhom-agent 0.142.0 selftest=os-update vmid=9201 ring=1 enabled=true guest-release=true host-release=true appliance=true ===
time=2026-10-04T16:42:05.041+02:00 level=INFO msg="osupdate: DONE" run=20261004T144154Z layer=guest vmid=9201 ring=1 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=10 reboot_needed=true wrapper_seconds=10.6
time=2026-10-04T16:42:15.090+02:00 level=INFO msg="osupdate: DONE" run=20261004T144154Z layer=host vmid=9201 ring=1 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=80 not_covered=80 restart_needed=34 reboot_needed=true wrapper_seconds=9.2
time=2026-10-04T16:42:15.108+02:00 level=INFO msg="osupdate: docker step skipped — ring 1 takes an engine set only inside a signed operator job (`11` §5.8)" run=20261004T144154Z vmid=9201 trigger=debug ring=1 enabled=true
docker step: skipped (see the log line above)
containerd.io 2.3.6-1~debian.13~trixie
docker-buildx-plugin 0.37.1-1~debian.13~trixie
docker-ce 5:29.8.2-1~debian.13~trixie
docker-ce-cli 5:29.8.2-1~debian.13~trixie
docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie
docker-compose-plugin 5.6.0-1~debian.13~trixie
@@ -0,0 +1,2 @@
signed: op=os_docker_step host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=cb4a902da3007b8f4c3a440d3f2d9a60 expires=2026-10-04T15:27:25Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,7 @@
time=2026-10-04T16:52:54.082+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=b7956e7339bfda31 op=os_docker_step key_id=felhom-op-1 nonce=cb4a902da3007b8f4c3a440d3f2d9a60
time=2026-10-04T16:52:55.697+02:00 level=INFO msg="osupdate: START" run=20261004T145254Z layer=docker vmid=9201 ring=1 trigger=signed enabled=true release=os-docker-20261004-142842
time=2026-10-04T16:53:07.878+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-docker-20261004-142842 layer=docker:9201 lane=slow mode=apply select=listed packages=6 authority=signed"
time=2026-10-04T16:53:07.878+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=6 not-installed=0 from-snapshot=0"
time=2026-10-04T16:53:07.878+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=0 upgraded=0 (nothing to do)"
time=2026-10-04T16:53:07.879+02:00 level=INFO msg="osupdate: DONE" run=20261004T145254Z layer=docker vmid=9201 ring=1 trigger=signed outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=10 reboot_needed=true wrapper_seconds=
time=2026-10-04T16:53:07.884+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=b7956e7339bfda31 op=os_docker_step
@@ -0,0 +1,2 @@
signed: op=os_docker_step host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=cb4a902da3007b8f4c3a440d3f2d9a60 expires=2026-10-04T15:38:36Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,3 @@
time=2026-10-04T16:52:54.082+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=b7956e7339bfda31 op=os_docker_step key_id=felhom-op-1 nonce=cb4a902da3007b8f4c3a440d3f2d9a60
time=2026-10-04T16:53:07.884+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=b7956e7339bfda31 op=os_docker_step
time=2026-10-04T17:07:54.007+02:00 level=WARN msg="signedjobs: REJECTED signed op — executor not called" job=dec707a162451c5b op=os_docker_step reason=rejected err="authz: replay (nonce already seen): nonce cb4a902da3007b8f4c3a440d3f2d9a60"
@@ -0,0 +1,2 @@
HTTP/1.1 303 See Other
Location: /system?flash=done
@@ -0,0 +1,16 @@
time=2026-10-04T16:39:43.360+02:00 level=INFO msg="audit: gate decision" class=os_docker_step host=demo-hp-bb76ea guest="" source=one_shot_job disposition=destructive allowed=true reason=signed key_id=felhom-op-1 nonce=f21cbf93… durable_id=""
time=2026-10-04T16:39:43.360+02:00 level=INFO msg="gate decision" class=os_docker_step guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
time=2026-10-04T16:39:43.360+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=c9f9efc098335bf9 op=os_docker_step key_id=felhom-op-1 nonce=f21cbf9302fb8b655b3ccf3e36f1cbea
os-apply: LIVE-RESTORE already on
time=2026-10-04T16:39:45.427+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: LIVE-RESTORE already on"
time=2026-10-04T16:39:45.427+02:00 level=INFO msg="osupdate: START" run=20261004T143943Z layer=docker vmid=9201 ring=0 trigger=signed enabled=true release=undo-to-29.7.2
os-apply: START release=undo-to-29.7.2 layer=docker:9201 lane=slow mode=apply select=listed packages=6 authority=signed UNDO
os-apply: PLAN upgrade=6 already=0 not-installed=0 from-snapshot=0
os-apply: DONE rc=0 seconds=16.7 upgraded=6 restart-needed=containerd-shim reboot-needed=no
time=2026-10-04T16:40:26.840+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=undo-to-29.7.2 layer=docker:9201 lane=slow mode=apply select=listed packages=6 authority=signed UNDO"
time=2026-10-04T16:40:26.840+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=6 already=0 not-installed=0 from-snapshot=0"
time=2026-10-04T16:40:26.840+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=16.7 upgraded=6 restart-needed=containerd-shim reboot-needed=no"
time=2026-10-04T16:41:02.662+02:00 level=INFO msg="osupdate: DONE" run=20261004T143943Z layer=docker vmid=9201 ring=0 trigger=signed outcome=applied healthy=true reason="" upgraded=6 pending=6 not_covered=0 restart_needed=1 reboot_needed=false wrapper_seconds=41.3
time=2026-10-04T16:41:02.716+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=c9f9efc098335bf9 op=os_docker_step
IDS IDENTICAL (24 containers)
after: engine=29.7.2 live=true
@@ -0,0 +1,8 @@
time=2026-10-04T16:42:55.380+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
guest nothing healthy=true
time=2026-10-04T16:43:10.212+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=0 already=0 not-installed=0 from-snapshot=0"
host nothing healthy=true
time=2026-10-04T16:43:58.115+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=6 already=0 not-installed=0 from-snapshot=0"
docker applied healthy=true
IDS IDENTICAL (24 containers)
engine=29.8.2
@@ -0,0 +1 @@
before: 24 containers engine=29.8.2
@@ -0,0 +1 @@
{"release_id":"undo-to-29.7.2","undo":true,"packages":[{"name":"containerd.io","version":"2.3.3-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-buildx-plugin","version":"0.36.1-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-ce","version":"5:29.7.2-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-ce-cli","version":"5:29.7.2-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-ce-rootless-extras","version":"5:29.7.2-1~debian.13~trixie","origin":"Docker CE"},{"name":"docker-compose-plugin","version":"5.5.0-1~debian.13~trixie","origin":"Docker CE"}]}
@@ -0,0 +1,2 @@
signed: op=os_docker_step host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=f21cbf9302fb8b655b3ccf3e36f1cbea expires=2026-10-04T15:15:26Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,29 @@
== crash 1 issued 2026-10-04T14:51:28Z
kernel.panic = 10
0defc390-c295-4e9f-aa6e-13d69a1f24c3
back: ping+ssh after 54s
7.0.14-20-pve
kernel.panic = 10
{
"armed": true,
"boot_id": "cc0aa40b-ac53-4524-a46f-6c5983ce5ed9",
"config": {
"LIMIT": 3,
"PANIC_SECONDS": 10,
"REARM_HOURS": 24,
"WINDOW_MINUTES": 60
},
"kernel_panic": 10,
"last_boot_at": "2026-10-04T14:52:08Z",
"last_boot_unclean": true,
"tripped": false,
"unclean_boots": [
"2026-10-04T14:52:08Z"
],
"unclean_boots_24h": 1,
"unclean_boots_in_window": 1,
"updated_at": "2026-10-04T14:52:18Z",
"version": 1
}
hub log since the boot:
2026/10/04 16:52:57 [INFO] host-report from demo-hp-bb76ea (1 guests, 4 storage targets, 0 backups, 2 restore-tests, 2 pbs-snapshots, 13538 bytes)
@@ -0,0 +1,33 @@
== crash 2 issued 2026-10-04T14:53:36Z
kernel.panic = 10
cc0aa40b-ac53-4524-a46f-6c5983ce5ed9
back: ping+ssh after 53s
7.0.14-20-pve
kernel.panic = 0
{
"armed": false,
"boot_id": "53f69dfb-93ef-44c3-a9ad-803d86dfb0b3",
"config": {
"LIMIT": 3,
"PANIC_SECONDS": 10,
"REARM_HOURS": 24,
"WINDOW_MINUTES": 60
},
"kernel_panic": 0,
"last_boot_at": "2026-10-04T14:54:33Z",
"last_boot_unclean": true,
"tripped": true,
"tripped_at": "2026-10-04T14:54:43Z",
"tripped_reason": "2 unclean boots within 60 minutes \u2014 the next crash leaves the box off (limit 3)",
"unclean_boots": [
"2026-10-04T14:52:08Z",
"2026-10-04T14:54:33Z"
],
"unclean_boots_24h": 2,
"unclean_boots_in_window": 2,
"updated_at": "2026-10-04T14:54:43Z",
"version": 1
}
hub log since the boot:
2026/10/04 16:52:57 [INFO] host-report from demo-hp-bb76ea (1 guests, 4 storage targets, 0 backups, 2 restore-tests, 2 pbs-snapshots, 13538 bytes)
2026/10/04 16:54:53 [INFO] host-report from demo-hp-bb76ea (1 guests, 4 storage targets, 0 backups, 2 restore-tests, 2 pbs-snapshots, 13395 bytes)
@@ -0,0 +1,4 @@
== crash 3 issued 2026-10-04T14:55:13Z
kernel.panic = 0
53f69dfb-93ef-44c3-a9ad-803d86dfb0b3
STAYED OFF for 184s (no ping)
@@ -0,0 +1,29 @@
== after the operator's power-on, ssh back 2026-10-04T15:17:41Z (27s after his word)
7.0.14-20-pve
kernel.panic = 0
{
"armed": false,
"boot_id": "182af911-22a9-41ec-97ac-7ee247a4d470",
"config": {
"LIMIT": 3,
"PANIC_SECONDS": 10,
"REARM_HOURS": 24,
"WINDOW_MINUTES": 60
},
"kernel_panic": 0,
"last_boot_at": "2026-10-04T15:17:26Z",
"last_boot_unclean": true,
"tripped": true,
"tripped_at": "2026-10-04T14:54:43Z",
"tripped_reason": "2 unclean boots within 60 minutes \u2014 the next crash leaves the box off (limit 3)",
"unclean_boots": [
"2026-10-04T14:52:08Z",
"2026-10-04T14:54:33Z",
"2026-10-04T15:17:26Z"
],
"unclean_boots_24h": 3,
"unclean_boots_in_window": 3,
"updated_at": "2026-10-04T15:17:36Z",
"version": 1
}
status: stopped
@@ -0,0 +1,12 @@
2026/10/04 17:21:00 [INFO] Operator email sent for demo-hp/app_start_failed
2026/10/04 17:26:30 [INFO] Operator email sent for demo-hp/app_stopped_unhealthy
2026/10/04 17:26:30 [INFO] Customer email sent to drill@felhom.eu for demo-hp/app_stopped_unhealthy
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_crash_restart
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_restarted_after_crash
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_crash_restart
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_restarted_after_crash
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_crash_restart
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_restarted_after_crash
2026/10/04 17:32:52 [WARN] host demo-hp-bb76ea crash guard: host_crash_guard_tripped
2026/10/04 17:32:52 [INFO] Operator email sent for demo-hp/host_crash_restart
2026/10/04 17:32:52 [INFO] Operator email sent for demo-hp/host_crash_guard_tripped
@@ -0,0 +1,240 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>System — Felhom Hub</title>
<link rel="stylesheet" href="/style.css?v=0.132.0">
<style>
.sys td, .sys th { white-space: nowrap; font-size: 0.82em; vertical-align: top; }
.sys .grp { border-left: 2px solid var(--border, #444); }
.c-warn { color: var(--warn); font-weight: 600; }
.c-bad { color: var(--danger, #e5534b); font-weight: 700; }
.sys form { display: inline; }
.rel-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(16rem, 1fr)); gap: 0.75rem; }
</style>
</head>
<body>
<svg xmlns="http://www.w3.org/2000/svg" style="display:none" aria-hidden="true">
<symbol id="i-triangle-alert" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m21.73 18-8-14a2 2 0 0 0-3.48 0l-8 14A2 2 0 0 0 4 21h16a2 2 0 0 0 1.73-3" /> <path d="M12 9v4" /> <path d="M12 17h.01" /></symbol>
<symbol id="i-check" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M20 6 9 17l-5-5" /></symbol>
<symbol id="i-server" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="20" height="8" x="2" y="2" rx="2" ry="2" /> <rect width="20" height="8" x="2" y="14" rx="2" ry="2" /> <line x1="6" x2="6.01" y1="6" y2="6" /> <line x1="6" x2="6.01" y1="18" y2="18" /></symbol>
<symbol id="i-settings" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M9.671 4.136a2.34 2.34 0 0 1 4.659 0 2.34 2.34 0 0 0 3.319 1.915 2.34 2.34 0 0 1 2.33 4.033 2.34 2.34 0 0 0 0 3.831 2.34 2.34 0 0 1-2.33 4.033 2.34 2.34 0 0 0-3.319 1.915 2.34 2.34 0 0 1-4.659 0 2.34 2.34 0 0 0-3.32-1.915 2.34 2.34 0 0 1-2.33-4.033 2.34 2.34 0 0 0 0-3.831A2.34 2.34 0 0 1 6.35 6.051a2.34 2.34 0 0 0 3.319-1.915" /> <circle cx="12" cy="12" r="3" /></symbol>
<symbol id="i-x" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M18 6 6 18" /> <path d="m6 6 12 12" /></symbol>
<symbol id="i-info" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 16v-4" /> <path d="M12 8h.01" /></symbol>
<symbol id="i-hard-drive" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M10 16h.01" /> <path d="M2.212 11.577a2 2 0 0 0-.212.896V18a2 2 0 0 0 2 2h16a2 2 0 0 0 2-2v-5.527a2 2 0 0 0-.212-.896L18.55 5.11A2 2 0 0 0 16.76 4H7.24a2 2 0 0 0-1.79 1.11z" /> <path d="M21.946 12.013H2.054" /> <path d="M6 16h.01" /></symbol>
<symbol id="i-cpu" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 20v2" /> <path d="M12 2v2" /> <path d="M17 20v2" /> <path d="M17 2v2" /> <path d="M2 12h2" /> <path d="M2 17h2" /> <path d="M2 7h2" /> <path d="M20 12h2" /> <path d="M20 17h2" /> <path d="M20 7h2" /> <path d="M7 20v2" /> <path d="M7 2v2" /> <rect x="4" y="4" width="16" height="16" rx="2" /> <rect x="8" y="8" width="8" height="8" rx="1" /></symbol>
<symbol id="i-clock" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="10" /> <path d="M12 6v6l4 2" /></symbol>
<symbol id="i-boxes" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M2.97 12.92A2 2 0 0 0 2 14.63v3.24a2 2 0 0 0 .97 1.71l3 1.8a2 2 0 0 0 2.06 0L12 19v-5.5l-5-3-4.03 2.42Z" /> <path d="m7 16.5-4.74-2.85" /> <path d="m7 16.5 5-3" /> <path d="M7 16.5v5.17" /> <path d="M12 13.5V19l3.97 2.38a2 2 0 0 0 2.06 0l3-1.8a2 2 0 0 0 .97-1.71v-3.24a2 2 0 0 0-.97-1.71L17 10.5l-5 3Z" /> <path d="m17 16.5-5-3" /> <path d="m17 16.5 4.74-2.85" /> <path d="M17 16.5v5.17" /> <path d="M7.97 4.42A2 2 0 0 0 7 6.13v4.37l5 3 5-3V6.13a2 2 0 0 0-.97-1.71l-3-1.8a2 2 0 0 0-2.06 0l-3 1.8Z" /> <path d="M12 8 7.26 5.15" /> <path d="m12 8 4.74-2.85" /> <path d="M12 13.5V8" /></symbol>
<symbol id="i-users" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M16 21v-2a4 4 0 0 0-4-4H6a4 4 0 0 0-4 4v2" /> <path d="M16 3.128a4 4 0 0 1 0 7.744" /> <path d="M22 21v-2a4 4 0 0 0-3-3.87" /> <circle cx="9" cy="7" r="4" /></symbol>
</svg>
<script>
function felhomConfirm(el,question,onYes){
if(!el||el.dataset.fcOpen)return;
el.dataset.fcOpen='1';
var wrap=document.createElement('span');wrap.className='inline-confirm';
var q=document.createElement('span');q.className='inline-confirm-q';q.textContent=question;
var yes=document.createElement('button');yes.type='button';yes.className='btn btn-sm btn-danger';yes.textContent='Igen';
var no=document.createElement('button');no.type='button';no.className='btn btn-sm btn-outline';no.textContent='Mégse';
wrap.appendChild(q);wrap.appendChild(yes);wrap.appendChild(no);
el.style.display='none';el.parentNode.insertBefore(wrap,el.nextSibling);
function close(){wrap.remove();el.style.display='';delete el.dataset.fcOpen;}
no.addEventListener('click',close);
yes.addEventListener('click',function(){close();onYes();});
}
document.addEventListener('click',function(e){
var btn=e.target.closest?e.target.closest('[data-confirm]'):null;
if(!btn)return;
e.preventDefault();
felhomConfirm(btn,btn.getAttribute('data-confirm'),function(){
var form=btn.closest('form');
if(form){if(form.requestSubmit)form.requestSubmit(btn);else form.submit();}
});
});
</script>
<div class="container">
<header>
<h1>Felhom <span>Hub</span></h1>
<nav class="nav-links">
<a href="/" class="nav-link">Dashboard</a>
<a href="/configs" class="nav-link">Customers</a>
<a href="/apps" class="nav-link">Apps</a>
<a href="/hosts" class="nav-link">Hosts</a>
<a href="/system" class="nav-link active">System</a>
<a href="/offsite" class="nav-link">Offsite</a>
<a href="/configuration" class="nav-link">Configuration</a>
</nav>
</header>
<h2 style="margin-bottom: 1rem;">System — versions and OS updates</h2>
<section class="card" style="margin-bottom: 1.5rem;">
<h3 style="margin-top: 0;">Approved releases</h3>
<div class="rel-grid">
<div><strong>guest</strong>: <code>os-guest-20261004-123933</code><br>
<span class="text-muted">272 packages · 2026-10-04 12:39 UTC · by auto</span></div>
<div><strong>host</strong>: <code>os-host-20261004-124133</code><br>
<span class="text-muted">607 packages · 2026-10-04 12:41 UTC · by auto</span></div>
<div><strong>docker</strong>: <code>os-docker-20261004-142842</code><br>
<span class="text-muted">6 packages · 2026-10-04 14:28 UTC · by operator</span></div>
</div>
<h3>What ring 0 runs now</h3>
<div class="rel-grid">
<div><strong>guest</strong>:
272 packages, first seen 2026-10-04 11:07 UTC<br>
<span class="text-muted">approved as os-guest-20261004-123933</span>
</div>
<div><strong>host</strong>:
607 packages, first seen 2026-10-04 12:24 UTC<br>
<span class="text-muted">approved as os-host-20261004-124133</span>
</div>
<div><strong>docker</strong>:
6 packages, first seen 2026-10-04 14:28 UTC<br>
<span class="text-muted">approved as os-docker-20261004-142842</span>
</div>
</div>
<form method="POST" action="/os/approve-now" style="margin-top: 0.8rem;">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<button type="submit" class="btn btn-sm btn-danger" data-confirm="Approve the guest and host sets ring 0 runs NOW, without the 24 h + 1 night wait? Every ring-1 box installs them at its next night run.">Approve now (guest + host)</button>
<span class="text-muted" style="font-size: 0.85em;">An urgent fix only — normally the hub approves after 24 h and one night.</span>
</form>
</section>
<section class="card" style="padding: 0; overflow-x: auto;">
<table class="data-table sys">
<thead>
<tr>
<th>Box</th><th>Ring / updates</th><th>Tunnel</th>
<th class="grp">Proxmox</th><th>Kernel (running)</th><th>Kernel (next boot)</th><th>Debian</th><th>Felhom release</th><th>Pending</th><th>Not covered</th><th>Held</th><th>Reboot needed</th><th>kernel.panic</th><th>Oops</th><th>Crash restarts 24 h</th><th>Crash guard</th>
<th class="grp">Guest Debian</th><th>Felhom release</th><th>Pending</th><th>Restart needed</th>
<th class="grp">Docker</th><th>containerd</th><th>live-restore</th><th>Docker release</th>
<th class="grp">Last OS leg</th>
</tr>
<tr class="text-muted"><th></th><th></th><th></th><th class="grp" colspan="13">host</th><th class="grp" colspan="4">guest</th><th class="grp" colspan="4">Docker engine</th><th class="grp"></th></tr>
</thead>
<tbody>
<tr>
<td><a href="/hosts/demo-felhom-8363b5">demo-felhom-8363b5</a><br><span class="text-muted">Demo Ügyfél</span>
</td>
<td>
ring 0
<form method="POST" action="/os/ring/demo-felhom-8363b5">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="1"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make demo-felhom-8363b5 a normal (ring 1) box? It then installs only approved releases.">→ normal</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/demo-felhom-8363b5">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for demo-felhom-8363b5? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td>running</td>
<td class="grp ">9.2.2</td>
<td>7.0.2-6-pve</td><td>7.0.2-6-pve</td><td>13.7</td>
<td>os-host-20261004-124133</td><td>80</td><td>0</td>
<td>none</td><td class="c-warn">since 2026-10-04</td><td>10 s</td><td>no</td>
<td>0</td><td>armed</td>
<td class="grp ">13.7</td>
<td>os-guest-20261004-123933</td><td>0</td><td>10</td>
<td class="grp ">29.8.2</td>
<td>2.3.6-1~debian.13~trixie</td><td>on</td><td>os-docker-20261004-142842</td>
<td class="grp " title="last successful leg: 51 min ago">40 min ago · nothing · 12 s</td>
</tr>
<tr>
<td><a href="/hosts/demo-hp-bb76ea">demo-hp-bb76ea</a><br><span class="text-muted">Demo HP</span>
</td>
<td>
ring 0
<form method="POST" action="/os/ring/demo-hp-bb76ea">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="1"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make demo-hp-bb76ea a normal (ring 1) box? It then installs only approved releases.">→ normal</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/demo-hp-bb76ea">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for demo-hp-bb76ea? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td>running</td>
<td class="grp ">9.2.2</td>
<td>7.0.14-20-pve</td><td>7.0.14-20-pve</td><td>13.7</td>
<td>ring0-20261004T144238Z</td><td>78</td><td>0</td>
<td>none</td><td>no</td><td class="c-warn">0 (stays off)</td><td>no</td>
<td class="c-warn">3</td><td class="c-bad" title="2 unclean boots within 60 minutes — the next crash leaves the box off (limit 3)">TRIPPED 2026-10-04T14:54:43Z</td>
<td class="grp ">13.7</td>
<td>ring0-20261004T144238Z</td><td>6</td><td>1</td>
<td class="grp ">29.8.2</td>
<td>2.3.6-1~debian.13~trixie</td><td>on</td><td>ring0-20261004T144238Z</td>
<td class="grp " title="last successful leg: 50 min ago">48 min ago · applied · 45 s</td>
</tr>
<tr>
<td><a href="/hosts/drill-r50-0a4f9a">drill-r50-0a4f9a</a><br><span class="text-muted">drill-r50</span>
<br><span class="c-warn" style="font-weight: normal;">no versions reported (agent older than v0.142.0)</span></td>
<td>
ring 1
<form method="POST" action="/os/ring/drill-r50-0a4f9a">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="ring" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Make drill-r50-0a4f9a a DEMO (ring 0) box? It then installs every new fix first and takes unsigned Docker steps if its root-owned ring-0 mark allows.">→ demo</button>
</form><br>
updates <strong>ON</strong>
<form method="POST" action="/os/enabled/drill-r50-0a4f9a">
<input type="hidden" name="_csrf" value=""><input type="hidden" name="return" value="/system">
<input type="hidden" name="on" value="0"><button type="submit" class="btn btn-sm btn-outline" data-confirm="Switch OS updates OFF for drill-r50-0a4f9a? It keeps reporting and installs nothing.">switch off</button>
</form>
</td>
<td class="c-bad">inactive</td>
<td class="grp c-warn">unknown</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td>
<td>—</td><td>0</td><td>0</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>no</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>no</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn">not installed</td>
<td class="grp c-warn">unknown</td>
<td>—</td><td>0</td><td>0</td>
<td class="grp c-warn">unknown</td>
<td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td class="c-warn" title="the box could not read it (agent older than v0.142.0, or the guest is down)">unknown</td><td>—</td>
<td class="grp " title="last successful leg: never">never · — · 0 s</td>
</tr>
</tbody>
</table>
</section>
<p class="text-muted" style="font-size: 0.85em;">Amber: worth a look. Red: an operator alarm fires (`08` §6.3). "unknown": the box could not read the value — never a guess.</p>
<footer style="margin-top: 2rem; color: var(--text-muted); font-size: 0.8rem; text-align: center;">
Felhom Hub <span style="font-family: var(--font-mono)">0.132.0</span>
</footer>
</div>
</body>
</html>
@@ -0,0 +1,3 @@
crash-guard: RE-ARMED by operator (was tripped: True); kernel.panic=10
kernel.panic = 10
{'armed': True, 'tripped': False, 'rearmed_at': '2026-10-04T15:33:25Z', 'rearmed_by': 'operator', 'unclean_boots_in_window': 0, 'unclean_boots_24h': 3, 'last_trip': {'at': '2026-10-04T14:54:43Z', 'reason': '2 unclean boots within 60 minutes — the next crash leaves the box off (limit 3)'}}
+13
View File
@@ -26,6 +26,19 @@
---
## 2026-10-04 (evening) — System page, Docker slow lane, crash guard (agent v0.142.0, hub v0.132.0, installer 1.30.0)
> Evidence: `audits/os-docker-crash-2026-10-04/`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-852** | **The operator could see no box's Debian, Proxmox, kernel or Docker version anywhere.** The box reports them (`system` stanza: Proxmox + kernel from the API, the wrapper's read-only `facts`); hub v0.132.0 shows them on the new **System** page (with the ring / switch / approve buttons) and a Proxmox / kernel column on Hosts. **Reasoning kept: a value the box could not read says `unknown`, amber, with the reason — never empty, never guessed.** | CLOSED 2026-10-04 — FIXED (`09` decision 89) | `partA/` |
| **R-835** | **Turning Docker's live-restore OFF by a restart stops every container and starts none.** Resolved by design (decision 87): live-restore is ON everywhere — the golden bakes it, installed boxes get it once by a RELOAD (measured: same ids on 9202 6/6, demo-hp 24/24, demo-felhom 5/5), and the wrapper refuses a Docker step while it is off (R15). **Rule kept: never turned off by a plain restart.** | CLOSED 2026-10-04 — RESOLVED BY DESIGN | `partB/b1..b4` |
| **R-848** | **A held host package was invisible to the hub.** The facts report `held`; the System page shows it amber. | CLOSED 2026-10-04 — FIXED | `partA/` |
| **R-849** | **The guest's "restart needed since" never cleared.** The wrapper scans the guest on every pass too (agent v0.142.0). | CLOSED 2026-10-04 — FIXED | `partB/agent-redproofs.txt` |
| **R-851** | **A host that panicked stayed stopped (`kernel.panic = 0`).** The crash guard (decision 88): `kernel.panic = 10`; the 3rd unclean stop within 60 minutes leaves the box off; re-arms after 24 h or by `felhom-crash-guard rearm`. MEASURED on demo-hp: crash 1 and 2 restarted by themselves (54 s, 53 s), the guard tripped, crash 3 stayed off until the operator switched it on; the hub mailed the trip. **Reasoning kept: a crash, a power cut and a hard reset cannot be told apart on these boxes (pstore saved nothing for a real panic) — the guard counts every unclean stop.** | CLOSED 2026-10-04 — FIXED | `partC/` |
| **R-854** | **The first Docker steps were stored as GUEST reports** — they ran (agent rc) before hub v0.132.0 was live, and the older hub maps an unknown layer to `guest`; the guest candidate briefly read 0 packages. One-time; a pass with the released agent re-reported both layers. Nothing to fix. | CLOSED 2026-10-04 — ONE-TIME, NO ACTION | `partB/b5-*` |
## 2026-10-04 (afternoon) — OS updates, host fast lane + fleet view + alarms (agent v0.141.0/v0.141.1, hub v0.131.0/v0.131.1, controller v0.292.0)
> Evidence: `audits/os-host-lane-2026-10-04/`.
+8 -10
View File
@@ -302,13 +302,13 @@ stopping line that lies.
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
## Box system & updates — 21 rows (P2 4, P3 14, P4 3)
## Box system & updates — 20 rows (P2 4, P3 13, P4 3)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator |
| **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC |
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
| **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/<path>` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC |
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
@@ -323,13 +323,12 @@ stopping line that lies.
| **R-184** | Box system & updates | P4 | **Nothing prevents the hub from vouching an agent version that was never released.** The R-115 gate proves every RELEASED version is installable, but it works from git tags — so a hub artifact-manifest entry naming a version with no tag and no package is invisible to it. The installer would then die at step 5 on a virgin machine, as root | **READY (S) — NEW 2026-08-03** | — | **Filed BECAUSE the R-115 gate deliberately does not cover it, rather than leaving the gap unstated.** CI cannot check it: the hub's `/api/v1/artifacts/<customer>` answers **401** without a per-customer retrieval passphrase and Gitea's package **listing** api answers **401** without a token (both measured 2026-08-03, P-C), so a credential-free gate can ask *"is this version installable"* but never *"which version is vouched"*. **Two shapes, and the second is better:** (a) give CI a hub credential — expands what CI can reach, and is the operator's call not a gate author's; (b) **validate at vouch time, in the hub**: the operator UI's Day-0 artifact form refuses a version whose package is not downloadable. (b) fails closed at the moment of the decision, needs no new credential anywhere, and puts the check where the mistake is actually made. **Exposure is low and should be said so:** vouching is a deliberate operator action against a version they have just released, and R-115's release path now makes released-but-unpublished nearly impossible. This is the residue, not the main risk | CC |
| **R-194** | Box system & updates | P4 | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
| **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC |
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC |
| **R-851** | Box system & updates | P3 | **A host that panics stays stopped: `kernel.panic = 0`.** READ 2026-10-04 on demo-hp (os-host-lane Part E): `kernel.panic = 0`, `kernel.panic_on_oops = 0`. The brief assumed the host restarts after a panic; it does not — the box stays down until a person power-cycles it, and only `softdog` (dead in a panic) runs. A fix (`kernel.panic = 10` set by the installer, maybe `panic_on_oops`) changes how every box behaves and what a household sees, so it is the operator's call, together with the kernel lane. `audits/os-host-lane-2026-10-04/partE/e0-readonly-demo-hp.txt` | **READY — operator decision (with the kernel lane); owner: operator** | — | — | operator |
| **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC |
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC.** 2026-10-04 (evening): the demo boxes got the v0.142.0 wrapper, the crash guard and the two root-owned trust files BY HAND again (`audits/os-docker-crash-2026-10-04/partD/`). **RULED 2026-10-04 ~12:20 (`09` §3 decision 82): not built now — "There are no older boxes"; no installed box exists outside the two demo boxes, which take wrapper changes by hand. Kept open, not re-ranked. Reviewer's note (as written): every box installed from now on becomes an "older box" the first time the wrapper or the sudoers line changes again (the 2026-10-04 host-lane brief changes the wrapper); R-840 must be solved before the first box that cannot be reached by hand.** | — | — | CC |
## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6)
## Monitoring & notifications — 25 rows (P2 2, P3 16, P4 7)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -357,8 +356,9 @@ stopping line that lies.
| **R-337** | Monitoring & notifications | P4 | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC |
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator |
## Hub & operator — 25 rows (P2 1, P3 8, P4 16)
## Hub & operator — 23 rows (P2 1, P3 7, P4 15)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
@@ -384,9 +384,7 @@ stopping line that lies.
| **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator |
| **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator |
| **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC |
| **R-848** | Hub & operator | P4 | **A held host package is invisible to the hub.** MEASURED 2026-10-04 on demo-hp (undo runbook proof): with `tzdata` held after a by-hand undo, the wrapper's `pending` stayed 78 — apt's simulation leaves held packages out, so the fleet view shows nothing and no alarm can see a hold that was forgotten. The hold lives only in the incident's register row (`runbooks/os-updates-host-undo.md`). Fix direction: the wrapper reports `apt-mark showhold` and the fleet line shows it. | **READY — owner: CC** | — | — | CC |
| **R-849** | Hub & operator | P4 | **The fleet view's GUEST "reboot needed since" never clears.** 2026-10-04: the guest is scanned only after an install (R-845, one `pct exec`), so a guest restart is never seen; the guest line keeps the date of the last install that said "needed". No alarm reads the guest line (the reboot alarm is host-only), so it is display only. Fix direction: scan the guest on every pass too (one `pct exec`, ~1 s) or hide the guest date. `audits/os-host-lane-2026-10-04/partC/live/fleet-during-ring1-test.json` | **READY — owner: CC** | — | — | CC |
| **R-852** | Hub & operator | P3 | **The operator cannot see any box's Debian, Proxmox, kernel or Docker version anywhere.** FOUND 2026-10-04 (operator question: *"Where should I be able to see the OS/Proxmox versions of the boxes?"*): no box reports them — the agent's host report carries no Debian version, no `pveversion`, no running or next-boot kernel, no guest Debian or Docker engine version (`felhom-agent/internal/proxmox/types.go:35` has a `PVEVersion` field that nothing reads); the OS fleet view is `GET /os/fleet` JSON only, with no page, no tab and no buttons for ring, switch or "approve now"; its "release" is Felhom's OS release id, not a Debian or Proxmox version. `09` §3 decision 89. | **READY — owner: CC (this session, Part A)** | — | — | CC |
| **R-855** | Hub & operator | P4 | **The hub's start log prints "after -1 healthy ring-0 night(s)" for the TEST override `OS_DOCKER_APPROVE_NIGHTS=0`** (the internal "none" value; the WARN line before it is right). Cosmetic, TEST configuration only. Fix: print 0. `audits/os-docker-crash-2026-10-04/partB/` (hub log 16:27:42) | **READY — owner: CC** | — | — | CC |
## Business & legal — 7 rows (P2 4, P4 3)
+19
View File
@@ -0,0 +1,19 @@
# The crash guard — read it, re-arm it
> **Design:** `architecture/11-os-updates.md` §5.9 (decision 88). **When:** the hub mailed `host_crash_guard_tripped`,
> or the System page shows **TRIPPED** for a box.
A tripped guard means: the box stopped uncleanly twice within an hour (a crash, a power cut or a hard reset), so it set
`kernel.panic = 0` — **the next crash leaves it off** until someone switches it on. It re-arms by itself after 24 h of
normal running.
1. **Read why** (as root on the box's host): `felhom-crash-guard status` — `unclean_boots`, `tripped_at`,
`tripped_reason`. Then the end of each crashed boot: `journalctl --list-boots` and `journalctl -b -1 -n 50` (a crashed
boot ends with no shutdown lines). `journalctl -k -b -1 | grep -iE "panic|oops|BUG:"` for a kernel message.
2. **Fix the cause first** if you found one (a bad kernel → boot the previous one; a power problem → the PSU, the cable).
3. **Re-arm** (operator's choice): `felhom-crash-guard rearm` → `kernel.panic = 10` now; the history stays; a fresh
60-minute window starts. The hub announces `host_crash_guard_rearmed` with the next report that carries the facts
(up to ~15 min, R-853).
4. **Check:** `sysctl kernel.panic` = 10; the System page shows "armed".
The numbers live in `/etc/felhom/crash-guard.conf` (LIMIT, WINDOW_MINUTES, PANIC_SECONDS, REARM_HOURS).
@@ -0,0 +1,54 @@
# Put a box's Docker engine back one set (OS updates, Docker slow lane)
> **When:** the hub mailed `os_update_health_failed` for the **docker** layer, or the System page shows a box whose
> apps broke after a Docker engine step.
> **Who:** the operator, or CC with the operator's word (CC may sign until the first paying customer, R-530 ruling).
> **Owner design:** `architecture/11-os-updates.md` §5.8. **Proved** on demo-hp 2026-10-04 —
> evidence `audits/os-docker-crash-2026-10-04/partB/undo/`.
The undo is an ordinary **signed `os_docker_step` with `"undo": true`** that names the previous engine set. The root
wrapper re-verifies the signature itself (against the root-owned `/etc/felhom/operator-signers`), allows the downgrade
only because the signed job says `undo`, and refuses unless `live-restore` is on — so the undo, like the step, restarts
no app. Nothing is done by hand on the box.
## 1. Find the previous set
- The box's previous Docker report (System page → Docker release column, or the hub's `os_reports` for the box, layer
`docker`): its `installed` list before the bad step. Or, on the host: `pct exec <vmid> -- grep -A3 "Start-Date"
/var/log/apt/history.log | tail` — each line `name:amd64 (old, new)`.
- All **six** names, each with its old version. Docker's repository keeps old versions (`11` C2), so no snapshot is
needed.
## 2. Sign and queue it
```bash
cd /mnt/5_hdd/felhom.eu/git/felhom-agent && go build -o /tmp/felhom-opsign ./cmd/felhom-opsign
P='{"release_id":"undo-to-<engine>","undo":true,"packages":[
{"name":"containerd.io","version":"<old>","origin":"Docker CE"},
{"name":"docker-buildx-plugin","version":"<old>","origin":"Docker CE"},
{"name":"docker-ce","version":"<old>","origin":"Docker CE"},
{"name":"docker-ce-cli","version":"<old>","origin":"Docker CE"},
{"name":"docker-ce-rootless-extras","version":"<old>","origin":"Docker CE"},
{"name":"docker-compose-plugin","version":"<old>","origin":"Docker CE"}]}'
KEY=$(sudo kubectl -n felhom-system get secret report-api -o jsonpath='{.data.REPORT_API_KEY}' | base64 -d)
/tmp/felhom-opsign -op os_docker_step -host <host_id> -key-id felhom-op-1 \
-key /mnt/5_hdd/felhom.eu/felhom-op-operational -params "$P" -ttl 45m \
-upload http://<hub ClusterIP>:8080 -hub-key "$KEY"
unset KEY
```
The box takes the job at its next poll (up to 15 minutes), under the heavy-op gate (never beside a backup).
## 3. Check
- Agent journal: `signedjobs: AUTHORIZED signed op` → `os-apply: START … layer=docker:… authority=signed UNDO` →
`signedjobs: signed op COMPLETED`. A `REFUSED: R3` means the signature, host, time window, nonce or package list did
not match; `R15` means live-restore is off.
- The hub: a docker-layer report, `outcome applied`, `healthy`, `undo: true`, the old engine.
- **The same container ids before and after** (the health rule checks it; a changed id is `health_failed`).
## Afterwards
On a **ring-0** box the next night step installs the newest set again — switch the box's updates OFF on the System page
first if that set is the bad one, and record it in the incident's register row. A **ring-1** box takes a set only by a
signed job, so it stays on the undone set until someone signs the next one.