Day 2026-10-07: Parts A-H done — hub 0.142.0 + agent 0.151.0 delivered, Proxmox lane proven on demo-felhom, kernel spike recorded; 130 -> 126
gates / gates (push) Successful in 3m3s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-07 14:19:58 +02:00
parent 374213f279
commit b0276c7197
23 changed files with 836 additions and 33 deletions
+7
View File
@@ -16,6 +16,13 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-07 (day) — the operator's answers built; the kernel spike (`09` §3 162–171).** Hub 0.142.0 + agent 0.151.0
> (+ bundle) on the three boxes: the Proxmox package lane (R-812 A, proven on demo-felhom), the controller-image verb
> (R-861 a+b), RESET's purge (R-32; the one-time clean-up waits for a main-account login), the other-key line (R-366),
> the retired DR fields (R-105). Agent vouch refused by the hub until the golden catches up. Kernel spike: ESP one-shot
> and BootNext pass on all three boxes; a systemd-armed watchdog fails (freeze needs a person) — design
> `audits/kernel-spike-2026-10-07/`. Register 130 → 126.
> **2026-10-07 (morning) — read-back B + D, the releases, the Tester 1 proof (`09` §3 160–161).** R-518 read back (night
> stop 91 s on demo-hp, one press 80 s; the page text holds; the two-tier night ~2026-10-08 is left). Released + delivered
> to demo-hp, demo-felhom, Tester 1: hub 0.141.0, agent 0.150.0 (+ bundle, probe 68/68), controller 0.302.0 (floors with
+29 -24
View File
@@ -1,33 +1,38 @@
# REPORT — the read-back, the releases and the Tester 1 proof (2026-10-07 morning)
Operator instruction 2026-10-07 08:40: Parts B and D of `TASK-night-readback-offsite-gap-tester1-2026-10-07.md`, then
release hub 0.141.0 → agent 0.150.0 → controller → the held catalog branch, deliver to demo-hp, demo-felhom and the
Tester 1 box, then the Tester 1 update test with the operator present; rule 11 in the shared rule file. Recorded as
`09` §3 decisions 160–161. **The task file is not on DooPlex** (searched the disk and git); Parts B and D were taken from
the evening report and the R-518 row.
# REPORT — the operator's answers built; the kernel lane spiked with real reboots (2026-10-07 day)
| Part | Result |
|---|---|
| **B** — the night read-back (R-518) | demo-hp (9 apps) stop **~91 s** (was 5 min 47 s); demo-felhom (1 app) ~11 s; local tier only, off-site not due; two channels each. `audits/readback-2026-10-07/RESULT-B-D.md` |
| **D** — one press, measured | demo-hp: **80 s** from the press to the last app (per app 39–79 s); the copy finished 4 min later with the apps up; only the local tier ran; the page's „kb. 1–1,5 perc" holds — no text change |
| hub 0.141.0 | built, manifest `011a481a`, Synced/Healthy at HEAD, image 0.141.0, `starting` log, healthz + System 200 at 06:57:42Z (40 s after the sync) |
| agent 0.150.0 | released (sha `a23d1c90…`, bundle `88456b38…`, tag `v0.150.0`), vouched (golden 0.301.0 and MinAgent 0.131.0 unchanged); signed `agent_update` → all three on 0.150.0 by 07:07Z; signed `agent_config_update` → `BUNDLE DONE written=1 same=24 self-check=ok`, probe 68/68 on all three by 07:21Z |
| controller 0.302.0 | MinAgent 0.131.0; floors 0.302.0 with declared MinAgent for demo-hp, demo-felhom, tester-1 → all three `0.302.0 (healthy)` by 07:10Z. Global floor and Tester 2 not touched. Golden waiver issued to 2026-10-13 (the weekly bake's date) |
| catalog | `night-held-2026-10-06` merged (`872039d`); synced on all three boxes (read in their templates) |
| **Tester 1 proof (R-892)** | vaultwarden 1.36.0 → 1.37.4 through the guarded Update: **proven, 14.4 s**, seed read back before and after; removed with its volume; app list equal; pointer restored byte-identical; drill reset. `audits/night-burndown-2026-10-06/s3/box/` |
| rule 11 | „Every helper prompt carries the brief's fences in full" — in all five copies (one md5) |
| **A** — rulings | `09` §3 decisions 162–170 recorded first, plus 171 (reboots without a per-reboot word, never DooPlex). R-822 closed by ruling. R-836 → P2. Shared rule file: no hub build/deploy without the operator (all five copies identical) |
| **B** — Proxmox package lane (R-812 A) | built (agent `pve` layer + `/etc/pve` write gate; hub candidate/approval/System page), delivered, **proven live on demo-felhom**: one signed step, 65 packages in 70 s, `pve-manager` 9.2.2 → 9.2.21, healthy, guest untouched, app 31/31; hub logged it. R-812 closed |
| **C** — agent image check (R-861 (a) A1, (b) B2) | built, delivered to the three boxes; on demo-hp: no `tee` grant, the verb is the route, an `alpine` ref refused (rc 3, file unchanged), felhom-op's pct lines exact. **Not seen: a full managed swap** (no newer controller today) |
| **D** — RESET purge + clean-up (R-32) | purge through the sub-account's own login built and delivered (hub 0.142.0). **The one-time clean-up STOPPED:** ~2.7 GB belongs to no live customer, reachable only by the pool box's main account — no such login on DooPlex |
| **E** — other-key archives line (R-366 slice 2) | built and delivered; R-366 closed |
| **F** — retire the empty fields (R-105) | built and delivered (columns kept, unread); `05`/`06` corrected; R-105 closed |
| **G** — kernel spike (R-836) | 24 reboots + 2 power cycles on Tester 1 (VM), demo-felhom, demo-hp (Secure Boot on). **ESP one-shot flag: works on all three. UEFI BootNext: works on all three. A watchdog armed by systemd: fails on all three** (a reset clears the timer; a frozen kernel needs a person). Design with two operator questions: `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`. Boxes left clean on a healthily booted default |
| **H** — Tester 2 (read only) | the household got 0 mails while the box was off (operator 10); no row |
| Rows before | Rows after | Opened | Closed |
|---|---|---|---|
| **137** | **130** | **0** | **7** (R-892, R-894, R-542, R-777, R-612, R-805, R-806) |
| **130** | **126** | **0** | **4** (R-822, R-812, R-366, R-105) |
**Fixed without a row:** none. **One slip:** the catalog merge was pushed in the same command as its unit-test run,
before reading the result (standing rule 1). Read right after: 156 tests, the single known error (`test_pg_conversion`
is a script, not a test module — the same on `main` before the merge); `catalog_gates.py --fast` was green before the push.
**Releases (one per repo changed):** hub **0.142.0** (live, healthz 200 within 50 s); agent **0.151.0** + bundle on demo-hp,
demo-felhom, Tester 1 (probe 68/68 each). **The agent vouch was refused by the hub** (`golden_behind_fleet`: golden 0.301.0
is behind the fleet's controller 0.302.0), so new installs keep agent 0.150.0 until the weekly bake. No controller release
(no controller code changed). CI: one lost job (felhom.eu run 1497, every step failed with no log) re-run once → success
(R-887 count +1).
**Instruction-file edits:** `.claude/rules/unprompted-work.md` §3 item 11 added in all five copies (operator ruling,
decision 160).
**Helpers:** each prompt carried the brief's fences in full (rule 11). **Security note** on `hub/internal/api/handler.go`
(raised by a background review, no detail): the last three changes log box/customer ids and counts only — nothing fixed.
**Teardown:** Tester 1 — vaultwarden removed through the product (no container, no volume), the saved
`controller.yaml.pre-r892` shredded, pointer restored identical; demo-hp — the one press made a normal local copy (kept
by retention), nothing else; hub — nothing provisioned. The night report stays in `REPORT-night-burndown-2026-10-06.md`.
**Teardown:** Tester 1 — spike entries, ESP flag, BootNext loader/entry and watchdog config removed; VM 341's test
watchdog device removed; default 7.0.2-6 (running). demo-felhom — the same; default 7.0.2-6; Proxmox userspace now 9.2.21.
demo-hp — the same; default 7.0.14-20. 7.0.14-20 is installed on Tester 1 and demo-felhom (booted healthily in the
spike) but not their default. Hub — nothing provisioned.
## Two decisions for you
1. **The kernel lane's promise:** may a box restart at night for a kernel update, and do we tell the household? My pick:
yes, with one line in the household's mail the day before. If you do nothing: kernels stay manual.
2. **The pool box clean-up (~2.7 GB):** give me the pool box's main-account login for one session, or delete the
folders yourself in the Hetzner console. My pick: you delete them in the console (no new credential on DooPlex). If
you do nothing: the leftovers stay; nothing reads them.
+18 -2
View File
@@ -2,8 +2,24 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-07 09:35: hub 0.141.0; demo-hp, demo-felhom and Tester 1 run controller 0.302.0, agent 0.150.0. The
open-items list is at 130. Report: `REPORT.md`.**
**Updated 2026-10-07 14:30: hub 0.142.0; demo-hp, demo-felhom and Tester 1 run agent 0.151.0 and controller 0.302.0.
The open-items list is at 126. Report: `REPORT.md`.**
## Day (2026-10-07): your answers built, and the kernel test with real reboots
- **Proxmox now updates its own programs** after your approval of each set: tested on demo-felhom (65 packages, 70 s,
the apps never stopped).
- **The box's helper can start only our own controller image** now; a strange image is refused.
- **RESET now deletes the household's off-site folder.** The old leftovers (~2.7 GB) are NOT deleted yet: that needs the
pool box's main login, which this server does not have.
- **After a reinstall you get one line** when old backups use another key. Two never-built hub fields are removed.
- **Kernel test, 24 reboots and your 2 power cycles:** a new kernel that crashes falls back to the old one by itself, on
all three boxes (two ways work). A kernel that FREEZES needs a person — the watchdog did not help on any box.
**Needs you (none urgent):**
1. May a box restart at night for a kernel update, and do we tell the household? My pick: yes, with a mail the day
before. If nothing: kernels stay manual.
2. The pool box leftovers: delete them in the Hetzner console, or give me the main login once. If nothing: they stay.
## Morning (2026-10-07): read back, released, delivered — and the Tester 1 test passed
@@ -0,0 +1,80 @@
Inst frr [10.6.1-1+pve2] (10.6.1-1+pve3 Proxmox Debian Repository:stable [amd64])
Inst shim-signed-common [1.48+pmx1+16.1-1+pmx1] (1.51+pmx1+16.1-2+pmx1 Proxmox Debian Repository:stable [all])
Inst shim-unsigned [16.1-1+pmx1] (16.1-2+pmx1 Proxmox Debian Repository:stable [amd64])
Inst shim-helpers-amd64-signed [1+16.1+1+pmx1] (1+16.1+2+pmx1 Proxmox Debian Repository:stable [amd64])
Inst shim-signed [1.48+pmx1+16.1-1+pmx1] (1.51+pmx1+16.1-2+pmx1 Proxmox Debian Repository:stable [amd64])
Inst librgw2 [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst libradosstriper1 [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst ceph-common [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst librbd1 [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst librados2 [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst python3-cephfs [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst libcephfs2 [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst python3-rgw [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst python3-rados [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64]) []
Inst python3-ceph-argparse [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [all]) []
Inst python3-ceph-common [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [all]) []
Inst python3-rbd [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64])
Inst ceph-fuse [19.2.3-pve4] (19.2.6-pve4 Proxmox Debian Repository:stable [amd64])
Inst chrony [4.6.1-3+deb13u1] (4.8-4~bpo13+2 Proxmox Debian Repository:stable [amd64])
Inst libcorosync-common4 [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst libcfg7 [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst libcmap4 [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst libcpg4 [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst libknet1t64 [1.31-pve1] (1.35-pve2 Proxmox Debian Repository:stable [amd64])
Inst libnozzle1t64 [1.31-pve1] (1.35-pve2 Proxmox Debian Repository:stable [amd64])
Inst libquorum5 [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst libvotequorum8 [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst corosync [3.1.10-pve2] (3.1.10-pve3 Proxmox Debian Repository:stable [amd64])
Inst frr-pythontools [10.6.1-1+pve2] (10.6.1-1+pve3 Proxmox Debian Repository:stable [all])
Inst libjs-extjs [7.0.0-5] (7.0.0-7 Proxmox Debian Repository:stable [all])
Inst libnvpair3linux [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [amd64])
Inst libproxmox-acme-plugins [1.7.1] (1.7.2 Proxmox Debian Repository:stable [all])
Inst libproxmox-backup-qemu0 [2.0.2] (2.0.3 Proxmox Debian Repository:stable [amd64])
Inst pve-qemu-kvm [11.0.0-3] (11.0.3-4 Proxmox Debian Repository:stable [amd64])
Inst libpve-notify-perl [9.1.5] (9.1.6 Proxmox Debian Repository:stable [all]) []
Inst libpve-cluster-api-perl [9.1.5] (9.1.6 Proxmox Debian Repository:stable [all]) []
Inst libpve-cluster-perl [9.1.5] (9.1.6 Proxmox Debian Repository:stable [all]) []
Inst pve-cluster [9.1.5] (9.1.6 Proxmox Debian Repository:stable [amd64]) []
Inst libpve-access-control [9.1.1] (9.1.2 Proxmox Debian Repository:stable [all]) []
Inst libpve-apiclient-perl [3.4.2] (3.4.3 Proxmox Debian Repository:stable [all]) []
Inst librados2-perl [1.5.0] (1.5.1 Proxmox Debian Repository:stable [amd64]) []
Inst proxmox-backup-client [4.2.0-1] (4.2.8-1 Proxmox Debian Repository:stable [amd64]) []
Inst proxmox-backup-file-restore [4.2.0-1] (4.2.8-1 Proxmox Debian Repository:stable [amd64]) []
Inst pve-manager [9.2.2] (9.2.21 Proxmox Debian Repository:stable [all]) []
Inst libproxmox-acme-perl [1.7.1] (1.7.2 Proxmox Debian Repository:stable [all]) []
Inst libpve-common-perl [9.1.12] (9.2.3 Proxmox Debian Repository:stable [all]) []
Inst libpve-guest-common-perl [6.0.3] (6.0.5 Proxmox Debian Repository:stable [all]) []
Inst qemu-server [9.1.15] (9.2.10 Proxmox Debian Repository:stable [amd64]) []
Inst libpve-storage-perl [9.1.5] (9.1.12 Proxmox Debian Repository:stable [all]) []
Inst pve-edk2-firmware-legacy [4.2025.05-2] (4.2026.08-1 Proxmox Debian Repository:stable [all]) []
Inst pve-edk2-firmware-ovmf [4.2025.05-2] (4.2026.08-1 Proxmox Debian Repository:stable [all]) []
Inst libpve-network-api-perl [1.6.5] (1.6.7 Proxmox Debian Repository:stable [all]) []
Inst libpve-network-perl [1.6.5] (1.6.7 Proxmox Debian Repository:stable [all]) []
Inst proxmox-firewall-data (0.1 Proxmox Debian Repository:stable [all]) []
Inst pve-firewall [6.0.4] (6.0.6 Proxmox Debian Repository:stable [amd64]) []
Inst pve-container [6.1.10] (6.1.14 Proxmox Debian Repository:stable [all]) []
Inst pve-ha-manager [5.2.4] (5.2.5 Proxmox Debian Repository:stable [amd64]) []
Inst novnc-pve [1.7.0-1] (1.7.0-2 Proxmox Debian Repository:stable [all]) []
Inst proxmox-enterprise-support-keyring [1.0] (1.1 Proxmox Debian Repository:stable [all]) []
Inst proxmox-mini-journalreader [1.6] (1.7 Proxmox Debian Repository:stable [amd64]) []
Inst proxmox-widget-toolkit [5.2.2] (5.2.10 Proxmox Debian Repository:stable [all]) []
Inst pve-docs [9.2.1] (9.2.13 Proxmox Debian Repository:stable [all])
Inst pve-i18n [3.7.4] (3.10.0 Proxmox Debian Repository:stable [all])
Inst pve-xtermjs [6.0.0-1] (6.0.0-2 Proxmox Debian Repository:stable [all])
Inst pve-yew-mobile-i18n [3.7.4] (3.10.0 Proxmox Debian Repository:stable [all])
Inst pve-yew-mobile-gui [0.7.0] (0.8.0 Proxmox Debian Repository:stable [amd64])
Inst libuutil3linux [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [amd64])
Inst libzfs7linux [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [amd64])
Inst libzpool7linux [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [amd64])
Inst proxmox-first-boot [9.2.5] (9.2.8 Proxmox Debian Repository:stable [amd64])
Inst pve-firmware [3.18-3] (3.18-7 Proxmox Debian Repository:stable [all])
Inst proxmox-kernel-7.0.14-22-pve-signed (7.0.14-22 Proxmox Debian Repository:stable [amd64])
Inst proxmox-kernel-7.0 [7.0.2-6] (7.0.14-22 Proxmox Debian Repository:stable [amd64])
Inst proxmox-kernel-helper [9.1.0+fde2] (9.2.0 Proxmox Debian Repository:stable [all])
Inst pve-edk2-firmware-aarch64 [4.2025.05-2] (4.2026.08-1 Proxmox Debian Repository:stable [all])
Inst pve-edk2-firmware [4.2025.05-2] (4.2026.08-1 Proxmox Debian Repository:stable [all])
Inst zfs-initramfs [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [all]) []
Inst zfsutils-linux [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [amd64])
Inst zfs-zed [2.4.2-pve1] (2.4.4-pve1 Proxmox Debian Repository:stable [amd64])
Inst tailscale [1.102.2] (1.102.5 Tailscale:pkgs.tailscale.com [amd64])
@@ -0,0 +1,336 @@
{
"release_id": "pve-2026-10-07-demo-felhom",
"packages": [
{
"name": "frr",
"version": "10.6.1-1+pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "librgw2",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "libradosstriper1",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "ceph-common",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "librbd1",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "librados2",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "python3-cephfs",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "libcephfs2",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "python3-rgw",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "python3-rados",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "python3-ceph-argparse",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "python3-ceph-common",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "python3-rbd",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "ceph-fuse",
"version": "19.2.6-pve4",
"origin": "Proxmox Debian Repository"
},
{
"name": "chrony",
"version": "4.8-4~bpo13+2",
"origin": "Proxmox Debian Repository"
},
{
"name": "libcorosync-common4",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libcfg7",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libcmap4",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libcpg4",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libknet1t64",
"version": "1.35-pve2",
"origin": "Proxmox Debian Repository"
},
{
"name": "libnozzle1t64",
"version": "1.35-pve2",
"origin": "Proxmox Debian Repository"
},
{
"name": "libquorum5",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libvotequorum8",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "corosync",
"version": "3.1.10-pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "frr-pythontools",
"version": "10.6.1-1+pve3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libjs-extjs",
"version": "7.0.0-7",
"origin": "Proxmox Debian Repository"
},
{
"name": "libnvpair3linux",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
},
{
"name": "libproxmox-acme-plugins",
"version": "1.7.2",
"origin": "Proxmox Debian Repository"
},
{
"name": "libproxmox-backup-qemu0",
"version": "2.0.3",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-qemu-kvm",
"version": "11.0.3-4",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-notify-perl",
"version": "9.1.6",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-cluster-api-perl",
"version": "9.1.6",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-cluster-perl",
"version": "9.1.6",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-cluster",
"version": "9.1.6",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-access-control",
"version": "9.1.2",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-apiclient-perl",
"version": "3.4.3",
"origin": "Proxmox Debian Repository"
},
{
"name": "librados2-perl",
"version": "1.5.1",
"origin": "Proxmox Debian Repository"
},
{
"name": "proxmox-backup-client",
"version": "4.2.8-1",
"origin": "Proxmox Debian Repository"
},
{
"name": "proxmox-backup-file-restore",
"version": "4.2.8-1",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-manager",
"version": "9.2.21",
"origin": "Proxmox Debian Repository"
},
{
"name": "libproxmox-acme-perl",
"version": "1.7.2",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-common-perl",
"version": "9.2.3",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-guest-common-perl",
"version": "6.0.5",
"origin": "Proxmox Debian Repository"
},
{
"name": "qemu-server",
"version": "9.2.10",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-storage-perl",
"version": "9.1.12",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-network-api-perl",
"version": "1.6.7",
"origin": "Proxmox Debian Repository"
},
{
"name": "libpve-network-perl",
"version": "1.6.7",
"origin": "Proxmox Debian Repository"
},
{
"name": "proxmox-firewall-data",
"version": "0.1",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-firewall",
"version": "6.0.6",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-container",
"version": "6.1.14",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-ha-manager",
"version": "5.2.5",
"origin": "Proxmox Debian Repository"
},
{
"name": "novnc-pve",
"version": "1.7.0-2",
"origin": "Proxmox Debian Repository"
},
{
"name": "proxmox-enterprise-support-keyring",
"version": "1.1",
"origin": "Proxmox Debian Repository"
},
{
"name": "proxmox-mini-journalreader",
"version": "1.7",
"origin": "Proxmox Debian Repository"
},
{
"name": "proxmox-widget-toolkit",
"version": "5.2.10",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-docs",
"version": "9.2.13",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-i18n",
"version": "3.10.0",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-xtermjs",
"version": "6.0.0-2",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-yew-mobile-i18n",
"version": "3.10.0",
"origin": "Proxmox Debian Repository"
},
{
"name": "pve-yew-mobile-gui",
"version": "0.8.0",
"origin": "Proxmox Debian Repository"
},
{
"name": "libuutil3linux",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
},
{
"name": "libzfs7linux",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
},
{
"name": "libzpool7linux",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
},
{
"name": "zfs-initramfs",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
},
{
"name": "zfsutils-linux",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
},
{
"name": "zfs-zed",
"version": "2.4.4-pve1",
"origin": "Proxmox Debian Repository"
}
],
"vmid": 9201
}
@@ -0,0 +1,14 @@
shim-signed-common 1.51+pmx1+16.1-2+pmx1 [Proxmox Debian Repository]
shim-unsigned 16.1-2+pmx1 [Proxmox Debian Repository]
shim-helpers-amd64-signed 1+16.1+2+pmx1 [Proxmox Debian Repository]
shim-signed 1.51+pmx1+16.1-2+pmx1 [Proxmox Debian Repository]
pve-edk2-firmware-legacy 4.2026.08-1 [Proxmox Debian Repository]
pve-edk2-firmware-ovmf 4.2026.08-1 [Proxmox Debian Repository]
proxmox-first-boot 9.2.8 [Proxmox Debian Repository]
pve-firmware 3.18-7 [Proxmox Debian Repository]
proxmox-kernel-7.0.14-22-pve-signed 7.0.14-22 [Proxmox Debian Repository] NEW
proxmox-kernel-7.0 7.0.14-22 [Proxmox Debian Repository]
proxmox-kernel-helper 9.2.0 [Proxmox Debian Repository]
pve-edk2-firmware-aarch64 4.2026.08-1 [Proxmox Debian Repository]
pve-edk2-firmware 4.2026.08-1 [Proxmox Debian Repository]
tailscale 1.102.5 [Tailscale]
@@ -0,0 +1,7 @@
pve-manager/9.2.2/b9984c6d90a4bd80 (running kernel: 7.0.2-6-pve)
felhom-controller 2026-10-07T10:44:54.860562348Z
opengist 2026-10-07T10:44:53.653308875Z
cloudflared 2026-10-07T10:44:53.660792998Z
filebrowser 2026-10-07T10:44:53.663466912Z
traefik 2026-10-07T10:44:53.648261309Z
status: running
@@ -0,0 +1,232 @@
12:00:34 302
12:00:39 302
12:00:44 302
12:00:49 302
12:00:54 302
12:00:59 302
12:01:04 302
12:01:09 302
12:01:14 302
12:01:19 302
12:01:24 302
12:01:29 302
12:01:34 302
12:01:39 302
12:01:44 302
12:01:49 302
12:01:54 302
12:01:59 302
12:02:04 302
12:02:09 302
12:02:14 302
12:02:19 302
12:02:24 302
12:02:29 302
12:02:34 302
12:02:39 302
12:02:44 302
12:02:49 302
12:02:54 302
12:03:00 302
12:03:05 302
12:03:10 302
12:03:15 302
12:03:20 302
12:03:25 302
12:03:30 302
12:03:35 302
12:03:40 302
12:03:45 302
12:03:50 302
12:03:55 302
12:04:00 302
12:04:05 302
12:04:10 302
12:04:15 302
12:04:20 302
12:04:25 302
12:04:30 302
12:04:35 302
12:04:40 302
12:04:45 302
12:04:50 302
12:04:55 302
12:05:00 302
12:05:05 302
12:05:10 302
12:05:15 302
12:05:20 302
12:05:25 302
12:05:30 302
12:05:35 302
12:05:40 302
12:05:46 302
12:05:51 302
12:05:56 302
12:06:01 302
12:06:06 302
12:06:11 302
12:06:16 302
12:06:21 302
12:06:26 302
12:06:31 302
12:06:36 302
12:06:41 302
12:06:46 302
12:06:51 302
12:06:56 302
12:07:01 302
12:07:06 302
12:07:11 302
12:07:16 302
12:07:21 302
12:07:26 302
12:07:31 302
12:07:36 302
12:07:41 302
12:07:46 302
12:07:51 302
12:07:56 302
12:08:01 302
12:08:06 302
12:08:11 302
12:08:16 302
12:08:21 302
12:08:26 302
12:08:31 302
12:08:37 302
12:08:42 302
12:08:47 302
12:08:52 302
12:08:57 302
12:09:02 302
12:09:07 302
12:09:12 302
12:09:17 302
12:09:22 302
12:09:27 302
12:09:32 302
12:09:37 302
12:09:42 302
12:09:47 302
12:09:52 302
12:09:57 302
12:10:02 302
12:10:07 302
12:10:12 302
12:10:17 302
12:10:22 302
12:10:27 302
12:10:32 302
12:10:37 302
12:10:42 302
12:10:47 302
12:10:52 302
12:10:57 302
12:11:02 302
12:11:07 302
12:11:12 302
12:11:17 302
12:11:22 302
12:11:28 302
12:11:33 302
12:11:38 302
12:11:43 302
12:11:48 302
12:11:53 302
12:11:58 302
12:12:03 302
12:12:08 302
12:12:13 302
12:12:18 302
12:12:23 302
12:12:28 302
12:12:33 302
12:12:38 302
12:12:43 302
12:12:48 302
12:12:53 302
12:12:58 302
12:13:03 302
12:13:08 302
12:13:13 302
12:13:18 302
12:13:23 302
12:13:28 302
12:13:33 302
12:13:38 302
12:13:43 302
12:13:48 302
12:13:53 302
12:13:58 302
12:14:03 302
12:14:08 302
12:14:14 302
12:14:19 302
12:14:24 302
12:14:29 302
12:14:34 302
12:14:39 302
12:14:44 302
12:14:49 302
12:14:54 302
12:14:59 302
12:15:04 302
12:15:09 302
12:15:14 302
12:15:19 302
12:15:24 302
12:15:29 302
12:15:34 302
12:15:39 302
12:15:44 302
12:15:49 302
12:15:54 302
12:15:59 302
12:16:04 302
12:16:09 302
12:16:14 302
12:16:19 302
12:16:24 302
12:16:29 302
12:16:34 302
12:16:39 302
12:16:44 302
12:16:49 302
12:16:54 302
12:16:59 302
12:17:05 302
12:17:10 302
12:17:15 302
12:17:20 302
12:17:25 302
12:17:30 302
12:17:35 302
12:17:40 302
12:17:45 302
12:17:50 302
12:17:55 302
12:18:00 302
12:18:05 302
12:18:10 302
12:18:15 302
12:18:20 302
12:18:25 302
12:18:30 302
12:18:35 302
12:18:40 302
12:18:45 302
12:18:50 302
12:18:55 302
12:19:00 302
12:19:05 302
12:19:10 302
12:19:15 302
12:19:20 302
12:19:25 302
12:19:30 302
12:19:35 302
12:19:40 302
12:19:45 302
12:19:50 302
12:19:55 302
@@ -0,0 +1,3 @@
== os_pve_step → demo-felhom (66 packages, /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/day-2026-10-07/B/B2-params.json), 2026-10-07T12:00:40Z
signed: op=os_pve_step host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=0a3d4def12ecdccd270ac79e78be17ba expires=2026-10-07T12:45:40Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,35 @@
pve-manager/9.2.21/4f6e0ac86f9e8c7f (running kernel: 7.0.2-6-pve)
felhom-controller 2026-10-07T10:44:54.860562348Z
opengist 2026-10-07T10:44:53.653308875Z
cloudflared 2026-10-07T10:44:53.660792998Z
filebrowser 2026-10-07T10:44:53.663466912Z
traefik 2026-10-07T10:44:53.648261309Z
status: running
Oct 07 14:15:13 demo-felhom felhom-agent[59282]: time=2026-10-07T14:15:13.072+02:00 level=INFO msg="audit: gate decision" class=os_pve_step host=demo-felhom-8363b5 guest="" source=one_shot_job disposition=destructive allowed=true reason=signed key_id=felhom-op-1 nonce=0a3d4def… durable_id=""
Oct 07 14:15:13 demo-felhom felhom-agent[59282]: time=2026-10-07T14:15:13.072+02:00 level=INFO msg="gate decision" class=os_pve_step guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
Oct 07 14:15:13 demo-felhom felhom-agent[59282]: time=2026-10-07T14:15:13.072+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=ae12ad4399ef57a1 op=os_pve_step key_id=felhom-op-1 nonce=0a3d4def12ecdccd270ac79e78be17ba
Oct 07 14:15:13 demo-felhom felhom-agent[59282]: time=2026-10-07T14:15:13.073+02:00 level=INFO msg="osupdate: pve step holds the /etc/pve write gate — the agent's own writes wait until it ends" run=20261007T121513Z layer=pve vmid=9201 trigger=signed
Oct 07 14:15:13 demo-felhom felhom-agent[59282]: time=2026-10-07T14:15:13.073+02:00 level=INFO msg="osupdate: START" run=20261007T121513Z layer=pve vmid=9201 ring=0 trigger=signed enabled=true release=pve-2026-10-07-demo-felhom
Oct 07 14:15:13 demo-felhom felhom-os-apply[91459]: os-apply: START release=pve-2026-10-07-demo-felhom layer=pve:9201 lane=slow mode=apply select=listed packages=66 authority=signed
Oct 07 14:15:16 demo-felhom felhom-os-apply[91551]: os-apply: REPAIR configured=0 journal=0 fixed=0
Oct 07 14:15:18 demo-felhom felhom-os-apply[91638]: os-apply: PLAN upgrade=65 already=0 not-installed=1 from-snapshot=0
Oct 07 14:15:19 demo-felhom felhom-os-apply[91691]: os-apply: NEW proxmox-firewall-data=0.1 (on the pve lane's allow-list)
Oct 07 14:16:29 demo-felhom felhom-os-apply[100237]: os-apply: CONFFILE updated /etc/apparmor.d/usr.sbin.chronyd (it was not changed locally)
Oct 07 14:16:29 demo-felhom felhom-os-apply[100238]: os-apply: CONFFILE updated /etc/network/if-post-down.d/chrony (it was not changed locally)
Oct 07 14:16:29 demo-felhom felhom-os-apply[100239]: os-apply: CONFFILE updated /etc/network/if-up.d/chrony (it was not changed locally)
Oct 07 14:16:29 demo-felhom felhom-os-apply[100240]: os-apply: CONFFILE updated /etc/ppp/ip-down.d/chrony (it was not changed locally)
Oct 07 14:16:29 demo-felhom felhom-os-apply[100241]: os-apply: CONFFILE updated /etc/ppp/ip-up.d/chrony (it was not changed locally)
Oct 07 14:16:30 demo-felhom felhom-os-apply[100514]: os-apply: DONE rc=0 seconds=69.6 upgraded=65 restart-needed=watchdog-mux reboot-needed=no
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=pve-2026-10-07-demo-felhom layer=pve:9201 lane=slow mode=apply select=listed packages=66 authority=signed"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 journal=0 fixed=0"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=65 already=0 not-installed=1 from-snapshot=0"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: NEW proxmox-firewall-data=0.1 (on the pve lane's allow-list)"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE updated /etc/apparmor.d/usr.sbin.chronyd (it was not changed locally)"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE updated /etc/network/if-post-down.d/chrony (it was not changed locally)"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE updated /etc/network/if-up.d/chrony (it was not changed locally)"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE updated /etc/ppp/ip-down.d/chrony (it was not changed locally)"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE updated /etc/ppp/ip-up.d/chrony (it was not changed locally)"
Oct 07 14:16:35 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:35.758+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=69.6 upgraded=65 restart-needed=watchdog-mux reboot-needed=no"
Oct 07 14:16:36 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:36.587+02:00 level=INFO msg="osupdate: DONE" run=20261007T121513Z layer=pve vmid=9201 ring=0 trigger=signed outcome=applied healthy=true reason="" upgraded=65 pending=5 not_covered=0 restart_needed=1 reboot_needed=false wrapper_seconds=82.6
Oct 07 14:16:36 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:36.672+02:00 level=INFO msg="osupdate: pve step released the /etc/pve write gate" run=20261007T121513Z layer=pve vmid=9201 trigger=signed
Oct 07 14:16:36 demo-felhom felhom-agent[59282]: time=2026-10-07T14:16:36.672+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=ae12ad4399ef57a1 op=os_pve_step
@@ -0,0 +1,6 @@
pve | : | — | a ring-0 box has not reported a Proxmox step | Approve now (guest + host) | An urgent fix only — normally the hub approves after 24 h and one night.
7.0.2-6-pve | 7.0.2-6-pve | 13.7 | os-host-20261004-124133 | 78 | 0
7.0.2-6-pve | 13.7 | os-host-20261004-124133 | 78 | 0 | none
7.0.2-6-pve | unknown | 13.7 | ring0-20261007T023651Z | 80 | 0
7.0.14-20-pve | unknown | 13.7 | ring0-20261007T024218Z | 78 | 0
7.0.2-6-pve | unknown | 13.7 | os-host-20261005-122429 | 79 | 0
@@ -0,0 +1 @@
2026/10/07 14:16:36 [INFO] osupdates: demo-felhom-8363b5 reported pve run 20261007T121513Z: ring=0 mode=apply outcome=applied healthy=true upgraded=65 pending=5 not-covered=0 restart-needed=1 reboot-needed=false wrapper=82.6s
@@ -0,0 +1,17 @@
# Part B live proof — the Proxmox package lane on demo-felhom (2026-10-07 14:15 CEST)
Agent 0.151.0 + its bundle (probe 68/68), hub 0.142.0. One signed `os_pve_step` (the product's ring-1 path, used by hand
outside 02:00–08:30; the debug selftest was NOT used because it also runs the Docker step the brief forbids). The set:
66 packages of origin "Proxmox Debian Repository" from a dry-run (`B1`, `B2-params.json`); 14 left out — shim, firmware,
`proxmox-first-boot`, the kernel names (`B2-skipped.txt`).
- Journal (`B6-after.txt`): `pve step holds the /etc/pve write gate` → `os-apply: START … layer=pve:9201 lane=slow
select=listed packages=66 authority=signed` → `PLAN upgrade=65 not-installed=1` → `NEW proxmox-firewall-data=0.1 (on the
pve lane's allow-list)` → `DONE rc=0 seconds=69.6 upgraded=65 restart-needed=watchdog-mux reboot-needed=no` →
`osupdate: DONE … outcome=applied healthy=true` → `released the /etc/pve write gate` → `signed op COMPLETED`.
- `pveversion`: 9.2.2 → **9.2.21**; running kernel unchanged (7.0.2-6).
- The guest kept running: every container's `StartedAt` identical before and after (`B3`, `B6`).
- opengist through traefik every 5 s: **31 of 31 answered 302** from 14:14:30 to 14:18:00 (`B4-sampler-5s.log`).
- Second channel, the hub's log: `demo-felhom-8363b5 reported pve run 20261007T121513Z … outcome=applied healthy=true
upgraded=65` (`B8-hub-log.txt`). The System page's pve row still reads „a ring-0 box has not reported a Proxmox step" —
by design: the candidate needs every ring-0 box's pve report, and demo-hp has run none (`osupdates/service.go:305-313`).
File diff suppressed because one or more lines are too long
@@ -0,0 +1,5 @@
Oct 07 13:45:01 demo-felhom felhom-agent[1312]: time=2026-10-07T13:45:01.163+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=56c40a58d980c080 op=agent_update
Oct 07 14:00:14 demo-felhom felhom-os-apply[75675]: os-apply: BUNDLE DONE agent=0.151.0 written=4 same=21 self-check=ok signers-created=False
Oct 07 14:00:14 demo-felhom felhom-agent[59282]: time=2026-10-07T14:00:14.074+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.151.0 written=4 same=21 self-check=ok signers-created=False"
Oct 07 14:00:14 demo-felhom felhom-agent[59282]: time=2026-10-07T14:00:14.646+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded=""
Oct 07 14:00:14 demo-felhom felhom-agent[59282]: time=2026-10-07T14:00:14.646+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=8e6752fa28b3d407 op=agent_config_update
@@ -0,0 +1,3 @@
== agent_config_update 0.151.0 → demo-felhom, 2026-10-07T11:45:28Z
signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=c8469f84b90c04e1bc746fca1066bc55 expires=2026-10-07T12:30:28Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,3 @@
== agent_config_update 0.151.0 (bundle bacd1d17…) → demo-hp, 2026-10-07T11:21:16Z
signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=99b8a24f5e7ed4b0c3b2484ac7d5a141 expires=2026-10-07T12:06:16Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,3 @@
== agent_config_update 0.151.0 → tester-1, 2026-10-07T11:51:38Z
signed: op=agent_config_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=5ee8899712b81aa50cdc92a41891b7d0 expires=2026-10-07T12:36:38Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,5 @@
hub 0.142.0: manifest 374213f2, sync 11:13:12Z, rollout ok, 'felhom-hub 0.142.0 starting' 13:13:18 local, healthz 200 + /system 200 at 11:14:02Z
agent 0.151.0: sha 0464354f…, bundle bacd1d17…, tag v0.151.0=dd7cdc0; vouch REFUSED by the hub (flash=golden_behind_fleet: golden 0.301.0 < fleet controller 0.302.0) — new installs keep agent 0.150.0 until the weekly bake
hp: felhom-agent 0.151.0 pve-manager/9.2.2
felhom-pve-lan: felhom-agent 0.151.0 pve-manager/9.2.21
root@192.168.0.154: felhom-agent 0.151.0 pve-manager/9.2.2
@@ -0,0 +1,5 @@
== agent_update 0.151.0 → demo-felhom, tester-1, 2026-10-07T11:37:23Z
signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=39432e033ce489b9f1046c5f7f99c9f9 expires=2026-10-07T12:22:23Z
uploaded signed op to the hub jobs queue
signed: op=agent_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=f5c91f341545b40ed059eed19ae26708 expires=2026-10-07T12:22:23Z
uploaded signed op to the hub jobs queue
@@ -0,0 +1,6 @@
== vouch 2026-10-07T11:15:06Z: agent 0.151.0, golden 0.301.0, min_agent 0.131.0 (unchanged)
HTTP/1.1 303 See Other
Location: /configuration?flash=golden_behind_fleet
== agent_update 0.151.0 → demo-hp only, 2026-10-07T11:15:47Z
signed: op=agent_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=af6bcd53853bc24468de902f287c87af expires=2026-10-07T12:00:47Z
uploaded signed op to the hub jobs queue
+3
View File
@@ -32,6 +32,9 @@ The full text of every row below: `git show 5ba1702fcf:documentation/backlog/OPE
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED (hub 0.142.0, agent 0.151.0; `09` §3 decision 169) | The never-built `hosts.dr_record_json` and `host_escrow.directive_json` have no writer and no reader any more (the columns stay, unread); the agent's `-directive` flag is gone; `05` §9/§11 and `06` §3.5 corrected. Tester 2: the operator believes it never set up off-site, so it has no escrow (decision 170). `audits/day-2026-10-07/F/`. |
| **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED (hub 0.142.0, agent 0.151.0; `09` §3 decision 168) | Slice 1 (the hub keeps the old key on a reinstall, hub 0.141.0) and slice 2 (the restore-test's skip of another key's archives becomes one edge-triggered operator line naming the box, the count and the date range) delivered to the three boxes; red-proofs `audits/day-2026-10-07/E/`. Option C (the recovery code at reinstall) declined by the operator. Not yet seen live: a reinstall that produces other-key archives. |
| **R-812** | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED AND PROVEN LIVE (agent 0.151.0, hub 0.142.0; `09` §3 decision 163) | Option A, the Proxmox package lane: on demo-felhom one signed step installed 65 Proxmox packages in 70 s (`pve-manager` 9.2.2 → 9.2.21), healthy, the /etc/pve write gate held, the guest untouched (StartedAt identical), the app answered 31/31; the hub logged the run. `audits/day-2026-10-07/B/RESULT.md`. The kernel is R-836 (the spike done, `audits/kernel-spike-2026-10-07/`). |
| **R-822** | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** (P2) | CLOSED 2026-10-07 — BY OPERATOR RULING (`09` §3 decision 166): the residual is accepted as a stated limit of decision 68 | Guard shipped controller v0.289.0; residual (past-dated fakes steering the keeps, a box-trusted window count) in `audits/night-burndown-2026-10-06/design-R-822.md`; the hub's own count is R-895, open. |
## 2026-10-07 (morning) — the read-back, the releases, the Tester 1 proof
+4 -7
View File
@@ -145,15 +145,13 @@ stopping line that lies.
| **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC |
| **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC |
## Backup & restore — 32 rows (P2 7, P3 11, P4 14)
## Backup & restore — 30 rows (P2 5, P3 11, P4 14)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-32.md`) — pick A: purge the repo + set-aside copies through the sub-account's own password login BEFORE deleting it (no main-account credential); the „repo data destroyed" comment at `hub/internal/offsite/offsite.go:276-280` is false for the shared tier; part (3) waits on a read-only measurement (`du` on port 23). Operator: agree to route A. **2026-10-07 07:58: `09` §3 decision 167 — option A yes, and delete the old RESET leftovers once.** **2026-10-07 (Part D step 2, read only):** the hub shows bx11 holding 3.3 GB of data while the four live customers account for ~0.6 GB → **~2.7 GB belongs to no live customer**; no live customer shows a set-aside folder, so the leftovers are in homes whose sub-account is gone, reachable only by the pool box's MAIN account. **The one-time clean-up is STOPPED:** no main-account login for bx11 is held on DooPlex (the API token in the credentials file sees no box). Needs: the main-account login, or the operator lists and deletes in the Hetzner console (`audits/day-2026-10-07/D/D2-listing-readonly.txt`). | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC |
| **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: TRACED in source — not a fault in a running path.** `hosts.dr_record_json` has no writer and no reader; `host_escrow.directive_json` is filled only by the by-hand `-directive` flag (the wizard's fixed argv carries none, so every wizard escrow stores `{}`), and its only routes (`/re-enroll`, `/restore-directive`) have no client. The built DR path reads the recipe, tenantsync and the escrow blob. Design with a pick (retire both and correct `05` §9/§11, `06` §3.5 — needs the operator's word): `audits/night-burndown-2026-10-06/design-R-105.md`. The `drives` third was not re-measured (the hub-DB read was refused). **2026-10-07 07:58: `09` §3 decisions 169–170 — option A yes (remove both fields, correct `05` §9); Tester 2: the operator believes it never set up off-site, so it has no escrow.** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC |
| **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-32.md`) — pick A: purge the repo + set-aside copies through the sub-account's own password login BEFORE deleting it (no main-account credential); the „repo data destroyed" comment at `hub/internal/offsite/offsite.go:276-280` is false for the shared tier; part (3) waits on a read-only measurement (`du` on port 23). Operator: agree to route A. **2026-10-07 07:58: `09` §3 decision 167 — option A yes, and delete the old RESET leftovers once.** **2026-10-07 (Part D step 2, read only):** the hub shows bx11 holding 3.3 GB of data while the four live customers account for ~0.6 GB → **~2.7 GB belongs to no live customer**; no live customer shows a set-aside folder, so the leftovers are in homes whose sub-account is gone, reachable only by the pool box's MAIN account. **The one-time clean-up is STOPPED:** no main-account login for bx11 is held on DooPlex (the API token in the credentials file sees no box). Needs: the main-account login, or the operator lists and deletes in the Hetzner console (`audits/day-2026-10-07/D/D2-listing-readonly.txt`). **2026-10-07: option A DELIVERED (hub 0.142.0):** RESET purges the repo and every `.orphaned-*` through the sub-account's own login before deleting it, and keeps the sub-account if the purge fails (red-proved, `audits/day-2026-10-07/D/`). Not yet seen live (no scratch RESET today). **The one-time clean-up of the ~2.7 GB of leftovers is STOPPED:** it needs the pool box's main-account login (see above). | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC |
| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator |
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC |
| **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** **2026-10-06 night: the wrong verdict was already gone** (agent v0.138.0, R-727 skips another key's archives); demo-hp's August archives are pruned by ep0's keep-last 2 (inferred, ep0 fenced). **NEW, the real gap:** the hub retained an escrow only on a restic-password change, so a reinstall's new backup key K overwrote the only copy of the old one — **fixed on felhom.eu main (hub, unreleased; red-proved `audits/night-burndown-2026-10-06/r366/`)**. Design `audits/night-burndown-2026-10-06/design-R-366.md` (pick B: retain + tell the operator; C, key continuity, is the operator's). Not closed: the operator signal (archives made with another key) is the design's second slice. **2026-10-07 07:58: `09` §3 decision 168 — keep the old key (hub 0.141.0) yes; option C no; build slice 2 (tell the operator when old copies use another key).** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC |
| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC |
| **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** **2026-10-07 07:58: kept open for later (`09` §3 decision 166).** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC |
@@ -197,7 +195,7 @@ stopping line that lies.
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. **2026-10-07 07:58: `09` §3 decision 165 — (a) A1 yes before the first paying customer; (b) B3 accept + B2 hygiene in the same bundle; (c) C2 accept.** | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC |
| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. **2026-10-07 07:58: `09` §3 decision 165 — (a) A1 yes before the first paying customer; (b) B3 accept + B2 hygiene in the same bundle; (c) C2 accept.** **2026-10-07: (a) A1 and (b) B2 DELIVERED (agent 0.151.0 + bundle on demo-hp, demo-felhom, Tester 1; probe 68/68).** Live on demo-hp: no `tee` grant left in `sudo -l -U felhom-agent`; the verb `felhom-priv-apply ^controller-image [0-9]+$` is the route; a hand-fed `docker.io/library/alpine:latest` → `REFUSED [I1]` rc 3, the guest's image file unchanged; the old `pct exec … tee` asks for a password; felhom-op's pct lines anchored (`audits/day-2026-10-07/C/C-live-demo-hp.txt`). **LEFT:** one managed controller swap seen through the verb — no newer controller existed today; the next controller release shows it. (c) accepted (decision 165). | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC |
| **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor |
| **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with "<domain>"` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC |
| **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC |
@@ -212,11 +210,10 @@ stopping line that lies.
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
## Box system & updates — 7 rows (P2 2, P3 5)
## Box system & updates — 6 rows (P2 1, P3 5)
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|---|---|---|---|---|---|---|---|
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 (evening) — the guest's DOCKER engine slow lane is BUILT and proven live (agent v0.142.0, hub v0.132.0; `11` §5.8): live-restore on everywhere, ring 0 steps under a root-owned mark, ring 1 and undo only by a signed job the wrapper re-verifies; the operator approves each engine set on the System page. LEFT: the kernel lane (R-836); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 (afternoon) — the HOST's Debian fast lane is BUILT and proven live too (agent v0.141.1, hub v0.131.1; `11` §8.2, `audits/os-host-lane-2026-10-04/`): appliances only, never kernel/boot/firmware, after a healthy guest step; fleet view and four alarms (§8.3). LEFT: the Docker and kernel slow lanes (R-836); existing boxes (R-840).** Earlier: NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-812.md`) — what is left is `11` §8 step 6: the host's Proxmox-origin packages and its kernel are never updated (read tonight: demo-hp 77 and demo-felhom 78 Proxmox packages pending on the 2026-10-05 lists, `pve-manager` 9.2.2 → 9.2.21; demo-felhom still runs its install-time kernel 7.0.2-6). Pick: a `pve` slow lane for the Proxmox packages without the kernel, approved per set like the Docker engine; the kernel stays an operator-run step until R-836's fallback is measured. Proposed split: close R-812 when that lane ships; R-836 carries the kernel. Two operator questions in the design. **2026-10-07 07:58: `09` §3 decisions 163–164 — option A (the Proxmox package lane, per-set approval) YES; option B (the kernel lane) YES before the first paying customer, starting with a reboot spike carried by R-836.** | — | — | CC + operator |
| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot <new>`: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. | — | — | CC |
| **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator |
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |