OS updates guest fast lane: records — 11 §8.1 BUILT, 03 cloudflared corrected, 07 §6.1 OS leg, 00 PARTIAL, monthly runbook infra pins, golden 0.291.0 record + vouch, register 331 -> 333 (R-837/838/726/843 closed; R-840/841/842/844/845 opened), STATUS, report
gates / gates (push) Successful in 31s
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -16,6 +16,15 @@
|
|||||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||||
|
|
||||||
|
|
||||||
|
> **2026-10-04 (afternoon) — OS updates, guest fast lane BUILT (`11` §8.1).** Agent v0.140.0: `configs/felhom-os-apply`
|
||||||
|
> (R1–R13, repair first, snapshot.debian.org fallback; FELHOM_OSAPPLY sudoers), `internal/osupdate` (the leg after a
|
||||||
|
> successful primary backup, under the heavy-op gate; health baseline = start of the leg), `--selftest=os-update`.
|
||||||
|
> Hub v0.130.0: `internal/osupdates` (rings, switch default ON, candidate = what ALL ring-0 boxes run, approval 24 h +
|
||||||
|
> 1 night via `OS_APPROVE_AFTER` / `OS_APPROVE_NIGHTS`, `os_update` desired block, `POST /hosts/{id}/os-report`,
|
||||||
|
> operator `/os/*`). Controller v0.291.0: decision 78 (R-726), infra pins raised (R-838) + `scripts/check-infra-pins.py`.
|
||||||
|
> Installer 1.29.0 installs the wrapper. **No automatic undo** (R-837 measured → R-842). **Existing boxes need the
|
||||||
|
> wrapper + sudoers by hand** (R-840 — the demo boxes got them by hand). `REPORT-os-guest-lane-2026-10-04.md`.
|
||||||
|
|
||||||
> **Rulings 2026-10-04 (~10:17) — recorded before the work (guest fast lane build).** `09` §3 decisions **78** (R-726
|
> **Rulings 2026-10-04 (~10:17) — recorded before the work (guest fast lane build).** `09` §3 decisions **78** (R-726
|
||||||
> option A: a returning household's box sets the old off-site copy aside on night one), **79** (`snapshot.debian.org`
|
> option A: a returning household's box sets the old off-site copy aside on night one), **79** (`snapshot.debian.org`
|
||||||
> fallback for a replaced approved version — option A; the reviewer had picked B) and **80** (next build: the guest
|
> fallback for a replaced approved version — option A; the reviewer had picked B) and **80** (next build: the guest
|
||||||
|
|||||||
@@ -0,0 +1,60 @@
|
|||||||
|
# REPORT — OS updates build step 1: the guest's Debian fast lane; decision 78; the infrastructure images — 2026-10-04
|
||||||
|
|
||||||
|
Architecture read: `11-os-updates.md` (with C1–C12, §5.4.1, §7.1 — the design; it won wherever it differed from the
|
||||||
|
brief, see below), `03-host-agent.md`, `07` §6.1, `09` §3 decisions 11/12/15/18, `08`. Baselines (re-verified):
|
||||||
|
felhom.eu `1b74ddc0c9` (hub 0.129.0), agent `596238cc2e` (0.139.0), controller `99a1497560` (0.290.0), catalog
|
||||||
|
`917a779cca`. Register 331, highest R-839. Rulings recorded first: `09` §3 decisions 78–80 (`6ed79cd`). Evidence: `documentation/audits/os-guest-lane-2026-10-04/` (parts A–G).
|
||||||
|
|
||||||
|
## The Part table
|
||||||
|
|
||||||
|
| Part | Result | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| A — the snapshot undo first (R-837) | **done — and it FAILED: no snapshot is possible** | PVE refuses any snapshot not named `vzdump` of a guest with host-path binds (mp8/mp9), as the agent's token (which has `VM.Snapshot` + `VM.Snapshot.Rollback`) and as root. By the brief's rule: **no automatic undo built**; the decision is in STATUS (R-842). Steps 2–5 (apply, roll back, re-apply) had nothing to roll back to; 9201 was brought current by the product's own leg in Part G. Thin pool unchanged. |
|
||||||
|
| B — the wrapper | **done** | `felhom-os-apply` (Python 3 stdlib), R1–R13, repair first, snapshot.debian.org fallback, log lines; host layer and slow lane refused. 35 tests; **every refusal red-proved** (13/13). `visudo -cf` OK. **Changed:** Python not shell (a JSON plan cannot be parsed safely in sh — so "shellcheck clean" became `ast`/compile-checked + the suite); one sudoers entry with a plan `mode` instead of a separate `--repair-only`. **The route for existing boxes: none exists** (R-840, with a proposal). |
|
||||||
|
| C — the agent's leg | **done** | After a successful primary whole-guest backup, under the heavy-op gate (red-proved: the gate is held), once per 20 h, 90 s settle. Health rule written and pinned (`HealthVerdict`). Report: full installed set with origins, pending, not covered, restart-needed. Debug action `--selftest=os-update`. 7 leg red-proofs + 2 hook red-proofs. |
|
||||||
|
| D — the hub | **done** — hub v0.130.0 | Rings, per-box switch (default ON), the candidate/approval rule (24 h + 1 night, config), approve-now, events, fleet JSON. 5 approval red-proofs; the `os_update` wire golden byte-identical in both repos. |
|
||||||
|
| E — household line + decision 78 | **done** | Line = hub customer event `os_update_applied` (info: on the household's timeline, not mailed; hu/en in the bundle). **Changed:** there is no box-side event surface, so the hub event is it (R-844). Decision 78 built in controller v0.291.0, red-proved both ways. |
|
||||||
|
| F — infrastructure images (R-838) | **done** | traefik v3.7.13, cloudflared 2026.9.3, filebrowser 1.5.6-stable; breaking changes named (none we use). A release moves all three (9202: ≤ 1.9 s / ≤ 1.5 s; demo boxes: public gap ≤ 19.6 s / ≤ 14.7 s incl. the controller restart). `scripts/check-infra-pins.py` + runbook section. **Changed:** the standing brief `claude/MONTHLY-security-retest.md` lives in the claude.ai project, not the repo — the repo half is the runbook; the project file is the operator's to update. `03` corrected (3 lines). |
|
||||||
|
| G — live proof | **done, one part changed** | Ring 0 on both boxes (53 packages each, healthy); approval with a 2-minute TEST wait (272 packages, auto), then the ruled values back; ring 1 on demo-felhom (exactly the 3 approved versions, nothing newer); a failed health check → `health_failed`, operator mail, household line. **Changed:** "show the rollback" — there is none (Part A). Teardown: no snapshot, no plan files, test config gone, demo-felhom back to ring 0. |
|
||||||
|
| H — release, golden, records | **done** (see Teardown for the golden) | Agent 0.140.0 (signed per box, both demo boxes on it), hub 0.130.0, controller 0.291.0 (floor 0.291.0, MinAgent 0.131.0 declared), installer 1.29.0. `11` §8.1, `00`, `07` §6.1, `03` updated. |
|
||||||
|
|
||||||
|
## Claims in the brief that turned out wrong (named)
|
||||||
|
|
||||||
|
1. **"The agent's token can snapshot and roll back"** — it HAS the rights, but no snapshot of a customer guest is
|
||||||
|
possible at all (bind mounts). Neither the token nor root can.
|
||||||
|
2. **"A snapshot rollback leaves the thin pool clean"** — unmeasurable: there was no snapshot.
|
||||||
|
3. **"A new sudoers line can reach an installed box through the product"** — false. Only the installer writes it;
|
||||||
|
the signed agent update replaces the binary only (R-840). The demo boxes got the wrapper + sudoers BY HAND.
|
||||||
|
4. **"A controller release moves the infrastructure containers"** — TRUE for all three. (I first wrote the opposite
|
||||||
|
for the file browser and corrected it the same hour: its start-up mount sync renders the new image.)
|
||||||
|
5. **"An agent event can reach the household's timeline"** — only through the hub (a hub customer event); the box has
|
||||||
|
no timeline of its own (R-844).
|
||||||
|
6. **"Before each guest update, the box takes a snapshot"** (the one-page summary) — impossible (Part A).
|
||||||
|
7. `11` vs the brief: `11` §5.4.1's `--repair-only` flag was folded into the plan; `11`'s "a missed night waits" holds.
|
||||||
|
|
||||||
|
## Found and fixed live (before the release)
|
||||||
|
|
||||||
|
- `--selftest=os-update` was refused by the flag's allow-list — and so was `--selftest=wgtunnel`, since S3 (R-843,
|
||||||
|
opened and closed; a new test pins every dispatched mode).
|
||||||
|
- The wrapper logged an UPDATED conffile as "kept" (dpkg's two message shapes; fixed + tested).
|
||||||
|
- An app stopped between the inventory and the apply escaped the health check; the baseline is now the start of the
|
||||||
|
leg (fixed + red-proved).
|
||||||
|
- My stopped-app test also made the box mail one `app_start_failed` (a second one was held by the cooldown).
|
||||||
|
|
||||||
|
## Rows
|
||||||
|
|
||||||
|
Closed: **R-837** (measured), **R-838**, **R-726**, **R-843** (opened and closed). Opened: **R-840** (no product route
|
||||||
|
to installed boxes, P2), **R-841** (the agent's cloudflared probe reads a host unit that does not exist, P3), **R-842**
|
||||||
|
(the undo decision, waiting on the operator), **R-844** (household line only on the hub, P4), **R-845** (a pass takes
|
||||||
|
3–4 min, P4). Narrowed: **R-812**. Register **331 → 333**.
|
||||||
|
|
||||||
|
## Teardown, three layers
|
||||||
|
|
||||||
|
- **Machines:** no snapshot on either 9201; no plan files; privatebin restarted and healthy; demo-felhom back to ring 0;
|
||||||
|
both 9201s fully Debian-current (openssl at the approved u3). 9202 runs controller 0.291.0 (from Part F).
|
||||||
|
**Kept on purpose:** the wrapper + sudoers on both demo hosts (installed by hand; the old sudoers saved as
|
||||||
|
`/root/felhom-agent.sudoers.bak-pre-osapply`); agent 0.140.0 (signed update).
|
||||||
|
- **Host (DooPlex):** helper scripts in the scratchpad only; the hub password copy shredded at the end.
|
||||||
|
- **Hub:** v0.130.0 at the ruled 24 h + 1 night (the TEST override reverted and the start log shows no override);
|
||||||
|
both demo boxes ring 0, ON; release `os-20261004-091417` approved (it was approved under the TEST wait — ring 1 boxes
|
||||||
|
will install it; every version in it already runs on both demo boxes). Floor 0.291.0.
|
||||||
@@ -2,8 +2,40 @@
|
|||||||
|
|
||||||
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
|
**Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).**
|
||||||
|
|
||||||
**Updated 2026-10-04 (day): the off-site topic is closed; the OS-update test is done. Both demo boxes run controller
|
**Updated 2026-10-04 (afternoon): guest system updates are automatic on the demo boxes. Both demo boxes run
|
||||||
0.290.0 and host agent 0.139.0. Hub 0.129.0. New installs get golden 0.290.0; every box's floor is 0.290.0.**
|
controller 0.291.0 and host agent 0.140.0. Hub 0.130.0. New installs get golden 0.291.0.**
|
||||||
|
|
||||||
|
## Today (2026-10-04, afternoon): the guest's security fixes install themselves
|
||||||
|
|
||||||
|
**One decision for you (a safe default if you say nothing):**
|
||||||
|
|
||||||
|
1. **Undo for a guest update that goes wrong.** I tested it first, as you asked: Proxmox cannot take a snapshot of a
|
||||||
|
customer box at all (the box is linked to the household's drives, and Proxmox refuses). So there is no automatic undo.
|
||||||
|
Today a failed check stops, mails you, and last night's whole-box backup (minutes old) is the undo, by hand.
|
||||||
|
- **A (my pick):** keep it so. Nothing new to build. A restore takes about 1–3 minutes plus losing what apps wrote
|
||||||
|
since the backup.
|
||||||
|
- **B:** build our own disk snapshot under Proxmox. Automatic, but a new mechanism nobody has tested, with a risk to
|
||||||
|
the disk pool.
|
||||||
|
- **If you say nothing:** A stays.
|
||||||
|
|
||||||
|
**What I did:**
|
||||||
|
- **Your two choices are built.** A returning household's new box sets the old off-site copy aside on night one and
|
||||||
|
starts a new one (nothing deleted). An approved version that Debian already replaced comes from Debian's dated archive.
|
||||||
|
- **Guest security fixes now install themselves.** Each night, after the whole-box backup, the demo boxes install
|
||||||
|
Debian's fixes. When both demo boxes run the same versions healthy for 24 hours and one night, the hub approves that
|
||||||
|
set; every other box then installs exactly those versions. Docker, the host and the kernel are not touched.
|
||||||
|
- Proven live: each demo box installed 53 fixes and stayed healthy. With a 2-minute test wait the hub approved the
|
||||||
|
set; then I put back 24 hours. demo-felhom, made "ring 1" for the test, installed exactly the 3 approved versions it
|
||||||
|
lacked and nothing newer. A run where I stopped an app failed its check and mailed you (you got that mail).
|
||||||
|
- You can switch OS updates off per box in the hub. They are ON by default. The household sees one line in its
|
||||||
|
timeline: "System security fixes installed".
|
||||||
|
- **The tunnel and two other built-in programs are current** (cloudflared was 4 months old). A new monthly check
|
||||||
|
catches them falling behind. The tunnel came back by itself on both demo boxes within 20 seconds.
|
||||||
|
- **Three bugs found and fixed during the live test**, before release.
|
||||||
|
- **Two gaps found, now rows:** existing boxes cannot receive the new update tool through the product (only new
|
||||||
|
installs, or by hand — the demo boxes got it by hand); and the hub's "tunnel status" for every box always reads
|
||||||
|
"inactive" because it checks the wrong place.
|
||||||
|
- **Rows:** 4 closed (one opened and closed the same day), 5 opened. The list went from 331 to 333.
|
||||||
|
|
||||||
## Today (2026-10-04, day): off-site closed, operating-system updates measured
|
## Today (2026-10-04, day): off-site closed, operating-system updates measured
|
||||||
|
|
||||||
@@ -148,8 +180,8 @@ Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer ne
|
|||||||
|
|
||||||
## What needs you
|
## What needs you
|
||||||
|
|
||||||
0. **The two decisions at the top of today's section** (a returning household's first night; approved OS updates
|
0. **The undo choice at the top of today's section** (A: keep the backup as the undo; B: build our own snapshot).
|
||||||
when Debian has moved on). Each has a safe default if you say nothing. The Hetzner key change stays your call (3 steps, in the list).
|
If you say nothing, A stays. Earlier today's two decisions are built. The Hetzner key change stays your call.
|
||||||
1. **plant-it:** keep the hidden template as it is, or remove it entirely (its image no longer exists). **If you say
|
1. **plant-it:** keep the hidden template as it is, or remove it entirely (its image no longer exists). **If you say
|
||||||
nothing:** it stays hidden; nothing runs it.
|
nothing:** it stays hidden; nothing runs it.
|
||||||
2. **Send the SparkyFitness request, and ask the Tandoor authors** (the "Before the first paying customer" list).
|
2. **Send the SparkyFitness request, and ask the Tandoor authors** (the "Before the first paying customer" list).
|
||||||
|
|||||||
@@ -232,7 +232,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
|||||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | — | **MISSING** | `felhom-host-install.sh:2133-2136` ("No upgrades are run") | Nothing runs them after install → finding R-812, intention R-808 (added 2026-10-03). Design: `architecture/11-os-updates.md` (NOT RATIFIED, 2026-10-04). **Spike done 2026-10-04** (`audits/os-updates-spike-2026-10-04/`): measured, nothing built — still MISSING |
|
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.140.0, hub v0.130.0 | **PARTIAL — the GUEST's Debian fast lane is PROVEN-LIVE (2026-10-04); the host, Docker and the kernel are MISSING** | `audits/os-guest-lane-2026-10-04/` — ring 0 (both demo boxes) installed 53 Debian fixes each, healthy; the hub approved a 272-package release (TEST wait 2 min, then the ruled 24 h + 1 night restored); demo-felhom as ring 1 installed exactly the 3 approved versions; a stopped app → `health_failed` + operator mail. Design `architecture/11-os-updates.md` §8.1 | **No automatic undo** (a customer guest cannot be snapshotted, R-837 → R-842); existing boxes need the wrapper + sudoers by hand (R-840); host / Docker / kernel lanes not built (R-812, R-835, R-836); `felhom-host-install.sh:2133-2136` still runs no host upgrades |
|
||||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||||
|
|
||||||
|
|||||||
@@ -47,7 +47,7 @@ Owns:
|
|||||||
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
|
1. **Proxmox lifecycle** — create/start/stop/destroy guests, snapshots, storage allocation. Via a scoped Proxmox API token (the **`FelhomAgent` operator role** — `proxmox-platform.md` §3.6, validated Phase 3 B3) for everything the API covers; raw host ops only where unavoidable.
|
||||||
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
|
2. **Storage management** — attach/classify targets, reconcile the storage manifest, mount USB-by-UUID, present mounts into guests.
|
||||||
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
|
3. **Backup/restore orchestration** — vzdump to the tiers, PBS, snapshot management, and the **self-restore-test**.
|
||||||
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub.
|
4. **Host & tunnel monitoring** — host metrics, guest up/down, storage-target status, and `cloudflared` health; reports the host domain to the hub. **[FACT, 2026-10-04] The "cloudflared health" leg reads a host unit that does not exist:** `internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST, but cloudflared is a container in the guest (below), so every box reports `inactive` (`Unit cloudflared.service could not be found`, demo-hp). R-841.
|
||||||
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
|
5. **Provisioning** — provision a guest **by restoring the golden base image** (§9), deploy the controller into it, hand it its bootstrap config; also **build and refresh the golden base image** itself.
|
||||||
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
|
6. **Hub control loop** — poll for desired state + signed jobs, reconcile, execute, report, heartbeat.
|
||||||
7. **Local API** — the per-guest authorization gate the controller calls.
|
7. **Local API** — the per-guest authorization gate the controller calls.
|
||||||
@@ -64,7 +64,7 @@ Explicitly does **not**:
|
|||||||
|
|
||||||
- **Native Go binary, systemd service** on the host: boot-start, `Restart=always`, systemd watchdog (kill+restart on hang), journald logging, resource limits.
|
- **Native Go binary, systemd service** on the host: boot-start, `Restart=always`, systemd watchdog (kill+restart on hang), journald logging, resource limits.
|
||||||
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
|
- **Root-minimized (boundary settled — Phase 3 B3).** The agent runs as a **non-root** service user with the scoped `FelhomAgent` token for all API-covered work + a **narrow `sudoers` allowlist** for true host ops. Per Phase 3 (B3) the boundary is settled: the entire per-customer guest lifecycle — provision (by restore, §9), config, start/stop, snapshot, backup, **restore**, destroy — is token-covered. Genuine OS-root is confined to: (1) building/refreshing the **golden base image** (`keyctl` create is `root@pam`-only — one-time at enrollment + a maintenance cadence, §9); (2) **host mounts** (USB mount-by-UUID, systemd mount units / fstab); (3) **SMART / hardware sensors**. Root therefore never sits on the per-customer path. See `proxmox-platform.md` §3.6 for the role + boundary table.
|
||||||
- **`cloudflared` is a separate systemd service**, not embedded in the agent. This is what makes the data path survive control-plane death by construction. The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.
|
- ~~**`cloudflared` is a separate systemd service**, not embedded in the agent. … The agent **manages and health-watches** it (see §5) but the tunnel does not live or die with the agent process.~~ **[FACT, corrected 2026-10-04 — `11-os-updates.md` C8, R-838]** `cloudflared` is a **container in the customer guest** (`cloudflare/cloudflared:<pin>`), rendered and kept up by the in-guest controller (`felhom-controller/controller/internal/infra/infra.go`, `internal/stacks/infra.go`) and baked into the golden; there is no host systemd unit. It is still NOT embedded in the agent, so the data path survives the agent's death — and the agent neither manages it nor (see R-841) sees its health. Its version moves only by a controller release (the pin), on the monthly re-test (`runbooks/monthly-floating-retest.md` "Infrastructure pins").
|
||||||
|
|
||||||
## 4. Control model — reconcile + signed destructive ops
|
## 4. Control model — reconcile + signed destructive ops
|
||||||
|
|
||||||
@@ -143,7 +143,7 @@ notification) is the control.
|
|||||||
|
|
||||||
**Box-initiated poll.** The hub never connects inbound. Each poll cycle exchanges:
|
**Box-initiated poll.** The hub never connects inbound. Each poll cycle exchanges:
|
||||||
|
|
||||||
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health, agent + controller versions, audit-log tail.
|
- **Up:** heartbeat + a host-domain state report — host CPU/RAM/disk, per-guest up/down + spec, storage-target status (USB connected? NFS/CIFS reachable? PBS reachable?), last backup per target, last restore-test result, `cloudflared` health (a host-unit probe that always reads `inactive` — R-841), agent + controller versions, audit-log tail.
|
||||||
- **Down:** the current desired state, any pending signed one-shot jobs, and config (poll interval, update window, policy changes).
|
- **Down:** the current desired state, any pending signed one-shot jobs, and config (poll interval, update window, policy changes).
|
||||||
|
|
||||||
**Dead-man's-switch (essential, not optional).** In a box-initiated model the heartbeat
|
**Dead-man's-switch (essential, not optional).** In a box-initiated model the heartbeat
|
||||||
@@ -633,6 +633,13 @@ buildable until then; recorded here so the front-half built in slice 7 lands rea
|
|||||||
never — the binary is what flips. Report field `selfupdate_pending` surfaces a runs-but-never-commits
|
never — the binary is what flips. Report field `selfupdate_pending` surfaces a runs-but-never-commits
|
||||||
binary. v1 scope-outs: no hub-floor auto-update, no failed-update auto-retry (the operator re-signs),
|
binary. v1 scope-outs: no hub-floor auto-update, no failed-update auto-retry (the operator re-signs),
|
||||||
no pending-timeout auto-rollback.
|
no pending-timeout auto-rollback.
|
||||||
|
- **OS updates, guest fast lane (agent v0.140.0, `11-os-updates.md` §8.1).** A second root-owned wrapper,
|
||||||
|
`/usr/local/sbin/felhom-os-apply`, behind ONE sudoers entry (`FELHOM_OSAPPLY`: `felhom-os-apply --plan
|
||||||
|
/var/lib/felhom-agent/os/plan-*.json`). The agent has no `apt` grant of its own for this; every rule (no removal,
|
||||||
|
no downgrade, no new or unlisted package, Debian origin only, the box's own customer guest only) is in the wrapper,
|
||||||
|
red-proved per rule. The leg runs after a successful primary whole-guest backup, under the heavy-op gate.
|
||||||
|
**[FACT] The signed agent update does NOT carry the wrapper or the sudoers line** — only the installer installs
|
||||||
|
them (R-840).
|
||||||
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
|
- **Controller (the easy case — it's a guest).** The agent owns the controller's lifecycle,
|
||||||
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
|
so the **agent updates the controller**: snapshot-before-update (free rollback, because the
|
||||||
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
|
controller *is* a snapshottable guest) → pull new image → redeploy → health-check → rollback
|
||||||
|
|||||||
@@ -362,6 +362,12 @@ never moved by an update or an undo, and the unit restore accepts it.
|
|||||||
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
|
> **R-191 (2026-08-04) — this row was RIGHT and the configuration disagreed with it, weekly, for as long as R-89 has been in force.** The contract has not changed: offsite retention is ep0's, the box's token is write-only, and the box cannot delete its own history. What had not followed was the installer's `keep_last: 2` on the offsite tier, so every weekly run uploaded its snapshot successfully and then failed the whole JOB on a prune the token is refused — `whole_guest_backup_failed` in the operator's inbox about a backup that had already succeeded. Fixed in installer **1.25.0** (`keep_last: 0`) and on both live boxes; a gate now asserts it. **Verified before changing it:** ep0's two prune jobs have run every day since 2026-07-27, 18 tasks, all OK. A doc that states the contract does not enforce it — the gate does.
|
||||||
|
|
||||||
|
|
||||||
|
**[FACT, 2026-10-04 — agent v0.140.0, `11-os-updates.md` §8.1] The OS leg closes the night.** After the
|
||||||
|
whole-guest backup (the controller drives it, inside [W+2h, W+6h)) ends SUCCESSFULLY on the primary tier, the agent
|
||||||
|
waits 90 s and runs the guest's Debian fast lane — still holding the host-wide heavy-op gate, so it never overlaps a
|
||||||
|
backup or a restore-test; at most once per 20 h. The backup minutes old is the guest's undo (no snapshot is possible,
|
||||||
|
R-837). A failed or missed backup → no OS leg that night.
|
||||||
|
|
||||||
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
|
**[FACT]** The three nightly legs derive from **one** customer-settable window start W at fixed
|
||||||
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
|
offsets — db-dump at W, Tier-2 at W+60m, offsite at W+105m — so they can never be misordered
|
||||||
(`cmd/controller/main.go:604-607`). Both boxes run W = `02:30`.
|
(`cmd/controller/main.go:604-607`). Both boxes run W = `02:30`.
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
> | | |
|
> | | |
|
||||||
> |---|---|
|
> |---|---|
|
||||||
> | **Status** | **NOT RATIFIED — a PROPOSAL with one operator ruling, corrected by the 2026-10-04 spike (§7.1, corrections C1–C12 below).** Ratification is Viktor's review, not an editor's. |
|
> | **Status** | **NOT RATIFIED — a PROPOSAL with operator rulings, corrected by the 2026-10-04 spike (§7.1, C1–C12); §8 step 2 BUILT 2026-10-04 (§8.1).** Ratification is Viktor's review, not an editor's. |
|
||||||
> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. |
|
> | **Written** | 2026-10-04, by the reviewer (project Claude), before any spike. |
|
||||||
> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 |
|
> | **Verified against** | felhom.eu `d07a1a9` · felhom-controller `99a1497` (v0.290.0) · felhom-agent `d766666` (v0.138.0) · hub v0.128.0 |
|
||||||
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
|
> | **Freshness** | **CURRENT** as of 2026-10-04. The spike `TASK-backup-close-and-os-updates-spike-2026-10-04` adds measurements here as `[FACT]` and corrects every claim it disproves. Mark this file STALE when it falls behind what the product does. |
|
||||||
@@ -403,7 +403,8 @@ not covered: 79 Proxmox + 1 Tailscale on the host (slow lane / not ours), 6 Dock
|
|||||||
Each step returns to the operator for go or no-go.
|
Each step returns to the operator for go or no-go.
|
||||||
|
|
||||||
1. **Spike** (measure Q1–Q10; no product code).
|
1. **Spike** (measure Q1–Q10; no product code).
|
||||||
2. **Guest Debian, fast lane.** Lowest risk: a snapshot undo exists.
|
2. **Guest Debian, fast lane.** ~~Lowest risk: a snapshot undo exists.~~ **BUILT 2026-10-04** — agent v0.140.0, hub
|
||||||
|
v0.130.0, installer 1.29.0; §8.1. **There is no snapshot undo** (R-837, measured).
|
||||||
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
|
3. **Host Debian, fast lane** (no kernel, no Proxmox packages).
|
||||||
4. **Fleet view and alarms** (§5.7).
|
4. **Fleet view and alarms** (§5.7).
|
||||||
5. **Slow lane: Docker engine.**
|
5. **Slow lane: Docker engine.**
|
||||||
@@ -412,6 +413,47 @@ Each step returns to the operator for go or no-go.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
### 8.1 Step 2 as BUILT (2026-10-04) `[FACT]`
|
||||||
|
|
||||||
|
Evidence: `audits/os-guest-lane-2026-10-04/` (parts A–G). Brief: guest fast lane, decisions 78–80 (`09` §3).
|
||||||
|
|
||||||
|
- **The undo (R-837, measured first): none automatic.** PVE refuses ANY snapshot of a customer guest —
|
||||||
|
`PVE/AbstractConfig.pm:755-757` skips non-snapshot mounts only for a snapshot named `vzdump`, and every customer
|
||||||
|
guest carries the host-path binds mp8/mp9. As the agent's token and as root: `snapshot feature is not available`.
|
||||||
|
The token's role HAS `VM.Snapshot` / `VM.Snapshot.Rollback`. So §5.6's first row does not exist: a failed health
|
||||||
|
check stops, reports `health_failed`, and the hub mails the operator; the whole-guest backup taken minutes earlier
|
||||||
|
is the undo, by hand. The choice of a real undo is in STATUS (R-842).
|
||||||
|
- **The wrapper** `felhom-os-apply` (agent repo `configs/`), Python 3 stdlib, per §5.4.1 with refusals R1–R13 (R12:
|
||||||
|
the host layer; R13: dpkg still broken after the repair). **Changed from the draft** *(decided by CC unattended —
|
||||||
|
operator may reverse)*: one sudoers entry (`--plan <file>`); the plan's `mode` field (`inventory` / `apply` /
|
||||||
|
`health`) replaces a separate `--repair-only` (the repair runs first on every apply); Python, because a JSON plan
|
||||||
|
cannot be parsed safely in sh. The box's own customer guest is "the guest that binds `/mnt/felhom-drives`" (R10).
|
||||||
|
- **The leg** (agent `internal/osupdate`): after a SUCCESSFUL primary whole-guest backup, inside the backup's
|
||||||
|
goroutine before the host-wide heavy-op gate is released (so never beside a backup or a restore-test, C10), 90 s
|
||||||
|
after the backup, at most once per 20 h. *Decided by CC unattended — operator may reverse:* the 90 s settle and the
|
||||||
|
20 h gap; the controller's own self-update (04:30) is not detected — the 5-minute health wait absorbs a restart.
|
||||||
|
- **The health rule** (`HealthVerdict`, pinned): docker answers, the guest resolves `deb.debian.org`, the controller's
|
||||||
|
health check is `healthy`, and every container running at the START of the leg runs again (healthy if it was). The
|
||||||
|
baseline merges the inventory's reading with the apply's own — found live: an app stopped between them escaped the
|
||||||
|
first rule. Wait 5 min, poll 15 s.
|
||||||
|
- **Rings and the switch** (hub, per box): ring 0 = demo-hp + demo-felhom, everything else ring 1; switch ON by
|
||||||
|
default; OFF → the box reports, installs nothing. A box with no `os_update` block (older hub) = ring 1, ON, no
|
||||||
|
release.
|
||||||
|
- **The approval rule (the ruled "1–2 day wait")**: every Debian / Debian-Security package=version that ALL ring-0
|
||||||
|
boxes having it agree on; approved when, since that set was first seen, **24 h** passed with every ring-0 report
|
||||||
|
healthy and every ring-0 box completed **1** post-backup night run (`OS_APPROVE_AFTER`, `OS_APPROVE_NIGHTS`; an
|
||||||
|
override is logged as a TEST configuration). The approval time is the snapshot.debian.org timestamp (decision 79).
|
||||||
|
"Approve now" is an operator event. *The 24 h / 1 night numbers: decided by CC unattended within the ruled 1–2 days.*
|
||||||
|
- **The household's line**: hub event `os_update_applied` (info: on the household's hub timeline, never mailed;
|
||||||
|
hu/en in the bundle). There is no surface on the box itself (R-844).
|
||||||
|
- **Measured live**: ring 0 — 53 Debian packages on each demo box (18.7–31.7 s inside the wrapper; 174–226 s for the
|
||||||
|
whole leg incl. the inventory), healthy, 6 Docker updates "not covered"; approval — with a 2-minute TEST wait, a
|
||||||
|
272-package release approved automatically, then the ruled values restored; ring 1 — demo-felhom installed exactly
|
||||||
|
the 3 approved versions it lacked (269 already current) and left the newer Docker packages alone; a deliberately
|
||||||
|
stopped app → `health_failed` after 5 min, the operator mailed, the household line recorded.
|
||||||
|
- **Not delivered to existing boxes by the product**: the wrapper and the sudoers line reach a box only through the
|
||||||
|
installer; the signed agent update replaces the binary only (R-840). The demo boxes got them BY HAND.
|
||||||
|
|
||||||
## 9. Where the rest lives
|
## 9. Where the rest lives
|
||||||
|
|
||||||
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
- The finding: **R-812** (`backlog/OPEN-ITEMS.md`). The intention: **R-808** (`backlog/ROADMAP.md`).
|
||||||
|
|||||||
@@ -12,3 +12,7 @@ R10: 4 raise(s) removed -> suite rc=1; failing tests: test_R10_bind_only_in_a_sn
|
|||||||
R11: 9 raise(s) removed -> suite rc=1; failing tests: test_R11_bad_name, test_R11_bad_version_string, test_R11_duplicate; own test(s) failed: YES
|
R11: 9 raise(s) removed -> suite rc=1; failing tests: test_R11_bad_name, test_R11_bad_version_string, test_R11_duplicate; own test(s) failed: YES
|
||||||
R12: 1 raise(s) removed -> suite rc=1; failing tests: test_R12_host_layer; own test(s) failed: YES
|
R12: 1 raise(s) removed -> suite rc=1; failing tests: test_R12_host_layer; own test(s) failed: YES
|
||||||
R13: 1 raise(s) removed -> suite rc=1; failing tests: test_R13_repair_does_not_fix_it; own test(s) failed: YES
|
R13: 1 raise(s) removed -> suite rc=1; failing tests: test_R13_repair_does_not_fix_it; own test(s) failed: YES
|
||||||
|
== RP (wrapper conffile): the old wording (every conffile line reported as kept)
|
||||||
|
FAIL: test_updated_vs_kept (__main__.Conffiles.test_updated_vs_kept)
|
||||||
|
AssertionError: 'os-apply: CONFFILE updated /etc/debian_version (it was not changed locally)' not found in ['os-apply: START release=os-t1 layer=guest:9201 lane=fast mode=apply packages=2', 'os-apply: REPAIR configured=0 fixed=0', 'os-apply: PLAN upgrade=2 already=0 not-installed=0 from-snapshot=0', 'os-apply: CONFFILE kept /etc/debian_version', 'os-apply: CONFFILE kept /etc/ssh/sshd_config', 'os-apply: DONE rc=0 seconds=0.0 upgraded=2']
|
||||||
|
FAILED (failures=1)
|
||||||
|
|||||||
@@ -33,3 +33,9 @@ FAIL gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.563s
|
|||||||
== RP (local API hook): run the leg after a FAILED backup too
|
== RP (local API hook): run the leg after a FAILED backup too
|
||||||
afterbackup_test.go:60: the leg ran after a FAILED backup: [8200]
|
afterbackup_test.go:60: the leg ran after a FAILED backup: [8200]
|
||||||
FAIL (recorded from the run above)
|
FAIL (recorded from the run above)
|
||||||
|
== RP (selftest flag): drop the os-update case from selftestFlag.Set
|
||||||
|
selftest_flag_test.go:37: dispatched but refused by --selftest: [os-update]
|
||||||
|
FAIL gitea.dooplex.hu/admin/felhom-agent/cmd/felhom-agent 0.010s
|
||||||
|
== RP (leg): baseline = the apply's own before-reading only (the pre-fix rule)
|
||||||
|
leg_test.go:279: an app that stopped during the run passed: {RunID:20261004T040000Z Trigger:night Mode:apply Ring:0 ReleaseID:ring0-20261004T040000Z Outcome:applied Healthy:true HealthReason: VMID:9201 Upgraded:[{Name:libc6 Version: Origin:}] Installed:[] Pending:[] NotCovered:[] RestartNeeded:[] DockerRestartNeeded:false RebootNeeded:false Refused:[]}
|
||||||
|
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.195s
|
||||||
|
|||||||
@@ -0,0 +1,90 @@
|
|||||||
|
1791104489.428006469 302
|
||||||
|
1791104490.114048259 302
|
||||||
|
1791104490.867431883 302
|
||||||
|
1791104491.567710242 302
|
||||||
|
1791104492.259675188 302
|
||||||
|
1791104492.950311346 302
|
||||||
|
1791104493.658119300 302
|
||||||
|
1791104494.345122536 302
|
||||||
|
1791104495.027891977 302
|
||||||
|
1791104495.724212868 302
|
||||||
|
1791104496.422407509 302
|
||||||
|
1791104497.128727616 302
|
||||||
|
1791104497.800224130 302
|
||||||
|
1791104498.519248669 302
|
||||||
|
1791104499.203467396 302
|
||||||
|
1791104499.898045533 502
|
||||||
|
1791104500.634582730 502
|
||||||
|
1791104501.341338498 502
|
||||||
|
1791104502.054985899 502
|
||||||
|
1791104502.792687510 502
|
||||||
|
1791104503.493970117 502
|
||||||
|
1791104504.236100842 502
|
||||||
|
1791104504.980276912 502
|
||||||
|
1791104505.689072396 502
|
||||||
|
1791104506.471770326 502
|
||||||
|
1791104507.185250940 502
|
||||||
|
1791104508.316227102 502
|
||||||
|
1791104509.010368824 404
|
||||||
|
1791104509.690610573 302
|
||||||
|
1791104510.381269521 302
|
||||||
|
1791104511.048606319 302
|
||||||
|
1791104511.718435404 302
|
||||||
|
1791104512.462336250 502
|
||||||
|
1791104513.187968905 502
|
||||||
|
1791104513.871106880 302
|
||||||
|
1791104514.590549839 302
|
||||||
|
1791104515.281227748 302
|
||||||
|
1791104515.964558604 302
|
||||||
|
1791104516.720756559 302
|
||||||
|
1791104517.426584135 302
|
||||||
|
1791104518.115760697 302
|
||||||
|
1791104518.810201702 302
|
||||||
|
1791104519.507291973 302
|
||||||
|
1791104520.172275341 302
|
||||||
|
1791104520.866469463 302
|
||||||
|
1791104521.541290853 302
|
||||||
|
1791104522.219399888 302
|
||||||
|
1791104522.891993819 302
|
||||||
|
1791104523.560829339 302
|
||||||
|
1791104524.240801665 302
|
||||||
|
1791104524.946212414 302
|
||||||
|
1791104525.627388820 302
|
||||||
|
1791104526.329685610 302
|
||||||
|
1791104527.011843100 302
|
||||||
|
1791104527.684678463 302
|
||||||
|
1791104528.380214494 302
|
||||||
|
1791104529.058330821 302
|
||||||
|
1791104529.740540446 302
|
||||||
|
1791104530.436470651 302
|
||||||
|
1791104531.121601771 302
|
||||||
|
1791104531.809593181 302
|
||||||
|
1791104532.503149505 302
|
||||||
|
1791104533.173713227 302
|
||||||
|
1791104533.856540265 302
|
||||||
|
1791104534.544973430 302
|
||||||
|
1791104535.253316994 302
|
||||||
|
1791104535.932080467 302
|
||||||
|
1791104536.599649531 302
|
||||||
|
1791104537.294371521 302
|
||||||
|
1791104537.968883762 302
|
||||||
|
1791104538.662330307 302
|
||||||
|
1791104539.355172110 302
|
||||||
|
1791104540.014510174 302
|
||||||
|
1791104540.732988090 302
|
||||||
|
1791104541.409134391 302
|
||||||
|
1791104542.085884593 302
|
||||||
|
1791104542.786152655 302
|
||||||
|
1791104543.468040923 302
|
||||||
|
1791104544.129791932 302
|
||||||
|
1791104545.053502785 302
|
||||||
|
1791104545.747716347 302
|
||||||
|
1791104546.430378570 302
|
||||||
|
1791104547.089481515 302
|
||||||
|
1791104547.826344317 302
|
||||||
|
1791104548.491677889 302
|
||||||
|
1791104549.200576500 302
|
||||||
|
1791104549.873171847 302
|
||||||
|
1791104550.549674984 302
|
||||||
|
1791104551.223931502 302
|
||||||
|
1791104551.897581518 302
|
||||||
@@ -0,0 +1,93 @@
|
|||||||
|
1791104489.428473157 302
|
||||||
|
1791104490.154456409 302
|
||||||
|
1791104490.836711390 302
|
||||||
|
1791104491.512992898 302
|
||||||
|
1791104492.188921801 302
|
||||||
|
1791104492.892582650 302
|
||||||
|
1791104493.561025000 302
|
||||||
|
1791104494.243801578 302
|
||||||
|
1791104494.946931718 302
|
||||||
|
1791104495.612770541 302
|
||||||
|
1791104496.296371937 302
|
||||||
|
1791104496.959464984 302
|
||||||
|
1791104497.624069604 302
|
||||||
|
1791104498.287369613 302
|
||||||
|
1791104498.954218253 302
|
||||||
|
1791104499.625720367 302
|
||||||
|
1791104500.285823257 302
|
||||||
|
1791104500.962562788 302
|
||||||
|
1791104501.623569883 302
|
||||||
|
1791104502.280153375 502
|
||||||
|
1791104502.973783514 502
|
||||||
|
1791104503.699803496 502
|
||||||
|
1791104504.477656042 502
|
||||||
|
1791104505.219746840 502
|
||||||
|
1791104505.935741922 502
|
||||||
|
1791104506.696658409 502
|
||||||
|
1791104507.383980538 502
|
||||||
|
1791104508.099312305 502
|
||||||
|
1791104508.778478656 502
|
||||||
|
1791104509.485583013 502
|
||||||
|
1791104510.206577081 502
|
||||||
|
1791104510.894142842 502
|
||||||
|
1791104511.583745893 502
|
||||||
|
1791104512.301754435 502
|
||||||
|
1791104512.965882550 502
|
||||||
|
1791104513.674358904 502
|
||||||
|
1791104514.365515398 302
|
||||||
|
1791104515.084994873 302
|
||||||
|
1791104515.780146089 302
|
||||||
|
1791104516.467673569 302
|
||||||
|
1791104517.127579680 302
|
||||||
|
1791104517.806574655 302
|
||||||
|
1791104518.464909459 302
|
||||||
|
1791104519.137104416 302
|
||||||
|
1791104519.819783794 502
|
||||||
|
1791104520.535296664 530
|
||||||
|
1791104521.190614199 302
|
||||||
|
1791104521.864351794 302
|
||||||
|
1791104522.541227619 302
|
||||||
|
1791104523.210468864 302
|
||||||
|
1791104523.862415986 302
|
||||||
|
1791104524.545740539 302
|
||||||
|
1791104525.247498962 302
|
||||||
|
1791104525.907280086 302
|
||||||
|
1791104526.568936189 302
|
||||||
|
1791104527.226485539 302
|
||||||
|
1791104527.883240645 302
|
||||||
|
1791104528.572125595 302
|
||||||
|
1791104529.261611733 302
|
||||||
|
1791104529.990339048 302
|
||||||
|
1791104530.672950131 302
|
||||||
|
1791104531.335651483 302
|
||||||
|
1791104532.020457877 302
|
||||||
|
1791104532.681731917 302
|
||||||
|
1791104533.347279268 302
|
||||||
|
1791104533.999334640 302
|
||||||
|
1791104534.662664742 302
|
||||||
|
1791104535.326620624 302
|
||||||
|
1791104535.998590184 302
|
||||||
|
1791104536.651396219 302
|
||||||
|
1791104537.299040037 302
|
||||||
|
1791104537.937915214 302
|
||||||
|
1791104538.640295166 302
|
||||||
|
1791104539.329815017 302
|
||||||
|
1791104539.985113137 302
|
||||||
|
1791104540.672365246 302
|
||||||
|
1791104541.358812693 302
|
||||||
|
1791104542.035139317 302
|
||||||
|
1791104542.698601778 302
|
||||||
|
1791104543.351804369 302
|
||||||
|
1791104544.018269004 302
|
||||||
|
1791104544.686076951 302
|
||||||
|
1791104545.360390747 302
|
||||||
|
1791104546.033087077 302
|
||||||
|
1791104546.697562917 302
|
||||||
|
1791104547.355845574 302
|
||||||
|
1791104548.046312941 302
|
||||||
|
1791104548.728344788 302
|
||||||
|
1791104549.396103966 302
|
||||||
|
1791104550.084478791 302
|
||||||
|
1791104550.772128650 302
|
||||||
|
1791104551.425494194 302
|
||||||
|
1791104552.081996680 302
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
demo-hp (public URL through Cloudflare, 09:01:30-09:06:00Z): samples 92 codes {'302': 73, '502': 18, '530': 1}; first non-302 09:01:42Z, last 09:02:00Z; gap at most 19.6 s
|
||||||
|
demo-felhom (public URL through Cloudflare, 09:01:30-09:06:00Z): samples 89 codes {'302': 74, '502': 14, '404': 1}; first non-302 09:01:39Z, last 09:01:53Z; gap at most 14.7 s
|
||||||
@@ -0,0 +1,306 @@
|
|||||||
|
=== felhom-agent 0.140.0-rc2 selftest=os-update vmid=9201 ring=0 enabled=true release=false ===
|
||||||
|
time=2026-10-04T11:08:52.934+02:00 level=INFO msg="osupdate: START" run=20261004T090852Z vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T090852Z
|
||||||
|
time=2026-10-04T11:09:08.241+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T090852Z layer=guest:9201 lane=fast mode=inventory packages=0"
|
||||||
|
time=2026-10-04T11:11:46.297+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T090852Z layer=guest:9201 lane=fast mode=apply packages=53"
|
||||||
|
time=2026-10-04T11:11:46.297+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||||
|
time=2026-10-04T11:11:46.297+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=53 already=0 not-installed=0 from-snapshot=0"
|
||||||
|
time=2026-10-04T11:11:46.297+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE updated /etc/debian_version (it was not changed locally)"
|
||||||
|
time=2026-10-04T11:11:46.297+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=18.7 upgraded=53"
|
||||||
|
time=2026-10-04T11:11:46.297+02:00 level=INFO msg="osupdate: DONE" run=20261004T090852Z vmid=9201 ring=0 trigger=debug outcome=applied healthy=true reason="" upgraded=53 pending=6 not_covered=6 restart_needed=7
|
||||||
|
--- os-update report ---
|
||||||
|
{
|
||||||
|
"docker_restart_needed": true,
|
||||||
|
"health_reason": "",
|
||||||
|
"healthy": true,
|
||||||
|
"mode": "apply",
|
||||||
|
"not_covered": [
|
||||||
|
"docker-ce-cli",
|
||||||
|
"containerd.io",
|
||||||
|
"docker-ce",
|
||||||
|
"docker-buildx-plugin",
|
||||||
|
"docker-ce-rootless-extras",
|
||||||
|
"docker-compose-plugin"
|
||||||
|
],
|
||||||
|
"outcome": "applied",
|
||||||
|
"pending": 6,
|
||||||
|
"refused": null,
|
||||||
|
"release_id": "ring0-20261004T090852Z",
|
||||||
|
"restart_needed": [
|
||||||
|
"agetty",
|
||||||
|
"containerd",
|
||||||
|
"cron",
|
||||||
|
"dbus-daemon",
|
||||||
|
"dhclient",
|
||||||
|
"sshd",
|
||||||
|
"systemd-logind"
|
||||||
|
],
|
||||||
|
"ring": 0,
|
||||||
|
"run_id": "20261004T090852Z",
|
||||||
|
"upgraded": [
|
||||||
|
{
|
||||||
|
"name": "libc6",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "base-files",
|
||||||
|
"version": "13.8+deb13u7",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bash",
|
||||||
|
"version": "5.2.37-2+b10",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bsdutils",
|
||||||
|
"version": "1:2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "gzip",
|
||||||
|
"version": "1.13-1+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libperl5.40",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "perl",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "perl-base",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "perl-modules-5.40",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "liblastlog2-2",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bsdextrautils",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "util-linux-extra",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libblkid1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libmount1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libsmartcols1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "mount",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "fdisk",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libuuid1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "util-linux",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libfdisk1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libaudit-common",
|
||||||
|
"version": "1:4.0.2-2+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libaudit1",
|
||||||
|
"version": "1:4.0.2-2+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libsqlite3-0",
|
||||||
|
"version": "3.46.1-7+deb13u2",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libc-bin",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "login",
|
||||||
|
"version": "1:4.16.0-2+really2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "logsave",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libext2fs2t64",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "e2fsprogs",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libexpat1",
|
||||||
|
"version": "2.8.3-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl-provider-legacy",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libssl3t64",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "postfix",
|
||||||
|
"version": "3.10.13-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "python3.13",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libpython3.13-stdlib",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "python3.13-minimal",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libpython3.13-minimal",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "tzdata",
|
||||||
|
"version": "2026c-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcap2",
|
||||||
|
"version": "1:2.75-10+deb13u1+b3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libpcre2-8-0",
|
||||||
|
"version": "10.46-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "dhcpcd-base",
|
||||||
|
"version": "1:10.1.0-11+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bind9-dnsutils",
|
||||||
|
"version": "1:9.20.29-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bind9-host",
|
||||||
|
"version": "1:9.20.29-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bind9-libs",
|
||||||
|
"version": "1:9.20.29-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libc-l10n",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "locales",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libssh2-1t64",
|
||||||
|
"version": "1.11.1-1+deb13u2",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "curl",
|
||||||
|
"version": "8.14.1-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcurl4t64",
|
||||||
|
"version": "8.14.1-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcap2-bin",
|
||||||
|
"version": "1:2.75-10+deb13u1+b3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcom-err2",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcurl3t64-gnutls",
|
||||||
|
"version": "8.14.1-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libss2",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,306 @@
|
|||||||
|
=== felhom-agent 0.140.0-rc2 selftest=os-update vmid=9201 ring=0 enabled=true release=false ===
|
||||||
|
time=2026-10-04T11:04:21.448+02:00 level=INFO msg="osupdate: START" run=20261004T090421Z vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T090421Z
|
||||||
|
time=2026-10-04T11:04:42.234+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T090421Z layer=guest:9201 lane=fast mode=inventory packages=0"
|
||||||
|
time=2026-10-04T11:08:07.757+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T090421Z layer=guest:9201 lane=fast mode=apply packages=53"
|
||||||
|
time=2026-10-04T11:08:07.757+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||||
|
time=2026-10-04T11:08:07.757+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=53 already=0 not-installed=0 from-snapshot=0"
|
||||||
|
time=2026-10-04T11:08:07.757+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: CONFFILE kept /etc/debian_version"
|
||||||
|
time=2026-10-04T11:08:07.757+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=31.7 upgraded=53"
|
||||||
|
time=2026-10-04T11:08:07.758+02:00 level=INFO msg="osupdate: DONE" run=20261004T090421Z vmid=9201 ring=0 trigger=debug outcome=applied healthy=true reason="" upgraded=53 pending=6 not_covered=6 restart_needed=7
|
||||||
|
--- os-update report ---
|
||||||
|
{
|
||||||
|
"docker_restart_needed": true,
|
||||||
|
"health_reason": "",
|
||||||
|
"healthy": true,
|
||||||
|
"mode": "apply",
|
||||||
|
"not_covered": [
|
||||||
|
"docker-ce-cli",
|
||||||
|
"containerd.io",
|
||||||
|
"docker-ce",
|
||||||
|
"docker-buildx-plugin",
|
||||||
|
"docker-ce-rootless-extras",
|
||||||
|
"docker-compose-plugin"
|
||||||
|
],
|
||||||
|
"outcome": "applied",
|
||||||
|
"pending": 6,
|
||||||
|
"refused": null,
|
||||||
|
"release_id": "ring0-20261004T090421Z",
|
||||||
|
"restart_needed": [
|
||||||
|
"agetty",
|
||||||
|
"containerd",
|
||||||
|
"cron",
|
||||||
|
"dbus-daemon",
|
||||||
|
"dhclient",
|
||||||
|
"sshd",
|
||||||
|
"systemd-logind"
|
||||||
|
],
|
||||||
|
"ring": 0,
|
||||||
|
"run_id": "20261004T090421Z",
|
||||||
|
"upgraded": [
|
||||||
|
{
|
||||||
|
"name": "libc6",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "base-files",
|
||||||
|
"version": "13.8+deb13u7",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bash",
|
||||||
|
"version": "5.2.37-2+b10",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bsdutils",
|
||||||
|
"version": "1:2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "gzip",
|
||||||
|
"version": "1.13-1+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libperl5.40",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "perl",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "perl-base",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "perl-modules-5.40",
|
||||||
|
"version": "5.40.1-6+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "liblastlog2-2",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bsdextrautils",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "util-linux-extra",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libblkid1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libmount1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libsmartcols1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "mount",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "fdisk",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libuuid1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "util-linux",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libfdisk1",
|
||||||
|
"version": "2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libaudit-common",
|
||||||
|
"version": "1:4.0.2-2+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libaudit1",
|
||||||
|
"version": "1:4.0.2-2+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libsqlite3-0",
|
||||||
|
"version": "3.46.1-7+deb13u2",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libc-bin",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "login",
|
||||||
|
"version": "1:4.16.0-2+really2.41.5-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "logsave",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libext2fs2t64",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "e2fsprogs",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libexpat1",
|
||||||
|
"version": "2.8.3-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl-provider-legacy",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libssl3t64",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "postfix",
|
||||||
|
"version": "3.10.13-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "python3.13",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libpython3.13-stdlib",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "python3.13-minimal",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libpython3.13-minimal",
|
||||||
|
"version": "3.13.5-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "tzdata",
|
||||||
|
"version": "2026c-0+deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcap2",
|
||||||
|
"version": "1:2.75-10+deb13u1+b3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libpcre2-8-0",
|
||||||
|
"version": "10.46-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "dhcpcd-base",
|
||||||
|
"version": "1:10.1.0-11+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bind9-dnsutils",
|
||||||
|
"version": "1:9.20.29-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bind9-host",
|
||||||
|
"version": "1:9.20.29-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "bind9-libs",
|
||||||
|
"version": "1:9.20.29-1~deb13u1",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libc-l10n",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "locales",
|
||||||
|
"version": "2.41-12+deb13u4",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libssh2-1t64",
|
||||||
|
"version": "1.11.1-1+deb13u2",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "curl",
|
||||||
|
"version": "8.14.1-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcurl4t64",
|
||||||
|
"version": "8.14.1-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcap2-bin",
|
||||||
|
"version": "1:2.75-10+deb13u1+b3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcom-err2",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libcurl3t64-gnutls",
|
||||||
|
"version": "8.14.1-2+deb13u5",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libss2",
|
||||||
|
"version": "1.47.2-3+b12",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
=== felhom-agent 0.140.0-rc2 selftest=os-update vmid=9201 ring=1 enabled=true release=true ===
|
||||||
|
time=2026-10-04T11:15:59.945+02:00 level=INFO msg="osupdate: START" run=20261004T091559Z vmid=9201 ring=1 trigger=debug enabled=true release=os-20261004-091417
|
||||||
|
time=2026-10-04T11:16:10.634+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-20261004-091417 layer=guest:9201 lane=fast mode=inventory packages=0"
|
||||||
|
time=2026-10-04T11:20:04.641+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=os-20261004-091417 layer=guest:9201 lane=fast mode=apply packages=272"
|
||||||
|
time=2026-10-04T11:20:04.641+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||||
|
time=2026-10-04T11:20:04.641+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=3 already=269 not-installed=0 from-snapshot=0"
|
||||||
|
time=2026-10-04T11:20:04.641+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=9.3 upgraded=3"
|
||||||
|
time=2026-10-04T11:20:04.642+02:00 level=INFO msg="osupdate: DONE" run=20261004T091559Z vmid=9201 ring=1 trigger=debug outcome=applied healthy=true reason="" upgraded=3 pending=6 not_covered=6 restart_needed=10
|
||||||
|
--- os-update report ---
|
||||||
|
{
|
||||||
|
"docker_restart_needed": true,
|
||||||
|
"health_reason": "",
|
||||||
|
"healthy": true,
|
||||||
|
"mode": "apply",
|
||||||
|
"not_covered": [
|
||||||
|
"docker-ce-cli",
|
||||||
|
"containerd.io",
|
||||||
|
"docker-ce",
|
||||||
|
"docker-buildx-plugin",
|
||||||
|
"docker-ce-rootless-extras",
|
||||||
|
"docker-compose-plugin"
|
||||||
|
],
|
||||||
|
"outcome": "applied",
|
||||||
|
"pending": 6,
|
||||||
|
"refused": null,
|
||||||
|
"release_id": "os-20261004-091417",
|
||||||
|
"restart_needed": [
|
||||||
|
"agetty",
|
||||||
|
"containerd",
|
||||||
|
"cron",
|
||||||
|
"dbus-daemon",
|
||||||
|
"dhclient",
|
||||||
|
"sshd",
|
||||||
|
"systemd",
|
||||||
|
"systemd-journal",
|
||||||
|
"systemd-logind",
|
||||||
|
"systemd-network"
|
||||||
|
],
|
||||||
|
"ring": 1,
|
||||||
|
"run_id": "20261004T091559Z",
|
||||||
|
"upgraded": [
|
||||||
|
{
|
||||||
|
"name": "libssl3t64",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl-provider-legacy",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
=== felhom-agent 0.140.0-rc3 selftest=os-update vmid=9201 ring=0 enabled=true release=false ===
|
||||||
|
time=2026-10-04T11:22:44.828+02:00 level=INFO msg="osupdate: START" run=20261004T092244Z vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T092244Z
|
||||||
|
time=2026-10-04T11:22:59.130+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T092244Z layer=guest:9201 lane=fast mode=inventory packages=0"
|
||||||
|
time=2026-10-04T11:23:37.731+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T092244Z layer=guest:9201 lane=fast mode=apply packages=3"
|
||||||
|
time=2026-10-04T11:23:37.731+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: REPAIR configured=0 fixed=0"
|
||||||
|
time=2026-10-04T11:23:37.731+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: PLAN upgrade=3 already=0 not-installed=0 from-snapshot=0"
|
||||||
|
time=2026-10-04T11:23:37.731+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: DONE rc=0 seconds=3.8 upgraded=3"
|
||||||
|
time=2026-10-04T11:28:41.271+02:00 level=INFO msg="osupdate: DONE" run=20261004T092244Z vmid=9201 ring=0 trigger=debug outcome=health_failed healthy=false reason="privatebin was running and is not" upgraded=3 pending=6 not_covered=6 restart_needed=10
|
||||||
|
--- os-update report ---
|
||||||
|
{
|
||||||
|
"docker_restart_needed": true,
|
||||||
|
"health_reason": "privatebin was running and is not",
|
||||||
|
"healthy": false,
|
||||||
|
"mode": "apply",
|
||||||
|
"not_covered": [
|
||||||
|
"docker-ce-cli",
|
||||||
|
"containerd.io",
|
||||||
|
"docker-ce",
|
||||||
|
"docker-buildx-plugin",
|
||||||
|
"docker-ce-rootless-extras",
|
||||||
|
"docker-compose-plugin"
|
||||||
|
],
|
||||||
|
"outcome": "health_failed",
|
||||||
|
"pending": 6,
|
||||||
|
"refused": null,
|
||||||
|
"release_id": "ring0-20261004T092244Z",
|
||||||
|
"restart_needed": [
|
||||||
|
"agetty",
|
||||||
|
"containerd",
|
||||||
|
"cron",
|
||||||
|
"dbus-daemon",
|
||||||
|
"dhclient",
|
||||||
|
"sshd",
|
||||||
|
"systemd",
|
||||||
|
"systemd-journal",
|
||||||
|
"systemd-logind",
|
||||||
|
"systemd-network"
|
||||||
|
],
|
||||||
|
"ring": 0,
|
||||||
|
"run_id": "20261004T092244Z",
|
||||||
|
"upgraded": [
|
||||||
|
{
|
||||||
|
"name": "openssl-provider-legacy",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "libssl3t64",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "openssl",
|
||||||
|
"version": "3.5.7-1~deb13u3",
|
||||||
|
"origin": ""
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
=== felhom-agent 0.140.0 selftest=os-update vmid=9201 ring=0 enabled=true release=false ===
|
||||||
|
time=2026-10-04T11:38:36.515+02:00 level=INFO msg="osupdate: START" run=20261004T093836Z vmid=9201 ring=0 trigger=debug enabled=true release=ring0-20261004T093836Z
|
||||||
|
time=2026-10-04T11:38:50.533+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: START release=ring0-20261004T093836Z layer=guest:9201 lane=fast mode=inventory packages=0"
|
||||||
|
time=2026-10-04T11:38:50.534+02:00 level=INFO msg="osupdate: DONE" run=20261004T093836Z vmid=9201 ring=0 trigger=debug outcome=nothing healthy=true reason="" upgraded=0 pending=6 not_covered=6 restart_needed=10
|
||||||
|
--- os-update report ---
|
||||||
|
{
|
||||||
|
"docker_restart_needed": true,
|
||||||
|
"health_reason": "",
|
||||||
|
"healthy": true,
|
||||||
|
"mode": "inventory",
|
||||||
|
"not_covered": [
|
||||||
|
"docker-ce-cli",
|
||||||
|
"containerd.io",
|
||||||
|
"docker-ce",
|
||||||
|
"docker-buildx-plugin",
|
||||||
|
"docker-ce-rootless-extras",
|
||||||
|
"docker-compose-plugin"
|
||||||
|
],
|
||||||
|
"outcome": "nothing",
|
||||||
|
"pending": 6,
|
||||||
|
"refused": null,
|
||||||
|
"release_id": "ring0-20261004T093836Z",
|
||||||
|
"restart_needed": [
|
||||||
|
"agetty",
|
||||||
|
"containerd",
|
||||||
|
"cron",
|
||||||
|
"dbus-daemon",
|
||||||
|
"dhclient",
|
||||||
|
"sshd",
|
||||||
|
"systemd",
|
||||||
|
"systemd-journal",
|
||||||
|
"systemd-logind",
|
||||||
|
"systemd-network"
|
||||||
|
],
|
||||||
|
"ring": 0,
|
||||||
|
"run_id": "20261004T093836Z",
|
||||||
|
"upgraded": null
|
||||||
|
}
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
os_update_applied|System security fixes installed (3 package(s)).|hub|Oct 04 09:28
|
||||||
|
os_update_health_failed|OS update on demo-hp-bb76ea: the guest was NOT healthy after 3 package(s) were installed (privatebin was running and is not). Nothing was undone automatically (no guest snapshot is possible, R-837); last night's whole-guest backup is th
|
||||||
|
os_update_applied|System security fixes installed (3 package(s)).|hub|Oct 04 09:21
|
||||||
|
os_update_applied|System security fixes installed (53 package(s)).|hub|Oct 04 09:03
|
||||||
|
sent|OS update on demo-hp-bb76ea: the guest was NOT healthy after 3 package(s) were installed (privatebin was running and is not). Nothing was undone automatically (no guest snapshot is possible, R-837); last night's whole-guest backup is the undo.|Oct 04 09:2
|
||||||
|
skipped|OS update on demo-hp-bb76ea: the guest was NOT healthy after 3 package(s) were installed (privatebin was running and is not). Nothing was undone automatically (no guest snapshot is possible, R-837); last night's whole-guest backup is the undo.|Oct 04 0
|
||||||
@@ -26,6 +26,17 @@
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## 2026-10-04 (day) — OS updates, guest fast lane (agent v0.140.0, hub v0.130.0, controller v0.291.0)
|
||||||
|
|
||||||
|
> Evidence: `audits/os-guest-lane-2026-10-04/`.
|
||||||
|
|
||||||
|
| Row | What | Closed | Evidence |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **R-837** | **The guest snapshot undo was unmeasured.** MEASURED: impossible — PVE refuses any non-`vzdump` snapshot of a guest with host-path binds (mp8/mp9), as the agent's token (which has the rights) and as root (`PVE/AbstractConfig.pm:755-757`). The automatic undo was not built (the brief's stop rule); the decision is R-842. **Rule kept: a guest with host binds has no PVE snapshot; the night's backup is its undo.** | CLOSED 2026-10-04 — MEASURED | `partA/README.md` |
|
||||||
|
| **R-838** | **The infrastructure images never moved** (cloudflared four months behind). Controller v0.291.0: traefik v3.7.13, cloudflared 2026.9.3, filebrowser 1.5.6-stable (release notes read; nothing we use breaks). A controller release DOES move all three (bring-up + start-up mount sync; measured on 9202 and both demo boxes, public gap ≤ 19.6 s); `scripts/check-infra-pins.py` + the runbook's "Infrastructure pins" section put them on the monthly re-test. | CLOSED 2026-10-04 — FIXED controller v0.291.0 | `partF/` |
|
||||||
|
| **R-726** | **A returning household's new box made no off-site copy on night one.** Decision 78 built (controller v0.291.0): a claimed box that has never made an off-site copy sets the orphaned old copy aside and starts a new one; nothing deleted; a box that has made copies still asks. Red-proved both ways. | CLOSED 2026-10-04 — FIXED controller v0.291.0 | `partE/r726-redproof.txt` |
|
||||||
|
| **R-843** | **`--selftest=wgtunnel` (since S3) and `--selftest=os-update` were dispatched but refused by the flag's allow-list.** Found live (the OS debug action could not run); both accepted in agent v0.140.0; `TestSelftestFlag_AcceptsEveryDispatchedMode` pins every dispatched mode. **Rule kept: a dispatch case without a flag case is dead code — the test reads both.** | CLOSED 2026-10-04 — FIXED agent v0.140.0 (opened and closed the same day) | `partC/agent-leg-redproofs.txt` |
|
||||||
|
|
||||||
## 2026-10-04 (day) — the off-site topic closed (hub v0.129.0, agent v0.139.0)
|
## 2026-10-04 (day) — the off-site topic closed (hub v0.129.0, agent v0.139.0)
|
||||||
|
|
||||||
> Evidence: `audits/backup-close-2026-10-04/`.
|
> Evidence: `audits/backup-close-2026-10-04/`.
|
||||||
|
|||||||
@@ -191,7 +191,7 @@ stopping line that lies.
|
|||||||
| **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC |
|
| **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC |
|
||||||
| **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator |
|
| **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator |
|
||||||
|
|
||||||
## Backup & restore — 55 rows (P2 11, P3 23, P4 21)
|
## Backup & restore — 54 rows (P2 10, P3 23, P4 21)
|
||||||
|
|
||||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||||
|---|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|---|
|
||||||
@@ -204,7 +204,6 @@ stopping line that lies.
|
|||||||
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller)** | — | — | CC |
|
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller)** | — | — | CC |
|
||||||
| **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | — | — | CC |
|
| **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | — | — | CC |
|
||||||
| **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC |
|
| **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC |
|
||||||
| **R-726** | Backup & restore | P2 | **[P2-MEDIUM] A new box for a customer who had one before makes NO off-site copy on night one: the old repository is found orphaned, and the fix is a button nobody pointed the household to.** MEASURED 2026-09-30 00:15 UTC on the new-household drill box (`tester-1`, whose previous box was deleted 2026-09-17): `[offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset` → `offbox_repo_orphaned` (warning) to the household's timeline and an operator mail; `offsite-integrity` then checked nothing. The household's page is honest („A távoli tároló másik kulccsal készült mentéseket tartalmaz … Új távoli mentés indítása…", old history set aside, never deleted), but the evening before, the recovery-code ceremony and the off-site page raised nothing, and the guide does not mention it. The hub re-issued the off-site credentials on re-enroll by itself; it could have known the repository would orphan. Customer data was never at risk (the old repository is untouched); the household simply has no off-site copy until someone presses the button. **Fix direction:** offer the reset at the recovery-code ceremony when the repository already holds another key's snapshots, or the hub's re-enroll re-issue sets the old history aside the same way (it is the same move-aside), and the guide says so. | **WAITING-ON-OPERATOR 2026-10-04 — two options in STATUS (A: a fresh box sets the old history aside by itself on its first night, as an unclaimed box already does; B: ask at the recovery-code evening). CC's pick: A. Nothing built.** | — | — | CC + operator |
|
|
||||||
| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC |
|
| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC |
|
||||||
| **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC |
|
| **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC |
|
||||||
| **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC |
|
| **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC |
|
||||||
@@ -268,7 +267,7 @@ stopping line that lies.
|
|||||||
| **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks/<n>/deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC |
|
| **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks/<n>/deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC |
|
||||||
| **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC |
|
| **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC |
|
||||||
|
|
||||||
## Security & access — 31 rows (P2 4, P3 24, P4 3)
|
## Security & access — 30 rows (P2 3, P3 24, P4 3)
|
||||||
|
|
||||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||||
|---|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|---|
|
||||||
@@ -302,15 +301,14 @@ stopping line that lies.
|
|||||||
| **R-134** | Security & access | P4 | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC |
|
| **R-134** | Security & access | P4 | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC |
|
||||||
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
|
| **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC |
|
||||||
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
|
| **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator |
|
||||||
| **R-838** | Security & access | P2 | **`cloudflared` — the internet-facing tunnel on every box — never moves: pinned at `2026.6.0` since 2026-06-11; upstream is `2026.9.3`.** MEASURED 2026-10-04 (`11` C8, Q9): it is not a host package but a container in the guest, pinned in `controller/internal/infra/infra.go:26` and baked into the golden; traefik and filebrowser are pinned the same way. Nothing re-tests or raises these pins on a schedule; the app-image update arc (`09`) covers catalog apps, not these. Fix direction: put the infrastructure pins on the same monthly re-test as app images, raised by a controller release. `audits/os-updates-spike-2026-10-04/README.md` (Q9) | **READY — owner: CC (pin raise) + operator (cadence)** | — | — | CC |
|
|
||||||
|
|
||||||
## Box system & updates — 20 rows (P2 3, P3 14, P4 3)
|
## Box system & updates — 22 rows (P2 4, P3 15, P4 3)
|
||||||
|
|
||||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||||
|---|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|---|
|
||||||
| **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator |
|
| **R-530** | Box system & updates | P2 | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. **NARROWED 2026-09-16 (operator ruling 1):** the keys stay on DooPlex owner-only and CC may sign `agent_update` until the first PAYING customer (testers excluded) — recorded in `CONTEXT.md` + `04-control-plane-authorization.md` §3.1. Both demo boxes now run agent 0.131.0 (demo-hp 2026-09-15, demo-felhom 2026-09-16, each by a per-box signed job; Peti's box untouched, still 0.130.0). **What remains:** a fleet rollout step — signing per box does not scale past a handful, and nothing lists which boxes are behind. **2026-09-25:** Peti's box was RETIRED (operator ruling) — it no longer counts as a box behind. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** | — | — | operator |
|
||||||
| **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC |
|
| **R-604** | Box system & updates | P2 | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | — | — | CC |
|
||||||
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
|
| **R-812** | Box system & updates | P2 | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** SEARCHED 2026-10-03 (read-only): `felhom-controller`, `felhom-agent`, `app-catalog-felhom.eu` and `felhom.eu` hold no `apt-get upgrade`, `apt full-upgrade`, `unattended-upgrades`, `pveupgrade` or `needrestart` that runs on a box. The installer aligns the host's Proxmox repositories to no-subscription *"so the box can pull security updates"* and then says plainly *"No upgrades are run"* (`scripts/felhom-host-install.sh:2133-2136`). The guest's Docker engine is installed when the golden is BAKED (`felhom-agent/configs/build-golden.sh:124-125`), so a fresh install gets that week's engine and an installed box keeps it forever. The only `apt full-upgrade` in the project is a by-hand step for the off-site endpoint ep0 (`documentation/runbooks/offsite-endpoint.md:41`), not a box. App images ARE updated (the update arc); the layer under them is not. The intention, with its scope, is **R-808** in `ROADMAP.md`. | **NARROWED AGAIN 2026-10-04 (day) — the GUEST's Debian fast lane is BUILT and proven live (agent v0.140.0, hub v0.130.0, installer 1.29.0; `11` §8.1, `audits/os-guest-lane-2026-10-04/`). LEFT: the host, Docker and kernel lanes; the undo (R-842); existing boxes (R-840).** Earlier: NARROWED 2026-10-04 — the SPIKE is done (`11` §7.1, corrections C1–C12, `audits/os-updates-spike-2026-10-04/`); no product code yet. LEFT: the build steps of `11` §8, each with the operator's go; the §5.3 snapshot question is in STATUS; preconditions R-835, R-836, R-837.** | — | — | CC + operator |
|
||||||
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
|
| **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC |
|
||||||
| **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/<path>` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC |
|
| **R-50b** | Box system & updates | P3 | **[P2] A root-owned privileged host artifact is delivered unversioned from `main` — "which wrapper is on this host?" is unanswerable.** `configs/felhom-pbs-apply` installs to `/usr/local/sbin/felhom-pbs-apply` (0755 root:root) and is the pinned sudoers vector for `create\ |reconcile\|grant` against `/etc/pve/priv/storage`. It is fetched by `felhom-host-install.sh:1914` via `fetch_raw`, which hits `raw/branch/main/<path>` — **no tag, no pin, no checksum, and no record in the Day-0 artifact manifest**, unlike the agent binary (sha256-vouched) and the golden image. Three consequences: (1) two hosts installed a week apart can carry different privileged wrapper code while both reporting the same agent version; (2) a host hotfixed in place (felhom-pve, 2026-07-18) is indistinguishable from one that fetched the same content — the fleet has no inventory of it; (3) an accidental push to `main` reaches the next install of every host with no review gate between commit and root-owned deployment. | **NARROWED** — **(a) SHIPPED 2026-07-21; (b)/(c) open** — moved from `ROADMAP.md` 2026-10-03: it states a checkable fact about the shipped product, so it is a FINDING (the sorting rule). **Re-ranked 2026-10-03: [P2] → P3 — operator-only; leg (a) shipped.** | — | **Surfaced 2026-07-21 while stopping the R-39 v0.90.1 publish** (`felhom-controller/REPORT.md` §5): the publish was cancelled precisely because the version number would have claimed to carry a fix that in fact rides this unversioned channel. Candidate shapes, in increasing cost: (a) record the wrapper's sha256 in the Day-0 artifact manifest beside the agent binary and have the agent report the installed file's hash, so drift is at least *visible*; (b) `fetch_raw` takes a pinned ref (tag or commit) supplied by the manifest rather than `main`; (c) the wrapper becomes a published generic-registry artifact with the same gate ladder as the agent binary. **(a) is the cheap honest first step and would have caught this class already.** Pairs with R-39 (whose remaining fleet half is specced separately) **(a) SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** `ArtifactManifest.WrapperSHA256` + an operator field; agents report the installed wrapper's sha256 each cycle and the host page surfaces a mismatch. **An unknown on EITHER side reads as quiet, never as drift** — lighting every host amber on rollout day is how a warning becomes background noise. Live confirmation of exactly the problem: felhom-pve's July-18 in-place hotfix hashed `2888f2ea…`, matching **no commit anyone could name**; it now reports `104db0a4…` against a vouchable manifest value. **(b)/(c) REMAIN OPEN:** the wrapper is still fetched unversioned from `raw/branch/main` — this makes drift *visible*, it does not fix the channel. Also recorded: the 0440 sudoers file is not agent-readable, so its drift stays invisible. **Flips (2026-10-03):** `00` §A "The installer is PUBLISHED, not pushed" — the same discipline for the privileged wrappers. **Re-ranked 2026-10-03:** [P2] → P3: operator-only; (a) shipped 2026-07-21 (the report carries the wrapper sha256), (b)/(c) open. **Finding-shaped** — an R-424 instance; check against today's product before building. | CC |
|
||||||
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
|
| **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |
|
||||||
@@ -327,8 +325,10 @@ stopping line that lies.
|
|||||||
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
| **R-373** | Box system & updates | P4 | **`SysDataGrowGB` is the intended lever for the system-data volume, it works, and nothing sets it.** Written down 2026-08-02 in `audits/SPIKE-recovery-unit-space-2026-08-02.md:230-232`, under an explicit *"### Not filed"* heading: the 20 G / 50 G mismatch was ruled a tier-sizing decision rather than a defect, *"`SysDataGrowGB` is the intended lever and it works; nothing sets it."* A lever with no caller is the same shape as R-368's comment — a setting that names a behaviour nothing invokes. **Age when filed: 20 days.** | **OPEN — LOW** | R-368 (same shape) | Either wire it to something an operator can reach, or remove it and record the sizing decision where a reader will meet it. | CC |
|
||||||
| **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC |
|
| **R-835** | Box system & updates | P3 | **Turning Docker's `live-restore` OFF with a restart stops every running container and starts none.** MEASURED 2026-10-04 on scratch 9202: `live-restore` on (via `systemctl reload docker`, which does enable it) kept all 6 containers running across two engine steps; `systemctl reload` with the baked `daemon.json` did NOT turn it off; a `systemctl restart docker` did — and the new daemon stopped every container (`Exited (0)`, `Removing stale sandbox … isRestore=false`) and restarted none, though all are `unless-stopped`. Nothing brought them back for 3.5 min. A precondition for the Docker slow lane (`11` C5): if `live-restore` ships, turning it off must be a guarded act (stop apps first), never a plain restart. `audits/os-updates-spike-2026-10-04/partG/` | **READY — design input, owner: CC** | — | — | CC |
|
||||||
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **READY — measure before the kernel slow lane; owner: CC + operator (reboots)** | — | — | CC |
|
| **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin <new> --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **READY — measure before the kernel slow lane; owner: CC + operator (reboots)** | — | — | CC |
|
||||||
| **R-837** | Box system & updates | P3 | **The guest snapshot undo (`11` §5.6) is unmeasured.** 2026-10-04: customer guests sit on LVM-thin and can snapshot, but scratch 9202 sits on `dir` storage (`snapshot feature is not available`), so the spike measured a backup-and-restore undo instead (73 s down, all apps back, libc back). The snapshot + rollback time on LVM-thin, and what it does to the thin pool (R-672's shape), were not measured — no fenced LVM-thin scratch guest this session. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — before the guest fast lane is built; owner: CC** | — | needs an LVM-thin scratch guest on a demo host | CC |
|
|
||||||
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
| **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC |
|
||||||
|
| **R-840** | Box system & updates | P2 | **A new root wrapper or a new sudoers line cannot reach a box that is already installed — there is no product route.** FOUND 2026-10-04 (guest OS lane, Part B 1): the agent's signed `agent_update` replaces ONLY the binary (`configs/felhom-selfupdate-guarded` swaps `/usr/local/bin/felhom-agent`); wrappers (`/usr/local/sbin/felhom-*`) and `/etc/sudoers.d/felhom-agent` are written only by `felhom-host-install.sh` step 5 on a new box. So agent v0.140.0's OS leg is DEAD on every existing box (the capability probe will show `osapply-run` missing) until someone installs `felhom-os-apply` + the FELHOM_OSAPPLY line by hand — which is what the demo boxes got this session. The `wireguard-tools` line (S3) reached the fleet the same way: by reinstall or by hand. **Proposal:** a signed `agent_config_update` op (operator-signed like `agent_update`) whose params pin the agent TAG and the sha256 of a config bundle (sudoers + wrappers); the self-update wrapper installs it as root after `visudo -cf` and a syntax check, keeps the previous copies, and the capability probe confirms. `audits/os-guest-lane-2026-10-04/README.md` | **READY — design + operator go; owner: CC** | — | — | CC |
|
||||||
|
| **R-841** | Box system & updates | P3 | **The agent's `cloudflared` health probe reads a host systemd unit that does not exist — every box reports its tunnel `inactive`.** FOUND 2026-10-04: `felhom-agent/internal/hub/cloudflared.go` runs `systemctl is-active cloudflared` on the HOST; cloudflared is a container in the GUEST (`11` C8), so demo-hp answers `inactive` / `Unit cloudflared.service could not be found`, and the hub stores that in `cloudflared_status` for every box. The field is equally consistent with "tunnel down" and "never checked" (R-96 rule 3). Fix direction: read the guest's container state (the agent already may `pct exec * -- docker inspect -f *`), or drop the field; and say so in `03` (corrected 2026-10-04). | **READY — owner: CC** | — | — | CC |
|
||||||
|
| **R-842** | Box system & updates | P3 | **The guest OS update has no automatic undo: a customer guest cannot be snapshotted.** MEASURED 2026-10-04 (R-837, closed): PVE refuses any snapshot not named `vzdump` when a guest has host-path binds (mp8, mp9) — as the agent's token (which holds `VM.Snapshot` / `VM.Snapshot.Rollback`) and as root. Today a failed health check stops, reports `health_failed` and mails the operator; the whole-guest backup taken minutes earlier is the undo, by hand. **Options (STATUS):** A — keep it so (no new mechanism; the backup is minutes old; a restore costs ~1–3 min down plus app data written since); B — a root-wrapper LVM-thin snapshot of rootfs + mp0 behind PVE's back, rolled back with the guest stopped (a new mechanism nobody has measured; thin-pool risk, R-672's shape). `audits/os-guest-lane-2026-10-04/partA/README.md` | **WAITING-ON-OPERATOR 2026-10-04 — CC's pick: A** | — | — | operator |
|
||||||
|
|
||||||
## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6)
|
## Monitoring & notifications — 24 rows (P2 2, P3 16, P4 6)
|
||||||
|
|
||||||
@@ -359,7 +359,7 @@ stopping line that lies.
|
|||||||
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC |
|
||||||
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
| **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC |
|
||||||
|
|
||||||
## Hub & operator — 21 rows (P2 1, P3 7, P4 13)
|
## Hub & operator — 23 rows (P2 1, P3 7, P4 15)
|
||||||
|
|
||||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||||
|---|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|---|
|
||||||
@@ -384,6 +384,8 @@ stopping line that lies.
|
|||||||
| **R-688** | Hub & operator | P4 | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** **Re-ranked 2026-10-03: P3→P4: the dialog no longer promises it; what is left is operator comfort.** | — | — | CC + operator |
|
| **R-688** | Hub & operator | P4 | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** **Re-ranked 2026-10-03: P3→P4: the dialog no longer promises it; what is left is operator comfort.** | — | — | CC + operator |
|
||||||
| **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator |
|
| **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator |
|
||||||
| **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator |
|
| **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator |
|
||||||
|
| **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC |
|
||||||
|
| **R-845** | Hub & operator | P4 | **One OS-leg pass takes 3–4 minutes even when it installs 3 packages**: the wrapper's inventory (an `apt-get update`, `apt-cache policy` over every installed package, a `/proc/*/maps` scan for restart-needed, two simulations) dominates; measured 174–245 s per pass on the demo boxes vs 3.8–31.7 s for the install itself. It runs at night under the heavy-op gate, so it delays a restore-test by minutes, nothing worse. Fix direction: one `apt-cache policy` per run and the restart scan only after an install. `audits/os-guest-lane-2026-10-04/partG/` | **READY — owner: CC** | — | — | CC |
|
||||||
|
|
||||||
## Business & legal — 7 rows (P2 4, P4 3)
|
## Business & legal — 7 rows (P2 4, P4 3)
|
||||||
|
|
||||||
|
|||||||
@@ -55,6 +55,23 @@ summary names each app DONE or STOPPED with its reason. Order: database/redis li
|
|||||||
**Monthly cost (measured 2026-10-01):** see "What it cost" below. linuxserver images (bookstack, radarr, sonarr,
|
**Monthly cost (measured 2026-10-01):** see "What it cost" below. linuxserver images (bookstack, radarr, sonarr,
|
||||||
code-server) are rebuilt upstream weekly under the same tag, so most months they come up.
|
code-server) are rebuilt upstream weekly under the same tag, so most months they come up.
|
||||||
|
|
||||||
|
## 4a. Infrastructure pins (R-838, 2026-10-04)
|
||||||
|
|
||||||
|
The box's three built-in containers — traefik, cloudflared, filebrowser — are pinned in
|
||||||
|
`felhom-controller/controller/internal/infra/infra.go` (`TraefikImage`, `CloudflaredImage`, `FileBrowserImage`). They are
|
||||||
|
**not** catalog templates, so `retest-floating.py` never sees them. Each month:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd felhom-controller/controller && python3 scripts/check-infra-pins.py # exit 1 = at least one BEHIND (report only)
|
||||||
|
```
|
||||||
|
|
||||||
|
For each BEHIND pin: read the upstream release notes between the pinned and the newest version (same channel:
|
||||||
|
traefik v3.x, cloudflared YYYY.M.P, filebrowser N.N.N-stable — never a beta), name any breaking change against what we
|
||||||
|
configure, raise the constant, release the controller, prove it on 9202 then on both demo boxes (the containers are
|
||||||
|
recreated within ~20 s of the new controller starting: traefik/cloudflared by the base-infra bring-up, filebrowser by
|
||||||
|
the start-up mount sync), and time the public gap through the tunnel. Measured 2026-10-04: ≤ 19.6 s on demo-hp, ≤ 14.7 s
|
||||||
|
on demo-felhom, counting the controller's own restart (`audits/os-guest-lane-2026-10-04/partF/`).
|
||||||
|
|
||||||
## 5. Teardown (three layers, stated)
|
## 5. Teardown (three layers, stated)
|
||||||
|
|
||||||
Machine: 9202 back on the live catalog (`repoint.py restore`), apps the run installed are removed by it. Host: `pct
|
Machine: 9202 back on the live catalog (`repoint.py restore`), apps the run installed are removed by it. Host: `pct
|
||||||
|
|||||||
@@ -0,0 +1,6 @@
|
|||||||
|
Round-trip of the published golden 0.291.0 (2026-10-04, from DooPlex, anonymous GET)
|
||||||
|
URL: https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.291.0/golden.tar.zst
|
||||||
|
HTTP: 200, size_download=648807981 bytes
|
||||||
|
sha256 (downloaded): 64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb
|
||||||
|
GOLDEN_SHA256 (bake.log line 328): 64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb
|
||||||
|
Result: MATCH
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
# Golden 0.291.0 — vouch + floor, 2026-10-04 (guest-OS-lane brief Part H, by CC)
|
||||||
|
Floor: POST /configuration/global-floor min_controller_version=0.291.0 min_agent=0.131.0 → 303
|
||||||
|
hub log 11:01:32 CEST: Global controller-version floor set to "0.291.0" (declared MinAgent "0.131.0")
|
||||||
|
both demo boxes self-updated to controller 0.291.0 within ~10 s (09:01:40Z / 09:01:42Z).
|
||||||
|
Vouch: POST /configuration/artifacts agent_version=0.140.0 golden_version=0.291.0 min_agent=0.131.0
|
||||||
|
agent_sha256=ae2d60b794869c51d6b063c8e31e2da98ecdbd36c6b75febd64e2fb4266c1250
|
||||||
|
golden_sha256=64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb → 303
|
||||||
|
hub log 11:40:14 CEST: Artifact manifest set: agent=0.140.0 golden=0.291.0 min_agent="0.131.0" wrapper_sha=false
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
# Golden 0.291.0 — bake + publish, 2026-10-04
|
||||||
|
|
||||||
|
Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–4, in the drill VM
|
||||||
|
on DooPlex. Step 5 (vouching in the hub) was **not** done; it is the operator's act.
|
||||||
|
|
||||||
|
- Controller image: `gitea.dooplex.hu/admin/felhom-controller:0.291.0` (manifest present in the
|
||||||
|
registry before the bake).
|
||||||
|
- Build script: `felhom-agent/configs/build-golden.sh` v3.0.0, agent repo `main` = `a55eedcf2cc1`
|
||||||
|
(tree clean, HEAD == origin/main after `git fetch`); sha256 of the copy in the VM matched the repo
|
||||||
|
file (`e4c9ede772e7…`, the same script bytes as the 0.290.0 bake).
|
||||||
|
- Drill VM: reverted to `virgin`, cold-booted per §4.0; `pveversion` = `pve-manager/9.2.2`.
|
||||||
|
- Step 2: `pveam update` → `update successful`; template `debian-13-standard_13.6-1_amd64.tar.zst`
|
||||||
|
(the only `_amd64` debian-13 entry), downloaded with checksum verified.
|
||||||
|
- Pre-gate: `GET …/generic/felhom-golden/0.291.0/golden.tar.zst` → **404** before the bake.
|
||||||
|
- Token: copied file → file (`scp`); launched via the in-VM runner script as transient unit
|
||||||
|
`golden-bake`. `systemctl show golden-bake -p Environment -p ExecStart | grep -c -F <token>` = **0**
|
||||||
|
(control: same output with the token appended = **1**).
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
```
|
||||||
|
GOLDEN_VERSION=0.291.0
|
||||||
|
GOLDEN_SHA256=64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb
|
||||||
|
```
|
||||||
|
|
||||||
|
## Pass markers (quoted verbatim from `bake.log`)
|
||||||
|
|
||||||
|
```
|
||||||
|
83: docker OK (overlay2; data-root /var/lib/docker)
|
||||||
|
319:INFO: including mount point rootfs ('/') in backup
|
||||||
|
320:INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||||
|
325:[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||||
|
326:[golden] upload OK (HTTP 201)
|
||||||
|
```
|
||||||
|
|
||||||
|
`grep -E 'excluding|FATAL' bake.log` → no matches.
|
||||||
|
|
||||||
|
## Token-leak grep
|
||||||
|
|
||||||
|
On the copy in this directory (the one that would be committed):
|
||||||
|
`grep -c -F "$(cat ~/.gitea-token)" bake.log` = **0**.
|
||||||
|
Positive control: a throwaway copy with the token appended → **1**; the copy was `shred -u`'d.
|
||||||
|
Same 0 / 1 result on the scratch copy pulled straight from the VM.
|
||||||
|
|
||||||
|
## Round trip
|
||||||
|
|
||||||
|
See `02-round-trip.txt`: the published package downloaded anonymously (HTTP 200, 648807981 bytes)
|
||||||
|
hashes to `64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb` — **matches**
|
||||||
|
GOLDEN_SHA256.
|
||||||
|
|
||||||
|
## Teardown state
|
||||||
|
|
||||||
|
- `pct destroy 9100 --purge` → rc 0 (both LVs removed); `pct list` empty.
|
||||||
|
- `/root/.gitea-token`, `/root/bake-run.sh`, `/root/bake.log` in the VM: `shred -u`, confirmed absent.
|
||||||
|
- VM powered off; no `qemu-system-x86` process remained.
|
||||||
|
- `qemu-img snapshot -a virgin drill.qcow2` → OK; snapshot list shows only `virgin`.
|
||||||
|
- Hub, k3s, demo boxes, ep0, PBS, Storage Box: not touched.
|
||||||
|
- `df -h`: `/mnt/5_hdd` 37 %, `/` 53 % (before and after).
|
||||||
@@ -0,0 +1,330 @@
|
|||||||
|
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.291.0
|
||||||
|
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
|
||||||
|
Logical volume "vm-9100-disk-0" created.
|
||||||
|
Logical volume pve/vm-9100-disk-0 changed.
|
||||||
|
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||||
|
Filesystem UUID: 169cbd9e-cd43-4c1b-9084-a5d1393901df
|
||||||
|
Superblock backups stored on blocks:
|
||||||
|
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||||
|
4096000, 7962624
|
||||||
|
Logical volume "vm-9100-disk-1" created.
|
||||||
|
Logical volume pve/vm-9100-disk-1 changed.
|
||||||
|
Creating filesystem with 6291456 4k blocks and 1572864 inodes
|
||||||
|
Filesystem UUID: 9c3933d5-01fe-491e-8cc2-065f403c1b3c
|
||||||
|
Superblock backups stored on blocks:
|
||||||
|
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||||
|
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
|
||||||
|
Total bytes read: 553512960 (528MiB, 104MiB/s)
|
||||||
|
Detected container architecture: amd64
|
||||||
|
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
|
||||||
|
done: SHA256:ZOYAxm1HMLEOJnDNfKK0MqicitZbUFFAN+BErHjSorw root@felhom-golden
|
||||||
|
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
|
||||||
|
done: SHA256:7uGzlVgJldBrBZLh+MMtKQ1zQ46jFiBCSeAd8MLjfD0 root@felhom-golden
|
||||||
|
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
|
||||||
|
done: SHA256:2n1/Q6PXXJCBu5BYPo8czMOY92VC0a6oj7GHkMNjklg root@felhom-golden
|
||||||
|
[golden] starting + installing Docker (official repo, trixie channel) …
|
||||||
|
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||||
|
perl: warning: Setting locale failed.
|
||||||
|
perl: warning: Please check that your locale settings:
|
||||||
|
LANGUAGE = (unset),
|
||||||
|
LC_ALL = (unset),
|
||||||
|
LC_CTYPE = (unset),
|
||||||
|
LC_NUMERIC = (unset),
|
||||||
|
LC_COLLATE = (unset),
|
||||||
|
LC_TIME = (unset),
|
||||||
|
LC_MESSAGES = (unset),
|
||||||
|
LC_MONETARY = (unset),
|
||||||
|
LC_ADDRESS = (unset),
|
||||||
|
LC_IDENTIFICATION = (unset),
|
||||||
|
LC_MEASUREMENT = (unset),
|
||||||
|
LC_PAPER = (unset),
|
||||||
|
LC_TELEPHONE = (unset),
|
||||||
|
LC_NAME = (unset),
|
||||||
|
LANG = "en_US.UTF-8"
|
||||||
|
are supported and installed on your system.
|
||||||
|
perl: warning: Falling back to the standard locale ("C").
|
||||||
|
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||||
|
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||||
|
perl: warning: Setting locale failed.
|
||||||
|
perl: warning: Please check that your locale settings:
|
||||||
|
LANGUAGE = (unset),
|
||||||
|
LC_ALL = (unset),
|
||||||
|
LC_CTYPE = (unset),
|
||||||
|
LC_NUMERIC = (unset),
|
||||||
|
LC_COLLATE = (unset),
|
||||||
|
LC_TIME = (unset),
|
||||||
|
LC_MESSAGES = (unset),
|
||||||
|
LC_MONETARY = (unset),
|
||||||
|
LC_ADDRESS = (unset),
|
||||||
|
LC_IDENTIFICATION = (unset),
|
||||||
|
LC_MEASUREMENT = (unset),
|
||||||
|
LC_PAPER = (unset),
|
||||||
|
LC_TELEPHONE = (unset),
|
||||||
|
LC_NAME = (unset),
|
||||||
|
LANG = "en_US.UTF-8"
|
||||||
|
are supported and installed on your system.
|
||||||
|
perl: warning: Falling back to the standard locale ("C").
|
||||||
|
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||||
|
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||||
|
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
|
||||||
|
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
|
||||||
|
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
|
||||||
|
Unable to find image 'hello-world:latest' locally
|
||||||
|
latest: Pulling from library/hello-world
|
||||||
|
4f55086f7dd0: Pulling fs layer
|
||||||
|
4f55086f7dd0: Verifying Checksum
|
||||||
|
4f55086f7dd0: Download complete
|
||||||
|
4f55086f7dd0: Pull complete
|
||||||
|
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
|
||||||
|
Status: Downloaded newer image for hello-world:latest
|
||||||
|
docker OK (overlay2; data-root /var/lib/docker)
|
||||||
|
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
|
||||||
|
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
|
||||||
|
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
|
||||||
|
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.291.0 (no registry cred at deploy) …
|
||||||
|
|
||||||
|
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
|
||||||
|
Configure a credential helper to remove this warning. See
|
||||||
|
https://docs.docker.com/go/credential-store/
|
||||||
|
|
||||||
|
0.291.0: Pulling from admin/felhom-controller
|
||||||
|
774043ccc8cc: Pulling fs layer
|
||||||
|
ab6b448d4be9: Pulling fs layer
|
||||||
|
23a5bfa58353: Pulling fs layer
|
||||||
|
862a57157567: Pulling fs layer
|
||||||
|
2a5fbd09d7cf: Pulling fs layer
|
||||||
|
61606e64d0b9: Pulling fs layer
|
||||||
|
862a57157567: Waiting
|
||||||
|
2a5fbd09d7cf: Waiting
|
||||||
|
61606e64d0b9: Waiting
|
||||||
|
774043ccc8cc: Verifying Checksum
|
||||||
|
774043ccc8cc: Download complete
|
||||||
|
23a5bfa58353: Verifying Checksum
|
||||||
|
23a5bfa58353: Download complete
|
||||||
|
862a57157567: Verifying Checksum
|
||||||
|
862a57157567: Download complete
|
||||||
|
2a5fbd09d7cf: Verifying Checksum
|
||||||
|
2a5fbd09d7cf: Download complete
|
||||||
|
61606e64d0b9: Verifying Checksum
|
||||||
|
61606e64d0b9: Download complete
|
||||||
|
ab6b448d4be9: Verifying Checksum
|
||||||
|
ab6b448d4be9: Download complete
|
||||||
|
774043ccc8cc: Pull complete
|
||||||
|
ab6b448d4be9: Pull complete
|
||||||
|
23a5bfa58353: Pull complete
|
||||||
|
862a57157567: Pull complete
|
||||||
|
2a5fbd09d7cf: Pull complete
|
||||||
|
61606e64d0b9: Pull complete
|
||||||
|
Digest: sha256:a8f814aef30616de67b4a726a6fcdb7f675bfe784ebefc288d0882eb02f1ba39
|
||||||
|
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.291.0
|
||||||
|
gitea.dooplex.hu/admin/felhom-controller:0.291.0
|
||||||
|
[golden] asking the controller which infra images it manages …
|
||||||
|
[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
|
||||||
|
v3.7.13: Pulling from library/traefik
|
||||||
|
e2de96513ba9: Pulling fs layer
|
||||||
|
b686a4f73445: Pulling fs layer
|
||||||
|
78cb21c375ca: Pulling fs layer
|
||||||
|
acb2f33459b1: Pulling fs layer
|
||||||
|
acb2f33459b1: Waiting
|
||||||
|
b686a4f73445: Download complete
|
||||||
|
e2de96513ba9: Verifying Checksum
|
||||||
|
e2de96513ba9: Download complete
|
||||||
|
acb2f33459b1: Verifying Checksum
|
||||||
|
acb2f33459b1: Download complete
|
||||||
|
78cb21c375ca: Verifying Checksum
|
||||||
|
78cb21c375ca: Download complete
|
||||||
|
e2de96513ba9: Pull complete
|
||||||
|
b686a4f73445: Pull complete
|
||||||
|
78cb21c375ca: Pull complete
|
||||||
|
acb2f33459b1: Pull complete
|
||||||
|
Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0
|
||||||
|
Status: Downloaded newer image for traefik:v3.7.13
|
||||||
|
docker.io/library/traefik:v3.7.13
|
||||||
|
2026.9.3: Pulling from cloudflare/cloudflared
|
||||||
|
2cc7ee286bf3: Pulling fs layer
|
||||||
|
c172f21841df: Pulling fs layer
|
||||||
|
218cf840d0d9: Pulling fs layer
|
||||||
|
f6069939f718: Pulling fs layer
|
||||||
|
d6b1b89eccac: Pulling fs layer
|
||||||
|
2780920e5dbf: Pulling fs layer
|
||||||
|
7c12895b777b: Pulling fs layer
|
||||||
|
3214acf345c0: Pulling fs layer
|
||||||
|
52630fc75a18: Pulling fs layer
|
||||||
|
dd64bf2dd177: Pulling fs layer
|
||||||
|
b839dfae01f6: Pulling fs layer
|
||||||
|
ebddc55facdc: Pulling fs layer
|
||||||
|
c4bc6f35ff5e: Pulling fs layer
|
||||||
|
b96fe2995f90: Pulling fs layer
|
||||||
|
58c0c263dc73: Pulling fs layer
|
||||||
|
bd8962e29291: Pulling fs layer
|
||||||
|
cac2ae0193cb: Pulling fs layer
|
||||||
|
f0383d5ebc47: Pulling fs layer
|
||||||
|
52630fc75a18: Waiting
|
||||||
|
dd64bf2dd177: Waiting
|
||||||
|
b839dfae01f6: Waiting
|
||||||
|
ebddc55facdc: Waiting
|
||||||
|
c4bc6f35ff5e: Waiting
|
||||||
|
b96fe2995f90: Waiting
|
||||||
|
58c0c263dc73: Waiting
|
||||||
|
bd8962e29291: Waiting
|
||||||
|
cac2ae0193cb: Waiting
|
||||||
|
f0383d5ebc47: Waiting
|
||||||
|
d6b1b89eccac: Waiting
|
||||||
|
2780920e5dbf: Waiting
|
||||||
|
7c12895b777b: Waiting
|
||||||
|
3214acf345c0: Waiting
|
||||||
|
f6069939f718: Waiting
|
||||||
|
c172f21841df: Verifying Checksum
|
||||||
|
c172f21841df: Download complete
|
||||||
|
2cc7ee286bf3: Download complete
|
||||||
|
218cf840d0d9: Verifying Checksum
|
||||||
|
218cf840d0d9: Download complete
|
||||||
|
f6069939f718: Download complete
|
||||||
|
d6b1b89eccac: Verifying Checksum
|
||||||
|
d6b1b89eccac: Download complete
|
||||||
|
2780920e5dbf: Verifying Checksum
|
||||||
|
2780920e5dbf: Download complete
|
||||||
|
2cc7ee286bf3: Pull complete
|
||||||
|
7c12895b777b: Verifying Checksum
|
||||||
|
7c12895b777b: Download complete
|
||||||
|
3214acf345c0: Verifying Checksum
|
||||||
|
3214acf345c0: Download complete
|
||||||
|
52630fc75a18: Download complete
|
||||||
|
dd64bf2dd177: Verifying Checksum
|
||||||
|
dd64bf2dd177: Download complete
|
||||||
|
b839dfae01f6: Verifying Checksum
|
||||||
|
b839dfae01f6: Download complete
|
||||||
|
ebddc55facdc: Verifying Checksum
|
||||||
|
ebddc55facdc: Download complete
|
||||||
|
c172f21841df: Pull complete
|
||||||
|
c4bc6f35ff5e: Verifying Checksum
|
||||||
|
c4bc6f35ff5e: Download complete
|
||||||
|
58c0c263dc73: Verifying Checksum
|
||||||
|
58c0c263dc73: Download complete
|
||||||
|
b96fe2995f90: Verifying Checksum
|
||||||
|
b96fe2995f90: Download complete
|
||||||
|
bd8962e29291: Verifying Checksum
|
||||||
|
bd8962e29291: Download complete
|
||||||
|
cac2ae0193cb: Verifying Checksum
|
||||||
|
cac2ae0193cb: Download complete
|
||||||
|
218cf840d0d9: Pull complete
|
||||||
|
f0383d5ebc47: Verifying Checksum
|
||||||
|
f0383d5ebc47: Download complete
|
||||||
|
f6069939f718: Pull complete
|
||||||
|
d6b1b89eccac: Pull complete
|
||||||
|
2780920e5dbf: Pull complete
|
||||||
|
7c12895b777b: Pull complete
|
||||||
|
3214acf345c0: Pull complete
|
||||||
|
52630fc75a18: Pull complete
|
||||||
|
dd64bf2dd177: Pull complete
|
||||||
|
b839dfae01f6: Pull complete
|
||||||
|
ebddc55facdc: Pull complete
|
||||||
|
c4bc6f35ff5e: Pull complete
|
||||||
|
b96fe2995f90: Pull complete
|
||||||
|
58c0c263dc73: Pull complete
|
||||||
|
bd8962e29291: Pull complete
|
||||||
|
cac2ae0193cb: Pull complete
|
||||||
|
f0383d5ebc47: Pull complete
|
||||||
|
Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c
|
||||||
|
Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3
|
||||||
|
docker.io/cloudflare/cloudflared:2026.9.3
|
||||||
|
1.5.6-stable: Pulling from gtstef/filebrowser
|
||||||
|
55afa1ecc21d: Pulling fs layer
|
||||||
|
8ed8f35f8d4f: Pulling fs layer
|
||||||
|
989b226a579c: Pulling fs layer
|
||||||
|
660aeead31d5: Pulling fs layer
|
||||||
|
4f4fb700ef54: Pulling fs layer
|
||||||
|
adce24567e4c: Pulling fs layer
|
||||||
|
f17ea56b313b: Pulling fs layer
|
||||||
|
6b6f3b3efe88: Pulling fs layer
|
||||||
|
4ed1ca4f3fce: Pulling fs layer
|
||||||
|
e6fc9c6a5757: Pulling fs layer
|
||||||
|
d47782d1182a: Pulling fs layer
|
||||||
|
f17ea56b313b: Waiting
|
||||||
|
6b6f3b3efe88: Waiting
|
||||||
|
4ed1ca4f3fce: Waiting
|
||||||
|
e6fc9c6a5757: Waiting
|
||||||
|
d47782d1182a: Waiting
|
||||||
|
660aeead31d5: Waiting
|
||||||
|
4f4fb700ef54: Waiting
|
||||||
|
adce24567e4c: Waiting
|
||||||
|
55afa1ecc21d: Verifying Checksum
|
||||||
|
55afa1ecc21d: Download complete
|
||||||
|
660aeead31d5: Verifying Checksum
|
||||||
|
660aeead31d5: Download complete
|
||||||
|
4f4fb700ef54: Verifying Checksum
|
||||||
|
4f4fb700ef54: Download complete
|
||||||
|
8ed8f35f8d4f: Verifying Checksum
|
||||||
|
8ed8f35f8d4f: Download complete
|
||||||
|
f17ea56b313b: Verifying Checksum
|
||||||
|
f17ea56b313b: Download complete
|
||||||
|
55afa1ecc21d: Pull complete
|
||||||
|
989b226a579c: Verifying Checksum
|
||||||
|
989b226a579c: Download complete
|
||||||
|
6b6f3b3efe88: Verifying Checksum
|
||||||
|
6b6f3b3efe88: Download complete
|
||||||
|
4ed1ca4f3fce: Verifying Checksum
|
||||||
|
4ed1ca4f3fce: Download complete
|
||||||
|
adce24567e4c: Verifying Checksum
|
||||||
|
adce24567e4c: Download complete
|
||||||
|
e6fc9c6a5757: Verifying Checksum
|
||||||
|
e6fc9c6a5757: Download complete
|
||||||
|
d47782d1182a: Verifying Checksum
|
||||||
|
d47782d1182a: Download complete
|
||||||
|
8ed8f35f8d4f: Pull complete
|
||||||
|
989b226a579c: Pull complete
|
||||||
|
660aeead31d5: Pull complete
|
||||||
|
4f4fb700ef54: Pull complete
|
||||||
|
adce24567e4c: Pull complete
|
||||||
|
f17ea56b313b: Pull complete
|
||||||
|
6b6f3b3efe88: Pull complete
|
||||||
|
4ed1ca4f3fce: Pull complete
|
||||||
|
e6fc9c6a5757: Pull complete
|
||||||
|
d47782d1182a: Pull complete
|
||||||
|
Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8
|
||||||
|
Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable
|
||||||
|
docker.io/gtstef/filebrowser:1.5.6-stable
|
||||||
|
1.1.0: Pulling from admin/felhom-samba
|
||||||
|
897d797d2723: Pulling fs layer
|
||||||
|
3051591aa250: Pulling fs layer
|
||||||
|
ce57a3f93416: Pulling fs layer
|
||||||
|
fb94eeec2fe1: Pulling fs layer
|
||||||
|
fb94eeec2fe1: Waiting
|
||||||
|
ce57a3f93416: Verifying Checksum
|
||||||
|
ce57a3f93416: Download complete
|
||||||
|
fb94eeec2fe1: Verifying Checksum
|
||||||
|
fb94eeec2fe1: Download complete
|
||||||
|
897d797d2723: Verifying Checksum
|
||||||
|
897d797d2723: Download complete
|
||||||
|
3051591aa250: Verifying Checksum
|
||||||
|
3051591aa250: Download complete
|
||||||
|
897d797d2723: Pull complete
|
||||||
|
3051591aa250: Pull complete
|
||||||
|
ce57a3f93416: Pull complete
|
||||||
|
fb94eeec2fe1: Pull complete
|
||||||
|
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
|
||||||
|
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||||
|
gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||||
|
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
|
||||||
|
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
|
||||||
|
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
|
||||||
|
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
|
||||||
|
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
|
||||||
|
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
|
||||||
|
[golden] identity-clean + minimize …
|
||||||
|
[golden] stop + archive …
|
||||||
|
INFO: including mount point rootfs ('/') in backup
|
||||||
|
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||||
|
INFO: archive file size: 618MB
|
||||||
|
INFO: Finished Backup of VM 9100 (00:00:29)
|
||||||
|
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_04-11_36_50.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
|
||||||
|
[golden] publishing golden (648807981 bytes, sha256 64c2fe5d58d65706…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.291.0/golden.tar.zst
|
||||||
|
[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||||
|
[golden] upload OK (HTTP 201)
|
||||||
|
GOLDEN_VERSION=0.291.0
|
||||||
|
GOLDEN_SHA256=64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb
|
||||||
|
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.291.0 / 64c2fe5d58d65706185023b359c42843ecd2d215298dd29febfc983ad43e11fb
|
||||||
|
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
|
||||||
Reference in New Issue
Block a user