diff --git a/CONTEXT.md b/CONTEXT.md index b07567f9..469a4f79 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,11 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-07 (late evening) — pre-night fixes, the night moved to 7→8 (`09` §3 174–176).** Agent v0.153.0 (ring 0 stages +> exactly the told kernel — R-898 closed), controller v0.303.0 (`backup/driveready.go`: captures wait ≤10 min after start +> for a live drive mount — R-897 closed), hub v0.143.1 (`reply_to` = operator on every household mail). Armed: demo-felhom +> 7.0.14-20 (told 18:12), demo-hp 7.0.14-22 (told 18:38). Read back 2026-10-08 morning. Register 128 → 126. + > **2026-10-07 (evening) — the kernel lane BUILT (R-836, `09` §3 decisions 172–173; `11` §5.11).** Agent v0.152.0 (+ step > bundle `0.152.0-step1`, + bundle) on demo-hp, demo-felhom, Tester 1; hub v0.143.0 (deployed attended). One-shot flag on > the ESP, option C on the one-shot entry, default pinned to the running kernel and moved only by a healthy boot (host diff --git a/REPORT.md b/REPORT.md index 36f51b90..1ba80027 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,64 +1,34 @@ -# REPORT — the kernel lane, 2026-10-07 (evening) +# REPORT — three fixes before the first real kernel night, 2026-10-07 (evening) + +The brief said 2026-10-08; at 18:34 it was 2026-10-07. Asked; the operator answered *"Start today, and do both boxes this +night"* (`09` §3 decision 176). So the first real kernel night is **7→8**, read back on the morning of 2026-10-08. | Part | Result | |---|---| -| Rulings | Decisions 172 (kernel lane yes: A + C, night restarts, household mailed the day before, a freeze needing a person accepted) and 173 (R-32 leftovers left; P2 → P4, blocked on a safe main-account path) recorded FIRST (`0c5dbfbf`). | -| A — the one-shot boot | **Done.** Wrapper layer `kernel` (agent v0.152.0) + the two GRUB generators in the bundle (option C on the one-shot entry only). Red tests first-class: a box without the ESP flag refused (R20), a kernel outside the set refused (R23), the snippet clears the flag before booting, an install never moves the default. Red-proof `audits/kernel-lane-2026-10-07/A/redproof.txt` (16 mutations, each red). | -| B — boot good, the self-revert | **Done.** Health = the host rule + the hub reached, 20 min (measured: all healthy 68 s after a reboot on demo-felhom, 272 s on demo-hp; the spike's > 6 min Tailscale delay did not recur — no row). A box that never returns raises the hub's existing `host_stale` at 45 min and `host_down` at 90 min. One self-revert per step. Crash guard: a step adds at most ONE unclean boot (a panic before userspace adds none) — pinned by `KernelStepCannotLeaveTheBoxOff`, and seen live (0 counted after the panic). | -| C — night slot, approval, System page | **Done.** The kernel step ends the night leg (after a healthy host step, after the backup, trigger night, once per night); ring 1 stages only by a signed job. Hub v0.143.0 (deployed with you attending): due boxes, "Approve kernel set", two System page cells. `11` §5.11 + §8 step 6 written. | -| D — the household mail | **Done.** One mail the day before, 09:00–20:00 Budapest, its language, registered address; `tonight` only after the mail service accepted it (no mail, no step). Sent for real to demo-felhom at 18:12 (Gmail). Parity green (both locales; tests). | -| E 1 — Tester 1 by hand | **PASS ×3** (`audits/kernel-lane-2026-10-07/E/RESULT.md`): forced panic → back on the old kernel by itself (screens); held guest → judged 20 min → one self-revert; healthy → 7.0.14-22 the default. Tester 1 left on 7.0.14-22 (healthy, default). | -| E 2 — option C | **Unmeasured:** the Proxmox kernel has no lockup test module (`CONFIG_TEST_LOCKUP` not set). No tool built. | -| E 3 — ring-0 night run | **Not yet — next session.** Tonight demo-felhom was told about 7.0.14-20 (from last night's report) but its sources now offer 7.0.14-22: the step will refuse (R23) before any change (R-898). demo-hp's lists were a day old. Both demo boxes are due for 7.0.14-22 on the night 2026-10-08→09. | -| E 4 — Tester 1 as ring 1 | **Not yet** (after E3; Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the next kernel). | -| Small | R-861 (a): no controller release in this arc — not read back. Tailscale delay: not reproduced → no row. | +| Rulings | Decisions 174 (mail text stays + a reply address), 175 (window 09–20, max 3 mails: keep), 176 (today, both boxes tonight) recorded FIRST (`a0ff737e`). | +| A — the told kernel boots (R-898) | **Done, delivered, closed.** Agent v0.153.0: ring 0 stages EXACTLY the told kernel (select `listed`, the set from the kver). Told 20 / sources offer 22 → 20 installs (wrapper test); 20 gone → R7 before any change (wrapper test); the hub keeps R7 temporary and tells the household again for the newer kernel (hub test). Red-proofs `audits/kernel-night-2026-10-07/A/redproof.txt`. | +| B — first backup after a restart waits (R-897) | **Done, delivered, closed.** The reviewer's pick, taken: controller v0.303.0 — for 10 min after the controller starts, a capture for an app on a drive runs only once that drive is a LIVE mount in the controller's own namespace (`/proc/self/mountinfo`, the signal the startup app gate already uses); the 5-min refresh skips, a data run waits; after 10 min it runs and logs once. The agent's bind order is unchanged. Red test: not-bound → no capture; bound → capture. | +| C — a reply reaches a person | **Built and delivered (hub v0.143.1); the header read-back is NOT done.** Every mail to a household now carries `reply_to` = `admin@felhom.eu` (all household templates invite contact: the kernel notice, the event sign-off "contact your operator", the setup/link mails); a mail to the operator carries none. Test pinned and red-proved. One test mail sent through the mail service with the hub's exact fields to `tester1@felhom.eu` — it arrived in Gmail (16:54 UTC). **Read-back tried:** the Gmail tool returns no headers (METADATA_ONLY, no RAW); the mail service's read API refused (`restricted_api_key` — send-only, correctly). Note: the demo households' own address IS `admin@felhom.eu`, so their mails carry no Reply-To by design. | +| D — deliver, arm the night | **Done.** Hub 0.143.1 deployed 18:54 local (operator attending), agent 0.153.0 on all three boxes 19:03 (binary only — no root file changed), controller 0.303.0 on all three 18:50 (per-customer floors; global untouched). **Armed:** demo-felhom → **7.0.14-20-pve**, household mail **18:12**; demo-hp → **7.0.14-22-pve**, household mail **18:38**. No reboot by hand. | -**Rows: 126 before → 128 after. Opened 2 (R-897, R-898). Closed 0.** +**Rows: 128 before → 126 after. Opened 0. Closed 2 (R-897, R-898).** -**Releases:** agent v0.152.0 (tag `d03ab7f`; binary `95ff4220…`, bundle `f0c2cec3…`, step bundle `0.152.0-step1` -`0b71d32b…`), delivered by signed jobs to demo-hp, demo-felhom and Tester 1 (binary → step bundle → bundle). Hub v0.143.0 -(`ab3b7ea2`, deployed `0ece2ff7` 15:39, Synced/Healthy). The installer's uninstall now knows the two GRUB files -(unreleased, no tag cut). Nothing on ep0, Tester 2 or DooPlex's own system; no Docker engine step; no delivery after 02:00. +**How demo-hp got told tonight:** its apt lists were a day old, so its last report named no new kernel. I switched its OS +updates OFF on the hub for ~1 minute, ran the agent's own report-only pass (`--selftest=os-update`; with the switch off it +installs nothing and runs no Docker/Proxmox/kernel step), and switched it back ON (`kernel-night-2026-10-07/demo-hp-inventory-pass.txt`). +The hub mailed 20 s after the host report arrived (18:37:51 → 18:38:11). -**Read for this report:** `11-os-updates.md` (§5.3, §5.8–5.10, §8.2 — §5.11 new), `03-host-agent.md`, `07` §6.1, -`audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`. +**Releases:** agent v0.153.0 (binary `b204ebe6…`, tag at `2d1e5d0`); controller v0.303.0 (`29bebbb`, MinAgent 0.131.0); +hub v0.143.1 (`87765bfa`, deployed `d1457892`, Synced/Healthy). CI green on every push, checked by commit. No change on +ep0, Tester 2 or DooPlex's own system. -**Also seen:** the hub image was built while one audit text file (outside `hub/`) was uncommitted; the build copies -`hub/` only, and `hub/` equalled the pushed commit (checked by diff). R-897: after demo-hp's timing reboot, one app's -backup failed in the second the agent re-bound the drive. - -## The household mail (both languages, as sent) - -**Hungarian** — subject: `[Felhom] Ma éjjel újraindul a Felhom dobozod` - -> Szia! -> -> Ma éjjel a Felhom dobozod egy biztonsági frissítés miatt újraindul. Ilyenkor az alkalmazásaid néhány percig nem -> érhetők el, utána maguktól visszajönnek. Neked nem kell semmit tenned. -> -> Ha reggelre a doboz mégsem érhető el: húzd ki a tápkábelét, várj 10 másodpercet, és dugd vissza. A doboz ilyenkor -> magától a korábbi, bevált rendszerrel indul el. -> -> Ha ez sem segít, válaszolj erre a levélre, és segítünk. -> -> Üdv, Felhom.eu - -**English** — subject: `[Felhom] Your Felhom box restarts tonight` - -> Hi, -> -> Tonight your Felhom box restarts for a security update. Your apps will be away for a few minutes, then come back by -> themselves. You do not need to do anything. -> -> If the box is not back by morning: unplug its power cable, wait 10 seconds, and plug it back in. The box then starts -> its earlier, proven system by itself. -> -> If that does not help, reply to this e-mail and we will help. -> -> Best, Felhom.eu +**What the morning read-back must check:** each demo box on its told kernel as the default; the household apps back and +the minutes they were down; no false backup failure after the restart (R-897's fix); the hub's kernel events. The two +boxes run DIFFERENT kernels after tonight, so "Approve kernel set" waits until both boot the same one. ## Decisions for you -1. **Is the mail text right?** My pick: keep it. If you do nothing: it goes out as written. -2. **Mail time window 09:00–20:00, at most 3 mails per kernel.** My pick: keep (decided by me, you may reverse). If you - do nothing: it stays. +1. **Check the reply address in one click:** open „TEST — Reply-To check" in Gmail and press Reply — it should go to + admin@felhom.eu. My pick: do it once. If you do nothing: the code and its test say it works; nobody has seen it. +2. **The demo households' address is your own**, so their kernel mails carry no Reply-To. My pick: leave it. If you do + nothing: it stays (a reply to them lands in your catch-all anyway). diff --git a/STATUS.md b/STATUS.md index 362d2f94..ce73654b 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,17 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-07 18:40: hub 0.143.0; demo-hp, demo-felhom and Tester 1 run agent 0.152.0 and controller 0.302.0. -The open-items list is at 128. Report: `REPORT.md`.** +**Updated 2026-10-07 19:10: hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller 0.303.0. +The open-items list is at 126. Report: `REPORT.md`.** + +## Tonight (2026-10-07 → 08): the first real kernel night + +- **Both demo boxes restart tonight** for a new kernel; both households were mailed (18:12 and 18:38). +- **Three fixes went in first:** the box takes exactly the kernel named in the mail; the first backup after a restart + waits for the drive; a household's reply goes to your address. +- **Tomorrow morning:** read back the night — new kernels, apps back, no false backup failure. + +**Needs you:** press Reply on the „TEST — Reply-To check" mail once (it should go to admin@). If nothing: untested by eye. ## Evening (2026-10-07): the kernel lane is built diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 39ffb454..e4b5350b 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -35,6 +35,8 @@ The full text of every row below: `git show 5ba1702fcf:documentation/backlog/OPE | **R-105** | **Three hub-held DR records are empty on the entire live fleet.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED (hub 0.142.0, agent 0.151.0; `09` §3 decision 169) | The never-built `hosts.dr_record_json` and `host_escrow.directive_json` have no writer and no reader any more (the columns stay, unread); the agent's `-directive` flag is gone; `05` §9/§11 and `06` §3.5 corrected. Tester 2: the operator believes it never set up off-site, so it has no escrow (decision 170). `audits/day-2026-10-07/F/`. | | **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED (hub 0.142.0, agent 0.151.0; `09` §3 decision 168) | Slice 1 (the hub keeps the old key on a reinstall, hub 0.141.0) and slice 2 (the restore-test's skip of another key's archives becomes one edge-triggered operator line naming the box, the count and the date range) delivered to the three boxes; red-proofs `audits/day-2026-10-07/E/`. Option C (the recovery code at reinstall) declined by the operator. Not yet seen live: a reinstall that produces other-key archives. | | **R-812** | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED AND PROVEN LIVE (agent 0.151.0, hub 0.142.0; `09` §3 decision 163) | Option A, the Proxmox package lane: on demo-felhom one signed step installed 65 Proxmox packages in 70 s (`pve-manager` 9.2.2 → 9.2.21), healthy, the /etc/pve write gate held, the guest untouched (StartedAt identical), the app answered 31/31; the hub logged the run. `audits/day-2026-10-07/B/RESULT.md`. The kernel is R-836 (the spike done, `audits/kernel-spike-2026-10-07/`). | +| **R-897** | **After a host restart the guest's first backup capture could run in the second the agent re-bound the drive and fail (`permission denied`).** (P3) | CLOSED 2026-10-07 — DELIVERED: controller v0.303.0 — for 10 minutes after the controller starts, a capture for an app on a drive runs only once the drive is a live mount in the controller's namespace (refresh skips, data run waits; after the window it runs and logs once). The agent's bind order unchanged | controller `internal/backup/driveready.go`, `TestDriveReady_*` (red-proved). Delivered by per-customer floors to demo-hp, demo-felhom, Tester 1. First live check: the kernel night 2026-10-07→08 | +| **R-898** | **Ring 0's told kernel could be a day old: the step staged whatever was pending and refused (R23) a kernel newer than the one the household was told about.** (P3) | CLOSED 2026-10-07 — DELIVERED: agent v0.153.0 stages EXACTLY the told kernel (select `listed`, `KernelSet(kver)`); a told version no longer installable is refused before any change (R7) and the hub (v0.143.1, test-pinned) treats R7 as temporary and tells the household again for the newer kernel | `audits/kernel-night-2026-10-07/A/redproof.txt`; agent `TestKernel_Ring0ToldNightStagesThenReboots`, wrapper `test_ring0_listed_installs_the_told_kernel_not_the_newest`, `test_ring0_told_kernel_gone_is_refused_before_any_change`; hub `TestKernel_ToldKernelGoneIsTemporaryAndTheNewerIsOffered`. First live use: the night 2026-10-07→08 (demo-felhom told 7.0.14-20) | | **R-822** | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** (P2) | CLOSED 2026-10-07 — BY OPERATOR RULING (`09` §3 decision 166): the residual is accepted as a stated limit of decision 68 | Guard shipped controller v0.289.0; residual (past-dated fakes steering the keeps, a box-trusted window count) in `audits/night-burndown-2026-10-06/design-R-822.md`; the hub's own count is R-895, open. | ## 2026-10-07 (morning) — the read-back, the releases, the Tester 1 proof diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index adb5eebc..20371906 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -164,7 +164,6 @@ stopping line that lies. | **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** **2026-10-05 (burn-down night): NEEDS A DESIGN (and money).** A selection rule needs a multi-box configuration and a second box. Next: an operator pick of the rule and its trigger (e.g. add box 2 at 70 %). | — | — | CC | | **R-548** | Backup & restore | P3 | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. **-- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27** („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's `local` holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). `audits/pg-last-six-2026-09-30/C/`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** **Operator ruling 2026-10-05 18:23: kept OPEN as a known risk to a household's restore; owner the operator; not worked on in the burn-down.** | — | — | operator | -| **R-897** | Backup & restore | P3 | **After a host restart, the guest's first backup run can start while the agent re-binds the drive, and that app's backup fails.** MEASURED 2026-10-07 on demo-hp (a plain reboot for the kernel lane's wait, `audits/kernel-lane-2026-10-07/B/demo-hp/`): the agent bound `/mnt/felhom-drives/hdd_1` at 15:16:07 (prior_binds=0), the guest started after it and could not see that bind, so at 15:16:28 the reconcile NORMALIZED it (umount + mount, prior_binds=1, `internal/localapi/intermediary.go`); in that same second the controller's backup refresh failed for calibre-web: `recovery_unit_capture_failed` — `mkdir /mnt/felhom-drives/hdd_1/backups: permission denied` — and `backup_run_failures` (1 of 9). Every later 20 s tick was quiet; the app stayed healthy. **Why it matters now:** the kernel lane restarts a box at night (`09` §3 decision 172), so each kernel step can raise a false backup failure the morning after. Fix direction (a design, not taken here): the controller's post-boot backup run waits until the drive reads bound-and-live (`BoundUnderParent`), or the agent's first bind waits for the guest to run so no normalize is needed. | **OPEN — filed 2026-10-07 (kernel-lane arc).** | — | — | CC | | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk **Checked from source 2026-10-05 (burn-down round 2):** Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word. | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | @@ -215,8 +214,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). | — | — | CC | -| **R-898** | Box system & updates | P3 | **Ring 0's told kernel can be a day old: the hub names the kernel from the box's LAST host report, the box's sources may offer a newer one by night, and the step then refuses (R23) — one night and one household mail lost.** SEEN 2026-10-07 18:12: demo-felhom's household was mailed for 7.0.14-20-pve (from the night report) while the box's apt already offered 7.0.14-22; the night leg would stage `pending-kernel` → 22 ≠ the told 20 → R23 before any change (safe; read from the code, `felhom-os-apply` Kernel.stage). Fix direction (needs an agent + hub release, so not in this session): ring 0 stages EXACTLY the told kernel (select `listed`, the set derived from the kver — Proxmox keeps old kernel versions), and the hub treats a told-kernel mismatch as transient, not „ended for good". | **OPEN — filed 2026-10-07 (kernel-lane arc); tonight's demo-felhom run shows it.** | — | — | CC | +| **R-836** | Box system & updates | P2 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** **2026-10-07 07:58: raised P3 → P2 (`09` §3 decision 164 — the kernel lane before the first paying customer); the spike runs with the operator present.** **2026-10-07 SPIKE DONE (operator present; 24 reboots, 2 power cycles; `audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`).** On Tester 1 (VM), demo-felhom and demo-hp (Secure Boot on): **a one-shot flag in a GRUB env block on the ESP (vfat) WORKS** — new kernel once, then the old one; a panicking new kernel (`panic=10`) falls back with no person. **UEFI `BootNext` WORKS too.** **A hardware watchdog armed by systemd during the reboot FAILS on all three** (i6300ESB, Intel TCO, AMD SP5100): the reset clears the timer, so a FROZEN new kernel needs a power cycle. Pick: the ESP flag + lockup-to-panic kernel options (unmeasured). Two operator questions in the design (night restarts and the household; a freeze needs a person). Boxes left clean on a healthily booted default. **2026-10-07 14:43 — `09` §3 decision 172: the kernel lane is YES — option A (the ESP one-shot flag) with option C (lockup → panic) on the one-shot entry; a box may restart at night; the household is mailed the day before (what to do if it is not back: unplug, plug back in); each kernel set is operator-approved after ring 0; a frozen new kernel needing a power cycle is ACCEPTED for the first customers. BUILD STARTED 2026-10-07 (the kernel-lane arc).** **2026-10-07 evening: BUILT and proven by hand** — agent v0.152.0 + hub v0.143.0 (`11` §5.11), delivered to demo-hp, demo-felhom and Tester 1. Tester 1 (`audits/kernel-lane-2026-10-07/E/RESULT.md`): a forced panic on the one-shot → back on the old kernel by itself (fell_back, 0 unclean boots counted); a held guest → judged 20 min → ONE self-revert → old kernel; a healthy boot → 7.0.14-22 the default. The household mail went out for demo-felhom 18:12 (Gmail). Option C UNMEASURED (no test_lockup module). **WHERE IT STOPPED (next session):** the ring-0 NIGHT run (Part E3) — tonight demo-felhom refuses the told 7.0.14-20 (R-898, safe); 7.0.14-22 is due on both demo boxes for the night 2026-10-08→09 (mail 2026-10-08 daytime); read back the morning of 2026-10-09; then approve the set on the System page and deliver to the Tester 1 box as ring 1 by a signed os_kernel_step (Part E4 — Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the NEXT kernel). **2026-10-07 evening (decision 176): the night run moved to TONIGHT 7→8.** Pre-night fixes delivered (agent v0.153.0 — ring 0 stages exactly the told kernel, R-898 closed; controller v0.303.0 — the first capture after a boot waits for the drive, R-897 closed; hub v0.143.1 — Reply-To on household mails). Armed at 19:05: demo-felhom due 7.0.14-20 (household told 18:12), demo-hp due 7.0.14-22 (told 18:38; demo-hp's day-old apt lists refreshed by a report-only pass with its updates switched off for ~1 min). **WHERE IT STOPPED: read back the night on 2026-10-08 morning** (each box on its kernel as default, the apps back, minutes down, no false backup failure), then approve the set (the two boxes ran DIFFERENT kernels — the button waits until both boot the same one) and Part E4. | — | — | CC | | **R-862** | Box system & updates | P3 | **Tester 2 cannot take the config bundle until one by-hand bootstrap is done: its `felhom-os-apply` (agent 0.142.0) predates the bundle mode, and no signed job can write a root file on it.** FOUND 2026-10-04 (R-840 build, Part C): the route reaches every box installed from installer 1.31.0 on, and the demo boxes (bootstrapped by CC); Tester 2 has every root file of agent 0.142.0 (installer 1.30.0) and lacks only the R-858 wrapper fix, which matters only for a Docker step it gets solely from a signed job. CC has no route to Tester 2 (its door admits only the operator's WireGuard peer; `felhom-op` cannot become root). The steps: `runbooks/config-bundle.md` "Tester 2". Then CC signs the bundle and reads it back. | **WAITING-ON-OPERATOR** — the operator said (2026-10-04 ~19:05) he will try through his tunnel **2026-10-07 07:58 (`09` §3 decision 170): the operator believes Tester 2 never set up the off-site backup (no escrow); the laptop may stay off for days; the tester plans to add an HDD.** **2026-10-07 (Part H, read only):** in Tester 2's last 10 notifications (2026-10-05 09:16 → 2026-10-07 06:58) the household got 0 mails, the operator 10 (host_down ×9, expected_dbdump_missed ×1); control: the household channel shows on the other three customers. So the household is not mailed daily while the box is off; no row (`audits/day-2026-10-07/H/`). | the operator's WireGuard tunnel | the operator runs the three bootstrap commands; CC sends the bundle | operator | | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC |