pre-night fixes delivered; R-897, R-898 closed; night 7->8 armed (demo-felhom 7.0.14-20, demo-hp 7.0.14-22); 128 -> 126
gates / gates (push) Successful in 3m9s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-07 19:05:21 +02:00
parent d14578922f
commit 0829f0346c
5 changed files with 43 additions and 59 deletions
+5
View File
@@ -16,6 +16,11 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-07 (late evening) — pre-night fixes, the night moved to 7→8 (`09` §3 174–176).** Agent v0.153.0 (ring 0 stages
> exactly the told kernel — R-898 closed), controller v0.303.0 (`backup/driveready.go`: captures wait ≤10 min after start
> for a live drive mount — R-897 closed), hub v0.143.1 (`reply_to` = operator on every household mail). Armed: demo-felhom
> 7.0.14-20 (told 18:12), demo-hp 7.0.14-22 (told 18:38). Read back 2026-10-08 morning. Register 128 → 126.
> **2026-10-07 (evening) — the kernel lane BUILT (R-836, `09` §3 decisions 172–173; `11` §5.11).** Agent v0.152.0 (+ step
> bundle `0.152.0-step1`, + bundle) on demo-hp, demo-felhom, Tester 1; hub v0.143.0 (deployed attended). One-shot flag on
> the ESP, option C on the one-shot entry, default pinned to the running kernel and moved only by a healthy boot (host
+24 -54
View File
@@ -1,64 +1,34 @@
# REPORT — the kernel lane, 2026-10-07 (evening)
# REPORT — three fixes before the first real kernel night, 2026-10-07 (evening)
The brief said 2026-10-08; at 18:34 it was 2026-10-07. Asked; the operator answered *"Start today, and do both boxes this
night"* (`09` §3 decision 176). So the first real kernel night is **7→8**, read back on the morning of 2026-10-08.
| Part | Result |
|---|---|
| Rulings | Decisions 172 (kernel lane yes: A + C, night restarts, household mailed the day before, a freeze needing a person accepted) and 173 (R-32 leftovers left; P2 → P4, blocked on a safe main-account path) recorded FIRST (`0c5dbfbf`). |
| A — the one-shot boot | **Done.** Wrapper layer `kernel` (agent v0.152.0) + the two GRUB generators in the bundle (option C on the one-shot entry only). Red tests first-class: a box without the ESP flag refused (R20), a kernel outside the set refused (R23), the snippet clears the flag before booting, an install never moves the default. Red-proof `audits/kernel-lane-2026-10-07/A/redproof.txt` (16 mutations, each red). |
| B — boot good, the self-revert | **Done.** Health = the host rule + the hub reached, 20 min (measured: all healthy 68 s after a reboot on demo-felhom, 272 s on demo-hp; the spike's > 6 min Tailscale delay did not recur — no row). A box that never returns raises the hub's existing `host_stale` at 45 min and `host_down` at 90 min. One self-revert per step. Crash guard: a step adds at most ONE unclean boot (a panic before userspace adds none) — pinned by `KernelStepCannotLeaveTheBoxOff`, and seen live (0 counted after the panic). |
| C — night slot, approval, System page | **Done.** The kernel step ends the night leg (after a healthy host step, after the backup, trigger night, once per night); ring 1 stages only by a signed job. Hub v0.143.0 (deployed with you attending): due boxes, "Approve kernel set", two System page cells. `11` §5.11 + §8 step 6 written. |
| D — the household mail | **Done.** One mail the day before, 09:00–20:00 Budapest, its language, registered address; `tonight` only after the mail service accepted it (no mail, no step). Sent for real to demo-felhom at 18:12 (Gmail). Parity green (both locales; tests). |
| E 1 — Tester 1 by hand | **PASS ×3** (`audits/kernel-lane-2026-10-07/E/RESULT.md`): forced panic → back on the old kernel by itself (screens); held guest → judged 20 min → one self-revert; healthy → 7.0.14-22 the default. Tester 1 left on 7.0.14-22 (healthy, default). |
| E 2 — option C | **Unmeasured:** the Proxmox kernel has no lockup test module (`CONFIG_TEST_LOCKUP` not set). No tool built. |
| E 3 — ring-0 night run | **Not yet — next session.** Tonight demo-felhom was told about 7.0.14-20 (from last night's report) but its sources now offer 7.0.14-22: the step will refuse (R23) before any change (R-898). demo-hp's lists were a day old. Both demo boxes are due for 7.0.14-22 on the night 2026-10-08→09. |
| E 4 — Tester 1 as ring 1 | **Not yet** (after E3; Tester 1 already runs 7.0.14-22, so the ring-1 proof needs the next kernel). |
| Small | R-861 (a): no controller release in this arc — not read back. Tailscale delay: not reproduced → no row. |
| Rulings | Decisions 174 (mail text stays + a reply address), 175 (window 09–20, max 3 mails: keep), 176 (today, both boxes tonight) recorded FIRST (`a0ff737e`). |
| A — the told kernel boots (R-898) | **Done, delivered, closed.** Agent v0.153.0: ring 0 stages EXACTLY the told kernel (select `listed`, the set from the kver). Told 20 / sources offer 22 → 20 installs (wrapper test); 20 gone → R7 before any change (wrapper test); the hub keeps R7 temporary and tells the household again for the newer kernel (hub test). Red-proofs `audits/kernel-night-2026-10-07/A/redproof.txt`. |
| B — first backup after a restart waits (R-897) | **Done, delivered, closed.** The reviewer's pick, taken: controller v0.303.0 — for 10 min after the controller starts, a capture for an app on a drive runs only once that drive is a LIVE mount in the controller's own namespace (`/proc/self/mountinfo`, the signal the startup app gate already uses); the 5-min refresh skips, a data run waits; after 10 min it runs and logs once. The agent's bind order is unchanged. Red test: not-bound → no capture; bound → capture. |
| C — a reply reaches a person | **Built and delivered (hub v0.143.1); the header read-back is NOT done.** Every mail to a household now carries `reply_to` = `admin@felhom.eu` (all household templates invite contact: the kernel notice, the event sign-off "contact your operator", the setup/link mails); a mail to the operator carries none. Test pinned and red-proved. One test mail sent through the mail service with the hub's exact fields to `tester1@felhom.eu` — it arrived in Gmail (16:54 UTC). **Read-back tried:** the Gmail tool returns no headers (METADATA_ONLY, no RAW); the mail service's read API refused (`restricted_api_key` — send-only, correctly). Note: the demo households' own address IS `admin@felhom.eu`, so their mails carry no Reply-To by design. |
| D — deliver, arm the night | **Done.** Hub 0.143.1 deployed 18:54 local (operator attending), agent 0.153.0 on all three boxes 19:03 (binary only — no root file changed), controller 0.303.0 on all three 18:50 (per-customer floors; global untouched). **Armed:** demo-felhom → **7.0.14-20-pve**, household mail **18:12**; demo-hp → **7.0.14-22-pve**, household mail **18:38**. No reboot by hand. |
**Rows: 126 before → 128 after. Opened 2 (R-897, R-898). Closed 0.**
**Rows: 128 before → 126 after. Opened 0. Closed 2 (R-897, R-898).**
**Releases:** agent v0.152.0 (tag `d03ab7f`; binary `95ff4220…`, bundle `f0c2cec3…`, step bundle `0.152.0-step1`
`0b71d32b…`), delivered by signed jobs to demo-hp, demo-felhom and Tester 1 (binary → step bundle → bundle). Hub v0.143.0
(`ab3b7ea2`, deployed `0ece2ff7` 15:39, Synced/Healthy). The installer's uninstall now knows the two GRUB files
(unreleased, no tag cut). Nothing on ep0, Tester 2 or DooPlex's own system; no Docker engine step; no delivery after 02:00.
**How demo-hp got told tonight:** its apt lists were a day old, so its last report named no new kernel. I switched its OS
updates OFF on the hub for ~1 minute, ran the agent's own report-only pass (`--selftest=os-update`; with the switch off it
installs nothing and runs no Docker/Proxmox/kernel step), and switched it back ON (`kernel-night-2026-10-07/demo-hp-inventory-pass.txt`).
The hub mailed 20 s after the host report arrived (18:37:51 → 18:38:11).
**Read for this report:** `11-os-updates.md` (§5.3, §5.8–5.10, §8.2 — §5.11 new), `03-host-agent.md`, `07` §6.1,
`audits/kernel-spike-2026-10-07/DESIGN-kernel-lane.md`.
**Releases:** agent v0.153.0 (binary `b204ebe6…`, tag at `2d1e5d0`); controller v0.303.0 (`29bebbb`, MinAgent 0.131.0);
hub v0.143.1 (`87765bfa`, deployed `d1457892`, Synced/Healthy). CI green on every push, checked by commit. No change on
ep0, Tester 2 or DooPlex's own system.
**Also seen:** the hub image was built while one audit text file (outside `hub/`) was uncommitted; the build copies
`hub/` only, and `hub/` equalled the pushed commit (checked by diff). R-897: after demo-hp's timing reboot, one app's
backup failed in the second the agent re-bound the drive.
## The household mail (both languages, as sent)
**Hungarian** — subject: `[Felhom] Ma éjjel újraindul a Felhom dobozod`
> Szia!
>
> Ma éjjel a Felhom dobozod egy biztonsági frissítés miatt újraindul. Ilyenkor az alkalmazásaid néhány percig nem
> érhetők el, utána maguktól visszajönnek. Neked nem kell semmit tenned.
>
> Ha reggelre a doboz mégsem érhető el: húzd ki a tápkábelét, várj 10 másodpercet, és dugd vissza. A doboz ilyenkor
> magától a korábbi, bevált rendszerrel indul el.
>
> Ha ez sem segít, válaszolj erre a levélre, és segítünk.
>
> Üdv, Felhom.eu
**English** — subject: `[Felhom] Your Felhom box restarts tonight`
> Hi,
>
> Tonight your Felhom box restarts for a security update. Your apps will be away for a few minutes, then come back by
> themselves. You do not need to do anything.
>
> If the box is not back by morning: unplug its power cable, wait 10 seconds, and plug it back in. The box then starts
> its earlier, proven system by itself.
>
> If that does not help, reply to this e-mail and we will help.
>
> Best, Felhom.eu
**What the morning read-back must check:** each demo box on its told kernel as the default; the household apps back and
the minutes they were down; no false backup failure after the restart (R-897's fix); the hub's kernel events. The two
boxes run DIFFERENT kernels after tonight, so "Approve kernel set" waits until both boot the same one.
## Decisions for you
1. **Is the mail text right?** My pick: keep it. If you do nothing: it goes out as written.
2. **Mail time window 09:00–20:00, at most 3 mails per kernel.** My pick: keep (decided by me, you may reverse). If you
do nothing: it stays.
1. **Check the reply address in one click:** open „TEST — Reply-To check" in Gmail and press Reply — it should go to
admin@felhom.eu. My pick: do it once. If you do nothing: the code and its test say it works; nobody has seen it.
2. **The demo households' address is your own**, so their kernel mails carry no Reply-To. My pick: leave it. If you do
nothing: it stays (a reply to them lands in your catch-all anyway).
+11 -2
View File
@@ -2,8 +2,17 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-07 18:40: hub 0.143.0; demo-hp, demo-felhom and Tester 1 run agent 0.152.0 and controller 0.302.0.
The open-items list is at 128. Report: `REPORT.md`.**
**Updated 2026-10-07 19:10: hub 0.143.1; demo-hp, demo-felhom and Tester 1 run agent 0.153.0 and controller 0.303.0.
The open-items list is at 126. Report: `REPORT.md`.**
## Tonight (2026-10-07 → 08): the first real kernel night
- **Both demo boxes restart tonight** for a new kernel; both households were mailed (18:12 and 18:38).
- **Three fixes went in first:** the box takes exactly the kernel named in the mail; the first backup after a restart
waits for the drive; a household's reply goes to your address.
- **Tomorrow morning:** read back the night — new kernels, apps back, no false backup failure.
**Needs you:** press Reply on the „TEST — Reply-To check" mail once (it should go to admin@). If nothing: untested by eye.
## Evening (2026-10-07): the kernel lane is built
+2
View File
@@ -35,6 +35,8 @@ The full text of every row below: `git show 5ba1702fcf:documentation/backlog/OPE
| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED (hub 0.142.0, agent 0.151.0; `09` §3 decision 169) | The never-built `hosts.dr_record_json` and `host_escrow.directive_json` have no writer and no reader any more (the columns stay, unread); the agent's `-directive` flag is gone; `05` §9/§11 and `06` §3.5 corrected. Tester 2: the operator believes it never set up off-site, so it has no escrow (decision 170). `audits/day-2026-10-07/F/`. |
| **R-366** | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED (hub 0.142.0, agent 0.151.0; `09` §3 decision 168) | Slice 1 (the hub keeps the old key on a reinstall, hub 0.141.0) and slice 2 (the restore-test's skip of another key's archives becomes one edge-triggered operator line naming the box, the count and the date range) delivered to the three boxes; red-proofs `audits/day-2026-10-07/E/`. Option C (the recovery code at reinstall) declined by the operator. Not yet seen live: a reinstall that produces other-key archives. |
| **R-812** | **[P2] A box never receives operating-system security updates — not the Proxmox host, not the guest's Debian, not its Docker engine.** (P2) — full text `git show 374213f279:documentation/backlog/OPEN-ITEMS.md` | CLOSED 2026-10-07 — DELIVERED AND PROVEN LIVE (agent 0.151.0, hub 0.142.0; `09` §3 decision 163) | Option A, the Proxmox package lane: on demo-felhom one signed step installed 65 Proxmox packages in 70 s (`pve-manager` 9.2.2 → 9.2.21), healthy, the /etc/pve write gate held, the guest untouched (StartedAt identical), the app answered 31/31; the hub logged the run. `audits/day-2026-10-07/B/RESULT.md`. The kernel is R-836 (the spike done, `audits/kernel-spike-2026-10-07/`). |
| **R-897** | **After a host restart the guest's first backup capture could run in the second the agent re-bound the drive and fail (`permission denied`).** (P3) | CLOSED 2026-10-07 — DELIVERED: controller v0.303.0 — for 10 minutes after the controller starts, a capture for an app on a drive runs only once the drive is a live mount in the controller's namespace (refresh skips, data run waits; after the window it runs and logs once). The agent's bind order unchanged | controller `internal/backup/driveready.go`, `TestDriveReady_*` (red-proved). Delivered by per-customer floors to demo-hp, demo-felhom, Tester 1. First live check: the kernel night 2026-10-07→08 |
| **R-898** | **Ring 0's told kernel could be a day old: the step staged whatever was pending and refused (R23) a kernel newer than the one the household was told about.** (P3) | CLOSED 2026-10-07 — DELIVERED: agent v0.153.0 stages EXACTLY the told kernel (select `listed`, `KernelSet(kver)`); a told version no longer installable is refused before any change (R7) and the hub (v0.143.1, test-pinned) treats R7 as temporary and tells the household again for the newer kernel | `audits/kernel-night-2026-10-07/A/redproof.txt`; agent `TestKernel_Ring0ToldNightStagesThenReboots`, wrapper `test_ring0_listed_installs_the_told_kernel_not_the_newest`, `test_ring0_told_kernel_gone_is_refused_before_any_change`; hub `TestKernel_ToldKernelGoneIsTemporaryAndTheNewerIsOffered`. First live use: the night 2026-10-07→08 (demo-felhom told 7.0.14-20) |
| **R-822** | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** (P2) | CLOSED 2026-10-07 — BY OPERATOR RULING (`09` §3 decision 166): the residual is accepted as a stated limit of decision 68 | Guard shipped controller v0.289.0; residual (past-dated fakes steering the keeps, a box-trusted window count) in `audits/night-burndown-2026-10-06/design-R-822.md`; the hub's own count is R-895, open. |
## 2026-10-07 (morning) — the read-back, the releases, the Tester 1 proof
File diff suppressed because one or more lines are too long