kernel night 7->8 read back: demo-felhom passed (7.0.14-20 default, ~1.5 min), demo-hp no step (R-899 filed); 126 -> 127
gates / gates (push) Successful in 2m58s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-08 06:41:20 +02:00
parent 0829f0346c
commit de4a8d20ba
6 changed files with 142 additions and 28 deletions
+21 -25
View File
@@ -1,34 +1,30 @@
# REPORT — three fixes before the first real kernel night, 2026-10-07 (evening)
# REPORT — read-back of the first real kernel night (2026-10-07 → 08)
The brief said 2026-10-08; at 18:34 it was 2026-10-07. Asked; the operator answered *"Start today, and do both boxes this
night"* (`09` §3 decision 176). So the first real kernel night is **7→8**, read back on the morning of 2026-10-08.
| Part | Result |
| Check | Result |
|---|---|
| Rulings | Decisions 174 (mail text stays + a reply address), 175 (window 09–20, max 3 mails: keep), 176 (today, both boxes tonight) recorded FIRST (`a0ff737e`). |
| A — the told kernel boots (R-898) | **Done, delivered, closed.** Agent v0.153.0: ring 0 stages EXACTLY the told kernel (select `listed`, the set from the kver). Told 20 / sources offer 22 → 20 installs (wrapper test); 20 gone → R7 before any change (wrapper test); the hub keeps R7 temporary and tells the household again for the newer kernel (hub test). Red-proofs `audits/kernel-night-2026-10-07/A/redproof.txt`. |
| B — first backup after a restart waits (R-897) | **Done, delivered, closed.** The reviewer's pick, taken: controller v0.303.0 — for 10 min after the controller starts, a capture for an app on a drive runs only once that drive is a LIVE mount in the controller's own namespace (`/proc/self/mountinfo`, the signal the startup app gate already uses); the 5-min refresh skips, a data run waits; after 10 min it runs and logs once. The agent's bind order is unchanged. Red test: not-bound → no capture; bound → capture. |
| C — a reply reaches a person | **Built and delivered (hub v0.143.1); the header read-back is NOT done.** Every mail to a household now carries `reply_to` = `admin@felhom.eu` (all household templates invite contact: the kernel notice, the event sign-off "contact your operator", the setup/link mails); a mail to the operator carries none. Test pinned and red-proved. One test mail sent through the mail service with the hub's exact fields to `tester1@felhom.eu` — it arrived in Gmail (16:54 UTC). **Read-back tried:** the Gmail tool returns no headers (METADATA_ONLY, no RAW); the mail service's read API refused (`restricted_api_key` — send-only, correctly). Note: the demo households' own address IS `admin@felhom.eu`, so their mails carry no Reply-To by design. |
| D — deliver, arm the night | **Done.** Hub 0.143.1 deployed 18:54 local (operator attending), agent 0.153.0 on all three boxes 19:03 (binary only — no root file changed), controller 0.303.0 on all three 18:50 (per-customer floors; global untouched). **Armed:** demo-felhom → **7.0.14-20-pve**, household mail **18:12**; demo-hp → **7.0.14-22-pve**, household mail **18:38**. No reboot by hand. |
| demo-felhom — step ran | **PASS.** Whole-guest backup 04:35; guest/host/Proxmox steps 04:37–04:38; kernel **staged 04:39:04 as exactly the told 7.0.14-20** (the sources offered 7.0.14-22 — R-898's fix held); restart 04:39:08; boot on 7.0.14-20; judged healthy **04:40:38** (38 s after the agent started); **7.0.14-20 is the default**, flag empty, step `good`. |
| demo-felhom — apps | Away about **1.5 min** (restart 04:39:08 → every container healthy at the 04:40:38 verdict). |
| demo-felhom — alarms | None: no `host_stale`, no backup failure mail (demo-felhom has no drive apps, so R-897's fix was not exercised there). Crash guard armed, 0 unclean boots. Hub logged staged → judging → applied. |
| demo-hp — step ran | **NO.** No whole-guest backup that night: the last one was yesterday's morning press at 08:49:53; with the 24 h cadence it came due at 08:49 today, after the window closed (08:30). No backup → no night leg → no kernel step. Nothing changed: running and default 7.0.14-20. Filed **R-899**. |
| Reply-To | **Proven by you:** your reply to the test mail (2026-10-07 20:29) went to admin@felhom.eu. |
| Your inbox overnight | Only Tester 2's expected missed-backup mails (the laptop is off). |
**Rows: 128 before → 126 after. Opened 0. Closed 2 (R-897, R-898).**
**Rows: 126 before → 127 after. Opened 1 (R-899). Closed 0.**
**How demo-hp got told tonight:** its apt lists were a day old, so its last report named no new kernel. I switched its OS
updates OFF on the hub for ~1 minute, ran the agent's own report-only pass (`--selftest=os-update`; with the switch off it
installs nothing and runs no Docker/Proxmox/kernel step), and switched it back ON (`kernel-night-2026-10-07/demo-hp-inventory-pass.txt`).
The hub mailed 20 s after the host report arrived (18:37:51 → 18:38:11).
**What happens next, by itself:**
- demo-hp is still due 7.0.14-22. A new household mail is allowed after 14:38 today (20 h after the last); its
whole-guest backup is due inside tonight's window → the step should run on the night 8→9.
- demo-felhom now sees 7.0.14-22 pending → due; its household gets a mail after 09:00; it should step to 22 tonight too.
- After that both run 7.0.14-22, and "Approve kernel set" can appear. Read back the morning of 2026-10-09.
**Releases:** agent v0.153.0 (binary `b204ebe6…`, tag at `2d1e5d0`); controller v0.303.0 (`29bebbb`, MinAgent 0.131.0);
hub v0.143.1 (`87765bfa`, deployed `d1457892`, Synced/Healthy). CI green on every push, checked by commit. No change on
ep0, Tester 2 or DooPlex's own system.
**Seen, not fixed:** the after-boot "judging" report carries ring 1 (the agent reads its ring from the hub's block, which
it has not fetched yet in the first second after a boot). Cosmetic: the hub's approval reads its own ring list, not
this field. Not fixed today because it needs an agent release and delivery for a label.
**What the morning read-back must check:** each demo box on its told kernel as the default; the household apps back and
the minutes they were down; no false backup failure after the restart (R-897's fix); the hub's kernel events. The two
boxes run DIFFERENT kernels after tonight, so "Approve kernel set" waits until both boot the same one.
Evidence: `documentation/audits/kernel-night-2026-10-07/readback/`.
## Decisions for you
1. **Check the reply address in one click:** open „TEST — Reply-To check" in Gmail and press Reply — it should go to
admin@felhom.eu. My pick: do it once. If you do nothing: the code and its test say it works; nobody has seen it.
2. **The demo households' address is your own**, so their kernel mails carry no Reply-To. My pick: leave it. If you do
nothing: it stays (a reply to them lands in your catch-all anyway).
1. **R-899 — a daytime "back up now" press skips the next night's backup (and its updates).** My pick: count nights,
not 24 hours (the backup is due when the last one is older than ~20 h at the window's start). If you do nothing: each
daytime press costs one night, and the kernel step waits one more day.