Files
felhom.eu/REPORT-day4-2026-10-09.md
T
2026-10-09 09:00:05 +02:00

5.0 KiB
Raw Blame History

REPORT — 2026-10-09: the kernel night read back, the release, the ep0-copy job, the live proofs

Operator present for the hub deploy (05:22Z) and the ep0-copy job (05:24–05:33Z). Architecture read: 07 §6.1 (backups), 11 §5.11 (the kernel lane), 09 §3 decisions 172, 181, 185–193.

Part table

Part Result
Kernel night 8→9 read-back demo-felhom: told 7.0.14-22, staged 04:44:42, restart 04:44:46, healthy 39 s after the agent, default now 7.0.14-22; apps down ~70 s (04:44:46 → 04:45:55). demo-hp: no kernel — the Proxmox step was refused R6 (proxmox-secure-boot-support pulled shim-signed-common), so the kernel step was skipped; still 7.0.14-20. „Approve kernel set": not shown (demo-hp has not booted the new kernel). One false backup alarm: demo-felhom „host tier: newest backup is 48h old" at 05:00 local — the backup ran 04:40; the restart emptied the agent's in-memory list (two nights running).
Two agent fixes (found in the read-back) The host report adds the saved last success per tier (R-894's file); the Secure Boot meta-package joins the boot-chain lane. Both red-proved. Live: demo-felhom's report now carries the 02:40Z backup after a restart.
Release agent 0.154.0 (db25b46, sha 36a54ad2…, bundle 93487989…), controller 0.304.0 (d7bfdcc), hub 0.144.0 (8dbdac5f, deployed 702e19ee). One release per repo.
Delivery (before 20:00) demo-hp, demo-felhom, Tester 1: agent 0.154.0 (signed agent_update, 07:16–07:18 local), bundle (68/68 capability probe each), controller 0.304.0 (floors set 07:05 local, MinAgent 0.131.0). Tester 2 not touched. Global floor not touched. Agent NOT vouched for fresh installs (the vouch also needs a golden ≥ the fleet's controller; no golden today).
Hub deploy Synced/Healthy at 702e19ee, pod image 0.144.0, log „felhom-hub 0.144.0 starting" (operator present).
D9 ep0-copy job Installed with the operator present: new local PBS key root@pam!ep0-copy-gc (DatastoreAdmin on /datastore/ep0-copy only), three root-only files, never printed. First dry run found a parser defect (ep0 answers {"data": [...]}; the job read none of the customers) — the mass-absence guard stopped it, nothing recorded or deleted. Fixed, red-proved, pushed, reinstalled; dry run 5 of 5; timer on; first real run 08:00 local: 5 of 5, deleted 0.
D1 hub buttons Proven on Tester 1 (operator's choice: 9202 has no hub): run_job fill-watch and offsite_backup_now both „done" on the host page; box log controls. Countdown buttons not proven (no countdown exists).
D2 box off Proven on Tester 1: whole VM off → „last connected" at +114 s, tick open at 6 min, ingress log control. A stopped guest alone is restarted by the agent in 36 s.
D3 health mail + dashboard Proven: household mail (Gmail) has no { and has the dashboard line; operator mail keeps the details; 9202's banners English/Hungarian by language.
D4 sign-in survives restart Proven on 9202, including the password-change revocation.
D6 Cloudflare keys Done later with the operator present: the hub's own check run read-only against all 4 stored keys — all PASS (one zone, their own domain); a fake key is refused. A same-key re-save would not have checked anything (the edit path checks only a new key or domain), so the check ran outside the save.
D7, D8 Not run: D7 needs a window (Tester 1's next ~2026-10-12); D8 needs a scratch off-site repo.

Open-items list: 141 when the register work started, 135 after. Opened 0, closed 6 (R-279, R-30, R-79, R-35, R-901, R-138). The three small findings (false backup alarm, Secure Boot meta-package, ep0 job parser) were fixed in the session, not filed.

What changed on the boxes during the proofs (all put back)

  • Tester 1: the guest stopped once (36 s, restarted by the agent); the whole VM shut down 05:53–05:58Z; filebrowser stopped twice (~5 min each); the household mail address set to the catch-all for one mail, then cleared (empty address + no events; the hub keeps the old address by its no-clobber guard, with no events, so no mail goes out).
  • 9202: controller image 0.303.0 → 0.304.0 by hand; the password changed and set back; language en and back to hu.
  • DooPlex: /etc/felhom/ep0-copy-gc/ (3 root-only files), /usr/local/sbin/felhom-ep0-copy-gc, the service + timer, one PBS token and its ACL. No Docker cleanup, no reboot.

Other session

A parallel session had an unstaged deletion of REPORT-facebook-cover-r919.md in the shared clone; an autostash met its newer commit. I put the working tree back to exactly its state (the file deleted, unstaged) and dropped the stash.

Instruction-file edits

None. A stale fact was corrected in documentation/operations/nodes.md (not an instruction file): Tester 1 SSH by key from DooPlex now works (root@192.168.0.154); before it said the key was not authorized (2026-10-06).

Evidence: documentation/audits/kernel-night-2026-10-08/, documentation/audits/release-2026-10-09/.