Files
felhom.eu/REPORT.md
T
2026-10-06 19:38:44 +02:00

6.5 KiB
Raw Blame History

REPORT — night read-back brief, the night-free parts (2026-10-06 evening; Parts B and D follow after 08:30 on 2026-10-07)

The brief said „start after 08:30 on 2026-10-07". The operator then said „start A, C, E, F now". Parts B (the night read-back) and D (one press, outside 02:00–08:30) wait for tomorrow. Nothing was delivered to a box tonight, so the night of 2026-10-06→07 runs controller v0.301.0 and agent v0.149.0, and Part B reads a clean result. The hub was released (it is not a box).

Part Result
A — the rulings done — 09 §3 decisions 157 (R-528 option A + the Docker-approval check) and 158 (R-892: VM 341 on the HP box); R-892's "no Proxmox host known" corrected
B — the night read-back (R-518) waits for 08:30 2026-10-07
C — demo-hp's off-site copy diagnosed: the premise was wrong — no 10-day gap. Last copy 2026-10-01 20:15Z (ep0's own listing, verify ok); every box within 7 days. Two real findings: an off-site tier read DUE after an agent restart while its storage was unreachable (R-894, filed), and the alarm's operator mail failed and was never retried (fixed, hub v0.140.0)
D — one press, measured waits for after 08:30 2026-10-07
E — the Tester 1 box in the update test (R-892) route built and identity matched; the live proof is BLOCKED — DooPlex's key is not authorized on VM 341, and the permission check refused fetching its vaulted password (the brief's rule: a refusal stops that item)
F — the Docker-approval memory-kill check (R-528) built: hub v0.140.0 LIVE (the approval waits for a passing check); the agent's wrapper merged, unreleased (ships as v0.150.0 after Part B); proven by hand on demo-hp's guest
small check — wger's 100 % peak file cache, not a kill (kill counter 0, restarts 0, anon 50.3 %); the memory hint stays
Rows before Rows after Opened Closed
137 138 1 (R-894) 0 (so far)

Part C — what happened on demo-hp, with times

  • My last report was wrong: it said „demo-hp had no successful off-site run in 10 days". I read a log view cut by tail -40 that started on 2026-10-04. The agent's full log and ep0's own listing both show a completed off-site copy on 2026-10-01 20:15Z (15.2 GB, verify ok). With the 7-day cadence the next one is due ~2026-10-08.
  • ep0, read only (C/C2-ep0-snapshot-listing.txt): demo-hp 2026-10-01T20:15Z, demo-felhom 2026-10-06T04:21Z, tester-1 2026-10-04T19:57Z, Tester-2 2026-10-04T16:31Z — every box within 7 days.
  • 2026-10-05 04:25Z (06:25 local), demo-hp: the off-site storage answered Can't connect to 10.77.0.1:8007; the agent had restarted at 02:57Z (04:57 local), 1.5 hours before; its per-tier backup record is in memory only (internal/backup/store.go, R-348), so the due-check fell back to an EMPTY record and read the tier DUE although it was not (internal/localapi/server.go newestArchiveOn → unknown → in-memory). The controller asked; vzdump failed (could not activate storage 'felhom-pbs'). By design the controller stops the apps before it asks; whether it did that night is not known (the controller's log was lost to a later restart). → R-894.
  • Who was told: the hub recorded whole_guest_backup_failed (error) at 04:27Z; the household was not mailed (by design: operator only); the operator's mail FAILED (Resend: context deadline exceeded) and was never retried. It is the only failed operator mail of 692 since 2026-02-16 (C/C4-hub-failed-mails.txt).
  • Fixed (hub v0.140.0): a failed operator mail is tried again after 1, 5 and 15 minutes; each try is a row; giving up is an ERROR line. Red-proof C/red-operator-mail-retry.txt („the failed mail was never sent again (tries=1)").
  • Not read: the household's backups page on demo-hp (no session tonight).

Part E — the Tester 1 box

  • Identity (E1-identity-match.txt): VM 341 night1004-tester1 on demo-hp, at 192.168.0.154 (found by its MAC); its certificate CN=felhom.enkicsifelhom.hu; the agent's own hub report for tester-1-d70be4 says host.node=felhom; its guest at 192.168.0.101 answers felhom.enkicsifelhom.hu with the Felhom login (200) and demo-hp's domain 404.
  • Built (catalog d63ea35): box_walk.py TARGETS (9202, 9201, tester-1 via -J demo-hp), BOX_ADMIN_SEED_GUESTS adds ("tester-1", "9201"); tests BoxWalkTargets (red-proved). operations/nodes.md has the box.
  • Blocked: ssh -J demo-hp root@192.168.0.154 → Permission denied (publickey,password). The VM has no guest agent; editing its disk needs a VM stop (a reboot — not allowed). I tried to read the hub's code for revealing the vaulted console password; the session's permission check refused it, and I did not try another way. R-892 now asks the operator.

Part F — the memory-kill check

  • Hub v0.140.0 (LIVE 17:34Z): DockerStatus needs, per ring-0 box, a passing oom_check with the set; failed, errored or missing blocks. Red-proof F/red-hub-docker-approval.txt; the System page hides the button and says why.
  • Agent (acccb66, unreleased): after a Docker step the wrapper runs a throwaway container from the running controller's image (--pull never, --network none, 64 MB cap, one 200 MB block); pass = OOMKilled=true AND the oom event; always removed. 12 red-proofs in F/. The helper also found 11 wrapper tests that never ran (a unittest.main() mid-file) — moved; all pass.
  • By hand on demo-hp's guest (F/F1, F/F2): OOMKilled=true, exit 137, 21 → 21 containers, none left. The oom event was MISSING when the events window ended in the same second as the run, and present (create attach start oom die) with the window ending a second later — the wrapper waits 2 s and reads to epoch + 1.
  • Until agent v0.150.0 reaches the ring-0 boxes, no Docker set can be approved (none is pending).

Instruction-file edits

None tonight.

CI

felhom.eu e0bdd52 → 1448 success (and eed1dbd 1446, 01c4a5d 1447); catalog d63ea35 → 1445; agent 7e82f32 → 1449; this commit checked after its push.

What is left for 2026-10-07 after 08:30

Part B (read both demo boxes' night under v0.301.0), Part D (one press on a demo box, measured; the page text), the agent v0.150.0 release + bundle and its delivery, the controller release for Part D's text, and Part E's live proof if the operator authorizes the key. Teardown tonight: none needed (no box provisioned; the hub DB copies deleted; the by-hand check containers removed — 0 left).