Decisions 157-158 recorded; R-892 corrected (VM 341), R-894 filed (unreadable off-site storage reads due after an agent restart); Part C/E/F evidence; nodes.md: the Tester 1 box; 11 §5.8: the memory-kill check
gates / gates (push) Successful in 2m47s
gates / gates (push) Successful in 2m47s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -915,6 +915,17 @@ its length, and both fixes cost something the household would notice — operato
|
||||
127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer.
|
||||
*Operator ruling 2026-10-05.* (R-861)
|
||||
|
||||
### 2026-10-06 (18:24) — two operator rulings (recorded before the work)
|
||||
|
||||
157. **R-528 — option A: the counter read and the „probably out of memory" text are NOT built.** The row stays open as
|
||||
a watch. **Added:** a new Docker engine set cannot be approved until it is shown to report a memory kill correctly
|
||||
(`OOMKilled=true` and the `oom` event) on the boxes that ran it. *Operator ruling 2026-10-06 18:24.* (It confirms
|
||||
the measurement that stopped decision 155's build: Docker 29.8.2 reported every kill tried.)
|
||||
158. **R-892 — yes: CC may reach the Tester 1 box by SSH.** It runs as **VM 341 on the HP box** (`ssh hp`; evidence
|
||||
`audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`). Disposable; decision 149 already allows the admin seed
|
||||
there. *Operator ruling 2026-10-06 18:24.* **The reviewer's error, recorded:** the previous brief named „the Tester 1
|
||||
box" without saying where it runs, although the reviewer knew; R-892 was filed because of that gap.
|
||||
|
||||
### 2026-10-06 (14:24) — three operator rulings and the reviewer's three design picks (recorded before the work)
|
||||
|
||||
151. **R-469 — closed.** The engine-major gate stays as decision 35's permanent per-app check. *Operator ruling
|
||||
|
||||
@@ -415,6 +415,14 @@ must never overlap a backup, a restore-test or a self-update.~~
|
||||
- **Approval**: the hub never approves a Docker set automatically; the System page's button works after every ring-0 box
|
||||
ran the set in 2 healthy night Docker steps (`OS_DOCKER_APPROVE_NIGHTS` TEST override, logged). An approval nudges no
|
||||
box. Undo: `runbooks/os-updates-docker-undo.md`.
|
||||
- **The engine must report a memory kill (decision 157, agent v0.150.0 + hub v0.140.0, 2026-10-06).** After a Docker step
|
||||
the wrapper runs `oom_check`: a throwaway container from the image the running controller uses (`--pull never`,
|
||||
`--network none`, no volume, label `felhom.oomcheck=1`, 64 MB cap) asks for one 200 MB block; pass = `OOMKilled=true` AND
|
||||
the `oom` event; the container is removed whatever happens; the result rides the report as `oom_check`. The approval
|
||||
button waits until every ring-0 box reported a PASSING check with the set, and refuses while any report of the set
|
||||
carries a failed or errored one. **Measured by hand on demo-hp's guest (29.8.2):** `OOMKilled=true`, exit 137, the
|
||||
event `create attach start oom die` — but only when the events window ends a second AFTER the run (a window closed in
|
||||
the same second missed it); the wrapper waits 2 s and reads to epoch + 1 (`audits/readback-2026-10-07/F/`).
|
||||
|
||||
**The design as written before the build:**
|
||||
|
||||
|
||||
@@ -0,0 +1,133 @@
|
||||
Sep 22 04:29:29 demo-hp felhom-agent[1195783]: time=2026-09-22T04:29:29.642+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-15T04:11:51Z) is already proven; local: newest settled archive (landed 2026-09-20T04:43:47Z) is already proven"
|
||||
Sep 22 06:12:46 demo-hp felhom-agent[1195783]: time=2026-09-22T06:12:46.231+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1790050365164112326
|
||||
Sep 22 06:16:22 demo-hp felhom-agent[1195783]: time=2026-09-22T06:16:22.067+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-22T04:12:45Z size_bytes=18418647089 uncovered_volumes=2
|
||||
Sep 22 06:16:22 demo-hp felhom-agent[1195783]: time=2026-09-22T06:16:22.067+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1790050365164112326 archive=felhom-pbs:backup/ct/9201/2026-09-22T04:12:45Z
|
||||
Sep 22 06:58:40 demo-hp felhom-agent[1195783]: time=2026-09-22T06:58:40.652+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_22-06_52_48.tar.zst size_bytes=6267269544 uncovered_volumes=2
|
||||
Sep 22 06:58:40 demo-hp felhom-agent[1195783]: time=2026-09-22T06:58:40.652+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790052768442108602 archive=local:backup/vzdump-lxc-9201-2026_09_22-06_52_48.tar.zst
|
||||
Sep 22 10:29:29 demo-hp felhom-agent[1195783]: time=2026-09-22T10:29:29.346+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/vzdump-lxc-9201-2026_09_21-06_47_30.tar.zst landed=2026-09-21T04:47:30Z reason="newest settled archive (landed 2026-09-21T04:47:30Z) has not been proven (last proven archive was a dif
|
||||
Sep 22 10:34:49 demo-hp felhom-agent[1195783]: time=2026-09-22T10:34:49.417+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=local:backup/vzdump-lxc-9201-2026_09_21-06_47_30.tar.zst duration_s=320.064619459
|
||||
Sep 22 16:29:29 demo-hp felhom-agent[1195783]: time=2026-09-22T16:29:29.663+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-15T04:11:51Z) is already proven; local: newest settled archive (landed 2026-09-21T04:47:30Z) is already proven"
|
||||
Sep 22 22:29:29 demo-hp felhom-agent[1195783]: time=2026-09-22T22:29:29.659+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-15T04:11:51Z) is already proven; local: newest settled archive (landed 2026-09-21T04:47:30Z) is already proven"
|
||||
Sep 23 04:29:29 demo-hp felhom-agent[1195783]: time=2026-09-23T04:29:29.655+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-15T04:11:51Z) is already proven; local: newest settled archive (landed 2026-09-21T04:47:30Z) is already proven"
|
||||
Sep 23 07:02:53 demo-hp felhom-agent[1195783]: time=2026-09-23T07:02:53.924+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst size_bytes=7417540996 uncovered_volumes=2
|
||||
Sep 23 07:02:53 demo-hp felhom-agent[1195783]: time=2026-09-23T07:02:53.924+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790139325067362818 archive=local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst
|
||||
Sep 23 10:29:29 demo-hp felhom-agent[1195783]: time=2026-09-23T10:29:29.358+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-22T04:12:45Z landed=2026-09-22T04:12:45Z reason="newest settled archive (landed 2026-09-22T04:12:45Z) has not been proven (last proven archive was a differen
|
||||
Sep 23 10:37:41 demo-hp felhom-agent[1195783]: time=2026-09-23T10:37:41.197+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-pbs:backup/ct/9201/2026-09-22T04:12:45Z duration_s=491.830022319
|
||||
Sep 23 16:29:29 demo-hp felhom-agent[1195783]: time=2026-09-23T16:29:29.364+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/vzdump-lxc-9201-2026_09_22-06_52_48.tar.zst landed=2026-09-22T04:52:48Z reason="newest settled archive (landed 2026-09-22T04:52:48Z) has not been proven (last proven archive was a dif
|
||||
Sep 23 16:35:15 demo-hp felhom-agent[1195783]: time=2026-09-23T16:35:15.142+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=local:backup/vzdump-lxc-9201-2026_09_22-06_52_48.tar.zst duration_s=345.775360191
|
||||
Sep 23 22:29:29 demo-hp felhom-agent[1195783]: time=2026-09-23T22:29:29.632+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-22T04:12:45Z) is already proven; local: newest settled archive (landed 2026-09-22T04:52:48Z) is already proven"
|
||||
Sep 24 04:29:29 demo-hp felhom-agent[1195783]: time=2026-09-24T04:29:29.641+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-22T04:12:45Z) is already proven; local: newest settled archive (landed 2026-09-22T04:52:48Z) is already proven"
|
||||
Sep 24 07:05:52 demo-hp felhom-agent[1195783]: time=2026-09-24T07:05:52.441+02:00 level=ERROR msg="local-api: backup job failed" vmid=9201 target=local job=backup-9201-1790225968108254908 err="backup: vzdump task vmid 9201: proxmox: task UPID:demo-hp:001C2729:11475EED:6AB4AE30:vzdump:9201:felhom-agent@pve!agent: failed: exitstatus \"job errors\""
|
||||
Sep 24 07:24:21 demo-hp felhom-agent[1195783]: time=2026-09-24T07:24:21.352+02:00 level=ERROR msg="local-api: backup job failed" vmid=9201 target=local job=backup-9201-1790227073555655065 err="backup: vzdump task vmid 9201: proxmox: task UPID:demo-hp:001D29F4:11490EBD:6AB4B281:vzdump:9201:felhom-agent@pve!agent: failed: exitstatus \"job errors\""
|
||||
Sep 24 07:49:16 demo-hp felhom-agent[1195783]: time=2026-09-24T07:49:16.829+02:00 level=ERROR msg="local-api: backup job failed" vmid=9201 target=local job=backup-9201-1790228574136769346 err="backup: vzdump task vmid 9201: proxmox: task UPID:demo-hp:001E7D73:114B58E7:6AB4B85E:vzdump:9201:felhom-agent@pve!agent: failed: exitstatus \"job errors\""
|
||||
Sep 24 08:22:57 demo-hp felhom-agent[1195783]: time=2026-09-24T08:22:57.653+02:00 level=ERROR msg="local-api: backup job failed" vmid=9201 target=local job=backup-9201-1790230975578595274 err="backup: vzdump task vmid 9201: proxmox: task UPID:demo-hp:0020B26F:114F02F7:6AB4C1BF:vzdump:9201:felhom-agent@pve!agent: failed: exitstatus \"job errors\""
|
||||
Sep 24 10:29:29 demo-hp felhom-agent[1195783]: time=2026-09-24T10:29:29.327+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst landed=2026-09-23T04:55:25Z reason="newest settled archive (landed 2026-09-23T04:55:25Z) has not been proven (last proven archive was a dif
|
||||
Sep 24 10:36:30 demo-hp felhom-agent[1195783]: time=2026-09-24T10:36:30.331+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst err="reconcile: restore-test start task: proxmox: task UPID:demo-hp:002789C4:115B3A8C:6AB4E106:vzstart:990000:felhom-agent@pve!agent: failed: exitstatus \"startup for container '990000' failed\""
|
||||
Sep 24 12:47:34 demo-hp felhom-agent[1195783]: time=2026-09-24T12:47:34.646+02:00 level=WARN msg="local-api: could not read the backup storage for the due-check — falling back to the in-memory record" vmid=9201 target=felhom-pbs err="proxmox: GET /nodes/demo-hp/storage/felhom-pbs/content -> HTTP 500: {\"data\":null,\"message\":\"felhom-pbs: error fetching datastores - 500 Can't connect to 10.77.
|
||||
Sep 24 13:10:04 demo-hp felhom-agent[3134463]: time=2026-09-24T13:10:04.268+02:00 level=INFO msg="stale-lock: removed dangling vzdump snapshot" vmid=9201
|
||||
Sep 24 22:06:23 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:23.065+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst size_bytes=8182759056 uncovered_volumes=2
|
||||
Sep 24 22:06:23 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:23.065+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790279965190292701 archive=local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst
|
||||
Sep 24 22:06:26 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:26.321+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1790280385256228451
|
||||
Sep 24 22:06:45 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:45.066+02:00 level=WARN msg="backup: ignoring an INCOMPLETE archive when computing tier freshness — it is not a successful backup" target=felhom-pbs vmid=9201 volid=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z size_bytes=1 reason="size 1 B is below the 1048576 B plausibility floor — an aborted/incomplete archive, not a succe
|
||||
Sep 24 22:07:45 demo-hp felhom-agent[756657]: time=2026-09-24T22:07:45.839+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=backup:felhom-pbs
|
||||
Sep 24 22:10:57 demo-hp felhom-agent[756657]: time=2026-09-24T22:10:57.339+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z size_bytes=23236593106 uncovered_volumes=2
|
||||
Sep 24 22:10:57 demo-hp felhom-agent[756657]: time=2026-09-24T22:10:57.339+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1790280385256228451 archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z
|
||||
Sep 25 22:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T22:57:53.343+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 25 22:57:54 demo-hp felhom-agent[995644]: time=2026-09-25T22:57:54.955+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=23108536763 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 25 22:57:54 demo-hp felhom-agent[995644]: time=2026-09-25T22:57:54.955+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.5 GiB"
|
||||
Sep 26 04:38:34 demo-hp felhom-agent[995644]: time=2026-09-26T04:38:34.211+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_26-04_30_46.tar.zst size_bytes=8384567204 uncovered_volumes=2
|
||||
Sep 26 04:38:34 demo-hp felhom-agent[995644]: time=2026-09-26T04:38:34.211+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790389845916353490 archive=local:backup/vzdump-lxc-9201-2026_09_26-04_30_46.tar.zst
|
||||
Sep 26 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-26T04:57:53.350+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 26 04:57:55 demo-hp felhom-agent[995644]: time=2026-09-26T04:57:55.007+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620981 avail_bytes=23079614940 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 26 04:57:55 demo-hp felhom-agent[995644]: time=2026-09-26T04:57:55.007+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.5 GiB"
|
||||
Sep 26 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-26T10:57:53.346+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 26 10:57:54 demo-hp felhom-agent[995644]: time=2026-09-26T10:57:54.951+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=23068046210 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 26 10:57:54 demo-hp felhom-agent[995644]: time=2026-09-26T10:57:54.951+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.5 GiB"
|
||||
Sep 26 16:57:53 demo-hp felhom-agent[995644]: time=2026-09-26T16:57:53.359+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 26 16:57:55 demo-hp felhom-agent[995644]: time=2026-09-26T16:57:55.016+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=23056477481 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 26 16:57:55 demo-hp felhom-agent[995644]: time=2026-09-26T16:57:55.017+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.5 GiB"
|
||||
Sep 26 22:57:53 demo-hp felhom-agent[995644]: time=2026-09-26T22:57:53.370+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 26 22:57:55 demo-hp felhom-agent[995644]: time=2026-09-26T22:57:55.013+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=23044908752 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 26 22:57:55 demo-hp felhom-agent[995644]: time=2026-09-26T22:57:55.013+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.5 GiB"
|
||||
Sep 27 04:43:29 demo-hp felhom-agent[995644]: time=2026-09-27T04:43:29.138+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst size_bytes=8380885161 uncovered_volumes=2
|
||||
Sep 27 04:43:29 demo-hp felhom-agent[995644]: time=2026-09-27T04:43:29.138+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790476546353426458 archive=local:backup/vzdump-lxc-9201-2026_09_27-04_35_47.tar.zst
|
||||
Sep 27 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-27T04:57:53.352+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 27 04:57:54 demo-hp felhom-agent[995644]: time=2026-09-27T04:57:54.985+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620981 avail_bytes=22663140685 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 27 04:57:54 demo-hp felhom-agent[995644]: time=2026-09-27T04:57:54.985+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 27 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-27T10:57:53.351+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different
|
||||
Sep 27 10:57:54 demo-hp felhom-agent[995644]: time=2026-09-27T10:57:54.968+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620981 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has
|
||||
Sep 27 10:57:54 demo-hp felhom-agent[995644]: time=2026-09-27T10:57:54.968+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.853+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a differen
|
||||
Sep 27 20:13:19 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:19.466+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620981 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, ha
|
||||
Sep 27 20:13:19 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:19.467+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 28 02:13:17 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:17.870+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a differen
|
||||
Sep 28 02:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:19.630+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, ha
|
||||
Sep 28 02:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:19.630+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 28 04:44:58 demo-hp felhom-agent[1370926]: time=2026-09-28T04:44:58.683+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_28-04_37_06.tar.zst size_bytes=8825181955 uncovered_volumes=2
|
||||
Sep 28 04:44:58 demo-hp felhom-agent[1370926]: time=2026-09-28T04:44:58.683+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790563025947283530 archive=local:backup/vzdump-lxc-9201-2026_09_28-04_37_06.tar.zst
|
||||
Sep 28 08:13:17 demo-hp felhom-agent[1370926]: time=2026-09-28T08:13:17.866+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a differen
|
||||
Sep 28 08:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T08:13:19.655+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=21801270353 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, ha
|
||||
Sep 28 08:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T08:13:19.655+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 20.3 GiB"
|
||||
Sep 28 14:13:17 demo-hp felhom-agent[1370926]: time=2026-09-28T14:13:17.848+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a differen
|
||||
Sep 28 14:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T14:13:19.532+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=25653657207 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, ha
|
||||
Sep 28 14:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T14:13:19.533+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 23.9 GiB"
|
||||
Sep 28 22:25:23 demo-hp felhom-agent[2355282]: time=2026-09-28T22:25:23.861+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a differen
|
||||
Sep 28 22:33:59 demo-hp felhom-agent[2355282]: time=2026-09-28T22:33:59.872+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z duration_s=516.008133181
|
||||
Sep 29 04:25:24 demo-hp felhom-agent[2355282]: time=2026-09-29T04:25:24.198+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven"
|
||||
Sep 29 04:50:53 demo-hp felhom-agent[2355282]: time=2026-09-29T04:50:53.815+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst size_bytes=9516315002 uncovered_volumes=2
|
||||
Sep 29 04:50:53 demo-hp felhom-agent[2355282]: time=2026-09-29T04:50:53.815+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790649730829915100 archive=local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst
|
||||
Sep 29 10:25:24 demo-hp felhom-agent[2355282]: time=2026-09-29T10:25:24.184+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven"
|
||||
Sep 29 16:25:24 demo-hp felhom-agent[2355282]: time=2026-09-29T16:25:24.219+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven"
|
||||
Sep 29 22:25:24 demo-hp felhom-agent[2355282]: time=2026-09-29T22:25:24.194+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven"
|
||||
Sep 30 04:25:24 demo-hp felhom-agent[2355282]: time=2026-09-30T04:25:24.194+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven"
|
||||
Sep 30 10:25:23 demo-hp felhom-agent[2355282]: time=2026-09-30T10:25:23.855+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst landed=2026-09-29T02:42:11Z reason="newest settled archive (landed 2026-09-29T02:42:11Z) has not been proven (last proven archive was a dif
|
||||
Sep 30 10:27:48 demo-hp felhom-agent[2355282]: time=2026-09-30T10:27:48.373+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=local:backup/vzdump-lxc-9201-2026_09_29-04_42_11.tar.zst duration_s=144.515446972
|
||||
Sep 30 13:47:12 demo-hp felhom-agent[2135127]: time=2026-09-30T13:47:12.666+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_30-13_37_38.tar.zst size_bytes=9978959837 uncovered_volumes=2
|
||||
Sep 30 13:47:12 demo-hp felhom-agent[2135127]: time=2026-09-30T13:47:12.666+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790768258239656621 archive=local:backup/vzdump-lxc-9201-2026_09_30-13_37_38.tar.zst
|
||||
Sep 30 16:40:32 demo-hp felhom-agent[2135127]: time=2026-09-30T16:40:32.449+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Sep 30 22:40:32 demo-hp felhom-agent[2135127]: time=2026-09-30T22:40:32.547+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 01 04:40:32 demo-hp felhom-agent[2135127]: time=2026-10-01T04:40:32.442+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 01 08:11:08 demo-hp felhom-agent[2135127]: time=2026-10-01T08:11:08.130+02:00 level=WARN msg="local-api: storage view unavailable for the backup-target presence check — assuming present" target=felhom-pbs err="storage: NodeStorage: proxmox: GET /nodes/demo-hp/storage: Get \"https://127.0.0.1:8006/api2/json/nodes/demo-hp/storage\": context canceled"
|
||||
Oct 01 08:11:08 demo-hp felhom-agent[2135127]: time=2026-10-01T08:11:08.130+02:00 level=WARN msg="local-api: could not read the backup storage for the due-check — falling back to the in-memory record" vmid=9201 target=felhom-pbs err="proxmox: GET /nodes/demo-hp/storage/felhom-pbs/content: Get \"https://127.0.0.1:8006/api2/json/nodes/demo-hp/storage/felhom-pbs/content\": context canceled"
|
||||
Oct 01 10:40:32 demo-hp felhom-agent[2135127]: time=2026-10-01T10:40:32.422+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 01 16:40:32 demo-hp felhom-agent[2135127]: time=2026-10-01T16:40:32.078+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/vzdump-lxc-9201-2026_09_30-13_37_38.tar.zst landed=2026-09-30T11:37:38Z reason="newest settled archive (landed 2026-09-30T11:37:38Z) has not been proven (last proven archive was a dif
|
||||
Oct 01 16:43:11 demo-hp felhom-agent[2135127]: time=2026-10-01T16:43:11.704+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=local:backup/vzdump-lxc-9201-2026_09_30-13_37_38.tar.zst duration_s=159.623154174
|
||||
Oct 01 22:15:25 demo-hp felhom-agent[2135127]: time=2026-10-01T22:15:25.563+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_01-22_10_48.tar.zst size_bytes=4681052311 uncovered_volumes=2
|
||||
Oct 01 22:15:25 demo-hp felhom-agent[2135127]: time=2026-10-01T22:15:25.563+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790885448142698430 archive=local:backup/vzdump-lxc-9201-2026_10_01-22_10_48.tar.zst
|
||||
Oct 01 22:15:30 demo-hp felhom-agent[2135127]: time=2026-10-01T22:15:30.441+02:00 level=INFO msg="local-api: backup reached snapshotted (app may resume)" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1790885728244905026
|
||||
Oct 01 22:18:31 demo-hp felhom-agent[2135127]: time=2026-10-01T22:18:31.284+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-10-01T20:15:29Z size_bytes=15223817181 uncovered_volumes=2
|
||||
Oct 01 22:18:31 demo-hp felhom-agent[2135127]: time=2026-10-01T22:18:31.284+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1790885728244905026 archive=felhom-pbs:backup/ct/9201/2026-10-01T20:15:29Z
|
||||
Oct 01 22:40:32 demo-hp felhom-agent[2135127]: time=2026-10-01T22:40:32.400+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 02 04:40:32 demo-hp felhom-agent[2135127]: time=2026-10-02T04:40:32.387+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 02 10:40:32 demo-hp felhom-agent[2135127]: time=2026-10-02T10:40:32.460+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 02 16:40:32 demo-hp felhom-agent[2135127]: time=2026-10-02T16:40:32.406+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-24T20:06:25Z) is already proven; local: no settled archive yet — nothing to prove (newborn or still settling)"
|
||||
Oct 02 22:40:32 demo-hp felhom-agent[2135127]: time=2026-10-02T22:40:32.065+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-10-01T20:15:29Z landed=2026-10-01T20:15:29Z reason="newest settled archive (landed 2026-10-01T20:15:29Z) has not been proven (last proven archive was a differen
|
||||
Oct 02 22:46:02 demo-hp felhom-agent[2135127]: time=2026-10-02T22:46:02.535+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-pbs:backup/ct/9201/2026-10-01T20:15:29Z duration_s=330.466115271
|
||||
Oct 03 04:38:20 demo-hp felhom-agent[2135127]: time=2026-10-03T04:38:20.266+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_03-04_33_33.tar.zst size_bytes=4674947408 uncovered_volumes=2
|
||||
Oct 03 04:38:20 demo-hp felhom-agent[2135127]: time=2026-10-03T04:38:20.266+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790994812809499660 archive=local:backup/vzdump-lxc-9201-2026_10_03-04_33_33.tar.zst
|
||||
Oct 03 04:40:32 demo-hp felhom-agent[2135127]: time=2026-10-03T04:40:32.407+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 03 10:40:32 demo-hp felhom-agent[2135127]: time=2026-10-03T10:40:32.425+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 03 16:40:32 demo-hp felhom-agent[2135127]: time=2026-10-03T16:40:32.455+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 03 22:40:32 demo-hp felhom-agent[2135127]: time=2026-10-03T22:40:32.411+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 04 04:39:41 demo-hp felhom-agent[2135127]: time=2026-10-04T04:39:41.809+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst size_bytes=4679944932 uncovered_volumes=2
|
||||
Oct 04 04:39:41 demo-hp felhom-agent[2135127]: time=2026-10-04T04:39:41.809+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1791081294394153713 archive=local:backup/vzdump-lxc-9201-2026_10_04-04_34_55.tar.zst
|
||||
Oct 04 04:40:32 demo-hp felhom-agent[2135127]: time=2026-10-04T04:40:32.412+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 05 01:49:22 demo-hp felhom-agent[447632]: time=2026-10-05T01:49:22.786+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 05 04:40:16 demo-hp felhom-agent[447632]: time=2026-10-05T04:40:16.647+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_05-04_35_24.tar.zst size_bytes=4882309520 uncovered_volumes=2
|
||||
Oct 05 04:40:16 demo-hp felhom-agent[447632]: time=2026-10-05T04:40:16.647+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1791167723831350602 archive=local:backup/vzdump-lxc-9201-2026_10_05-04_35_24.tar.zst
|
||||
Oct 05 06:25:36 demo-hp felhom-agent[2031788]: time=2026-10-05T06:25:36.366+02:00 level=WARN msg="local-api: could not read the backup storage for the due-check — falling back to the in-memory record" vmid=9201 target=felhom-pbs err="proxmox: GET /nodes/demo-hp/storage/felhom-pbs/content -> HTTP 500: {\"data\":null,\"message\":\"felhom-pbs: error fetching datastores - 500 Can't connect to 10.77.
|
||||
Oct 05 06:26:23 demo-hp felhom-agent[2031788]: time=2026-10-05T06:26:23.344+02:00 level=ERROR msg="local-api: backup job failed" vmid=9201 target=felhom-pbs job=backup-9201-felhom-pbs-1791174359051131871 err="backup: vzdump task vmid 9201: proxmox: task UPID:demo-hp:0022DC8C:00483400:6AC326DF:vzdump:9201:felhom-agent@pve!agent: failed: exitstatus \"could not activate storage 'felhom-pbs': felhom-p
|
||||
Oct 05 06:30:36 demo-hp felhom-agent[2031788]: time=2026-10-05T06:30:36.395+02:00 level=WARN msg="local-api: could not read the backup storage for the due-check — falling back to the in-memory record" vmid=9201 target=felhom-pbs err="proxmox: GET /nodes/demo-hp/storage/felhom-pbs/content -> HTTP 500: {\"message\":\"felhom-pbs: error fetching datastores - 500 Can't connect to 10.77.0.1:8007 (Conn
|
||||
Oct 05 06:35:36 demo-hp felhom-agent[2031788]: time=2026-10-05T06:35:36.377+02:00 level=WARN msg="local-api: could not read the backup storage for the due-check — falling back to the in-memory record" vmid=9201 target=felhom-pbs err="proxmox: GET /nodes/demo-hp/storage/felhom-pbs/content -> HTTP 500: {\"data\":null,\"message\":\"felhom-pbs: error fetching datastores - 500 Can't connect to 10.77.
|
||||
Oct 05 10:26:59 demo-hp felhom-agent[1323]: time=2026-10-05T10:26:59.170+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 05 11:24:01 demo-hp felhom-agent[1323]: time=2026-10-05T11:24:01.884+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_10_05-11_19_30.tar.zst size_bytes=4880501245 uncovered_volumes=2
|
||||
Oct 05 11:24:01 demo-hp felhom-agent[1323]: time=2026-10-05T11:24:01.884+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1791191969558324187 archive=local:backup/vzdump-lxc-9201-2026_10_05-11_19_30.tar.zst
|
||||
Oct 05 11:24:09 demo-hp felhom-agent[1323]: time=2026-10-05T11:24:09.634+02:00 level=INFO msg="local-api: backup refused — a heavy operation is already in flight" vmid=9201 requested_target=felhom-pbs busy=backup:local
|
||||
Oct 05 12:57:08 demo-hp felhom-agent[453394]: time=2026-10-05T12:57:08.743+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 05 18:57:09 demo-hp felhom-agent[453394]: time=2026-10-05T18:57:09.755+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 05 19:42:24 demo-hp felhom-agent[1625894]: time=2026-10-05T19:42:24.219+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 06 01:12:37 demo-hp felhom-agent[2556733]: time=2026-10-06T01:12:37.224+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 06 07:12:38 demo-hp felhom-agent[2556733]: time=2026-10-06T07:12:38.110+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="local: no settled archive yet — nothing to prove (newborn or still settling); felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven"
|
||||
Oct 06 12:21:17 demo-hp felhom-agent[297545]: time=2026-10-06T12:21:17.652+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/vzdump-lxc-9201-2026_10_05-11_19_30.tar.zst landed=2026-10-05T09:19:30Z reason="newest settled archive (landed 2026-10-05T09:19:30Z) has not been proven (last proven archive was a diff
|
||||
Oct 06 12:22:46 demo-hp felhom-agent[297545]: time=2026-10-06T12:22:46.987+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=local:backup/vzdump-lxc-9201-2026_10_05-11_19_30.tar.zst duration_s=89.331269652
|
||||
Oct 06 18:22:47 demo-hp felhom-agent[297545]: time=2026-10-06T18:22:47.786+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-10-01T20:15:29Z) is already proven; local: newest settled archive (landed 2026-10-05T09:19:30Z) is already proven"
|
||||
@@ -0,0 +1,14 @@
|
||||
Tester-2 ct 9201 2026-10-04T16:31:13Z 2096481498 ok
|
||||
<string>:5: DeprecationWarning: datetime.datetime.utcfromtimestamp() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.fromtimestamp(timestamp, datetime.UTC).
|
||||
demo-felhom ct 9201 2026-09-22T04:12:20Z 5840921438 ok
|
||||
demo-felhom ct 9201 2026-09-29T04:16:43Z 6956268440 ok
|
||||
demo-felhom ct 9201 2026-10-06T04:21:09Z 2749038678 ok
|
||||
<string>:5: DeprecationWarning: datetime.datetime.utcfromtimestamp() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.fromtimestamp(timestamp, datetime.UTC).
|
||||
<string>:5: DeprecationWarning: datetime.datetime.utcfromtimestamp() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.fromtimestamp(timestamp, datetime.UTC).
|
||||
demo-hp ct 9201 2026-09-24T20:06:25Z 23236593218 ok
|
||||
demo-hp ct 9201 2026-10-01T20:15:29Z 15223817295 ok
|
||||
operator host dooplex-hub 2026-10-05T14:23:25Z 369808250 ok
|
||||
operator host dooplex-hub 2026-10-06T00:32:21Z 372921211 ok
|
||||
<string>:5: DeprecationWarning: datetime.datetime.utcfromtimestamp() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.fromtimestamp(timestamp, datetime.UTC).
|
||||
tester-1 ct 9201 2026-10-04T19:57:18Z 4913374015 ok
|
||||
<string>:5: DeprecationWarning: datetime.datetime.utcfromtimestamp() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.fromtimestamp(timestamp, datetime.UTC).
|
||||
@@ -0,0 +1,48 @@
|
||||
events columns: ['id', 'customer_id', 'event_type', 'severity', 'message', 'details_json', 'source', 'created_at']
|
||||
('2026-10-05 04:27:31', 'whole_guest_backup_failed', 'error', 'Whole-guest backup FAILED on the felhom-pbs tier — retrying with backoff (next attempt in 15m0s)')
|
||||
('2026-10-05 05:16:22', 'controller_updated', 'info', 'Controller frissítve: 0.293.0 → 0.294.0')
|
||||
('2026-10-05 05:16:27', 'controller_started', 'info', 'Controller elindult (0.294.0)')
|
||||
('2026-10-05 06:09:05', 'os_update_applied', 'info', 'System security fixes installed on the box (13 package(s)).')
|
||||
('2026-10-05 06:11:02', 'os_update_applied', 'info', 'System security fixes installed on the box (13 package(s)).')
|
||||
('2026-10-05 06:16:39', 'controller_started', 'info', 'Controller elindult (0.294.0)')
|
||||
('2026-10-05 06:18:03', 'os_update_failed', 'error', 'OS update (guest) on demo-hp-bb76ea failed: {"dpkg_audit":"clean","rc":100,"tail":["E: dpkg was interrupted, you must manually run \'sudo dpkg --configure -a\' to correct the problem."]} no health readi')
|
||||
('2026-10-05 06:19:35', 'os_update_applied', 'info', 'System security fixes installed on the box (12 package(s)).')
|
||||
('2026-10-05 06:29:57', 'host_restarted_after_crash', 'info', 'Your box restarted by itself after an unexpected stop. You do not need to do anything.')
|
||||
('2026-10-05 06:29:57', 'host_crash_restart', 'warning', 'demo-hp-bb76ea restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T06:14:33Z — 4 such restart(s) in 24 h. kernel.panic now 10.')
|
||||
('2026-10-05 07:29:43', 'controller_updated', 'info', 'Controller frissítve: 0.294.0 → 0.295.0')
|
||||
('2026-10-05 07:29:48', 'controller_started', 'info', 'Controller elindult (0.295.0)')
|
||||
('2026-10-05 07:58:50', 'controller_started', 'info', 'Controller elindult (0.295.0)')
|
||||
('2026-10-05 08:00:42', 'os_update_applied', 'info', 'System security fixes installed on the box (12 package(s)).')
|
||||
('2026-10-05 08:12:06', 'host_restarted_after_crash', 'info', 'Your box restarted by itself after an unexpected stop. You do not need to do anything.')
|
||||
('2026-10-05 08:12:06', 'host_crash_restart', 'warning', 'demo-hp-bb76ea restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:56:41Z — 5 such restart(s) in 24 h. kernel.panic now 10.')
|
||||
('2026-10-05 09:16:15', 'recovery_credential_revealed', 'info', 'Az üzemeltető lekérte a géped konzolos hozzáférési jelszavát (távoli hibaelhárítás).')
|
||||
('2026-10-05 10:25:08', 'controller_updated', 'info', 'Controller frissítve: 0.295.0 → 0.296.0')
|
||||
('2026-10-05 10:25:13', 'controller_started', 'info', 'Controller elindult (0.296.0)')
|
||||
('2026-10-05 10:28:08', 'agent_capability_degraded', 'warning', 'Host demo-hp-bb76ea: agent privileged capability degraded — wg-conf-install (impairs: wg-felhom conf install (root content check, R-861))')
|
||||
('2026-10-05 11:13:08', 'agent_capability_recovered', 'info', 'Host demo-hp-bb76ea: agent privileged capabilities recovered (all required grants restored)')
|
||||
== notification_log demo-hp 2026-10-05:
|
||||
['id', 'customer_id', 'event_type', 'severity', 'message', 'status', 'error_message', 'created_at', 'channel']
|
||||
(1084, 'demo-hp', 'whole_guest_backup_failed', 'error', 'Whole-guest backup FAILED on the felhom-pbs tier — retrying with backoff (next attempt in 15m0s)', 'failed', 'sending request: Post "https://api.resend.com/emails": context deadline exceeded (Client.Timeout exceeded while awaiting headers)', '
|
||||
(1085, 'demo-hp', 'whole_guest_backup_failed', 'error', 'Whole-guest backup FAILED on the felhom-pbs tier — retrying with backoff (next attempt in 15m0s)', 'skipped', 'operator_only', '2026-10-05 04:27:41', 'customer')
|
||||
(1086, 'demo-hp', 'os_update_failed', 'error', 'OS update (guest) on demo-hp-bb76ea failed: {"dpkg_audit":"clean","rc":100,"tail":["E: dpkg was interrupted, you must manually run \'sudo dpkg --configure -a\' to correct the problem."]} no health reading', 'sent', '', '2026-10-05 06:18:03', 'operator'
|
||||
(1087, 'demo-hp', 'os_update_failed', 'error', 'OS update (guest) on demo-hp-bb76ea failed: {"dpkg_audit":"clean","rc":100,"tail":["E: dpkg was interrupted, you must manually run \'sudo dpkg --configure -a\' to correct the problem."]} no health reading', 'skipped', 'operator_only', '2026-10-05 06:18
|
||||
(1088, 'demo-hp', 'host_crash_restart', 'warning', 'demo-hp-bb76ea restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T06:14:33Z — 4 such restart(s) in 24 h. kernel.panic now 10.', 'sent', '', '2026-10-05 06:29:57', 'operator')
|
||||
(1089, 'demo-hp', 'host_crash_restart', 'warning', 'demo-hp-bb76ea restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T06:14:33Z — 4 such restart(s) in 24 h. kernel.panic now 10.', 'skipped', 'operator_only', '2026-10-05 06:29:57', 'customer')
|
||||
(1090, 'tester-1', 'node_stale', 'warning', 'No report received for 45m', 'sent', '', '2026-10-05 06:46:41', 'operator')
|
||||
(1091, 'tester-1', 'host_stale', 'warning', 'Host tester-1-d70be4: no report for 45m', 'sent', '', '2026-10-05 06:56:41', 'operator')
|
||||
(1092, 'tester-1', 'node_down', 'error', 'No report received for 1h30m', 'sent', '', '2026-10-05 07:31:41', 'operator')
|
||||
(1093, 'tester-1', 'host_recovered', 'info', 'Host tester-1-d70be4: reports resumed (was stale for 35m)', 'sent', '', '2026-10-05 07:31:42', 'operator')
|
||||
(1094, 'tester-1', 'host_crash_restart', 'warning', 'tester-1-d70be4 restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:31:20Z — 1 such restart(s) in 24 h. kernel.panic now 10.', 'sent', '', '2026-10-05 07:33:12', 'operator')
|
||||
(1095, 'tester-1', 'host_crash_restart', 'warning', 'tester-1-d70be4 restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:31:20Z — 1 such restart(s) in 24 h. kernel.panic now 10.', 'skipped', 'operator_only', '2026-10-05 07:33:12', 'customer')
|
||||
(1096, 'tester-1', 'node_recovered', 'info', 'Reports resumed (was down for 46m)', 'sent', '', '2026-10-05 07:33:41', 'operator')
|
||||
(1097, 'Tester-2', 'host_down', 'error', 'Host Tester-2-be8404: no report for 13h45m', 'sent', '', '2026-10-05 07:51:34', 'operator')
|
||||
(1098, 'demo-hp', 'host_crash_restart', 'warning', 'demo-hp-bb76ea restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:56:41Z — 5 such restart(s) in 24 h. kernel.panic now 10.', 'sent', '', '2026-10-05 08:12:06', 'operator')
|
||||
(1099, 'demo-hp', 'host_crash_restart', 'warning', 'demo-hp-bb76ea restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:56:41Z — 5 such restart(s) in 24 h. kernel.panic now 10.', 'skipped', 'operator_only', '2026-10-05 08:12:06', 'customer')
|
||||
(1100, 'tester-1', 'host_crash_guard_tripped', 'error', 'The crash guard of tester-1-d70be4 TRIPPED at 2026-10-05T07:57:17Z: 2 unclean boots within 60 minutes — the next crash leaves the box off (limit 3). The next crash leaves the box OFF until someone switches it on. Re-arm: `felhom-crash-guard re
|
||||
(1101, 'tester-1', 'host_crash_guard_tripped', 'error', 'The crash guard of tester-1-d70be4 TRIPPED at 2026-10-05T07:57:17Z: 2 unclean boots within 60 minutes — the next crash leaves the box off (limit 3). The next crash leaves the box OFF until someone switches it on. Re-arm: `felhom-crash-guard re
|
||||
(1102, 'tester-1', 'host_crash_restart', 'warning', 'tester-1-d70be4 restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:57:12Z — 2 such restart(s) in 24 h. kernel.panic now 0.', 'sent', '', '2026-10-05 08:12:32', 'operator')
|
||||
(1103, 'tester-1', 'host_crash_restart', 'warning', 'tester-1-d70be4 restarted by itself after an unexpected stop (a kernel crash, a power cut or a hard reset) at 2026-10-05T07:57:12Z — 2 such restart(s) in 24 h. kernel.panic now 0.', 'skipped', 'operator_only', '2026-10-05 08:12:32', 'customer')
|
||||
(1104, 'Tester-2', 'host_down', 'error', 'Host Tester-2-be8404: no report for 15h10m', 'sent', '', '2026-10-05 09:16:10', 'operator')
|
||||
(1105, 'demo-hp', 'agent_capability_degraded', 'warning', 'Host demo-hp-bb76ea: agent privileged capability degraded — wg-conf-install (impairs: wg-felhom conf install (root content check, R-861))', 'sent', '', '2026-10-05 10:28:08', 'operator')
|
||||
(1106, 'tester-1', 'agent_capability_degraded', 'warning', 'Host tester-1-d70be4: agent privileged capability degraded — wg-conf-install (impairs: wg-felhom conf install (root content check, R-861))', 'sent', '', '2026-10-05 10:28:09', 'operator')
|
||||
(1107, 'demo-felhom', 'agent_capability_degraded', 'warning', 'Host demo-felhom-8363b5: agent privileged capability degraded — wg-conf-install (impairs: wg-felhom conf install (root content check, R-861))', 'sent', '', '2026-10-05 10:39:09', 'operator')
|
||||
@@ -0,0 +1,6 @@
|
||||
all notification rows: (1112, '2026-02-16 18:46:11', '2026-10-06 09:56:51')
|
||||
by channel/status: [('customer', 'sent', 155), ('customer', 'skipped', 166), ('operator', 'failed', 1), ('operator', 'recorded', 22), ('operator', 'refused', 7), ('operator', 'sent', 691), ('operator', 'suppressed', 70)]
|
||||
failed OPERATOR mails, by severity: [('error', 1)]
|
||||
failed operator mails, error/critical, newest 15:
|
||||
('2026-10-05 04:27:41', 'demo-hp', 'whole_guest_backup_failed', 'error', 'sending request: Post "https://api.resend.com/emails": context deadline exceeded (Client.Timeout exceeded whil')
|
||||
failed CUSTOMER mails: (0,)
|
||||
@@ -0,0 +1,27 @@
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.077s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.092s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/api 5.785s
|
||||
? gitea.dooplex.hu/admin/felhom-hub/internal/assets [no test files]
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/claim 3.336s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/cloudflare 0.008s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/configgen 0.031s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.154s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/gitea 0.013s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/hetznerapi 0.053s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/i18n 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/intent 0.092s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/mailrelay 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/monitor 38.751s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/notify 1.624s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/offsite 0.761s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/offsiteheal 0.269s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/offsitekeys 0.245s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 1.325s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/pbsdrheal 0.622s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/poke 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/semver 0.007s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/store 4.224s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/sysfacts 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/tenantsync 0.042s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 25.654s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/wgsync 0.237s
|
||||
@@ -0,0 +1,12 @@
|
||||
$ go test ./internal/notify/ -run TestOperatorMailRetry -v (processOperator without the retryOperatorEmail call)
|
||||
=== RUN TestOperatorMailRetry_FailedOperatorMailIsRetried
|
||||
r895_operator_retry_test.go:34: the failed mail was never sent again (tries=1)
|
||||
--- FAIL: TestOperatorMailRetry_FailedOperatorMailIsRetried (0.03s)
|
||||
=== RUN TestOperatorMailRetry_RetriesAreBounded
|
||||
r895_operator_retry_test.go:69: tries=1, want 4
|
||||
--- FAIL: TestOperatorMailRetry_RetriesAreBounded (0.03s)
|
||||
=== RUN TestOperatorMailRetry_SuccessIsNotRetried
|
||||
--- PASS: TestOperatorMailRetry_SuccessIsNotRetried (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/notify 0.104s
|
||||
FAIL
|
||||
@@ -0,0 +1,7 @@
|
||||
host_reports row for tester-1-d70be4 at 2026-10-06 16:51:29
|
||||
host.node = felhom
|
||||
guest_net.guests.ip = 192.168.0.101
|
||||
customer tester-1 domain: []
|
||||
VM 341 (demo-hp) PVE certificate: CN=felhom.enkicsifelhom.hu, SAN IP 192.168.0.154, DNS felhom
|
||||
felhom.enkicsifelhom.hu -> 200
|
||||
felhom.enkisfelhom.hu -> 404
|
||||
@@ -0,0 +1,10 @@
|
||||
Tue Oct 6 17:12:17 UTC 2026
|
||||
containers before: 21
|
||||
image: gitea.dooplex.hu/admin/felhom-controller:0.301.0
|
||||
Killed
|
||||
run rc=137
|
||||
inspect: true 137
|
||||
oom events:
|
||||
removed felhom-oomcheck-byhand-3226729
|
||||
left: 0; containers after: 21
|
||||
engine 29.8.2
|
||||
@@ -0,0 +1,7 @@
|
||||
run rc=137
|
||||
inspect: true 137 id=0f65ef3dde661a4f19
|
||||
A container=NAME event=oom, until T1+1: [oom felhom-oomcheck-byhand2-3227326 ]
|
||||
B event=oom only: [oom felhom-oomcheck-byhand2-3227326 ]
|
||||
C every event of the container: [create attach start oom die ]
|
||||
removed
|
||||
left: 0
|
||||
@@ -0,0 +1,63 @@
|
||||
== R-528 agent tests, clone /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528 branch r528-oom-check base de812bc + working tree, 2026-10-06T17:13:46Z
|
||||
$ python3 -W ignore configs/test_felhom_os_apply.py
|
||||
----------------------------------------------------------------------
|
||||
Ran 99 tests in 0.605s
|
||||
|
||||
OK
|
||||
rc=0
|
||||
$ python3 -W ignore configs/test_felhom_config_bundle.py
|
||||
----------------------------------------------------------------------
|
||||
Ran 55 tests in 1.085s
|
||||
|
||||
OK
|
||||
rc=0
|
||||
$ go build ./...
|
||||
rc=0
|
||||
$ go vet ./...
|
||||
rc=0
|
||||
$ go test ./...
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/cmd/felhom-agent 0.154s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/cmd/felhom-opsign 0.007s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/authz 0.018s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/backup 0.193s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/capability 0.178s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/config 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/desired 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/dr 0.009s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/escrow 54.459s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/fasttick 0.004s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/felhomsshd 0.243s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/fstrim 0.013s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/guesthook 0.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/guestnet 0.025s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/httpx 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/hub 1.150s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/lanresolver 0.149s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi 1.611s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/log 0.006s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/mgmtplane 0.010s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 2.266s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/pbs 0.263s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/pbsdr 0.066s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/poke 5.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/provision 0.013s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/proxmox 0.061s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/reconcile 0.278s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/restorespace 0.010s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/selfheal 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate 0.018s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs 0.015s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/storage 0.452s
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/wgtunnel 0.225s
|
||||
rc=0
|
||||
$ python3 scripts/agent_gates.py --fast
|
||||
==============================================================================
|
||||
== summary
|
||||
==============================================================================
|
||||
reuse-refs OK (exit 0)
|
||||
instructions OK (exit 0)
|
||||
release-complete OK (exit 0)
|
||||
observations OK (exit 0)
|
||||
|
||||
all agent gates OK
|
||||
rc=0
|
||||
@@ -0,0 +1,27 @@
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.067s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.074s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/api 5.720s
|
||||
? gitea.dooplex.hu/admin/felhom-hub/internal/assets [no test files]
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/claim 3.295s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/cloudflare 0.007s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/configgen 0.031s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/dbsnap 0.151s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/gitea 0.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/hetznerapi 0.054s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/i18n 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/intent 0.091s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/mailrelay 0.005s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/monitor 38.738s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/notify 1.684s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/offsite 0.820s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/offsiteheal 0.322s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/offsitekeys 0.255s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 1.476s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/pbsdrheal 0.670s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/poke 0.012s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/semver 0.004s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/store 4.240s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/sysfacts 0.007s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/tenantsync 0.044s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/web 26.271s
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/wgsync 0.238s
|
||||
@@ -0,0 +1,17 @@
|
||||
R-528 red-proof red-go-kept: the copy of oom_check removed from internal/osupdate/unsent.go
|
||||
mutation: `rep.OOMCheck = rawOrNil(wr.OOMCheck)` -> `_ = rawOrNil`
|
||||
$ go test ./internal/osupdate/ -run ^TestR868_KeptCopyCarriesTheOOMCheck$ -count=1
|
||||
rc=1
|
||||
2026/10/06 19:13:38 INFO osupdate: sending a kept report late (the agent was stopped mid-pass or the hub was away, R-868) run=20261007T020000Z layer=docker vmid=9201 ring=0 trigger=night path=/tmp/TestR868_KeptCopyCarriesTheOOMCheck2908158053/001/report-20261007T020000Z-docker-apply.json
|
||||
2026/10/06 19:13:38 INFO osupdate: DONE run=20261007T020000Z layer=docker vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="sent late — kept on the box until the hub could take it (R-868, R-875)" upgraded=1 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=0
|
||||
--- FAIL: TestR868_KeptCopyCarriesTheOOMCheck (0.00s)
|
||||
unsent_test.go:171: oom_check lost or changed: ""
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.008s
|
||||
FAIL
|
||||
|
||||
=== control: restored source ===
|
||||
$ go test ./internal/osupdate/ -run ^TestR868_KeptCopyCarriesTheOOMCheck$ -count=1
|
||||
rc=0
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.009s
|
||||
VERDICT: RED as expected, GREEN restored
|
||||
@@ -0,0 +1,26 @@
|
||||
R-528 red-proof red-go-runlayer: the copy of oom_check removed from internal/osupdate/leg.go
|
||||
mutation: `rep.OOMCheck = rawOrNil(wr.OOMCheck)` -> `_ = rawOrNil`
|
||||
$ go test ./internal/osupdate/ -run ^TestDocker_OOMCheckReachesTheHubUnchanged$ -count=1
|
||||
rc=1
|
||||
2026/10/06 19:13:37 INFO osupdate: START run=20261004T040000Z layer=guest vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261004T040000Z
|
||||
2026/10/06 19:13:37 INFO osupdate: wrapper line="os-apply: DONE rc=0"
|
||||
2026/10/06 19:13:37 INFO osupdate: DONE run=20261004T040000Z layer=guest vmid=9201 ring=0 trigger=night outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=0
|
||||
2026/10/06 19:13:37 INFO osupdate: START run=20261004T040000Z layer=host vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261004T040000Z
|
||||
2026/10/06 19:13:37 INFO osupdate: wrapper line="os-apply: DONE rc=0"
|
||||
2026/10/06 19:13:37 INFO osupdate: DONE run=20261004T040000Z layer=host vmid=9201 ring=0 trigger=night outcome=nothing healthy=true reason="" upgraded=0 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=0
|
||||
2026/10/06 19:13:37 INFO osupdate: wrapper line="os-apply: DONE rc=0"
|
||||
2026/10/06 19:13:37 INFO osupdate: live-restore vmid=9201 result=null
|
||||
2026/10/06 19:13:37 INFO osupdate: START run=20261004T040000Z layer=docker vmid=9201 ring=0 trigger=night enabled=true release=ring0-20261004T040000Z
|
||||
2026/10/06 19:13:37 INFO osupdate: wrapper line="os-apply: DONE rc=0"
|
||||
2026/10/06 19:13:37 INFO osupdate: DONE run=20261004T040000Z layer=docker vmid=9201 ring=0 trigger=night outcome=applied healthy=true reason="" upgraded=1 pending=0 not_covered=0 restart_needed=0 reboot_needed=false wrapper_seconds=0
|
||||
--- FAIL: TestDocker_OOMCheckReachesTheHubUnchanged (0.00s)
|
||||
leg_test.go:540: no oom_check in the docker report: {"run_id":"20261004T040000Z","layer":"docker","trigger":"night","mode":"apply","ring":0,"release_id":"ring0-20261004T040000Z","outcome":"applied","healthy":true,"vmid":9201,"upgraded":[{"name":"docker-ce","version":"5:29.8.2-1~debian.13~trixie","origin":""}],"docker_engine":"29.8.2","authority":"ring0"}
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.010s
|
||||
FAIL
|
||||
|
||||
=== control: restored source ===
|
||||
$ go test ./internal/osupdate/ -run ^TestDocker_OOMCheckReachesTheHubUnchanged$ -count=1
|
||||
rc=0
|
||||
ok gitea.dooplex.hu/admin/felhom-agent/internal/osupdate 0.009s
|
||||
VERDICT: RED as expected, GREEN restored
|
||||
@@ -0,0 +1,34 @@
|
||||
R-528 red-proof red-health-mode: the check also runs in docker health mode
|
||||
mutation: replace
|
||||
---
|
||||
self.report["health"] = self.health()
|
||||
return 0
|
||||
---
|
||||
with
|
||||
---
|
||||
self.report["health"] = self.health()
|
||||
self.report["oom_check"] = self.oom_check()
|
||||
return 0
|
||||
---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-health-mode-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_never_on_guest_or_host_or_in_health_mode
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_never_on_guest_or_host_or_in_health_mode (__main__.OOMCheck.test_never_on_guest_or_host_or_in_health_mode)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1301, in test_never_on_guest_or_host_or_in_health_mode
|
||||
self.assertNotIn("oom_check", r)
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^
|
||||
AssertionError: 'oom_check' unexpectedly found in {'health': {'containers': {'app': {'health': 'healthy', 'id': 'bbb222', 'state': 'running'}, 'felhom-controller': {'health': 'healthy', 'id': 'aaa111', 'state': 'running'}}, 'controller': 'healthy', 'controller_docker_ok': True, 'docker_ok': True, 'network_ok': True}, 'layer': 'docker', 'mode': 'health', 'oom_check': {'detail': 'the engine reported the memory kill: OOMKilled=true and the oom event', 'exit_code': 137, 'image': 'gitea.dooplex.hu/admin/felhom-controller:0.300.0', 'oom_event': True, 'oom_killed': True, 'result': 'pass'}, 'pass_seconds': 0.0, 'refused': None, 'release_id': 'os-docker-t1', 'vmid': 9201}
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.016s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_never_on_guest_or_host_or_in_health_mode
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,23 @@
|
||||
$ go test ./internal/osupdates/ -run TestDockerApproval -v -count=1 (oomCheckWaiting returns "")
|
||||
=== RUN TestDockerApproval_AFailedMemoryKillCheckBlocks
|
||||
[INFO] [store] app-update event types added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
[INFO] [store] app_stopped_unhealthy added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
docker_test.go:106: a set whose engine missed a memory kill was approvable: <nil>
|
||||
--- FAIL: TestDockerApproval_AFailedMemoryKillCheckBlocks (0.04s)
|
||||
=== RUN TestDockerApproval_AMissingMemoryKillCheckBlocks
|
||||
[INFO] [store] app-update event types added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
[INFO] [store] app_stopped_unhealthy added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
docker_test.go:121: a set with no memory-kill check was approvable: <nil>
|
||||
--- FAIL: TestDockerApproval_AMissingMemoryKillCheckBlocks (0.04s)
|
||||
=== RUN TestDockerApproval_AnErroredMemoryKillCheckBlocks
|
||||
[INFO] [store] app-update event types added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
[INFO] [store] app_stopped_unhealthy added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
docker_test.go:137: a set whose memory-kill check errored was approvable
|
||||
--- FAIL: TestDockerApproval_AnErroredMemoryKillCheckBlocks (0.04s)
|
||||
=== RUN TestDockerApproval_APassingCheckOnEveryBoxAllows
|
||||
[INFO] [store] app-update event types added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
[INFO] [store] app_stopped_unhealthy added to 0 household(s)' notification prefs (one-time, add-only): []
|
||||
--- PASS: TestDockerApproval_APassingCheckOnEveryBoxAllows (0.04s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.160s
|
||||
FAIL
|
||||
@@ -0,0 +1,35 @@
|
||||
R-528 red-proof red-image-error: an unreadable image does not stop the check (the early 'error' return removed)
|
||||
mutation: replace
|
||||
---
|
||||
log(f"os-apply: OOM-CHECK error — {res['detail']}")
|
||||
return res
|
||||
---
|
||||
with
|
||||
---
|
||||
log(f"os-apply: OOM-CHECK error — {res['detail']}")
|
||||
---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-image-error-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_image_unreadable_is_error
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_image_unreadable_is_error (__main__.OOMCheck.test_image_unreadable_is_error)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1272, in test_image_unreadable_is_error
|
||||
self.assertEqual(oc["result"], "error", oc)
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
AssertionError: 'pass' != 'error'
|
||||
- pass
|
||||
+ error
|
||||
: {'detail': 'the engine reported the memory kill: OOMKilled=true and the oom event', 'exit_code': 137, 'image': '', 'oom_event': True, 'oom_killed': True, 'result': 'pass'}
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.016s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_image_unreadable_is_error
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,32 @@
|
||||
R-528 red-proof red-mode-only: mode oom-check does not stop after the check (falls through to apt)
|
||||
mutation: replace
|
||||
---
|
||||
self.report["oom_check"] = self.oom_check()
|
||||
return 0
|
||||
self.who---
|
||||
with
|
||||
---
|
||||
self.report["oom_check"] = self.oom_check()
|
||||
self.who---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-mode-only-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_mode_oom_check_runs_only_the_check
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_mode_oom_check_runs_only_the_check (__main__.OOMCheck.test_mode_oom_check_runs_only_the_check)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1311, in test_mode_oom_check_runs_only_the_check
|
||||
self.assertFalse(any("apt-get" in c[-1] or "dpkg-query" in c[-1] for c in f.calls), f.calls)
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
AssertionError: True is not false : [('host', ['/usr/sbin/pct', 'status', '9201']), ('guest', 9201, ['docker', 'inspect', '-f', '{{.Config.Image}}', 'felhom-controller']), ('guest', 9201, ['date', '+%s']), ('guest', 9201, ['docker', 'run', '--name', 'felhom-oomcheck-3263952-df7cff79', '--pull', 'never', '--network', 'none', '--memory', '64m', '--memory-swap', '64m', '--label', 'felhom.oomcheck=1', '--entrypoint', 'sh', 'gitea.dooplex.hu/admin/felhom-controller:0.300.0', '-c', 'dd if=/dev/zero of=/dev/null bs=200M count=1']), ('guest', 9201, ['date', '+%s']), ('guest', 9201, ['docker', 'inspect', '-f', '{{.State.OOMKilled}} {{.State.ExitCode}}', 'felhom-oomcheck-3263952-df7cff79']), ('guest', 9201, ['docker', 'events', '--since', '1791115199', '--until', '1791115204', '--filter', 'container=felhom-oomcheck-3263952-df7cff79', '--filter', 'event=oom', '--format', '{{.Action}}']), ('guest', 9201, ['docker', 'rm', '-f', 'felhom-oomcheck-3263952-df7cff79']), ('guest', 9201, ['fuser', '/var/lib/dpkg/lock-frontend', '/var/lib/dpkg/lock']), ('guest', 9201, ['docker', 'ps', '-a', '--no-trunc', '--format', '{{.Names}}\t{{.State}}\t{{.Status}}\t{{.ID}}']), ('guest', 9201, ['getent', 'hosts', 'deb.debian.org']), ('guest', 9201, ['docker', 'exec', 'felhom-controller', 'docker', 'version', '--format', '{{.Server.Version}}']), ('guest', 9201, ['env', 'DEBIAN_FRONTEND=noninteractive', 'APT_LISTCHANGES_FRONTEND=none', 'NEEDRESTART_MODE=l', 'LC_ALL=C', 'apt-get', '-q', 'update']), ('guest', 9201, ['dpkg-query', '-W', '-f', '${Package}\t${Version}\t${db:Status-Abbrev}\n']), ('guest', 9201, ['apt-cache', 'policy', 'bash', 'containerd.io', 'docker-ce', 'libc6', 'openssl']), ('guest', 9201, ['env', 'DEBIAN_FRONTEND=noninteractive', 'APT_LISTCHANGES_FRONTEND=none', 'NEEDRESTART_MODE=l', 'LC_ALL=C', 'apt-get', '-s', '-q', 'dist-upgrade']), ('guest', 9201, ['sh', '-c', 'for p in /proc/[0-9]*; do grep -q "(deleted)" $p/maps 2>/dev/null || continue; grep -q "docker" $p/cgroup 2>/dev/null && continue; echo "${p#/proc/} $(cat $p/comm 2>/dev/null)"; done']), ('guest', 9201, ['docker', 'version', '--format', '{{.Server.Version}}']), ('guest', 9201, ['docker', 'ps', '-a', '--no-trunc', '--format', '{{.Names}}\t{{.State}}\t{{.Status}}\t{{.ID}}']), ('guest', 9201, ['getent', 'hosts', 'deb.debian.org']), ('guest', 9201, ['docker', 'exec', 'felhom-controller', 'docker', 'version', '--format', '{{.Server.Version}}'])]
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.002s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_mode_oom_check_runs_only_the_check
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,32 @@
|
||||
R-528 red-proof red-no-event: pass decided on OOMKilled alone (the oom event not required)
|
||||
mutation: replace
|
||||
---
|
||||
if res["oom_killed"] and res["oom_event"]:---
|
||||
with
|
||||
---
|
||||
if res["oom_killed"]:---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-no-event-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_no_event_is_fail
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_no_event_is_fail (__main__.OOMCheck.test_no_event_is_fail)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1265, in test_no_event_is_fail
|
||||
self.assertEqual(oc["result"], "fail", oc)
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
AssertionError: 'pass' != 'fail'
|
||||
- pass
|
||||
+ fail
|
||||
: {'detail': 'the engine reported the memory kill: OOMKilled=true and the oom event', 'exit_code': 137, 'image': 'gitea.dooplex.hu/admin/felhom-controller:0.300.0', 'oom_event': False, 'oom_killed': True, 'result': 'pass'}
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.016s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_no_event_is_fail
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,33 @@
|
||||
R-528 red-proof red-oom-pass: the docker apply no longer runs the check (the key line in run() removed)
|
||||
mutation: replace
|
||||
---
|
||||
self.report["oom_check"] = self.oom_check()
|
||||
return 0
|
||||
---
|
||||
with
|
||||
---
|
||||
pass
|
||||
return 0
|
||||
---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-oom-pass-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_pass_when_oomkilled_and_the_event
|
||||
rc=1
|
||||
E
|
||||
======================================================================
|
||||
ERROR: test_pass_when_oomkilled_and_the_event (__main__.OOMCheck.test_pass_when_oomkilled_and_the_event)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1217, in test_pass_when_oomkilled_and_the_event
|
||||
oc = rep["oom_check"]
|
||||
~~~^^^^^^^^^^^^^
|
||||
KeyError: 'oom_check'
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.014s
|
||||
|
||||
FAILED (errors=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_pass_when_oomkilled_and_the_event
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,32 @@
|
||||
R-528 red-proof red-oomkilled-false: pass decided on the oom event alone (OOMKilled not required)
|
||||
mutation: replace
|
||||
---
|
||||
if res["oom_killed"] and res["oom_event"]:---
|
||||
with
|
||||
---
|
||||
if res["oom_event"]:---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-oomkilled-false-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_oomkilled_false_is_fail
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_oomkilled_false_is_fail (__main__.OOMCheck.test_oomkilled_false_is_fail)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1257, in test_oomkilled_false_is_fail
|
||||
self.assertEqual(oc["result"], "fail", oc)
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
AssertionError: 'pass' != 'fail'
|
||||
- pass
|
||||
+ fail
|
||||
: {'detail': 'the engine reported the memory kill: OOMKilled=true and the oom event', 'exit_code': 0, 'image': 'gitea.dooplex.hu/admin/felhom-controller:0.300.0', 'oom_event': True, 'oom_killed': False, 'result': 'pass'}
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.014s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_oomkilled_false_is_fail
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,31 @@
|
||||
R-528 red-proof red-other-layers: the check runs on every layer's apply (the docker-layer condition removed)
|
||||
mutation: replace
|
||||
---
|
||||
if self.layer == "docker" and self.mode == "apply":
|
||||
# R-528---
|
||||
with
|
||||
---
|
||||
if self.mode == "apply":
|
||||
# R-528---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-other-layers-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_never_on_guest_or_host_or_in_health_mode
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_never_on_guest_or_host_or_in_health_mode (__main__.OOMCheck.test_never_on_guest_or_host_or_in_health_mode)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1301, in test_never_on_guest_or_host_or_in_health_mode
|
||||
self.assertNotIn("oom_check", r)
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^
|
||||
AssertionError: 'oom_check' unexpectedly found in {'docker_restart_needed': False, 'health_after': {'containers': {'app': {'health': 'healthy', 'id': 'bbb222', 'state': 'running'}, 'felhom-controller': {'health': 'healthy', 'id': 'aaa111', 'state': 'running'}}, 'controller': 'healthy', 'controller_docker_ok': True, 'docker_ok': True, 'network_ok': True}, 'health_before': {'containers': {'app': {'health': 'healthy', 'id': 'bbb222', 'state': 'running'}, 'felhom-controller': {'health': 'healthy', 'id': 'aaa111', 'state': 'running'}}, 'controller': 'healthy', 'controller_docker_ok': True, 'docker_ok': True, 'network_ok': True}, 'installed': [{'name': 'bash', 'origin': 'Debian', 'version': '5.2.37-2+b9'}, {'name': 'libc6', 'origin': 'Debian', 'version': '2.41-12+deb13u4'}, {'name': 'openssl', 'origin': 'Debian', 'version': '3.5.7-1~deb13u3'}], 'layer': 'guest', 'mode': 'apply', 'oom_check': {'detail': 'the engine reported the memory kill: OOMKilled=true and the oom event', 'exit_code': 137, 'image': 'gitea.dooplex.hu/admin/felhom-controller:0.300.0', 'oom_event': True, 'oom_killed': True, 'result': 'pass'}, 'pass_seconds': 0.0, 'pending': [{'from': '5.2.37-2+b9', 'name': 'bash', 'origin': ['Debian'], 'to': '5.2.37-2+b10'}], 'plan': {'already': 0, 'from_snapshot': 0, 'not_installed': 0, 'upgrade': 2}, 'reboot_needed': False, 'reboot_scanned': True, 'refused': None, 'release_id': 'os-t1', 'repair': {'clean_after': True, 'fixed': 0, 'half_configured_before': 0, 'journal_before': 0}, 'restart_needed': [], 'seconds': 0.0, 'upgraded': [{'name': 'libc6', 'version': '2.41-12+deb13u4'}, {'name': 'openssl', 'version': '3.5.7-1~deb13u3'}], 'vmid': 9201}
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.020s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_never_on_guest_or_host_or_in_health_mode
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,33 @@
|
||||
R-528 red-proof red-rm-finally: the removal is no longer in a finally (skipped when the check raised)
|
||||
mutation: replace
|
||||
---
|
||||
finally:
|
||||
try:
|
||||
mrc---
|
||||
with
|
||||
---
|
||||
if res["result"] != "error":
|
||||
try:
|
||||
mrc---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-rm-finally-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_container_removed_even_when_inspect_raises
|
||||
rc=1
|
||||
E
|
||||
======================================================================
|
||||
ERROR: test_container_removed_even_when_inspect_raises (__main__.OOMCheck.test_container_removed_even_when_inspect_raises)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1281, in test_container_removed_even_when_inspect_raises
|
||||
self.assertEqual(len(f.oom_removed), 1, "the container must be removed even when inspect raised")
|
||||
^^^^^^^^^^^^^
|
||||
AttributeError: 'Fake' object has no attribute 'oom_removed'
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.017s
|
||||
|
||||
FAILED (errors=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_container_removed_even_when_inspect_raises
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,29 @@
|
||||
R-528 red-proof red-settle-wait: no wait after the run before the end epoch is read (the measured miss, F1/F2)
|
||||
mutation: replace
|
||||
---
|
||||
self.r.sleep(OOMCHECK_SETTLE) # the engine---
|
||||
with
|
||||
---
|
||||
pass # self.r.sleep(OOMCHECK_SETTLE) # the engine---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-settle-wait-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_events_window_ends_after_the_settle_wait
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_events_window_ends_after_the_settle_wait (__main__.OOMCheck.test_events_window_ends_after_the_settle_wait)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1246, in test_events_window_ends_after_the_settle_wait
|
||||
self.assertTrue(waits and waits[0][1] >= 2, f"no settle wait after the run: {getattr(f, 'sleeps', None)}")
|
||||
~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
AssertionError: [] is not true : no settle wait after the run: None
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.016s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_events_window_ends_after_the_settle_wait
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,29 @@
|
||||
R-528 red-proof red-until-bound: the events window ends at the epoch before the run, not after the wait
|
||||
mutation: replace
|
||||
---
|
||||
"--until", str(t1 + 1),---
|
||||
with
|
||||
---
|
||||
"--until", str(t0 + 1),---
|
||||
$ OSAPPLY_UNDER_TEST=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/r528-mut/red-until-bound-felhom-os-apply python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_events_window_ends_after_the_settle_wait
|
||||
rc=1
|
||||
F
|
||||
======================================================================
|
||||
FAIL: test_events_window_ends_after_the_settle_wait (__main__.OOMCheck.test_events_window_ends_after_the_settle_wait)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py", line 1250, in test_events_window_ends_after_the_settle_wait
|
||||
self.assertEqual(until, f.oom_epochs[-1] + 1, "until = the guest epoch read after the wait, + 1")
|
||||
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
AssertionError: 1791115201 != 1791115204 : until = the guest epoch read after the wait, + 1
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 1 test in 0.016s
|
||||
|
||||
FAILED (failures=1)
|
||||
|
||||
=== control: the same test against the real wrapper ===
|
||||
$ python3 -W ignore /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/da0e2df0-3072-4d3e-bf10-c475ac1e5ae4/scratchpad/agent-r528/configs/test_felhom_os_apply.py OOMCheck.test_events_window_ends_after_the_settle_wait
|
||||
rc=0
|
||||
OK
|
||||
VERDICT: RED as expected, GREEN on the real wrapper
|
||||
@@ -0,0 +1,63 @@
|
||||
$ python3 scripts/test_upgrade_bench.py (ssh_args without the -J jump)
|
||||
test_all_three_conditions_allow (__main__.BenchAdminSeedGuard.test_all_three_conditions_allow) ... ok
|
||||
test_each_condition_alone_refuses (__main__.BenchAdminSeedGuard.test_each_condition_alone_refuses) ... ok
|
||||
test_the_bench_seeds_through_the_admin_invite_and_never_shows_the_token (__main__.BenchAdminSeedGuard.test_the_bench_seeds_through_the_admin_invite_and_never_shows_the_token) ... /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:177: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/.felhom-h-gr4ipisf' mode='r' encoding='utf-8'>
|
||||
headers.append(open(extra[i + 1][1:]).read().strip())
|
||||
ResourceWarning: Enable tracemalloc to get the object allocation traceback
|
||||
ok
|
||||
test_the_bench_venue_names_itself (__main__.BenchAdminSeedGuard.test_the_bench_venue_names_itself) ... ok
|
||||
test_the_box_walk_never_signs_in_as_admin (__main__.BenchAdminSeedGuard.test_the_box_walk_never_signs_in_as_admin) ... ok
|
||||
test_a_household_shaped_box_is_never_seeded (__main__.BoxAdminSeedGuard.test_a_household_shaped_box_is_never_seeded) ... ok
|
||||
test_a_refused_invite_is_inconclusive (__main__.BoxAdminSeedGuard.test_a_refused_invite_is_inconclusive) ... ok
|
||||
test_all_conditions_allow (__main__.BoxAdminSeedGuard.test_all_conditions_allow) ... ok
|
||||
test_each_condition_alone_refuses (__main__.BoxAdminSeedGuard.test_each_condition_alone_refuses) ... ok
|
||||
test_only_a_drill_address_is_invited (__main__.BoxAdminSeedGuard.test_only_a_drill_address_is_invited) ... ok
|
||||
test_the_test_box_seeds_through_the_invite_inside_the_box (__main__.BoxAdminSeedGuard.test_the_test_box_seeds_through_the_invite_inside_the_box) ... ok
|
||||
test_the_tester_1_box_is_a_test_box_and_demo_hp_9201_is_not (__main__.BoxAdminSeedGuard.test_the_tester_1_box_is_a_test_box_and_demo_hp_9201_is_not) ... ok
|
||||
test_an_unknown_target_is_refused (__main__.BoxWalkTargets.test_an_unknown_target_is_refused) ... ok
|
||||
test_every_row_is_complete (__main__.BoxWalkTargets.test_every_row_is_complete) ... ok
|
||||
test_the_default_stays_scratch_9202 (__main__.BoxWalkTargets.test_the_default_stays_scratch_9202) ... ok
|
||||
test_the_old_guest_switch_still_selects_demo_hp_9201 (__main__.BoxWalkTargets.test_the_old_guest_switch_still_selects_demo_hp_9201) ... ok
|
||||
test_the_tester_1_box_is_reached_through_the_hp_box (__main__.BoxWalkTargets.test_the_tester_1_box_is_reached_through_the_hp_box) ... ERROR
|
||||
test_names_changed_added_removed_and_nothing_else (__main__.ChangedFiles.test_names_changed_added_removed_and_nothing_else) ... ok
|
||||
test_never_a_bare_root_never_outside (__main__.ClearScratch.test_never_a_bare_root_never_outside) ... /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:52: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/bench-scratch-b_aqb9ob/hdd/appdata/x/config/config.php' mode='w' encoding='utf-8'>
|
||||
open(p, "w").write("last run")
|
||||
ResourceWarning: Enable tracemalloc to get the object allocation traceback
|
||||
ok
|
||||
test_the_apps_own_folders_are_cleared_and_said (__main__.ClearScratch.test_the_apps_own_folders_are_cleared_and_said) ... /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:52: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/bench-scratch-msfdhygg/hdd/appdata/nextcloud/config/config.php' mode='w' encoding='utf-8'>
|
||||
open(p, "w").write("last run")
|
||||
ResourceWarning: Enable tracemalloc to get the object allocation traceback
|
||||
/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:52: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/bench-scratch-msfdhygg/hdd/appdata/immich/config/config.php' mode='w' encoding='utf-8'>
|
||||
open(p, "w").write("last run")
|
||||
ResourceWarning: Enable tracemalloc to get the object allocation traceback
|
||||
ok
|
||||
test_nothing_new_is_refused (__main__.DefinitionEdge.test_nothing_new_is_refused) ... ok
|
||||
test_the_edge_names_the_new_service (__main__.DefinitionEdge.test_the_edge_names_the_new_service) ... ok
|
||||
test_the_to_render_reads_the_new_definition (__main__.DefinitionEdge.test_the_to_render_reads_the_new_definition) ... ok
|
||||
test_answers_of_any_code_are_the_app_answering (__main__.LoadVerdict.test_answers_of_any_code_are_the_app_answering) ... ok
|
||||
test_every_request_errored_is_inconclusive (__main__.LoadVerdict.test_every_request_errored_is_inconclusive) ... ok
|
||||
test_under_half_is_inconclusive_and_none_is_inconclusive (__main__.LoadVerdict.test_under_half_is_inconclusive_and_none_is_inconclusive) ... ok
|
||||
test_a_listed_name_that_grew_or_vanished_still_marks (__main__.MarkerIgnore.test_a_listed_name_that_grew_or_vanished_still_marks) ... ok
|
||||
test_a_moved_tree_the_file_walk_cannot_name_keeps_the_mark (__main__.MarkerIgnore.test_a_moved_tree_the_file_walk_cannot_name_keeps_the_mark) ... ok
|
||||
test_an_unlisted_changed_file_still_marks (__main__.MarkerIgnore.test_an_unlisted_changed_file_still_marks) ... ok
|
||||
test_every_entry_has_a_reason (__main__.MarkerIgnore.test_every_entry_has_a_reason) ... ok
|
||||
test_the_list_belongs_to_its_app (__main__.MarkerIgnore.test_the_list_belongs_to_its_app) ... ok
|
||||
test_the_six_measured_markers_do_not_mark_and_each_carries_a_reason (__main__.MarkerIgnore.test_the_six_measured_markers_do_not_mark_and_each_carries_a_reason) ... ok
|
||||
test_env_is_0600_and_shredded (__main__.SecretHygiene.test_env_is_0600_and_shredded) ... ok
|
||||
test_evidence_files_are_redacted (__main__.SecretHygiene.test_evidence_files_are_redacted) ... ok
|
||||
test_redact_longest_first (__main__.SecretHygiene.test_redact_longest_first) ... ok
|
||||
test_secret_values_are_the_generated_fields_and_the_seed_password (__main__.SecretHygiene.test_secret_values_are_the_generated_fields_and_the_seed_password) ... ok
|
||||
|
||||
======================================================================
|
||||
ERROR: test_the_tester_1_box_is_reached_through_the_hp_box (__main__.BoxWalkTargets.test_the_tester_1_box_is_reached_through_the_hp_box)
|
||||
----------------------------------------------------------------------
|
||||
Traceback (most recent call last):
|
||||
File "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py", line 432, in test_the_tester_1_box_is_reached_through_the_hp_box
|
||||
self.assertEqual(a[a.index("-J") + 1], "demo-hp")
|
||||
~~~~~~~^^^^^^
|
||||
ValueError: '-J' is not in list
|
||||
|
||||
----------------------------------------------------------------------
|
||||
Ran 36 tests in 0.043s
|
||||
|
||||
FAILED (errors=1)
|
||||
@@ -0,0 +1,9 @@
|
||||
wger's ladder step (catalog cf1ed43) reads memory_cgroup_peak_pct 100.0. From the bench's own verdict
|
||||
(audits/design-build-2026-10-06/F/bench/evidence/MV-wger/verdict.json), read 2026-10-06:
|
||||
wger: limit 402653184, cgroup peak 402653184 (1.0), anon peak 202338304 (0.503), oom_kills 0 (the kernel's
|
||||
memory.events counter), restarts 0, OOMKilled flag false
|
||||
the memory-watch log: no sample with kills>0 or rs>0
|
||||
=> FILE CACHE, not a kill: the entrypoint's collectstatic writes ~283 MB of static files at every start
|
||||
(F/bench-static-check.txt: "static volume: 283M; files: 22725"); page cache counts in the cgroup peak
|
||||
(memory note harness-memory-peak-includes-page-cache). The step's mark is judged on anon (50.3 %, decision 22).
|
||||
wger's memory hint (416M) stays.
|
||||
@@ -157,6 +157,7 @@ stopping line that lies.
|
||||
| **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
|
||||
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. | — | Read back the first night runs on demo-hp and demo-felhom under v0.301.0 (the agent's `snapshotted` line, the apps' StartedAt, the controller's window lines); close with that evidence. If nothing: the row stays open | CC |
|
||||
| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC |
|
||||
| **R-894** | Backup & restore | P3 | **After an agent restart, an UNREADABLE off-site storage makes the off-site tier look DUE, so the box asks for a copy that cannot be made.** MEASURED 2026-10-05 on demo-hp, read 2026-10-06 (`audits/readback-2026-10-07/C/`): the last off-site copy was 2026-10-01 20:15Z (ep0's own listing, verify `ok`), so the 7-day tier was NOT due; the agent had restarted at 04:57 local; at 06:25 `GET …/storage/felhom-pbs/content` answered 500 *Can't connect to 10.77.0.1:8007*; `newestArchiveOn` returned `unknown` and fell back to the in-memory record (`internal/localapi/server.go`), which a restart empties (`internal/backup/store.go` — memory only, R-348); so the tier read DUE, the controller requested it, and vzdump failed (*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks (`07` §6.4); whether it did that night is NOT KNOWN (the controller's log was lost to a later restart). The hub got `whole_guest_backup_failed` (error); its operator mail then failed (fixed in the hub this session: a failed operator mail is retried). The code's rule is deliberate: an unreadable storage must not suppress a backup. | **OPEN — filed 2026-10-06** | a design: keep the newest success per tier on disk (as `RestoreTestState` does) so the fallback is the last known copy, not "never"; or report an unreachable storage so no app is stopped for it | Decide the fallback; build it in the agent with a test that restarts the agent and then cannot read the storage; measure once | CC |
|
||||
| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC |
|
||||
| **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC |
|
||||
| **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** **2026-10-06: leg (a) PUSHED** to the live catalog (`c265b37`). **2026-10-06 (burn-down night, later): leg (b) NEEDS A DESIGN.** `07` §7.4 sets no direction; refusing the restore without the DB password, or `ALTER USER` after it, each change restore behaviour on customer data. | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC |
|
||||
@@ -232,7 +233,7 @@ stopping line that lies.
|
||||
| ID | Category | Sev | What | State | Blocked on | Next action | Owner |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-243** | Monitoring & notifications | P2 | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | — | — | operator |
|
||||
| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. **2026-10-06: STOPPED BEFORE ANY CODE — the measurement contradicted the design** (decision 155 said A then C). On scratch 9202, Docker 29.8.2, the `OOMKilled` flag was TRUE in all four shapes tried: a child process killed while the container kept running (`oom_kill` 0 → 3, flag true) and three main-process kills (exit 137, flag true); an `oom` event each time. The false flags of 2026-09-15 were on Docker 29.8.0. Cost read for the record: one exec read of 21 containers on demo-hp = 1.6 s. `audits/design-build-2026-10-06/`C/. | — | Operator: (A) drop A + C — today's engine reports the flag; keep the row as a watch for a box that reports a false flag; (B) build A + C anyway as insurance (≈ 1.6 s per 5-min scan on demo-hp). If nothing: nothing is built; the alarm keeps reading the flag | operator |
|
||||
| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. **2026-10-06: STOPPED BEFORE ANY CODE — the measurement contradicted the design** (decision 155 said A then C). On scratch 9202, Docker 29.8.2, the `OOMKilled` flag was TRUE in all four shapes tried: a child process killed while the container kept running (`oom_kill` 0 → 3, flag true) and three main-process kills (exit 137, flag true); an `oom` event each time. The false flags of 2026-09-15 were on Docker 29.8.0. Cost read for the record: one exec read of 21 containers on demo-hp = 1.6 s. `audits/design-build-2026-10-06/`C/. **2026-10-06 18:24: operator ruling, option A (decision 157):** nothing is built; the row stays open as a WATCH for a box whose Docker reports a false flag. **Added:** a Docker engine set may not be approved until it reports a memory kill correctly on the boxes that ran it (built in this row's next session). | — | Watch; build the Docker-approval memory-kill check (decision 157) | CC |
|
||||
| **R-79** | Monitoring & notifications | P3 | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the issues and warnings travel to the hub as sentences; changing them is the two-repo spike the row itself names. | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC |
|
||||
| **R-211** | Monitoring & notifications | P3 | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC |
|
||||
| **R-333** | Monitoring & notifications | P3 | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) |
|
||||
@@ -282,7 +283,7 @@ stopping line that lies.
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC |
|
||||
| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=<sha>` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin/<repo>/actions/runs/<id>/rerun`. Keep the 2026-10-12 check | operator |
|
||||
| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config`, no entry in `operations/nodes.md` and no Proxmox host known to this workspace — releases reach it only through the hub's signed jobs and floors. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** | the Tester 1 box's host and an SSH route | Operator: name where the Tester 1 box runs and whether CC may reach it by SSH; then `box_walk.py` gains a target table (host, guest, base URL) and `BOX_ADMIN_SEED_GUESTS` its row. If nothing: the walk stays on scratch 9202 | CC |
|
||||
| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** | — | Build the route (a jump through `hp`, the target in `box_walk.py`'s own table), prove one step there, add the box to `operations/nodes.md` | CC |
|
||||
| **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC |
|
||||
| **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC |
|
||||
| **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator |
|
||||
|
||||
@@ -34,6 +34,17 @@ rather than trusting these** (`ip -br addr show vmbr0`; the N100 read `.162` on
|
||||
tailnet addresses are the stable ones — use those. Direct LAN literals are **not** reachable from
|
||||
DooPlex while the boxes are away (`felhom-pve-lan` → `No route to host`, 2026-07-30).
|
||||
|
||||
## The Tester 1 box — a VM on the HP box (`09` §3 decision 158)
|
||||
|
||||
| | `tester-1-d70be4` |
|
||||
|---|---|
|
||||
| What it is | **QEMU VM 341 `night1004-tester1` on demo-hp** (installed from the ISO on 2026-10-04, `audits/night-2026-10-04/tester1/`); disposable (operator) |
|
||||
| PVE node name | `felhom` (its certificate: `CN=felhom.enkicsifelhom.hu`) |
|
||||
| LAN address | **DHCP** — read 2026-10-06 as `192.168.0.154` (MAC `bc:24:11:ac:e3:f9`; find it with `ip neigh` after a ping) |
|
||||
| Customer / guest | `tester-1`; guest LXC 9201 at `192.168.0.101`, dashboard `felhom.enkicsifelhom.hu` |
|
||||
| SSH | through the HP box: `ssh -J demo-hp root@<VM address>` — **DooPlex's key is NOT authorized there yet** (2026-10-06: `Permission denied (publickey,password)`); `box_walk.py` `TARGET=tester-1` holds the route |
|
||||
| Identity checked | 2026-10-06: the agent's own report `host.node=felhom` = the VM's certificate; the guest answers its domain (200) and not demo-hp's (404) — `audits/readback-2026-10-07/E1-identity-match.txt` |
|
||||
|
||||
**Which box is safe to break, and what may be done to each:
|
||||
[`../runbooks/target-selection.md`](../runbooks/target-selection.md).** This page is *what the hardware
|
||||
is*; that page is *what you may do to it*.
|
||||
|
||||
Reference in New Issue
Block a user