Day 4 (2026-10-09): kernel night read back, release delivered, ep0-copy job live, D1-D4 proven; R-279/R-30/R-79/R-35/R-901 closed
gates / gates (push) Successful in 4m46s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 08:37:18 +02:00
parent b9073e8fb6
commit 0cb169923c
15 changed files with 281 additions and 16 deletions
@@ -0,0 +1,3 @@
demo-felhom 2026-10-09 05:31:17 agent 0.154.0 backups: [('felhom-backup', '2026-10-09T02:40:16Z', True)]
demo-hp 2026-10-09 05:31:59 agent 0.154.0 backups: [('felhom-pbs', '2026-10-08T20:34:26Z', True), ('local', '2026-10-08T20:20:43Z', True)]
tester-1 2026-10-09 05:33:29 agent 0.154.0 backups: [('local', '2026-10-09T02:40:30Z', True)]
@@ -0,0 +1,15 @@
2026-10-09T05:33:47Z
## demo-hp
Oct 09 07:32:01 demo-hp felhom-agent[2730121]: time=2026-10-09T07:32:01.215+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)"
Oct 09 07:32:01 demo-hp felhom-agent[2730121]: time=2026-10-09T07:32:01.215+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.154.0 written=1 same=26 self-check=ok signers-created=False"
Oct 09 07:32:02 demo-hp felhom-agent[2730121]: time=2026-10-09T07:32:02.138+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded=""
## felhom-pve
Oct 09 07:31:16 demo-felhom felhom-agent[141055]: time=2026-10-09T07:31:16.948+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)"
Oct 09 07:31:16 demo-felhom felhom-agent[141055]: time=2026-10-09T07:31:16.948+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.154.0 written=1 same=26 self-check=ok signers-created=False"
Oct 09 07:31:17 demo-felhom felhom-agent[141055]: time=2026-10-09T07:31:17.520+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded=""
## root@192.168.0.154
Oct 09 07:33:30 felhom felhom-agent[3010618]: time=2026-10-09T07:33:30.409+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE WROTE /usr/local/sbin/felhom-os-apply (replaced)"
Oct 09 07:33:30 felhom felhom-agent[3010618]: time=2026-10-09T07:33:30.409+02:00 level=INFO msg="osupdate: wrapper" line="os-apply: BUNDLE DONE agent=0.154.0 written=1 same=26 self-check=ok signers-created=False"
Oct 09 07:33:31 felhom felhom-agent[3010618]: time=2026-10-09T07:33:31.207+02:00 level=WARN msg="osupdate: capability probe after the config bundle" ok=68 total=68 degraded=""
[exited with code 0]
@@ -0,0 +1,7 @@
2026-10-09T05:36:41Z demo-hp: agent 0.154.0
gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 31 minutes (healthy)
2026-10-09T05:36:42Z felhom-pve: agent 0.154.0
gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 31 minutes (healthy)
2026-10-09T05:36:43Z root@192.168.0.154: agent 0.154.0
gitea.dooplex.hu/admin/felhom-controller:0.304.0 Up 31 minutes (healthy)
Tester 2: not touched
@@ -4,3 +4,8 @@ rc=2
-- copy namespaces on disk: Tester-2 demo-felhom demo-hp operator tester-1
-- state file:
cat: /var/lib/felhom-ep0-copy-gc/absent.json: No such file or directory
== DRY RUN 2 2026-10-09T05:31:42Z (installed from felhom.eu 69fa9cf7, the shape fix)
ep0-copy-gc: done: 5 copy namespace(s), 5 on ep0, 0 absent, 0 to delete (dry run), 0 failed
rc=0
-- state file:
{}
@@ -0,0 +1,11 @@
== TIMER ON 2026-10-09T05:32:39Z (operator's word, present)
enabled
active
NEXT LEFT LAST PASSED UNIT ACTIVATES
Fri 2026-10-09 08:00:00 CEST 27min - - felhom-ep0-copy-gc.timer felhom-ep0-copy-gc.service
== first timer run (real mode, --apply)
Oct 09 08:00:00 dooplex systemd[1]: Starting felhom-ep0-copy-gc.service - Felhom: remove a deleted customer's namespace from ep0-copy (R-901, decision 181)...
Oct 09 08:00:00 dooplex felhom-ep0-copy-gc[1693084]: ep0-copy-gc: done: 5 copy namespace(s), 5 on ep0, 0 absent, 0 to delete, 0 failed
Oct 09 08:00:00 dooplex systemd[1]: felhom-ep0-copy-gc.service: Deactivated successfully.
Oct 09 08:00:00 dooplex systemd[1]: Finished felhom-ep0-copy-gc.service - Felhom: remove a deleted customer's namespace from ep0-copy (R-901, decision 181).
ExecMainStatus=0
@@ -0,0 +1 @@
es. Use the request buttons above — the box delivers on its next cycle. Operator Actions The box acts on its next report reply (usually seconds, at most one report interval) and answers on the report after that. Each press runs once. Recorded in the box's own log and in Events. Unanswered after a day: expired. Run off-site backup now disk-health-check fill-watch offsite-integrity offsite-proof Run check now Stop deletion countdown Extend countdown (days from now) # Action Requested By Outcome 2 offsite_backup_now 3 min ago operator CLI (basic auth) from 10.42.0.1 done the off-site backup finished 1 run_job fill-watch 3 min ago operator CLI (basic auth) from 10.42.0.1 done the job fill-watch ran to its end — this does not say what it found; its finding is in its own log line and alarms Network WireGuard 10.77.0.4 confirmed Interface Address vmbr0 192.168.0.154/24 Every routable address the box holds, as the kernel sees it. Loopback and link-local are excluded — including the 169.254.253.1 local-API island, which is identical on every box. The PVE web console is at https://<the LAN address>:8006 . Guest network healthy Guest State Address Route dhclient Repairs (1h) 9201 healthy 192.
@@ -0,0 +1,8 @@
== D1 on tester-1-d70be4 2026-10-09T05:37:15Z: POST /hosts/<id>/operator-action (the endpoint the host page's buttons call; Basic + X-Felhom-Operator)
action=run_job&arg=fill-watch -> 303
action=offsite_backup_now -> 303
refusal control (action=delete_everything) -> Refused: unknown action "delete_everything"
2026/10/09 07:37:15 [INFO] operator action #1 run_job fill-watch requested for tester-1 (host tester-1-d70be4) by operator CLI (basic auth) from 10.42.0.1 — the box acts on its next report
2026/10/09 07:37:15 [INFO] operator action #2 offsite_backup_now requested for tester-1 (host tester-1-d70be4) by operator CLI (basic auth) from 10.42.0.1 — the box acts on its next report
2026/10/09 07:37:15 [INFO] operator action refused for host tester-1-d70be4: unknown action "delete_everything"
@@ -0,0 +1,91 @@
== D2 on tester-1 2026-10-09T05:40:39Z
before: connected now
status: stopped
guest stopped at 2026-10-09T05:40:44Z
05:40:44 +4s: connected now
05:41:00 +20s: connected now
05:41:15 +35s: connected now
05:41:30 +50s: connected now
05:41:45 +65s: connected now
05:42:01 +81s: connected now
05:42:16 +96s: connected now
05:42:31 +111s: connected now
05:42:46 +126s: connected now
05:43:02 +142s: connected now
05:43:17 +157s: connected now
05:43:32 +172s: connected now
05:43:47 +187s: connected now
05:44:03 +203s: connected now
05:44:18 +218s: connected now
05:44:33 +233s: connected now
05:44:48 +248s: connected now
05:45:04 +264s: connected now
05:45:19 +279s: connected now
05:45:34 +294s: connected now
05:45:50 +310s: connected now
05:46:06 +326s: connected now
05:46:21 +341s: connected now
05:46:36 +356s: connected now
05:46:51 +371s: connected now
05:47:07 +387s: connected now
05:47:22 +402s: connected now
05:47:37 +417s: connected now
05:47:53 +433s: connected now
05:48:08 +448s: connected now
05:48:23 +463s: connected now
05:48:38 +478s: connected now
05:48:54 +494s: connected now
05:49:09 +509s: connected now
05:49:24 +524s: connected now
05:49:39 +539s: connected now
05:49:55 +555s: connected now
05:50:10 +570s: connected now
05:50:25 +585s: connected now
05:50:40 +600s: connected now
05:50:56 +616s: connected now
NOTE: run 1 aborted at +600s — the host agent's guest-power watchdog STARTED the guest again at 05:41:20Z (36 s after the stop, by design: onboot guest). A stopped guest on a live host is not a switched-off box. Run 2 switches off the whole Tester 1 VM.
== D2 run 2 2026-10-09T05:51:23Z: qm shutdown 341 (the whole Tester 1 machine) on the HP host
before: connected now
status: stopped
machine off at 2026-10-09T05:53:27Z
05:53:34 +7s: connected now
05:53:49 +22s: connected now
05:54:04 +37s: connected now
05:54:20 +53s: connected now
05:54:35 +68s: connected now
05:54:50 +83s: connected now
05:55:06 +99s: connected now
05:55:21 +114s: last connected 2026-10-09 05:51 UTC (4 min ago)
== GET /hosts/tester-1-d70be4/delete-impact 2026-10-09T05:55:32Z (read only; nothing deleted)
{"deletable":false,"escrow_present":true,"guests":2,"log_bundles":0,"off_tick_required":false,"pbs_secret_present":true,"recovery_present":true,"reports":449,"status":"ok","wg_peer_bound":true}
== control: GET /hosts/demo-hp-bb76ea/delete-impact (box running)
{"deletable":false,"escrow_present":true,"guests":2,"log_bundles":0,"off_tick_required":false,"pbs_secret_present":true,"recovery_present":true,"reports":7688,"status":"ok","wg_peer_bound":true}
05:55:49 ok deletable=False off_tick_required=False
05:56:00 ok deletable=False off_tick_required=False
05:56:10 ok deletable=False off_tick_required=False
05:56:20 ok deletable=False off_tick_required=False
05:56:30 ok deletable=False off_tick_required=False
05:56:40 ok deletable=False off_tick_required=False
05:56:51 ok deletable=False off_tick_required=False
05:57:01 ok deletable=False off_tick_required=False
05:57:11 ok deletable=False off_tick_required=False
05:57:21 ok deletable=False off_tick_required=False
05:57:31 ok deletable=False off_tick_required=False
05:57:42 ok deletable=True off_tick_required=True
== control (ingress-nginx, a different channel from the hub's presence): the wait calls 05:45-05:58Z
192.168.0.101 [09/Oct/2026:05:49:33 /api/v1/wait?gen=2
192.168.0.155 [09/Oct/2026:05:49:38 /api/v1/wait?gen=0
192.168.0.149 [09/Oct/2026:05:49:39 /api/v1/wait?gen=0
192.168.0.101 [09/Oct/2026:05:51:41 /api/v1/wait?gen=2
192.168.0.155 [09/Oct/2026:05:53:40 /api/v1/wait?gen=0
192.168.0.149 [09/Oct/2026:05:53:42 /api/v1/wait?gen=0
192.168.0.155 [09/Oct/2026:05:57:42 /api/v1/wait?gen=0
192.168.0.149 [09/Oct/2026:05:57:44 /api/v1/wait?gen=0
192.168.0.101 = Tester 1's guest 9201 (nodes.md). Its last call ended 05:51:41Z; the others kept calling.
== VM 341 started again 2026-10-09T05:57:58Z
05:58:36 back: connected now; delete-impact: ok False False
@@ -0,0 +1,42 @@
== D3 on tester-1 2026-10-09T06:04:10Z: docker stop filebrowser (a protected container) in guest 9201
filebrowser
2026/10/09 08:08:32 [INFO] Event from tester-1: health_critical (error) — Rendszer állapot kritikus (volt: ok)
2026/10/09 08:08:33 [INFO] Operator email sent for tester-1/health_critical
household mail: none — tester-1 has no customer_notifications row (no household address set), so the dispatcher returns before sending (dispatcher.go, prefs == nil). Read from a hub DB copy with its WAL.
== docker start filebrowser 2026-10-09T06:09:28Z
filebrowser
== 2026-10-09T06:09:57Z Tester 1 dashboard POST /settings/notifications: household address tester1@felhom.eu (our catch-all), the same 13 default events (was: no address)
-> 200
== run 2 2026-10-09T06:13:48Z: docker stop filebrowser again (household address now set)
filebrowser
2026/10/09 08:18:32 [INFO] Event from tester-1: health_critical (error) — Rendszer állapot kritikus (volt: ok)
2026/10/09 08:18:32 [INFO] Operator email suppressed for tester-1/health_critical — cooldown (key=tester-1:health_critical)
2026/10/09 08:18:33 [INFO] Customer email sent to (catch-all) for tester-1/health_critical
== docker start filebrowser 2026-10-09T06:18:44Z
== the household mail, read from the catch-all (Gmail, a different channel from the hub log), 2026-10-09T06:18:33Z to tester1@felhom.eu
Subject: [Felhom] Hiba: A rendszer állapota kritikus!
Kedves Ügyfél! / A Felhom rendszered a következő értesítést küldte: / A rendszer állapota kritikus! / Részletek: - Szerver: tester-1
- Időpont: 2026-10-09 08:18 - Szint: Hiba - Típus: health_critical - Üzenet: Rendszer állapot kritikus (volt: ok)
A részleteket a vezérlőpultodon látod. / Ha kérdésed van, vedd fel a kapcsolatot az üzemeltetővel.
CHECK: no "{" in the body (yes); the dashboard line is there (yes).
== the operator mail (run 1), 2026-10-09T06:08:33Z to admin@felhom.eu: "[Felhom] 🔴 tester-1: health_critical" — it still carries
Details: {"previous_status":"ok",...} (the note stays on the operator's side). The run-2 operator mail was held by its 6 h cooldown.
== 2026-10-09T06:19:23Z household address cleared (empty address + no events: the dashboard's guard refuses an empty address with events ticked): 200
2026/10/09 08:19:23 [INFO] Notification preferences updated for tester-1: email=tester1@felhom.eu, events=[]
hub row after the clear: address kept by the hub's empty-email no-clobber guard (handler.go, v0.71.0, by design), events=[] -> no household mail is sent (isEventEnabled). Behaviour as before the proof: no household mail.
== D3 dashboard half on scratch 9202 2026-10-09T06:19:59Z
language -> en: 302
stopped: filebrowser
06:30:11 dashboard (English): positive 'A protected system service is not running: filebrowser' -> 0; negative (Hungarian) 'rendszerszolgaltatas'/'nem fut' ASCII fragment -> 0; html lang -> lang="en"
note: on 9202 only traefik is a protected container (no tunnel, no samba), so stopping filebrowser raised no issue there; the controller's base-stack self-heal started it again at 06:24:24. The household-language check reads 9202's EXISTING health banners instead.
06:31:01 English dashboard (lang=en), the health banners:
- Storage not reachable: /mnt/felhom-drives/scratch_hdd/userdata/grimmory System monitor →
- Storage not reachable: /mnt/felhom-drives/scratch_hdd/userdata/metube System monitor →
- The hub connection is off — central monitoring is not running System monitor →
language -> hu (as found): 302
control, Hungarian dashboard, the same banners:
- Adattároló nem elérhető: /mnt/felhom-drives/scratch_hdd/userdata/grimmory Rendszermonitor →
- Adattároló nem elérhető: /mnt/felhom-drives/scratch_hdd/userdata/metube Rendszermonitor →
- Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor →
@@ -0,0 +1,18 @@
== D4 on scratch 9202 (controller 0.304.0) 2026-10-09T05:37:58Z
login: HTTP/2 302 , cookie present: yes
before restart: GET /api/settings -> 404
control, no cookie: GET /api/settings -> 401
before restart: /api/stacks with cookie -> 200, without -> 401
started 2026-10-09T05:37:26.038905595Z
started 2026-10-09T05:38:12.928521354Z healthy
after restart: /api/stacks with the SAME cookie -> 200, without -> 401
after restart: /settings with cookie -> 200
== part 2 2026-10-09T05:39:10Z: change the password (temp value kept out of the record), then restart
csrf scraped: yes
change: 302 -> https://192.168.0.114/login?flash=flash.login.password_changed
right after the change: old cookie A -> 401, cookie B -> 401
restarted 2026-10-09T05:39:24.066966039Z healthy
after the restart: old cookie A -> 401, cookie B -> 401
positive control: a new sign-in with the new password -> /api/stacks 200
password set back: 302
the original password signs in again -> /api/stacks 200
+12
View File
@@ -26,6 +26,18 @@
---
## 2026-10-09 — the release day: D1–D4 proven live, the ep0-copy job installed
The full text of every row below: `git show b9073e8fb6:documentation/backlog/OPEN-ITEMS.md`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-279** | **There is no operator-triggerable off-site backup.** (P4) | CLOSED 2026-10-09 — DELIVERED and proven live (hub 0.144.0 + controller 0.304.0): the host page's „Run off-site backup now" on the Tester 1 box (`09` §3 decision 185; the operator moved the hub-side proofs from 9202, which has no hub, to Tester 1 on 2026-10-09). The hub row read „done — the off-site backup finished"; control from the box's own log: `[offbox] backup OK: 3 app(s) backed up, 18 snapshot(s), 47s`. | `audits/release-2026-10-09/proofs/`D1/ |
| **R-35** | **Config-apply should not end the customer's session.** (P3) | CLOSED 2026-10-09 — DELIVERED and proven live on scratch 9202 (controller 0.304.0): signed in, restarted the controller, the same cookie read 200 (no cookie: 401); after a password change both old cookies read 401 at once and after the next restart; a new sign-in worked; the password was set back. | `audits/release-2026-10-09/proofs/`D4/ |
| **R-79** | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** (P3) | CLOSED 2026-10-09 — DELIVERED and proven live (hub 0.144.0 + controller 0.304.0): a `health_critical` on the Tester 1 box (a protected container stopped) → the household mail in the catch-all has no `{` and ends „A részleteket a vezérlőpultodon látod."; the operator mail still carries the details. Dashboard half on 9202: the health banners read English for `lang=en` and Hungarian for `lang=hu`. | `audits/release-2026-10-09/proofs/`D3/ |
| **R-30** | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** (P3) | CLOSED 2026-10-09 — DELIVERED and measured live (hub 0.144.0): the whole Tester 1 machine shut down 05:53:27Z; the host page read „last connected 05:51 UTC" at +114 s; the „box is off" tick opened at 05:57:42Z (6 min after the last call, as designed); control from the ingress log: Tester 1's guest (192.168.0.101) made its last `/api/v1/wait` at 05:51:41Z while the other boxes kept calling. A stopped GUEST on a live host is started again by the agent in 36 s (guest-power watchdog), so only a switched-off box opens the tick. demo-hp was never tick-deleted. | `audits/release-2026-10-09/proofs/`D2/ |
| **R-901** | **Two kinds of a household's data outlive the deletion of the customer, and no document says when they go.** (P2) | CLOSED 2026-10-09 — both halves LIVE: the hub's 1-year deletion of a deleted customer's audit rows ships in hub 0.144.0 (daily prune; nothing is a year old yet); the DooPlex `ep0-copy` job installed with the operator present — key root-only, dry run, then the 08:00 timer; the first dry run found a parser defect (ep0 answers `{"data": [...]}`) that the mass-absence guard stopped (nothing recorded, nothing deleted), fixed `69fa9cf7`, red-proved; the first real run read 5 of 5 namespaces, deleted 0. Both times are for the privacy-notice draft (R-813). | `audits/release-2026-10-09/ep0-copy-gc/` |
## 2026-10-08 (evening) — the operator's decision sheet D1–D10
The full text of every row below: `git show 4ef2fee090:documentation/backlog/OPEN-ITEMS.md`.
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -42,7 +42,7 @@ DooPlex while the boxes are away (`felhom-pve-lan` → `No route to host`, 2026-
| PVE node name | `felhom` (its certificate: `CN=felhom.enkicsifelhom.hu`) |
| LAN address | **DHCP** — read 2026-10-06 as `192.168.0.154` (MAC `bc:24:11:ac:e3:f9`; find it with `ip neigh` after a ping) |
| Customer / guest | `tester-1`; guest LXC 9201 at `192.168.0.101`, dashboard `felhom.enkicsifelhom.hu` |
| SSH | through the HP box: `ssh -J demo-hp root@<VM address>` — **DooPlex's key is NOT authorized there yet** (2026-10-06: `Permission denied (publickey,password)`); `box_walk.py` `TARGET=tester-1` holds the route |
| SSH | **`ssh root@192.168.0.154` from DooPlex works BY KEY** (measured 2026-10-09: `BatchMode=yes`, used for the release read-back); the old route `ssh -J demo-hp root@<VM address>` stays the fallback; `box_walk.py` `TARGET=tester-1` holds the route |
| Identity checked | 2026-10-06: the agent's own report `host.node=felhom` = the VM's certificate; the guest answers its domain (200) and not demo-hp's (404) — `audits/readback-2026-10-07/E1-identity-match.txt` |
**Which box is safe to break, and what may be done to each: