register: 9 closed (R-493/495/496/512/513/514/515/517/523), 9 opened (R-525..R-533), R-509/510/511/518 narrowed; ISO 1.27.1 publish record; website changelog; P1-fixes evidence
gates / gates (push) Successful in 20s

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-15 11:32:18 +02:00
parent 5572216506
commit 0a6cf60bf8
20 changed files with 334 additions and 17 deletions
@@ -0,0 +1,22 @@
secret keys: REPORT_API_KEY
## 2026-09-15T08:34:38Z before: felhom-agent 0.130.0
signed: op=agent_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=35882fe2913235671df3341b308ad68a expires=2026-09-15T09:04:39Z
wrote envelope to /tmp/claude-1000/agent-update-demo-hp.json
uploaded signed op to the hub jobs queue
## 2026-09-15T08:39:50Z after: felhom-agent 0.130.0
## 2026-09-15T08:44:17Z agent now: felhom-agent 0.131.0
## 2026-09-15T08:44:17Z final: felhom-agent 0.131.0
Sep 15 10:44:16 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:16.975+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-hp-bb76ea guest="" source=one_shot_job disposition=destructive allowed=true reason=signed key_id=felhom-op-1 nonce=35882fe2… durable_id=""
Sep 15 10:44:16 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:16.975+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
Sep 15 10:44:16 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:16.975+02:00 level=WARN msg="signedjobs: AUTHORIZED signed op — executing" job=b720fe262c748ea4 op=agent_update key_id=felhom-op-1 nonce=35882fe2913235671df3341b308ad68a
Sep 15 10:44:16 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:16.976+02:00 level=WARN msg="agent_update: downloading operator-signed binary" version=0.131.0 url=https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/0.131.0/felhom-agent sha256=1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c
Sep 15 10:44:17 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:17.388+02:00 level=WARN msg="agent_update: handing staged binary to the guarded wrapper" staged=/var/lib/felhom-agent/selfupdate/felhom-agent-0.131.0 version=0.131.0
Sep 15 10:44:17 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:17.537+02:00 level=WARN msg="agent_update: apply handed off; restart scheduled" version=0.131.0 wrapper=""
Sep 15 10:44:17 demo-hp felhom-agent[4703]: time=2026-09-15T10:44:17.537+02:00 level=WARN msg="signedjobs: signed op COMPLETED" job=b720fe262c748ea4 op=agent_update
felhom-agent 0.131.0
active
Sep 15 10:44:21 demo-hp felhom-agent[1526161]: time=2026-09-15T10:44:21.659+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s guests_dir=/var/lib/felhom-agent/guests
Sep 15 10:44:21 demo-hp felhom-agent[1526161]: time=2026-09-15T10:44:21.659+02:00 level=WARN msg="selfupdate: new version running — dwelling before commit" version=0.131.0 prev=0.130.0 dwell=1m0s
Sep 15 10:45:21 demo-hp sudo[1529408]: felhom-agent : PWD=/ ; USER=root ; COMMAND=/usr/local/sbin/felhom-selfupdate-guarded commit
Sep 15 10:45:21 demo-hp felhom-selfupdate-guarded[1529417]: committed (pending cleared, .prev retained)
Sep 15 10:45:21 demo-hp felhom-agent[1526161]: time=2026-09-15T10:45:21.684+02:00 level=WARN msg="selfupdate: update committed" version=0.131.0 wrapper=""
@@ -0,0 +1,6 @@
HTTP/1.1 303 See Other
Location: /customers/demo-hp?flash=floor_set
## 2026-09-15T08:46:59Z before: gitea.dooplex.hu/admin/felhom-controller:0.242.0
## 2026-09-15T08:47:18Z after: gitea.dooplex.hu/admin/felhom-controller:0.243.0 running
2026/09/15 10:46:59 [INFO] Customer demo-hp controller-version floor override set to "0.243.0" (declared MinAgent "0.131.0")
2026/09/15 10:47:02 [INFO] managed floor SERVED for demo-hp: floor 0.243.0, agent requirement "0.131.0" from declared (golden 0.242.0)
@@ -0,0 +1,4 @@
## 2026-09-15T09:14:41Z hub log lines 'Controller supervisor:'
2026/09/15 11:14:37 [INFO] Controller supervisor: controller_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
2026/09/15 11:14:37 [INFO] Controller supervisor: controller_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
2026/09/15 11:14:38 [INFO] Operator email sent for demo-hp/controller_crashloop
@@ -0,0 +1,9 @@
## A.4 kill 1 — IDLE, guest 9201 (demo-hp), agent 0.131.0, controller 0.243.0
pre health: 200
killed at 2026-09-15T08:53:27Z
dashboard health 200 again at 2026-09-15T08:54:26Z — 59 s after the kill (last non-200 seen: 1)
Sep 15 10:53:52 demo-hp felhom-agent[1526161]: time=2026-09-15T10:53:52.720+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status=exited seen=1 of=2
Sep 15 10:54:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:54:22.744+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exited unit=felhom-controller-bootstrap.service
Sep 15 10:54:24 demo-hp felhom-agent[1526161]: time=2026-09-15T10:54:24.016+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps"
Sep 15 10:54:24 demo-hp felhom-agent[1526161]: time=2026-09-15T10:54:24.016+02:00 level=INFO msg="controller-supervisor: alive" sweeps_since_boot=20 guests_evaluated=1 controllers_not_running=1
gitea.dooplex.hu/admin/felhom-controller:0.243.0 running started=2026-09-15T08:54:23.859136708Z policy=unless-stopped
@@ -0,0 +1,11 @@
session 79 csrf 64
deploy POST http=202 at 2026-09-15T09:04:55Z
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
killed at 2026-09-15T09:05:00Z (5 s into the deploy)
health 200 at 2026-09-15T09:09:08Z — 248 s after the kill
Sep 15 11:09:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:09:22.719+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:09:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:09:52.697+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:10:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:10:22.739+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
## teardown: remove http=502 at 2026-09-15T09:10:40Z
## CORRECTION 2026-09-15T09:11:57Z: the 'health 200 at 09:09:08Z — 248 s after the kill' line is WRONG. docker inspect: felhom-controller exited 09:05:02Z (exit 137, the kill) and was NOT restarted. What answered 200 once is unexplained. What happened instead, and it is the guard working as designed: the agent had restarted this controller at 08:54:24Z, 08:56:53Z and 08:58:24Z (my idle, unpark and failed-swap tests; the swap's own rollback at 09:01:57Z was correctly NOT counted), so on the second not-running sweep after this kill it logged 'CRASH-LOOP … restarts_in_window=3' at 09:05:52Z and paused restarts for 30 min. Mid-deploy restart timing is therefore NOT measured here; the resume after the pause is captured separately.
@@ -0,0 +1,15 @@
## A.4 kill 2 — PARKED, guest 9201
total 4
drwx------ 2 100000 100000 4096 Aug 21 18:00 bootstrap
-rw-r--r-- 1 root root 0 Sep 15 10:54 controller-parked
felhom-controller
parked + killed at 2026-09-15T08:54:49Z
## 2026-09-15T08:56:29Z after 100 s parked: health=502 container=exited
Sep 15 10:55:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:55:22.724+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
Sep 15 10:55:52 demo-hp felhom-agent[1526161]: time=2026-09-15T10:55:52.720+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
Sep 15 10:56:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:22.751+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
unparked at 2026-09-15T08:56:31Z
health 200 at 2026-09-15T08:56:56Z — 25 s after unpark
Sep 15 10:56:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:22.751+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
Sep 15 10:56:52 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:52.716+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exited unit=felhom-controller-bootstrap.service
Sep 15 10:56:53 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:53.962+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps"
@@ -0,0 +1,21 @@
token length: 43 endpoint: 169.254.253.1:8443
swap POST at 2026-09-15T08:57:20Z: 400
felhom-controller
## 2026-09-15T09:00:01Z swap status: Client sent an HTTP request to an HTTPS server.
Sep 15 10:57:52 demo-hp felhom-agent[1526161]: time=2026-09-15T10:57:52.700+02:00 level=INFO msg="controller-supervisor: controller observed not running — confirming on the next sweep" vmid=9201 status=exited seen=1 of=2
Sep 15 10:58:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:58:22.743+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exited unit=felhom-controller-bootstrap.service
Sep 15 10:58:24 demo-hp felhom-agent[1526161]: time=2026-09-15T10:58:24.047+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps"
gitea.dooplex.hu/admin/felhom-controller:0.243.0 running started=2026-09-15T08:58:23.90139374Z
## CORRECTION 2026-09-15T09:00:18Z: the block above is NOT a swap test — the endpoint has no scheme, curl spoke HTTP to the HTTPS local API, the agent answered 400 and no swap started. It is one more IDLE kill (killed 08:57:30Z, restarted 08:58:24Z). Re-run with https:// follows.
token length: 43 endpoint: https://169.254.253.1:8443
swap POST at 2026-09-15T09:00:18Z: 202
felhom-controller
## 2026-09-15T09:03:00Z swap status: {"ok":true,"data":{"current":"gitea.dooplex.hu/admin/felhom-controller:0.243.0","error":"new controller did not become healthy within timeout","in_flight":false,"previous":"gitea.dooplex.hu/admin/felhom-controller:0.243.0","state":"failed","target":"gitea.dooplex.hu/admin/felhom-controller:0.243.0"}}
Sep 15 11:00:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:00:22.591+02:00 level=INFO msg="controller-swap: image file written, restarting bootstrap" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.243.0
Sep 15 11:00:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:00:52.731+02:00 level=INFO msg="controller-supervisor: controller is not running during a controller SWAP — the swap owns it" vmid=9201 status=exited
Sep 15 11:01:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:22.734+02:00 level=INFO msg="controller-supervisor: controller is not running during a controller SWAP — the swap owns it" vmid=9201 status=exited
Sep 15 11:01:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:52.712+02:00 level=INFO msg="controller-supervisor: controller is not running during a controller SWAP — the swap owns it" vmid=9201 status=exited
Sep 15 11:01:55 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:55.784+02:00 level=WARN msg="controller-swap: health verdict negative — rolling back" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.243.0
Sep 15 11:01:55 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:55.784+02:00 level=WARN msg="controller-swap: rolling back" vmid=9201 previous=gitea.dooplex.hu/admin/felhom-controller:0.243.0 reason="new controller did not become healthy within timeout"
Sep 15 11:02:06 demo-hp felhom-agent[1526161]: time=2026-09-15T11:02:06.759+02:00 level=INFO msg="controller-swap: rolled back to previous, controller healthy" vmid=9201 previous=gitea.dooplex.hu/admin/felhom-controller:0.243.0
gitea.dooplex.hu/admin/felhom-controller:0.243.0 running started=2026-09-15T09:01:57.708483159Z
@@ -0,0 +1,11 @@
## 2026-09-15T08:47:43Z 9201 (hand-set by the operator 2026-09-15):
0
2026/09/15 08:47:09 infra.go:74: [WARN] [infra] filebrowser admin probe: Post "http://filebrowser:80/api/auth/login?username=admin": dial tcp: lookup filebrowser on 192.168.0.250:53: no such host (will retry)
## public, through Cloudflare from DooPlex:
files.enkisfelhom.hu login X-Password=admin -> 401
files.enkisfelhom.hu login X-Password=wrong-pw-x -> 401
## 2026-09-15T08:52:15Z state: "filebrowser_admin_state": "operator"
"filebrowser_admin_decided_at": "2026-09-15T08:52:10Z"
0
2026/09/15 08:47:09 infra.go:74: [WARN] [infra] filebrowser admin probe: Post "http://filebrowser:80/api/auth/login?username=admin": dial tcp: lookup filebrowser on 192.168.0.250:53: no such host (will retry)
2026/09/15 08:52:10 filebrowser_password.go:90: [INFO] [infra] filebrowser: admin/admin is refused (HTTP 401) — the password was set by someone; leaving it and recording "operator"
@@ -0,0 +1,74 @@
are supported and installed on your system.
## 2026-09-15T08:16:29Z 9202 BEFORE: controller=gitea.dooplex.hu/admin/felhom-controller:0.242.0
Unable to find image 'curlimages/curl:8.10.1' locally
8.10.1: Pulling from curlimages/curl
43c4264eed91: Pulling fs layer
b68d62cb323c: Pulling fs layer
4ca545ee6d5d: Pulling fs layer
43c4264eed91: Verifying Checksum
43c4264eed91: Download complete
4ca545ee6d5d: Verifying Checksum
4ca545ee6d5d: Download complete
b68d62cb323c: Verifying Checksum
b68d62cb323c: Download complete
43c4264eed91: Pull complete
b68d62cb323c: Pull complete
4ca545ee6d5d: Pull complete
Digest: sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b
Status: Downloaded newer image for curlimages/curl:8.10.1
admin/admin -> 200 wrong -> 401
## 2026-09-15T08:16:34Z switched to 0.243.0; waiting for the base-stack tick
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
sh: 1: Syntax error: Unterminated quoted string
## 2026-09-15T08:21:37Z settings:
admin/admin -> 401 wrong -> 401
2026/09/15 08:16:35 filebrowser_password.go:139: [INFO] [infra] filebrowser: admin/admin replaced by a generated password (value never logged); verified new=200 admin=401
2026/09/15 08:18:34 manager.go:558: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml
2026/09/15 08:20:34 manager.go:558: [DEBUG] [stacks] ScanStacks: found stack "filebrowser" deployed=false composePath=/opt/docker/stacks/filebrowser/docker-compose.yml
6d206504c99e filebrowser filebrowser gtstef/filebrowser:1.3.3-stable
2026/09/15 08:21:34 dbdump.go:169: [DEBUG] DiscoverDatabases: skipping container filebrowser (image=gtstef/filebrowser:1.3.3-stable, not a database)
gitea.dooplex.hu/admin/felhom-controller:0.243.0 running
state:
decided_at:
enc stored: 0
plaintext-looking 16-char field present: 0
settings files: /opt/docker/felhom-controller/data/settings.json
== /opt/docker/felhom-controller/data/settings.json
"filebrowser_admin_state": "generated"
"filebrowser_admin_decided_at": "2026-09-15T08:16:35Z"
1
## 2026-09-15T08:40:19Z via LAN https://192.168.0.114 (the endpoints the UI calls):
app page: card/reveal/admin counts as above
reveal -> password length 16 expected
files login: revealed=401 admin/admin=401 wrong=401
## CORRECTION 2026-09-15T08:40:31Z: the 08:40:19Z block above is INVALID — the dashboard password loaded as an empty string (read_credential.py call wrong), so there was no login, no session and no reveal; 'revealed=401' was an EMPTY password. Only admin/admin=401 and wrong=401 stand from that block. Re-run follows.
## RE-RUN 2026-09-15T08:40:48Z via LAN https://192.168.0.114, the endpoints the UI calls (login → /apps/filebrowser → POST reveal → FileBrowser login):
files login: revealed=200 admin/admin=401 wrong=401
@@ -0,0 +1,3 @@
## whole-guest section text:
Rendszermentés (teljes mentés) A teljes szerver — alkalmazások, beállítások és adatbázisok együtt — időszakos mentése, amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-15 06:26 (4 órája) 0 B Helyi tároló (local) Naprakész Következő mentés 4 órája — a mentési ablakon belül – Visszaállítás ellenőrizve Még nem futott Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-15 06:26 (4 órája) Naprakész Biztonsági szerver – külön hardver (PBS) ✓ Utolsó sikeres mentés: 2026-09-15 06:11 (4 órája) Naprakész Mentés most A mentés alatt az alkalmazások leállnak — általában néhány perc, nagyobb adatnál több.
## remote tile: lass="stat-label">Adatmentés aktív ✓ Távoli rendszermentés külön hardveren (PBS) <div class=
@@ -0,0 +1,22 @@
## 2026-09-15T08:33:09Z operator yes 2026-09-15; scope: datastore felhom-offsite, ns tester-1 ONLY
--- BEFORE (API list, ns tester-1):
"backup-id": "9201",
"backup-time": 1789401893,
"backup-type": "ct",
"size": 402
"size": 1927941845
"size": 877743
"size": 682
"owner": "felhom@pbs!tester-1",
"size": 1928820672,
--- FORGET ct/9201/1789401893 in ns tester-1:
null
rc=0
--- AFTER (API list, ns tester-1):
[]
--- AFTER (namespace dir):
drwxr-xr-x 4096 /mnt/pbs-datastore/ns/tester-1
drwxr-xr-x 4096 /mnt/pbs-datastore/ns/tester-1/ct
--- token still present (permissions):
Path: /datastore/felhom-offsite/tester-1
- Datastore.Backup (*)
@@ -0,0 +1,20 @@
## 2026-09-15T08:10:26Z ep0 read-only listing, scope: datastore felhom-offsite config + namespace tester-1 + token felhom@pbs!tester-1 ONLY
datastore felhom-offsite path: /mnt/pbs-datastore
--- namespace dir: /mnt/pbs-datastore/ns/tester-1
drwxr-xr-x 4096 2026-09-14T16:04:54.5042400850Z /mnt/pbs-datastore/ns/tester-1
drwxr-xr-x 4096 2026-09-14T16:04:54.5042400850Z /mnt/pbs-datastore/ns/tester-1/ct
drwxr-xr-x 4096 2026-09-14T16:04:54.5042400850Z /mnt/pbs-datastore/ns/tester-1/ct/9201
drwxr-xr-x 4096 2026-09-14T16:15:51.2697248430Z /mnt/pbs-datastore/ns/tester-1/ct/9201/2026-09-14T16:04:53Z
-rw-r--r-- 4216 2026-09-14T16:05:55.6363832670Z /mnt/pbs-datastore/ns/tester-1/ct/9201/2026-09-14T16:04:53Z/catalog.pcat1.didx
-rw-r--r-- 1180 2026-09-14T16:05:58.4043897160Z /mnt/pbs-datastore/ns/tester-1/ct/9201/2026-09-14T16:04:53Z/client.log.blob
-rw-r--r-- 682 2026-09-14T16:15:51.2697248430Z /mnt/pbs-datastore/ns/tester-1/ct/9201/2026-09-14T16:04:53Z/index.json.blob
-rw-r--r-- 402 2026-09-14T16:04:54.5642402260Z /mnt/pbs-datastore/ns/tester-1/ct/9201/2026-09-14T16:04:53Z/pct.conf.blob
-rw-r--r-- 27816 2026-09-14T16:05:55.5323830240Z /mnt/pbs-datastore/ns/tester-1/ct/9201/2026-09-14T16:04:53Z/root.pxar.didx
-rw-r--r-- 20 2026-09-14T16:04:54.5042400850Z /mnt/pbs-datastore/ns/tester-1/ct/9201/owner
--- apparent size of index files under the namespace (chunks are shared, deduplicated):
34316 /mnt/pbs-datastore/ns/tester-1
--- token permissions on /datastore/felhom-offsite/tester-1:
Privileges with (*) have the propagate flag set
Path: /datastore/felhom-offsite/tester-1
- Datastore.Backup (*)
@@ -0,0 +1,15 @@
session 79 csrf 64
2026-09-15T09:10:41Z deploy http=202
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
2026-09-15T09:10:53Z vaultwarden health: healthy
container env SIGNUPS_ALLOWED: false
stranger register via traefik (vault.enkisfelhom.hu): 422 body: <!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="color-scheme" content="light dark">
<title>422 Unprocessable Entity</titl
app page: invite fragment 'Invite User'=1 'meghiv'=1; negative control 'Hozd l'=0
## NOTE 2026-09-15T09:11:25Z: the 'stranger register … 422' line above is NOT a proof of refusal — 422 means the request body did not parse. Re-run with the spike's exact body follows.
## RE-RUN 2026-09-15T09:12:20Z stranger send-verification-email via traefik (vault.enkisfelhom.hu, catalog-deployed, SIGNUPS_ALLOWED=false): HTTP 400 message: Registration not allowed or user already exists
@@ -0,0 +1,19 @@
## 2026-09-15T09:21:32Z (a) restart case at 128M: docker oom events, last 15 min
container now: false restarting restarts=11
81: memory: 1280M
134: memory: 1280M
## 2026-09-15T09:25:05Z (b) cap restored: 1342177280 starting restarts=10 oom=false
memory hog exec rc=137 (137 = killed)
## 2026-09-15T09:25:10Z (c) child-process OOM inside the RUNNING container:
container: oomkilled=false status=restarting restarts=11
docker oom events since 2026-09-15T09:25:05Z:
## INTERPRETATION (2026-09-15)
- In this unprivileged LXC guest (Docker 29.8.0), NEITHER Docker signal reported the kills: State.OOMKilled stayed false in the
restart shape (128M cap, 11 restarts) and in the child-process shape (memory hog rc=137, container then restarting);
`docker events --filter event=oom` returned nothing for either window.
- The controller v0.243.0 check reads State.OOMKilled only. It covers the shape BIGNIGHT measured on VM 333
(oomkilled=true, restarts=0, container running) and is NOT proven live here. Filed as a row: a reliable OOM signal
(cgroup memory.events oom_kill read by the agent, or a restart-count trend) is a new mechanism.
- MISTAKE, stated: restoring the cap with `sed s/memory: 128M/memory: 1280M/` also rewrote the redis service's 128M cap
(compose line 134). Moot: the throwaway stack was removed right after (teardown-9202.txt); the catalog template in git
was never touched.
@@ -0,0 +1,18 @@
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
deploy http=202 at 2026-09-15T09:13:08Z
2026-09-15T09:14:52Z paperless-webserver health: healthy
workers=1 threads=1
api token length 40
## 20 posted in parallel at 2026-09-15T09:15:17Z
## 2026-09-15T09:16:18Z tasks: {'SUCCESS': 20}
documents count: 20
limit=1342177280 oomkilled=false restarts=0
cgroup: /sys/fs/cgroup/system.slice/docker-3e225319a21bae0017e58741f6dd3a949e2939575731e254c78d85d1a6f06636.scope
memory.peak=772370432 events: low 0 high 0 max 0 oom 0 oom_kill 0 oom_group_kill 0 sock_throttled 0
kernel OOM lines: 0
81: memory: 1280M
## 2026-09-15T09:17:08Z forced cap 128M, recreated
## 2026-09-15T09:19:39Z webserver: false running restarts=9 limit=134217728
## 2026-09-15T09:20:54Z controller log:
dashboard tag 'Memória elfogyott' count: 0; negative control 'Hiányzó tárhely' count: 0
@@ -0,0 +1,13 @@
2026-09-15T09:25:45Z remove paperless-ngx: {"ok":false,"error":"stack \"paperless-ngx\" is still running — stop it first before removing"}
http=409
2026-09-15T09:25:45Z remove vaultwarden: {"ok":false,"error":"stack \"vaultwarden\" is still running — stop it first before removing"}
http=409
remaining containers: paperless-webserver paperless-redis paperless-postgres vaultwarden <end>
2026-09-15T09:26:40Z stop paperless-ngx: {"ok":true,"message":"Stack paperless-ngx stop completed"}
http=200
2026-09-15T09:26:41Z stop vaultwarden: {"ok":true,"message":"Stack vaultwarden stop completed"}
http=200
2026-09-15T09:27:11Z remove paperless-ngx: {"ok":true,"data":{"removed":"paperless-ngx","volumes_removed":["paperless-ngx_paperless_data","paperless-ngx_paperless_postgres_data","paperless-ngx_paperless_redis_data"],"hdd_paths_removed":[],"hdd_paths_preserved":["
2026-09-15T09:27:11Z remove vaultwarden: {"ok":true,"data":{"removed":"vaultwarden","volumes_removed":["vaultwarden_vaultwarden_data"],"hdd_paths_removed":[],"hdd_paths_preserved":[]},"message":"Stack vaultwarden removed"}
http=200
2026-09-15T09:27:41Z remaining containers: <end>
+16
View File
@@ -26,6 +26,22 @@
---
## 2026-09-15 — the big night's P1 fixes (agent v0.131.0, controller v0.243.0, hub v0.114.0, catalog templates, ISO 1.27.1 published)
Nine rows closed. **Full original text: `git show <this commit>^ -- documentation/backlog/OPEN-ITEMS.md`.** Evidence folder: `documentation/audits/evidence-p1fixes-2026-09-15/`.
| ID | Title | Shipped | Evidence |
|---|---|---|---|
| **R-493** | [P1-HIGH] There are NO customer-facing install instructions, so a volunteer cannot begin — this blocks inviting anyone. | ISO 1.27.1 published + `felhom.eu/letoltes` live | 2026-09-15: round trip over `https://iso.felhom.eu` sha256 `25637007…c053`, 1 705 322 496 B; `felhom.eu/letoltes` 200 naming 1.27.1 (`documentation/tests/iso-release-1.27.1-2026-09-14/README.md`) |
| **R-495** | [P2-MEDIUM] The public installer asks a stranger four questions nothing answers, and REFUSES its own default on one of them. | answered by the guide; ISO 1.27.1 published 2026-09-15 | same as R-493 |
| **R-496** | [P2-MEDIUM] The box's console tells a stranger, in English and FIRST, to open the Proxmox admin page — and calls the owner passphrase „a jelszavadat”. | ISO 1.27.1 (console Felhom-only) published 2026-09-15 | same as R-493; G15 live proof on VM 332 |
| **R-512** | [P2-MEDIUM] Vaultwarden is installed with open registration, and the one control the page tells the customer to use to close it is read-only. | catalog template (no version change), 2026-09-15 | spike: stranger 400 / invite 200 / invited 200 (E1-vaultwarden-spike.txt); live on 9202 from the catalog: `SIGNUPS_ALLOWED=false`, stranger 400 „Registration not allowed" (E1-vaultwarden-9202-live.txt) |
| **R-513** | [P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.<domain>`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet. | controller v0.243.0 | 9202 (default login): generated, stored encrypted, reveal 200 (16 chars), revealed=200 admin=401 wrong=401; 9201 (hand-set): recorded `operator` 08:52:10Z, untouched; public `files.enkisfelhom.hu` admin → 401 (B1/B4 files) |
| **R-514** | [P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut". | catalog template: 1 worker × 1 thread, 1280M | live on 9202: 20 PDFs at once → 20/20 SUCCESS, memory.peak 772 370 432 B, no OOM (E2-paperless-9202-live.txt). The OOM-visibility half is NOT proven live → R-528 |
| **R-515** | [P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password. | catalog template, 2026-09-15 | `default_creds` removed; first steps point at „Automatikusan generált értékek" (app-catalog commit e6aa443) |
| **R-517** | [P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true. | controller v0.243.0 + agent v0.131.0 | live on 9201: per-tier rows „Helyi tároló (local) ✓ … Naprakész", „PBS ✓ … Naprakész", remote tick on a real PBS success (C4-backup-page-9201.txt). Found live and fixed after the release: an unknown size printed „0 B" → „–" (controller main d3eacbb, unreleased) |
| **R-523** | [P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule. | agent v0.131.0 + hub v0.114.0 (+ golden script `--restart always`, no bake) | measured first: kill leaves both `unless-stopped` and `always` exited (A1). Live on 9201: idle kill → dashboard 200 in 59 s; parked → stayed dead 100 s with the PARKED line, unpark → 200 in 25 s; kill during a swap → supervisor deferred ×3, swap rolled back itself; the crash-loop guard tripped for real after 3 test restarts and the hub mailed `controller_crashloop` (A4-* files). NOT measured: restart timing during a deploy (the budget was spent) → R-531 |
| **R-442** | **`remove_hdd_data: true` was INERT — the customer's data stayed on the drive while the API reported success (HTTP 200, `hdd_paths_removed: null`, 128 MB of Nextcloud left; demo-hp 2026-09-01).** Shipped in controller **v0.236.0** (2026-09-13). Removal resolved the drive from the GLOBAL `cfg.Paths.HDDPath` (no default, set on NO box — demo-hp AND demo-felhom both measured 0 `hdd_path` / 0 `FELHOM_PATHS_*`, so the fleet shares the shape); it now reads the app's OWN `app.yaml` `HDD_PATH` — the `07-backup-architecture.md` ~L437 rule that deploy, the start gate and the backup destination already implemented — and a data removal it cannot resolve, or whose drive is absent, is REFUSED (409, exact Hungarian sentence, typed `stacks.RemoveRefusedError`) BEFORE `compose down`, app kept. SSD app → `hdd_paths_removed: []`, never `null`, plus `hdd_note`. Missing folders stated in `hdd_paths_missing`. The backup-half refusal reaches the response (`backup_paths_refused`) and its base follows the same rule. **Reasoning kept:** *"declares no drive" ≠ "could not resolve the drive" — the first is a fact, the second a refusal*; *an app gone with its data left behind is unrecoverable from the UI — the customer cannot even re-run the removal*; *no fallback to the global — that silent fallback is the exact path this closes*. **Observation carried:** 8 of the 13 `needs_hdd` catalog apps bind ONLY `${USERDATA_PATH}` (the shared library) and no `${HDD_PATH}` folder, so for them "delete my data" correctly removes nothing on the drive and the modal shows no checkbox. | **CLOSED 2026-09-13 — shipped controller v0.236.0, proven live on demo-hp** | `audits/R442-2026-09-13/` — A: 63 MB written by Nextcloud ITSELF, gone after removal and listed with its size; C: 409 + sentence, all 69 files untouched, app still deployed, `[ERROR] … refused` logged; D: gokapi `[]` + note; ASCII controls (`llap`: C=2 A=0 D=0). 15 tests + two red-proofs in `felhom-controller/REPORT.md`. Full original text: `git show d6837d98ee24:documentation/backlog/OPEN-ITEMS.md`. |
| **R-449** | **UPDATE ARC SLICE 5 — an upgrade test that runs again. BUILT AND RUN 2026-09-06:** `app-catalog-felhom.eu/scripts/upgrade-test.py` + `upgrade_fixtures.py`, 7 edges across 3 apps, evidence in `audits/upgrade-spike-2026-09-06/`. **Reasoning kept — success is an APPLICATION-LEVEL READBACK, never file identity:** `survive2.py`'s sha256+inode rule is right for a redeploy and WRONG for an upgrade, because a migration is supposed to rewrite files and that rule would fail every correct upgrade. **Reasoning kept — nothing is ever seeded into a volume by hand (R-156);** an app with no non-browser route is recorded `inconclusive`, which is a result and not a licence to plant a file. **Reasoning kept — run the negative control FIRST:** C3's TO image exits immediately and came back `failed`; a harness that cannot fail a known-broken upgrade proves nothing with its greens. **WHAT IT MEASURED:** all five real catalog upgrades kept the customer's data; and **whether an upgrade can be undone is a property of the individual APP, not of upgrades** — docmost REFUSES (*"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*), privatebin does not, which independently reproduces the Nextcloud finding on a second app by a DIFFERENT mechanism and puts two measurements behind §4's ruling that "rollback" is the wrong word. **It also found a defect in our own catalog (R-459).** **What stays open, as its own rows rather than inside this one:** R-459 (the skipped MariaDB datadir upgrade), R-460 (bookstack's file half is unprovable headlessly), R-462 (the widening, costed). Full original text: `git show 417df06f3529:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — harness built, run, and proven by a red negative control** | `audits/SPIKE-upgrade-test-2026-09-06.md`; `audits/upgrade-spike-2026-09-06/evidence/`; catalog `0474ce387e6f` |
| **R-438** | **The catalog sync rewrote a DEPLOYED app's `docker-compose.yml` and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06).** `Syncer.copyTemplates` copied into every stack folder on a 15-minute cycle with **no deployed check**, so a deployed app's file and its running containers disagreed from that moment, and the next `compose up -d` from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s **with a pull**, and a boot reconciliation upgraded it **with nobody pressing anything**. **Reasoning kept — the distinction this row existed to protect:** `RestartStack`'s use of `up -d` to pick up template changes was **CHOSEN and written down in its own comment**, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. **Reasoning kept — one fear was measured SMALLER than stated:** a plain power cut does NOT upgrade anything, because Docker's `restart: unless-stopped` restores the containers on the old image and the reconciler logs `no boot-orphaned apps`; the unattended upgrade needs the narrower precondition *"and the app did not come back"*. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0** | `audits/SPIKE-app-update-2026-09-01.md` §2, §3, §8; `architecture/09-update-architecture.md`; `tests/VALIDATION-update-slice3-2026-09-06.md` |
+13 -13
View File
@@ -696,10 +696,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). **PARTLY SHIPPED in v0.242.0 (`d698ce3`), measured live on 9202 the same night:** the difference is computed and a fresh compose-created volume IS reported (`["opengist_opengist_data"]`), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (`docker volume create <name>`, `restore.go:154`) carries no labels — compose still removes it and the response says `[]` (`audits/v0242-2026-09-14/19-R489-cause.txt`). **Remaining fix:** list by the `<project>_` name prefix as well (union), or label the recreated volume as compose would. | **READY — rank P3-LOW; owner: CC (residual)** |
| **R-492** | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** |
| **R-493** | **[P1-HIGH] There are NO customer-facing install instructions, so a volunteer cannot begin — this blocks inviting anyone.** MEASURED 2026-09-14 (drill `audits/DRILL-fresh-install-0242-2026-09-14.md`, Phase 0.2), each source with what was searched: the website's 9 pages carry no mention of the installer, of `iso.felhom.eu`, or of any install step (`iso`, `letolt`, `telepit`); `https://iso.felhom.eu/` and `/index.html` both return **404** — only the exact object name `felhom-installer-1.26.1-pve9.2-1.iso` answers, so the file cannot be found without being told its name; `RUNBOOK-onboarding-draft-v4.md` is an operator-attended script; the R-11 tester one-pager was ruled 2026-07-21 and never written; no `email.md` exists in the workspace; the two hub mails a new customer receives (`Kösd össze a Felhom dobozodat`, `Elindult a Felhom szervered — beállító kód`) assume the box is already installed. **Stopgap written by the drill:** `documentation/runbooks/VOLUNTEER-first-hour.md` (Hungarian, the steps the product really needs, each addition over today listed at its top). **What it needs:** the operator to choose the channel and approve the text. **2026-09-14:** guide approved by the operator and aligned to ISO 1.27.1; download page planned at `felhom.eu/letoltes` (R-504). Stays open until the page and ISO are live. | **WAITING-ON-OPERATOR — publish yes; rank P1-HIGH** |
| **R-494** | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** |
| **R-495** | **[P2-MEDIUM] The public installer asks a stranger four questions nothing answers, and REFUSES its own default on one of them.** MEASURED 2026-09-14 on `felhom-installer-1.26.1` (drill screens `s04`–`s22`): every screen after the GRUB menu is **English** Proxmox (EULA, disk, locale, password, network, summary); **three disks are offered with no guidance which is the system disk** (the default happened to be right); the administrator e-mail is prefilled `mail@example.invalid`; the hostname is prefilled `pve.example.invalid` and pressing Next on it returns **„Invalid values: hostname does not look valid”** — a volunteer who accepts the defaults cannot continue; at the end, with the stick still in, the default-ticked auto-reboot boots back into the installer. None of this is wrong for Proxmox; all of it is unanswered for a Felhom household. **Fix shape (cheap first):** the volunteer instructions answer each question (done in `runbooks/VOLUNTEER-first-hour.md`); later, prefill hostname and e-mail from the ISO build. **ANSWERED 2026-09-14 by operator ruling + guide, not by code:** the interactive installer stays (the 2026-07-31 ruling re-affirmed); `VOLUNTEER-first-hour.md` answers every screen (disk, keyboard, password, e-mail, host name, stick) and release gate G14 pins the disk rule. The English screens remain by ruling. Closes when the guide and ISO are published. | **ANSWERED — awaiting publish; owner: operator (publish)** |
| **R-496** | **[P2-MEDIUM] The box's console tells a stranger, in English and FIRST, to open the Proxmox admin page — and calls the owner passphrase „a jelszavadat”.** MEASURED 2026-09-14 (drill screen `s29`): above the Hungarian pairing banner the console prints Proxmox's own `Welcome to the Proxmox Virtual Environment. Please use your web browser to configure this server - connect to: https://192.168.0.134:8006/` — the operator admin UI, which a household must never be sent to; the Felhom banner below says „add meg ezt a kódot és **a jelszavadat**”, while the self-bind mail and page name the same secret **„Tulajdonosi jelmondat”** (R-323 renamed the mail; `scripts/iso/felhom-bootstrap.sh:71-72` was not renamed). **Fix shape:** replace the Proxmox `/etc/issue` block on appliance installs; rename the console word. Both live in `felhom-bootstrap.sh`, a frozen ISO payload (G9), so it ships with the next ISO. **FIXED IN ISO 1.27.1, NOT YET PUBLISHED (2026-09-14).** 1.27.0 masked `pvebanner` at first boot and still showed the Proxmox block on that boot (VM 331 screen s20); 1.27.1 masks by symlink in the postinst and writes a no-ő/ű `/etc/issue`. Live on VM 332: first-boot console Felhom-only; after a proven reboot `pvebanner` masked, `/etc/issue` 0 × 8006. Harness 55/55, red first. Closes when 1.27.1 is published. | **FIXED — awaiting publish (operator yes); owner: CC** |
| **R-497** | **[P2-MEDIUM] No product channel ever gives the customer the „Tulajdonosi jelmondat”, yet the self-bind mail says they received it at setup.** MEASURED 2026-09-14: the 5-word phrase is minted at customer creation (`hub/internal/web/configs.go`, `RandomPassphrase(5)`) and shown **only on the operator's customer page**; the self-bind mail (`FormatSelfBindEmail`) says „amelyet a beállításkor kaptál” and the console asks for it; nothing sends it. Day-0 runbook A.2 says the operator must dictate or hand it over — a step that exists only in an operator document. A volunteer whose operator forgets cannot bind the box and is told they already have the phrase. **Fix shape:** the operator's customer-create confirmation says, in one line, to hand this phrase to the customer now; the volunteer instructions say where it comes from (done in `VOLUNTEER-first-hour.md`). **CLOSED 2026-09-14 — hub v0.113.0 deployed (ArgoCD Synced, image 0.113.0). Tests red first (`passphrase_handover_test.go`); live: `tester-1`'s page renders the hand-over sentence 1× (2× with `?flash=created`), negative control 0. The mail change is pinned by test; no mail could be read live (no mailbox).** Evidence: `audits/DOORSTEP-walk-1270-2026-09-14.md` §5. | **CLOSED — hub v0.113.0** |
| **R-498** | **[P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal `wiki.DOMAIN` — the placeholder is never filled in.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a **wiki.DOMAIN** címet a böngészőben", PrivateBin's „Nyisd meg a **paste.DOMAIN** címet". `grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml` → **52 of 53** templates carry it in `first_steps`; the controller renders the string as text (`internal/stacks/metadata.go`). A stranger reading their first instruction meets a word that is not an address. **Fix shape:** substitute the stack's real `SUBDOMAIN.DOMAIN` at render time (one place in the controller), with a render test per template that fails on a literal `DOMAIN`. | **READY — rank P3-LOW; owner: CC** |
| **R-499** | **[P2-MEDIUM] Every app without a data drive is told its data is „already in the full system backup (PBS)" and there is „nothing to do" — on a box with no PBS, whose only whole-box copy sits on the same disk.** MEASURED 2026-09-14 on a fresh 0.242.0 box with the DR tier off (drill step 7): `GET /stacks/bookstack/backup` renders „Ennek az alkalmazásnak az adatai a belső rendszerlemezen vannak, amelyek **már szerepelnek a teljes rendszermentésben (PBS)** … ehhez az alkalmazáshoz nincs külön teendő." The sentence sits under `{{if not .IsHDDApp}}` in `controller/internal/web/templates/tier2_config.html:20-26` and consults nothing about where the whole-guest backup goes. On the same box `/backups` says, correctly, „Helyi tároló (local)" and „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem." **Two pages of one product contradict each other, and the reassuring one is the false one.** **Fix shape:** branch the sentence on the box's actual whole-guest target (PBS vs local, same-disk flag the overview already computes); a render test per branch. | **READY — rank P2-MEDIUM; owner: CC** |
@@ -712,22 +709,25 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-506** | **[P3-LOW] `day0-install.md` A.1 says "the controller manages per-app hostnames itself via the tunnel" — it does not.** MEASURED 2026-09-14: no code under `felhom-controller/controller/internal` creates tunnel ingress, DNS records or tunnel configurations (`grep -i 'ingress\|cfd_tunnel\|/configurations\|dns_records'` → only comments saying cloudflared is deployed when a token exists; positive control: the geo-restriction CF API use IS found in `cmd/controller/main.go`). The controller's only Cloudflare act is geo-restriction; the tunnel runs `tunnel run` with the token, so its routes come from Cloudflare's remote config set by the operator. A reader following A.1 skips the one step that makes the dashboard reachable (R-505). **Fix shape:** A.1 names the public-hostname step and its service settings, copied from a working tunnel. | **READY — rank P3-LOW; owner: CC (doc), operator (the settings to copy)** |
| **R-507** | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** |
| **R-508** | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** |
| **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** |
| **R-510** | **[P1-HIGH] `tester-1`'s tunnel now has its route, and still gives a fresh box 502: the route sends traffic to `https://traefik` WITH certificate checking, and traefik answers the name `traefik` with its default certificate.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0), after the operator's R-505 fix: from DooPlex `https://felhom.enkicsifelhom.hu` → **502 ×3** (18:07:48Z, `server: cloudflare`). The box's `cloudflared` logs `Request failed … tls: failed to verify certificate: x509: certificate is valid for 544346c4….traefik.default, not traefik … ingressRule=0 originService=https://traefik`. Traefik itself holds a valid Let's Encrypt `CN=*.enkicsifelhom.hu` (openssl on 127.0.0.1:443 with SNI). **Control:** demo-hp's working tunnel config reads `{"hostname":"*.enkisfelhom.hu", "originRequest":{"noTLSVerify":true}, "service":"https://traefik"}` — the same route WITH `noTLSVerify`. So the Cloudflare-side public hostname for `*.enkicsifelhom.hu` lacks „No TLS Verify" (or an origin server name). The Cloudflare side is not visible to the session; the inference rests on the log line and the control. Intervention **I2** of the big night: the claim and every dashboard request go to the guest's LAN address with the name forced. **Fix:** operator ticks „No TLS Verify" on that public hostname, then day-0 A.1 names the setting beside the route. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (Cloudflare route), CC (day-0 A.1 wording after)** |
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-512** | **[P2-MEDIUM] Vaultwarden is installed with open registration, and the one control the page tells the customer to use to close it is read-only.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, controller 0.242.0, catalog vaultwarden 1.36.0-alpine): the deploy form sends `SIGNUPS_ALLOWED=true` (catalog default); after install, „Vaultwarden — Beállítások" says „Ez az alkalmazás már telepítve van. **Az alábbi beállítások csak olvashatók.**" and, below it, „Regisztráció engedélyezése — Igen / Nem — Új fiókok regisztrálásának engedélyezése. **Az első fiók létrehozása után állítsd 'Nem'-re.**" At 18:27:20Z a stranger with no invite and no login registered `idegen.probe@example.com` → **200** (`phase3/vaultwarden-stranger-signup.txt`). Once the dashboard is reachable through the tunnel, anyone who guesses `vault.<domain>` can open an account on the household's password server. A route does exist (the Vaultwarden admin panel's own setting, token under „Megjelenítés"), and no screen names it. **Fix shape:** default `SIGNUPS_ALLOWED=false` with an invite-first first step, or make that one field editable after install; the page must not instruct an act it forbids. | **READY — rank P2-MEDIUM; owner: CC (catalog + controller)** |
| **R-513** | **[P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.<domain>`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet.** MEASURED 2026-09-14 (BIGNIGHT): `POST /api/auth/login?username=admin` with `X-Password: admin` → **200 + a session token** on VM 333 (fresh ISO 1.27.1 install, controller 0.242.0) and on demo-hp guests **9201 and 9202** (loopback, `Host: files.enkisfelhom.hu`); negative control `admin` / wrong → **401** on all three. On VM 333 that token lists both sources — „Adatlemez" (the data drive's `userdata`: documents, media, photos…) and „Beolvasás" (with `paperless`) — `GET /api/users?id=self` 200. **Public exposure, measured by GET only:** `https://files.enkisfelhom.hu/` from DooPlex through Cloudflare → 200, FileBrowser Quantum, `passwordAvailable:true, noAuth:false` (`phase3/filebrowser-public-reachability-demo-hp.txt`); no login was attempted over the internet. The generated `config.yaml` sets no admin credential (FileBrowser's own default applies); no screen shows the customer any FileBrowser login. Geo-restriction narrows who can reach it, it does not authenticate. **Not changed tonight** (9201's standing state is fenced; no product code). **Fix shape:** the controller sets a generated admin password (or proxy auth behind the dashboard session) at stack creation and on every existing box, and shows it where the customer finds app credentials. | **READY — rank P1-HIGH; owner: CC (controller) · operator (rotate on live boxes first)** |
| **R-514** | **[P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, catalog paperless-ngx 2.20.15, `paperless-webserver` limit 805 306 368 B): 20 small 3-page PDFs posted through `/api/documents/post_document/` at 18:29:57Z. At 18:31:22Z the VM kernel logged `Memory cgroup out of memory: Killed process … (gs)` ×2, `([celeryd: celer)`, `([celery beat] -)` — constraint MEMCG of that container; `docker inspect` → `oomkilled=true restarts=0`. Paperless's own task list: **11 FAILURE (`WorkerLostError`), 1 STARTED, 8 PENDING**, unchanged 13 minutes later; **0 documents**. The controller shows the app running and healthy; no event, no alarm (`phase3/paperless-tasks.txt`, `paperless-oom-check.txt`). A household scanning a drawer of bills sees nothing arrive and no reason. **Fix shape:** raise the cap or set `PAPERLESS_TASK_WORKERS=1` / `PAPERLESS_THREADS_PER_WORKER=1` in the template, and let the dead-app/health check see an OOM-killed worker. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
| **R-515** | **[P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password.** MEASURED 2026-09-14 (BIGNIGHT, VM 333): `/apps/paperless-ngx` „Első lépések — Jelentkezz be: admin / admin" and „Alapértelmezett belépés admin / admin"; the catalog template generates `PAPERLESS_ADMIN_PASSWORD` (`password:16`) and the token call with the generated value succeeded (`phase3/seed-paperless.txt`). A household following the page is refused at its first login. Same card also sends documents to „FileBrowser … import/paperless" — the FileBrowser login is R-513. **Fix:** the card points at „Beállítások → Automatikusan generált értékek" as gokapi's does. | **READY — rank P3-LOW; owner: CC (catalog)** |
| **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). **SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL:** triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as `selfbind_link_sent` and shown on the Setup tab. Unit-proven with a red-proof (`TestSelfBind_EmailSetOnWaitingCustomerSendsLink`). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 `deprovision` + Cloudflare), which is fenced without the operator's word. **Closes on:** one real mail from either trigger. | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** |
| **R-510** | **[P1-HIGH] `tester-1`'s tunnel now has its route, and still gives a fresh box 502: the route sends traffic to `https://traefik` WITH certificate checking, and traefik answers the name `traefik` with its default certificate.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0), after the operator's R-505 fix: from DooPlex `https://felhom.enkicsifelhom.hu` → **502 ×3** (18:07:48Z, `server: cloudflare`). The box's `cloudflared` logs `Request failed … tls: failed to verify certificate: x509: certificate is valid for 544346c4….traefik.default, not traefik … ingressRule=0 originService=https://traefik`. Traefik itself holds a valid Let's Encrypt `CN=*.enkicsifelhom.hu` (openssl on 127.0.0.1:443 with SNI). **Control:** demo-hp's working tunnel config reads `{"hostname":"*.enkisfelhom.hu", "originRequest":{"noTLSVerify":true}, "service":"https://traefik"}` — the same route WITH `noTLSVerify`. So the Cloudflare-side public hostname for `*.enkicsifelhom.hu` lacks „No TLS Verify" (or an origin server name). The Cloudflare side is not visible to the session; the inference rests on the log line and the control. Intervention **I2** of the big night: the claim and every dashboard request go to the guest's LAN address with the name forced. **Fix:** operator ticks „No TLS Verify" on that public hostname, then day-0 A.1 names the setting beside the route. **2026-09-15, after the operator's tick:** three GETs from DooPlex 08:09:07–08:09:14Z → **530 ×3** (`server: cloudflare`) — no tunnel is connected, because tester-1 has no box since the BIGNIGHT teardown. The fix cannot be observed until a box exists. `day0-install.md` A.1 now names „No TLS Verify" beside the route, marked unproven. **Closes on:** 200 or the claim page through the tunnel on the next tester-1 box. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (Cloudflare route), CC (day-0 A.1 wording after)** |
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** |
| **R-517** | **[P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, controller 0.242.0, customer `tester-1` with the DR tier ticked but never provisioned — R-511): „Mentés most" at 19:03:23Z; the local tier succeeded (8 877 619 753 B, 362 s); the `felhom-pbs` tier then failed — agent: `could not activate storage 'felhom-pbs': storage 'felhom-pbs' does not exist`; `pvesm status` lists only `local` and `local-lvm`. At 19:15:36Z `/backups` read: „✗ · **Utolsó teljes mentés 2026-09-14 21:09 (5 perce) · 0 B · Biztonsági szerver – külön hardver (PBS) · Naprakész**" and „✓ **Távoli rendszermentés — külön hardveren (PBS)**" (`phase4/backup-pages-after.txt`). The successful 8.9 GB local backup is no longer shown; a 0-byte failed attempt is labelled up to date; the remote tier is ticked as present. The CLAUDE.md rule „presence is not success" in page form: an attempt's timestamp stands in for a result. The hub did raise a true `whole_guest_backup_failed (error)` to the operator; the customer's page says the opposite. **Fix shape:** the tile shows the newest SUCCESSFUL backup per tier, a failed tier as failed, and a tier whose storage does not exist as „nincs beállítva". **Also measured after F2's reboot (19:45:41Z):** the whole-system tile read „– · Utolsó teljes mentés – · Méret / cél · **Naprakész**" — no backup listed at all, still labelled up to date (the 21:03 CEST local success no longer shown). | **READY — rank P1-HIGH; owner: CC (controller)** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** |
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle (`qm reset` 22:08:17Z): the bootstrap started the controller at boot (22:10:23Z), dashboard 200 at +131 s, all 12 apps running at +229 s, the deployed homebox `running · deploying false · deployed true` — „Fut · Naprakész", **not stuck**. Hub `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z with its mail **also suppressed by cooldown**. `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
| **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** |
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-529** | **[P3-LOW] The agent-plane `host_stale` / `host_down` / `host_recovered` mails still wait out the one-hour quiet rule.** The 2026-09-15 ruling (decision A) named `node_*` only, and the task fenced „a design is not a defect". The same F9 silence can happen on the host plane. **What it needs:** the operator's word whether the ruling extends to `host_*`. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** |
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
| **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |
| **R-533** | **[P3-LOW] The operator signing keys were placed on DooPlex world-readable (mode 664) and sit outside any documented location.** OBSERVED 2026-09-15: `/mnt/5_hdd/felhom.eu/felhom-op-operational`, `felhom_op_ed25519`, `felhom-rec-recovery` arrived 664; CC set them to 600 (no other change). Their fingerprints match the signers demo-hp pins. **What it needs:** the operator decides where the keys live between sessions (hardware key, or a documented 0600 path) and records it in `operations/nodes.md`. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
@@ -1,7 +1,16 @@
# ISO release gate — felhom-installer-1.27.1 (2026-09-14)
**NOT PUBLISHED — awaiting the operator's yes.** Built and gated; the proof installs and the first-hour walk are Phase 4 of the
doorstep task; publishing waits for the operator's yes (task Phase 5 STOP).
**PUBLISHED 2026-09-15 08:45Z** on the operator's yes (decision 2026-09-15), by the P1-fixes task, from the exact file
gated here. Round trip over the real URL, 2026-09-15 08:45:55Z: `https://iso.felhom.eu/felhom-installer-1.27.1-pve9.2-1.iso`
→ **1 705 322 496 B, sha256 `25637007…c053`** (identical to the row below); `.sha256` → 200 (103 B); `.manifest.txt` →
200 (1 914 B). **Rollback is one link change:** `felhom-installer-1.26.1-pve9.2-1.iso` (+ `.sha256`, manifest) stays online
and was listed in the bucket before the upload. The download page `https://felhom.eu/letoltes` answered 200 naming 1.27.1
and the sha at 08:46:41Z. **G11 PASS** (above). **G12 NOT MEASURED:** the object-scoped token cannot read bucket settings
(`ListBuckets` 403s); `iso.felhom.eu/` and `/index.html` still 404, consistent with no public bucket listing — an
inference, not a measurement. **No git tag moved:** R-110 governs the host installer (`installer-v*`), which this ISO
does not change; nothing was re-tagged.
*Previously:* NOT PUBLISHED — awaiting the operator's yes (doorstep task Phase 5 STOP).
| | |
|---|---|
@@ -26,8 +35,8 @@ against the exact output file, stock ISO `proxmox-ve_9.2-1.iso` for the enumerat
| G8 postinst | **PASS** | forbidden systemctl verbs `0`, network tools `0`, `set -e` `0`, last line `exit 0` |
| G9 payload = HEAD | **PASS** | `felhom-bootstrap.sh` `973f0c8f…9dd6` = repo; `.service` `cf2e4678…9af3` = repo |
| G10 committed | **PASS** | `scripts/iso/` porcelain empty; `HEAD == origin/main` |
| G11 round trip | *publish-time* | — |
| G12 bucket private | *publish-time* | — |
| G11 round trip | **PASS (2026-09-15)** | downloaded bytes sha256 `25637007…c053`, 1 705 322 496 B = built file |
| G12 bucket private | **NOT MEASURED** | token is object-scoped (`ListBuckets` 403); root and `/index.html` 404 |
| G13 directories | **PASS** | `./etc/felhom/`, `./usr/local/sbin/`, `./lib/systemd/system/` each `1` |
| G14 a person chooses the disk | **static half PASS** | no answer file (G1), `filter.` keys `0`; the proof-install half is the walk |
| G15 console after first boot + reboot | **PASS (live, VM 332)** — first boot Felhom-only (screen `332-b12`); proven reboot (boot 16:14:11Z > `qm reboot` 16:13:56Z); `pvebanner` `masked` → `/dev/null`; `/etc/issue` 0 × `8006`; postinst log shows both acts | `systemctl mask pvebanner.service` present; issue text names `8006` `0`; see `audits/evidence-doorstep-walk-1270-2026-09-14/G15-live-332-iso1271.txt` |
+9
View File
@@ -1,5 +1,14 @@
# website/ changelog (newest on top)
## 2026-09-15 — New page: /letoltes (the installer download)
The volunteer guide sends people to `felhom.eu/letoltes`; the page did not exist. It now names installer
**1.27.1**, its size and SHA-256, how to check the sum on Linux/Mac and Windows, and — first — the disk
warning: the installer erases the disk you choose, never picks one, unplug the backup drive, call the
operator if unsure. `noindex` (it is for people holding the guide). Same skeleton as the other pages, no
nav entry; added to `sitemap.xml` and to `scripts/site_gates.py` PAGES; site gate OK. Published together
with the ISO (round trip verified first), so the link never pointed at a missing file.
## 2026-07-19 — Restored: the index grid background
The subtle grid behind the hero was **not** removed on purpose. It lived as a fixed `body::before` in