docs: R-691 (2) built (07 §6.5), R-701 (a) measured, R-702 claper default admin, R-703 calcom OOM; evidence
gates / gates (push) Successful in 25s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-28 10:37:42 +02:00
parent 03df4c8220
commit a8f790c0c6
39 changed files with 4023 additions and 4 deletions
@@ -96,3 +96,16 @@ claim page: 200
qemu processes: 0
disk reverted to virgin
DooPlex copy of the passphrase shredded
## 2026-09-28T08:35:05Z teardown layer 3 (hub): delete the scratch customer by the hub's own flow
HTTP/1.1 303 See Other
Location: /configs?flash=deleted
2026/09/28 10:35:05 [INFO] customer DELETE cascade started for drill-g0276 (journal #21, 1 host(s))
2026/09/28 10:35:05 [INFO] delete drill-g0276: host drill-g0276-b861fc deleted (escrow DEMOTED to retained custody)
2026/09/28 10:35:06 [INFO] tenantsync: deprovision ok for drill-g0276 (ns=drill-g0276, existed=false)
2026/09/28 10:35:06 [INFO] reset drill-g0276: PBS tenancy deprovisioned
2026/09/28 10:35:06 [INFO] [claim] reset to unclaimed for drill-g0276 (customer RESET) — next onboarding mints a fresh code
2026/09/28 10:35:06 [INFO] delete drill-g0276: residue purged (reports=1 app_telemetry=1 app_log_tails=0 log_tail_requests=0 notif_prefs=0 selfbind_tokens=0 appliance_registrations=0)
2026/09/28 10:35:06 [INFO] customer DELETE cascade COMPLETE for drill-g0276 (journal #21) — full teardown
-- after:
0
0
@@ -0,0 +1,34 @@
## 2026-09-28T08:04:59Z demo-hp 9201 paperless-ngx — read-only
paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 5 hours (healthy)
paperless-redis redis:7-alpine Up 5 hours (healthy)
paperless-postgres postgres:18-alpine Up 5 hours (healthy)
PG_VERSION in the running datadir: 18
server 18.6|documents 0|tags 0|correspondents 0|users 3
paperless-ngx_paperless_data
paperless-ngx_paperless_postgres_data
paperless-ngx_paperless_postgres_data.pre-update-20260928T022253Z
paperless-ngx_paperless_redis_data
-- controller log, conversion lines, last 2 nights:
2026/09/28 01:30:05 tier2.go:425: [INFO] [backup] Tier 2 copied paperless-ngx → /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx (399.8 MB, 1 leg(s), 3s) [SSD: state-only]
2026/09/28 02:22:44 unattended.go:320: [INFO] [update-leg] paperless-ngx: step pressed paperless-postgres=postgres:16-alpine, paperless-redis=redis:7-alpine, paperless-webserver=ghcr.io/paperless-ngx/paperless-ngx:2.20.15 → paperless-post
2026/09/28 02:22:44 update.go:786: [INFO] [stacks] update paperless-ngx: ladder — the last step (1 of 1) — the catalog's current definition
2026/09/28 02:22:44 update.go:800: [INFO] [stacks] update paperless-ngx: this step CONVERTS paperless-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume paperless-ngx_paperless_postgres_data
2026/09/28 02:22:44 update.go:814: [INFO] [stacks] update paperless-ngx: precondition met — Tier 2 (second drive) copy from 2026-09-28T00:30:01Z (1h53m0s old, limit 24h0m0s)
2026/09/28 02:22:46 undo.go:287: [INFO] [stacks] update paperless-ngx: the undo copy will hold 3 named volume(s), 398.1 MiB
2026/09/28 02:22:46 pgconvert.go:340: [INFO] [stacks] update paperless-ngx: conversion space — database volume paperless-ngx_paperless_postgres_data is 361 MB; the dump needs up to 451 MB + the 2 GB floor; 28.6 GB free beside the stack di
2026/09/28 02:22:46 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:53 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:53 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:57 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:57 update.go:1365: [INFO] [stacks] update paperless-ngx: phase converting
2026/09/28 02:22:57 pgconvert.go:423: [INFO] [stacks] update paperless-ngx: CONVERTING paperless-postgres PostgreSQL 16 → 18 (volume paperless-ngx_paperless_postgres_data, kept copy paperless-ngx_paperless_postgres_data.pre-update-2026092
2026/09/28 02:22:59 pgconvert.go:468: [INFO] [stacks] update paperless-ngx: pg_dumpall from 16 done in 591ms — 488 kB, completion line present; before: 2 database(s), 72 table(s), 1537 row(s)
2026/09/28 02:23:00 update.go:1365: [INFO] [stacks] update paperless-ngx: phase converting
2026/09/28 02:23:00 pgconvert.go:484: [INFO] [stacks] update paperless-ngx: database volume paperless-ngx_paperless_postgres_data emptied (its copy paperless-ngx_paperless_postgres_data.pre-update-20260928T022253Z holds the old datadir, mar
2026/09/28 02:23:07 pgconvert.go:513: [INFO] [stacks] update paperless-ngx: loaded into 18 in 1.68s (dropped the entrypoint's empty database(s) [paperless]; skipped CREATE ROLE for [paperless])
2026/09/28 02:23:07 pgconvert.go:529: [INFO] [stacks] update paperless-ngx: CONVERTED 16 → 18 in 9.93s — the check is equal (2 database(s), 72 table(s), 1537 row(s)), PG_VERSION 18
2026/09/28 02:23:49 update.go:1015: [INFO] [stacks] update paperless-ngx: KEEPING the pre-conversion datadir copy paperless-ngx_paperless_postgres_data.pre-update-20260928T022253Z (PostgreSQL 16) until a backup of the converted app is prove
2026/09/28 02:23:54 unattended.go:327: [INFO] [update-leg] paperless-ngx: step ended done after 70.1 s
@@ -0,0 +1 @@
## app list BEFORE Part E (deployed): ['adventurelog', 'bentopdf', 'bookstack', 'calibre-web', 'docmost', 'kimai', 'opengist', 'paperless-ngx', 'privatebin', 'romm']
@@ -0,0 +1,3 @@
2026-09-28T10:15:02 POST /backup/offbox/toggle app=nextcloud enabled=on -> HTTP/2 302 ['location: /backups/remote?flash=flash.offbox.app_setting_updated']
settings app_backup.nextcloud = {'enabled': False, 'offbox': True}
@@ -0,0 +1,10 @@
10:12:58 # kp deploy — 2026-09-28T10:12:55+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:13:01 before: deployed= False | appdata: paperless romm
10:13:04 deploy (no choice) -> 202 {"ok": true, "message": "Telepítés elindítva – az állapot a kártyán követhető"}
10:14:30 deployed after (85.5, 'running')
10:14:33 # kp seed — 2026-09-28T10:14:30+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:14:41 nextcloud: occ user:add :: The account "drill895c04" was created successfully Display name set to "drill895c04"
10:14:45 nextcloud: seeded user drill895c04
10:14:45 seeded (the account): True uid= drill895c04
10:14:50 PUT before-backup.txt -> 201
10:14:50 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
@@ -0,0 +1,10 @@
10:01:57 # kp deploy — 2026-09-28T10:01:53+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:02:00 before: deployed= False | appdata: paperless
10:02:03 deploy (no choice) -> 202 {"ok": true, "message": "Telepítés elindítva – az állapot a kártyán követhető"}
10:03:23 deployed after (80.4, 'running')
10:03:27 # kp seed — 2026-09-28T10:03:23+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:03:35 nextcloud: occ user:add :: The account "drill2bf734" was created successfully Display name set to "drill2bf734"
10:03:38 nextcloud: seeded user drill2bf734
10:03:38 seeded (the account): True uid= drill2bf734
10:03:44 PUT before-backup.txt -> 201
10:03:44 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
@@ -0,0 +1,100 @@
10:03:57 # kp window 10:06 — 2026-09-28T10:03:54+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:03:57 POST /backups/window window_start=10:06 -> 303
10:04:03 2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 11:06 (next run 2026-09-28 11:06 CEST)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 11:51 (next run 2026-09-28 11:51 CEST)
2026/09/28 08:03:57 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run
2026/09/28 08:03:57 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run
2026/09/28 08:03:57 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run
2026/09/28 08:03:57 backup_handlers.go:65: [INFO] [web] backup window set to 10:06 (legs 10:06/11:06/11:51)
10:07:00 # kp logs 15m — 2026-09-28T10:06:57+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:07:04 2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-29 02:30 CEST
2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: conversion-copy-release (every 1h0m0s)
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-29 03:30 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-29 04:15 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-29 05:10 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-29 06:00 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-29 05:30 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-29 04:00 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-29 03:30 CEST
2026/09/28 07:59:29 scheduler.go:206: [INFO] [scheduler] Starting scheduler with 17 jobs
2026/09/28 07:59:29 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 07:59:29 handlers.go:3622: [INFO] [web] FileBrowser: drive /mnt/felhom-drives/scratch_hdd/userdata/nextcloud not mounted — skipping userdata skeleton
2026/09/28 07:59:29 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 handlers.go:3622: [INFO] [web] FileBrowser: drive /mnt/felhom-drives/scratch_hdd/userdata/nextcloud not mounted — skipping userdata skeleton
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:01:29 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan
2026/09/28 08:01:29 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 64ms)
2026/09/28 08:02:03 router.go:431: [INFO] [api] Deploy requested for stack: nextcloud
2026/09/28 08:02:03 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:02:03 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:02:03 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:03 deploy.go:442: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH, DOMAIN, SUBDOMAIN, DB_PASSWORD]
2026/09/28 08:02:03 manager.go:1601: [INFO] [stacks] Deploying stack nextcloud — checking 3 images...
2026/09/28 08:02:09 healthprobe.go:190: [WARN] Health probe nextcloud: API GET :80/status.php → Get "http://nextcloud-redis:80/status.php": dial tcp: lookup nextcloud-redis on 127.0.0.11:53: no such host
2026/09/28 08:02:14 deploy.go:532: [INFO] [stacks] Stack nextcloud deployed successfully (took 10.9s)
2026/09/28 08:02:14 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:14 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:14 installed.go:420: [INFO] [stacks] installed-images nextcloud: recorded 3 service(s) (nextcloud=nextcloud:34.0.4-apache (sha256:a5ace30c695a…), nextcloud-db=mariadb:12.3 (sha256:805c8e104bd5…), nextcloud-redis=redis:7-alpine (sha256:858f009f9709…))
2026/09/28 08:02:14 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:14 pin.go:93: [INFO] [stacks] pin nextcloud: nextcloud=nextcloud:34.0.4-apache, nextcloud-db=mariadb:12.3, nextcloud-redis=redis:7-alpine
2026/09/28 08:02:17 manager.go:1562: [INFO] [stacks] Stack nextcloud post-start status:
2026/09/28 08:02:17 manager.go:1565: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
2026/09/28 08:02:17 manager.go:1565: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 13 seconds (healthy)
2026/09/28 08:02:17 manager.go:1565: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 13 seconds (healthy)
2026/09/28 08:03:29 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan
2026/09/28 08:03:29 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 103ms)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 10:06 (next run 2026-09-28 10:06 CEST)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 11:06 (next run 2026-09-28 11:06 CEST)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 11:51 (next run 2026-09-28 11:51 CEST)
4b4a913c4600 nextcloud-db nextcloud mariadb:12.3
4221e639f48b nextcloud-redis nextcloud redis:7-alpine
2026/09/28 08:03:57 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:03:57 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/09/28 08:04:29 scheduler.go:346: [INFO] [scheduler] Running job: backup-cache
2026/09/28 08:04:29 scheduler.go:346: [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/28 08:04:29 scheduler.go:363: [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
2026/09/28 08:04:29 scheduler.go:346: [INFO] [scheduler] Running job: system-health
4b4a913c4600 nextcloud-db nextcloud mariadb:12.3
4221e639f48b nextcloud-redis nextcloud redis:7-alpine
2026/09/28 08:04:30 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:04:30 scheduler.go:363: [INFO] [scheduler] Job system-health completed (took 151ms)
2026/09/28 08:04:30 scheduler.go:363: [INFO] [scheduler] Job backup-cache completed (took 156ms)
2026/09/28 08:05:29 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan
2026/09/28 08:05:29 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 104ms)
2026/09/28 08:06:00 scheduler.go:346: [INFO] [scheduler] Running job: db-dump
4b4a913c4600 nextcloud-db nextcloud mariadb:12.3
4221e639f48b nextcloud-redis nextcloud redis:7-alpine
2026/09/28 08:06:00 backup.go:571: [INFO] [backup] Discovered 2 database(s): nextcloud-db(mariadb), paperless-postgres(postgres)
2026/09/28 08:06:00 dbdump.go:410: [INFO] [backup] DB dump: nextcloud-db → nextcloud-mariadb.sql (852.5 KB, 471ms, 131 tables)
2026/09/28 08:06:00 backup.go:946: [INFO] [backup] Stopping nextcloud for safe volume dump
2026/09/28 08:06:00 manager.go:1305: [INFO] [stacks] Stopping stack: nextcloud
2026/09/28 08:06:02 manager.go:1314: [INFO] [stacks] Stack nextcloud stopped successfully (took 1.9s)
2026/09/28 08:06:03 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_db_data → 164.5 MB
2026/09/28 08:06:09 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 761.6 MB
2026/09/28 08:06:09 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 187.5 KB
2026/09/28 08:06:09 backup.go:956: [INFO] [backup] Restarting nextcloud after volume dump
2026/09/28 08:06:09 manager.go:1216: [INFO] [stacks] Starting stack: nextcloud
2026/09/28 08:06:21 manager.go:1231: [INFO] [stacks] Stack nextcloud started successfully (took 11.4s)
2026/09/28 08:06:24 manager.go:1562: [INFO] [stacks] Stack nextcloud post-start status:
2026/09/28 08:06:24 manager.go:1565: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
2026/09/28 08:06:24 manager.go:1565: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 14 seconds (healthy)
2026/09/28 08:06:24 manager.go:1565: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 14 seconds (healthy)
2026/09/28 08:06:43 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/09/28 08:06:43 scheduler.go:363: [INFO] [scheduler] Job db-dump completed (took 43.252s)
10:07:08 # kp window 02:30 — 2026-09-28T10:07:04+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:07:08 POST /backups/window window_start=02:30 -> 303
10:07:14 2026/09/28 08:07:08 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 11:06 → 03:30 (next run 2026-09-29 03:30 CEST)
2026/09/28 08:07:08 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 11:51 → 04:15 (next run 2026-09-29 04:15 CEST)
2026/09/28 08:07:08 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run
2026/09/28 08:07:08 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run
2026/09/28 08:07:08 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run
2026/09/28 08:07:08 backup_handlers.go:65: [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15)
@@ -0,0 +1,42 @@
10:07:26 # kp remove-keep — 2026-09-28T10:07:23+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:07:28 stop 200
10:08:11 remove (keep drive data, remove_backups=False) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": [], "hdd_paths_preserved": ["/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)"], "verified": true}, "message": "Stack nextcloud removed"}
10:08:22 after: deployed= False | appdata: nextcloud paperless
-rw-r--r-- 1 root root 5760 2026-09-28T08:06:43 manifest.json
drwxr-xr-x 2 root root 4096 2026-09-28T08:06:09 volume-dumps
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps:
total 864
drwxr-xr-x 2 root root 4096 2026-09-28T08:06:00 .
drwxr-xr-x 5 root root 4096 2026-09-28T08:06:43 ..
-rw-r--r-- 1 root root 872958 2026-09-28T08:06:00 nextcloud-mariadb.sql
ls: cannot access '/mnt/*/*/backups/secondary/nextcloud': No such file or directory
ls: cannot access '/mnt/*/backups/secondary/nextcloud': No such file or directory
10:08:25 # kp ask — 2026-09-28T10:08:22+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:08:28 deploy, no choice, lang=hu -> 409 use_offered=True use_desc='Az adatbázist innen töltjük vissza: saját mentés, 2026-09-28 10:06. A régi fájlokkal indítjuk az alkalmazást. Ami ezután változott, hiányozhat.' use_off=''
10:08:31 deploy, no choice, lang=en -> 409 use_offered=True use_desc='We load the database from its own backup, 2026-09-28 10:06 and start the app with the old files. Changes made after that time can be missing.' use_off=''
10:08:31 still not installed: False
10:08:34 # kp use — 2026-09-28T10:08:32+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:08:40 deploy kept_data=use -> 202 {"ok": true, "message": "A(z) Nextcloud megőrzött adatainak betöltése elindult."}
10:09:01 deployed after (20.1, 'degraded')
10:09:38 controller lines:
2026/09/28 08:08:40 kept_install.go:97: [INFO] [api] Deploy nextcloud: USE MY KEPT DATA — loading from /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (tier 1, 2026-09-28T08:06:00Z) under the kept files [/mnt/felhom-drives/scratch_hdd/appdata/nextcloud]
2026/09/28 08:08:40 restore_unit.go:538: [WARN] [backup] Restore nextcloud: generated replacement for [NEXTCLOUD_ADMIN_PASSWORD] — the credential was reset (old value unrecoverable); no data-encrypting key was involved, but a regenerated DATABASE password will not match the restored data directory's stored hash (R-12
2026/09/28 08:08:40 restore_unit.go:393: [INFO] [backup] Restoring nextcloud from recovery unit /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud: images=3, secrets recovered=2/3, data_keys=0
2026/09/28 08:08:41 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_db_data for nextcloud
2026/09/28 08:08:42 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_html for nextcloud
2026/09/28 08:08:49 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_redis_data for nextcloud
2026/09/28 08:08:49 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for nextcloud
2026/09/28 08:08:49 deploy.go:718: [INFO] [stacks] nextcloud: the restore generated [NEXTCLOUD_ADMIN_PASSWORD] — the app's own login came back with its data; the page will not show the new value as the password
2026/09/28 08:08:50 restore_db.go:87: [INFO] [backup] Restore nextcloud: replaying DB dump into nextcloud-db (mariadb)
2026/09/28 08:08:56 restore_db.go:97: [INFO] [backup] Restore nextcloud: replayed 1 DB dump(s)
2026/09/28 08:09:11 restore_unit.go:478: [INFO] [backup] Restore-from-unit completed: nextcloud — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
2026/09/28 08:09:11 kept_load.go:331: [INFO] [backup] kept load nextcloud from /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud done in 30s (volumes 3/3, dbs 1/1)
2026/09/28 08:09:12 kept.go:521: [INFO] [stacks] after_load nextcloud: nextcloud [php occ files:scan --all] in 656ms (err=<nil>): Starting scan for user 1 out of 2 (admin)
10:09:38 restore status: {"running": false, "op": "restore", "stack": "nextcloud", "started_at": "2026-09-28T08:08:40.940253282Z", "last": {"op": "restore", "stack": "nextcloud", "ok": true, "message": "A(z) Nextcloud a megőrzött adataival fut.", "finished_at": "2026-09-28T08:09:11.382296795Z"}, "last_recent": true}
10:09:41 # kp read — 2026-09-28T10:09:39+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:09:48 occ user:info — present ['drill2bf734'], absent [] (+ a uid that cannot exist, the control): {'drill2bf734': True, 'nobody85b783': False}
10:09:49 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
10:09:52 state: running images: nextcloud=nextcloud:34.0.4-apache nextcloud-redis=redis:7-alpine nextcloud-db=mariadb:12.3
@@ -0,0 +1,39 @@
10:10:08 # kp remove-keep delete-backups — 2026-09-28T10:10:04+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:10:10 stop 200
10:10:53 remove (keep drive data, remove_backups=True) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": [], "hdd_paths_preserved": ["/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)"], "backup_paths_removed": ["/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (928M)"], "verified": true}, "message": "Stack n
10:11:04 after: deployed= False | appdata: nextcloud paperless
ls: cannot access '/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud': No such file or directory
ls: cannot access '/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps': No such file or directory
ls: cannot access '/mnt/*/*/backups/secondary/nextcloud': No such file or directory
ls: cannot access '/mnt/*/backups/secondary/nextcloud': No such file or directory
10:11:07 # kp kept-delete — 2026-09-28T10:11:04+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:11:07 kept nextcloud items: ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud']
10:11:07 POST /kept-data/delete /mnt/felhom-drives/scratch_hdd/appdata/nextcloud -> HTTP/2 302 ['location: /kept-data?flash=A%28z%29+Nextcloud+meg%C5%91rz%C3%B6tt+adatai+t%C3%B6r%C3%B6lve.']
10:11:10 appdata now: ls: cannot access '/mnt/felhom-drives/scratch_hdd/kept': No such file or directory /mnt/felhom-drives/scratch_hdd/appdata: nextcloud paperless
ls: cannot access '/opt/docker/stacks/nextcloud/app.yaml': No such file or directory
0
0
/mnt/felhom-drives/scratch_hdd/backups/primary/:
/mnt/felhom-drives/scratch_hdd/userdata/:
calibre-web
documents
downloads
emby
immich
jellyfin
komga
media
navidrome
nextcloud
paperless-ngx
plex
radarr
romm
roms
sonarr
/mnt/felhom-drives/scratch_hdd/userdata/nextcloud
removed-empty-harness-dir
@@ -0,0 +1,3 @@
before: gitea.dooplex.hu/admin/felhom-controller:0.276.0
gitea.dooplex.hu/admin/felhom-controller:0.277.0
gitea.dooplex.hu/admin/felhom-controller:0.277.0 Up 25 seconds (healthy)
@@ -0,0 +1,9 @@
## 2026-09-28T08:11:50Z floor 0.276.0 -> 0.277.0 (min_agent 0.131.0 from the v0.277.0 header); golden vouched 0.276.0
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
name="min_controller_version" value="0.277.0"
## 2026-09-28T08:12:31Z both demo boxes, read from the boxes:
felhom-pve 9201: gitea.dooplex.hu/admin/felhom-controller:0.277.0 Up 28 seconds (healthy)
demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.277.0 Up 24 seconds (healthy)
2026/09/28 10:11:53 [INFO] managed floor SERVED for demo-felhom: floor 0.277.0, agent requirement "0.131.0" from declared (golden 0.276.0)
2026/09/28 10:11:53 [INFO] managed floor SERVED for demo-hp: floor 0.277.0, agent requirement "0.131.0" from declared (golden 0.276.0)
@@ -0,0 +1,7 @@
# after restoring the fix
--- PASS: TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest (0.00s)
--- PASS: TestR691_OffsiteLoadDownloadsJudgesThenRestores (0.00s)
--- PASS: TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone (0.01s)
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.024s
--- PASS: TestR691_InstallChoiceNamesTheOffsiteCopy (0.02s)
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.027s
@@ -0,0 +1,31 @@
# RP1 — KeptBestCopy returns KeptDBCopy alone (0.276.0's choice: the off-site copy is never looked at)
189,191c189
< local, lok := m.KeptDBCopy(app, drive)
< off, ook := m.KeptOffsiteCopy(ctx, app)
< return KeptNewer(local, lok, off, ook)
---
> return m.KeptDBCopy(app, drive) // RED-PROOF RP1
=== RUN TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest
r691_kept_offsite_test.go:114: only off-site copy: got {UnitDir: Tier:0 Time:0001-01-01 00:00:00 +0000 UTC DriveLabel: SnapshotID: UnitPath:} ok=false, want the off-site snapshot ab12cd34
--- FAIL: TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest (0.00s)
=== RUN TestR691_OffsiteLoadDownloadsJudgesThenRestores
r691_kept_offsite_test.go:162: no off-site copy offered
--- FAIL: TestR691_OffsiteLoadDownloadsJudgesThenRestores (0.00s)
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/data_version_not_recorded
r691_kept_offsite_test.go:210: no off-site copy offered
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/taken_of_another_drive
r691_kept_offsite_test.go:210: no off-site copy offered
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/definition_is_not_the_data's
r691_kept_offsite_test.go:210: no off-site copy offered
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/download_failed
r691_kept_offsite_test.go:210: no off-site copy offered
--- FAIL: TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.012s
=== RUN TestR691_InstallChoiceNamesTheOffsiteCopy
kept_install_test.go:147: en: use_offered=false use_desc="" use_off="There is no backup of the database, so the app cannot load the old files. You can look at the files under Kept data."
--- FAIL: TestR691_InstallChoiceNamesTheOffsiteCopy (0.02s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.024s
FAIL
@@ -0,0 +1,15 @@
# RP2 — the data-version refusal removed (a unit with no data block would be loaded blind)
240c240
< if man == nil || man.Data == nil {
---
> if man == nil { // RED-PROOF RP2
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/data_version_not_recorded
r691_kept_offsite_test.go:218: the load reported success
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/taken_of_another_drive
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/definition_is_not_the_data's
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/download_failed
--- FAIL: TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.019s
FAIL
@@ -0,0 +1,5 @@
# RP3 — observed, not staged: the FIRST draft of LoadKeptOffsite removed the downloaded copy in a `defer`, i.e. AFTER
# EndRestoreOp and after(ok) had already reported the outcome. The tests failed on it before the fix (copied from the run):
r691_kept_offsite_test.go:184: the downloaded copy was left behind at /tmp/TestR691_OffsiteLoadDownloadsJudgesThenRestores2898855198/002/backups/offsite-proof/nextcloud (err=<nil>)
r691_kept_offsite_test.go:233: the downloaded copy was left behind after a refusal (err=<nil>) (x3 sub-cases)
# Fix: cleanup() runs before fail()/EndRestoreOp on every path.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,190 @@
"""kp.py <step> [arg] — kept data + off-site restore, live, one step per call (2026-09-28, controller 0.277.0).
EVIDENCE, NOT PRODUCT. Every act goes through the product's own endpoints (the ones the pages call):
POST /api/stacks/nextcloud/deploy (with and without kept_data) · /stop · /remove · POST /backups/window
POST /kept-data/delete · GET /kept-data · the restore pages' form posts
and the app's OWN front doors for data: `occ user:add` inside nextcloud's container, and WebDAV as the
seeded user (R-156: nothing planted in a volume, no SQL).
Env: GUEST=9202|9201 (demo-hp), SC=<scratchpad> (seed tokens live there, never in the evidence tree),
OUT=<evidence file>. The throwaway app is nextcloud; the sub-domain is SUB (default c-nc).
"""
import json, os, re, subprocess, sys, time
import walk as w, fixtures
step = sys.argv[1]
arg = sys.argv[2] if len(sys.argv) > 2 else ""
APP = "nextcloud"
SUB = os.environ.get("SUB", "c-nc")
HDD = os.environ.get("HDD", {"9202": "/mnt/felhom-drives/scratch_hdd", "9201": "/mnt/felhom-drives/hdd_1"}[w.GUEST])
w.DRIVE = HDD + "/userdata" # the helper hard-codes 9202's drive for the folders a deploy needs
TOK = os.path.join(w.SC, f"seed-tokens-{w.GUEST}.json")
OUT = open(os.environ["OUT"], "a", buffering=1)
def say(*a):
w.say(*a)
OUT.write(time.strftime("%H:%M:%S ") + " ".join(map(str, a)) + "\n")
def g(cmd, timeout=600):
return w.guest(cmd, timeout=timeout)
def toks():
try:
return json.load(open(TOK))
except Exception:
return {}
def save_toks(t):
old = os.umask(0o077)
json.dump(t, open(TOK, "w"), default=str)
os.umask(old)
def deploy(choice=None, lang=""):
vals = w.deploy_values(APP, SUB)
vals["HDD_PATH"] = HDD
body = {"values": vals}
if choice:
body["kept_data"] = choice
return w.ctl("POST", f"/api/stacks/{APP}/deploy" + ("?lang=en" if lang else ""), body)
def wait_deployed(limit=900):
t0 = time.time()
while time.time() - t0 < limit:
st = w.stack(APP)
if st.get("deployed") and (st.get("app_config") or {}).get("pinned_images") and st.get("state") in ("running", "unhealthy", "degraded"):
return round(time.time() - t0, 1), st.get("state")
time.sleep(5)
return None, w.stack(APP).get("state")
def dav(mode, name=""):
"""A file through Nextcloud's OWN WebDAV as the seeded user."""
t = toks()["nextcloud"]
base = f"{w.BASE}/remote.php/dav/files/{t['uid']}/"
a = ["curl", "-sk", "--max-time", "60", "-u", f"{t['uid']}:{t['pw']}", "-H", f"Host: {SUB}.{w.DOMAIN}", "-w", "\n%{http_code}"]
if mode == "put":
r = subprocess.run(a + ["-X", "PUT", "--data-binary", f"kept-offsite {name} {time.strftime('%FT%T')}", base + name], capture_output=True, text=True)
return r.stdout.strip().splitlines()[-1] if r.stdout.strip() else "?"
r = subprocess.run(a + ["-X", "PROPFIND", "-H", "Depth: 1", base], capture_output=True, text=True)
lines = r.stdout.strip().splitlines()
code = lines[-1] if lines else "?"
names = sorted(set(re.findall(r"<d:href>[^<]*/([^/<]+)</d:href>", r.stdout)))
return code, names
def files_report(expect_present, expect_absent):
code, names = dav("ls")
ok = code == "207" and all(n in names for n in expect_present) and not any(n in names for n in expect_absent)
say(f"PROPFIND as the seeded user -> {code}; present {expect_present}: {[n in names for n in expect_present]}; "
f"absent {expect_absent}: {[n not in names for n in expect_absent]}; listing={names}")
return ok
def users_report(want_present, want_absent):
nc = fixtures.FIXTURES["nextcloud"]
res = {}
for u in want_present + want_absent + ["nobody" + os.urandom(3).hex()]:
got, _ = nc._info(w, u)
res[u] = got
ok = all(res[u] is True for u in want_present) and all(res[u] is False for u in want_absent)
say(f"occ user:info — present {want_present}, absent {want_absent} (+ a uid that cannot exist, the control): {res}")
return ok
w.login()
say(f"# kp {step} {arg} — {time.strftime('%FT%T%z')}; guest {w.GUEST}; controller {g('cat /etc/felhom-controller-image').strip()}")
if step == "deploy":
st = w.stack(APP)
say("before: deployed=", st.get("deployed"), "| appdata:", g(f"ls {HDD}/appdata 2>&1 | tr '\\n' ' '"))
code, d = deploy()
say(f"deploy (no choice) -> {code} {json.dumps(d, ensure_ascii=False)[:300]}")
say("deployed after", wait_deployed())
elif step == "seed":
t = toks()
t["nextcloud"] = fixtures.FIXTURES["nextcloud"].seed(w, SUB, say)
save_toks(t)
say("seeded (the account):", t["nextcloud"] is not None, "uid=", (t["nextcloud"] or {}).get("uid"))
say(f"PUT before-backup.txt -> {dav('put', 'before-backup.txt')}")
files_report(["before-backup.txt"], [])
elif step == "marker":
# Written AFTER the snapshot that will be restored: a second account + a file.
nc = fixtures.FIXTURES["nextcloud"]
uid = "marker" + os.urandom(3).hex()
out = g(f"docker exec -u www-data -e OC_PASS=Marker-{os.urandom(8).hex()} nextcloud php occ user:add --password-from-env {uid} 2>&1")
say(f"occ user:add {uid} :: {' '.join(out.split())[:160]}")
t = toks(); t["marker"] = uid; save_toks(t)
say(f"PUT after-snapshot.txt -> {dav('put', 'after-snapshot.txt')}")
users_report([t["nextcloud"]["uid"], uid], [])
files_report(["before-backup.txt", "after-snapshot.txt"], [])
elif step == "read":
t = toks()
present = [t["nextcloud"]["uid"]] + ([] if arg == "no-marker" else ([t["marker"]] if t.get("marker") else []))
absent = [t["marker"]] if (arg == "no-marker" and t.get("marker")) else []
users_report(present, absent)
files_report(["before-backup.txt"], [])
say("state:", w.stack(APP).get("state"), "images:", g(f"docker ps --filter label=com.docker.compose.project={APP} --format '{{{{.Names}}}}={{{{.Image}}}}' | tr '\\n' ' '"))
elif step == "ask":
for lang in ("", "en"):
code, d = deploy(None, lang)
dd = d.get("data") or {}
say(f"deploy, no choice, lang={lang or 'hu'} -> {code} use_offered={dd.get('use_offered')} use_desc={dd.get('use_desc')!r} use_off={dd.get('use_off')!r}")
say("still not installed:", w.stack(APP).get("deployed"))
elif step == "use":
since = g("date -u +%Y-%m-%dT%H:%M:%SZ").strip()
code, d = deploy("use")
say(f"deploy kept_data=use -> {code} {json.dumps(d, ensure_ascii=False)[:240]}")
say("deployed after", wait_deployed(1500))
for i in range(90):
if int((g(f"docker logs --since {since} felhom-controller 2>&1 | grep -c -E 'after_load|kept load .* FAILED'").strip() or "0")) > 0:
break
time.sleep(5)
time.sleep(15)
say("controller lines:\n" + g(f"docker logs --since {since} felhom-controller 2>&1 | grep -iE 'kept|after_load|restor|offbox|snapshot' | grep -v DEBUG | cut -c1-320 | head -50"))
say("restore status:", json.dumps(w.ctl("GET", "/api/backup/restore-status")[1].get("data"), ensure_ascii=False)[:400])
elif step == "remove-keep":
# arg: "keep-backups" (default) or "delete-backups"
code, d = w.ctl("POST", f"/api/stacks/{APP}/stop"); say("stop", code)
time.sleep(15)
rb = arg == "delete-backups"
code, d = w.ctl("POST", f"/api/stacks/{APP}/remove", {"remove_hdd_data": False, "remove_backups": rb})
say(f"remove (keep drive data, remove_backups={rb}) -> {code} {json.dumps(d, ensure_ascii=False)[:400]}")
time.sleep(8)
say("after: deployed=", w.stack(APP).get("deployed"), "| appdata:", g(f"ls {HDD}/appdata | tr '\\n' ' '; echo; ls -la --time-style=+%FT%T {HDD}/backups/primary/{APP} {HDD}/backups/primary/{APP}/db-dumps 2>&1 | tail -8; ls -d /mnt/*/*/backups/secondary/{APP} /mnt/*/backups/secondary/{APP} 2>&1"))
elif step == "window":
sess = open(f"{w.SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{w.SC}/csrf{os.getpid()}.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", f"window_start={arg}", f"{w.BASE}/backups/window"])
say(f"POST /backups/window window_start={arg} -> {r.stdout}")
time.sleep(3)
say(g("docker logs --since 1m felhom-controller 2>&1 | grep -E 'rescheduled|window set' | tail -6"))
elif step == "logs":
say(g(f"docker logs --since {arg} felhom-controller 2>&1 | grep -iE '{APP}|job|offbox|snapshot|kept' | grep -v DEBUG | cut -c1-300 | tail -80"))
elif step == "kept":
for lang in ("", "en"):
h = w.page("/kept-data" + ("?lang=en" if lang else ""))
t = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ", h))
i = t.find("extcloud")
say(f"GET /kept-data{'?lang=en' if lang else ''}: ...{t[max(0, i-200):i+500]}..." if i >= 0 else f"GET /kept-data{'?lang=en' if lang else ''}: no nextcloud row")
elif step == "kept-delete":
h = w.page("/kept-data")
paths = sorted(set(re.findall(r'name="path" value="([^"]*nextcloud[^"]*)"', h)))
say("kept nextcloud items:", paths)
sess = open(f"{w.SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{w.SC}/csrf{os.getpid()}.txt").read().strip()
for p in paths:
r = w.sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", f"path={p}", "--data-urlencode", "confirm=Nextcloud",
f"{w.BASE}/kept-data/delete"])
loc = [l for l in r.stdout.split("\n") if l.lower().startswith("location:")]
say(f"POST /kept-data/delete {p} -> {r.stdout.splitlines()[0] if r.stdout else '?'} {loc[:1]}")
say("appdata now:", g(f"ls {HDD}/appdata {HDD}/kept 2>&1 | tr '\\n' ' '"))
else:
sys.exit(f"unknown step {step}")
@@ -0,0 +1,507 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = os.environ.get('SC', '/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/a061fcea-2c08-478c-9fd6-991dc8e748f0/scratchpad')
EV = os.environ.get('EV', '/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/kept-offsite-2026-09-28')
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
import secrets as _s
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Serialised across concurrent walks: one git working tree, one lock."""
import fcntl
with open(f"{SC}/drill.lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
return _drill_bump(app, frm, to, service_hint)
def _drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
@@ -0,0 +1,9 @@
## 2026-09-28T08:33:16Z bench create — BEFORE
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
local-lvm lvmthin active 56487936 31384697 25103238 55.56%
nvme-scratch dir active 983379700 50609912 882743176 5.15%
total used free shared buff/cache available
Mem: 29 6 4 0 19 23
Configuration file 'nodes/demo-hp/lxc/9401.conf' does not exist
@@ -0,0 +1,17 @@
+ pveam download local debian-13-standard_13.6-1_amd64.tar.zst
+ tail -1
download of 'http://download.proxmox.com/images/system/debian-13-standard_13.6-1_amd64.tar.zst' to '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' finished
+ pct create 9401 local:vztmpl/debian-13-standard_13.6-1_amd64.tar.zst --hostname upgrade-harness --cores 6 --memory 10240 --swap 0 --rootfs nvme-scratch:40 --net0 name=eth0,bridge=vmbr0,ip=dhcp --unprivileged 1 --features nesting=1,keyctl=1 --onboot 0
+ tail -2
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:mDMV8ol8xGZRh95eFLQBtzrp/wlPiAmCskPzerTpe7Y root@upgrade-harness
+ pct start 9401
+ sleep 12
+ pct exec 9401 -- bash -c 'hostname; ip -4 -o addr show eth0 | awk "{print \$4}"; apt-get update -qq >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -qq docker.io docker-compose python3 python3-yaml curl >/dev/null 2>&1; docker --version; docker compose version; python3 --version; docker network create traefik-public >/dev/null && echo traefik-public-net; curl -s -o /dev/null -w net-%{http_code} https://ghcr.io/'
upgrade-harness
192.168.0.157/24
Docker version 26.1.5+dfsg1, build a72d7cd
Docker Compose version 2.26.1-4
Python 3.13.5
traefik-public-net
net-301
@@ -0,0 +1,3 @@
10:34:47 START claper ['claper-postgres=postgres:17-alpine']
10:34:56 DONE claper rc=1 verdict=None read_before=None read_after=None abort=None marks=None peaks={} detail='None' (9s)
10:34:56 QUEUE q1 FINISHED
@@ -0,0 +1 @@
10:35:51 START claper ['claper-postgres=postgres:17-alpine']
@@ -0,0 +1,45 @@
start ('200', {'ok': True, 'data': {'state': 'starting'}, 'message': 'Stack calcom start requested — state now: starting'})
calcom Up 2 seconds (health: starting)
calcom-postgres Up About a minute (healthy)
restarts=4 oom=false exit=0
📲 Updated app: 'zoom'
+ yarn start
• turbo 2.7.1
• Packages in scope: @calcom/web
• Running start in 1 packages
• Remote caching disabled
/calcom/node_modules/turbo/bin/turbo:273
throw e;
^
Error: Command failed: /calcom/node_modules/turbo-linux-64/bin/turbo run start --filter=@calcom/web
at genericNodeError (node:internal/errors:984:15)
at wrappedFn (node:internal/errors:538:14)
at checkExecSyncError (node:child_process:891:11)
at Object.execFileSync (node:child_process:927:15)
at Object.<anonymous> (/calcom/node_modules/turbo/bin/turbo:266:17)
at Module._compile (node:internal/modules/cjs/loader:1521:14)
at Module._extensions..js (node:internal/modules/cjs/loader:1623:10)
at Module.load (node:internal/modules/cjs/loader:1266:32)
at Module._load (node:internal/modules/cjs/loader:1091:12)
at Function.executeUserEntryPoint [as runMain] (node:internal/modules/run_main:164:12) {
status: null,
signal: 'SIGKILL',
output: [ null, null, null ],
pid: 137,
stdout: null,
stderr: null
}
Node.js v20.20.0
+ scripts/replace-placeholder.sh http://localhost:3000 https://c-cal.enkisfelhom.hu
Replacing all statically built instances of http://localhost:3000 with https://c-cal.enkisfelhom.hu.
+ scripts/wait-for-it.sh -- echo database is up
Error: you need to provide a host and port to test.
Usage:
scripts/wait-for-it.sh host:port|url [-t timeout] [-- command args]
-q | --quiet Do not output any status messages
-t TIMEOUT | --timeout=timeout Timeout in seconds, zero for no timeout
-- COMMAND ARGS Execute command with args after the test finishes
+ npx prisma migrate deploy --schema /calcom/packages/prisma/schema.prisma
@@ -0,0 +1,3 @@
10:17:34 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:21:00 [1] deployed, controller state=running, pinned={'calcom': 'calcom/cal.com:v6.2.0', 'calcom-postgres': 'postgres:16-alpine'}
deployed True
@@ -0,0 +1,33 @@
2026/09/28 08:17:34 router.go:431: [INFO] [api] Deploy requested for stack: calcom
2026/09/28 08:17:34 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:17:34 deploy.go:442: [INFO] [stacks] Deploying stack calcom with 5 env vars: [DOMAIN, SUBDOMAIN, NEXTAUTH_SECRET, CALENDSO_ENCRYPTION_KEY, DB_PASSWORD]
2026/09/28 08:17:34 manager.go:1601: [INFO] [stacks] Deploying stack calcom — checking 2 images...
2026/09/28 08:20:26 deploy.go:532: [INFO] [stacks] Stack calcom deployed successfully (took 171.7s)
2026/09/28 08:20:26 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:20:26 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:20:26 installed.go:420: [INFO] [stacks] installed-images calcom: recorded 2 service(s) (calcom=calcom/cal.com:v6.2.0 (sha256:ace3bb1219fb…), calcom-postgres=postgres:16-alpine (sha256:721873c34ceb…))
2026/09/28 08:20:26 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:20:26 pin.go:93: [INFO] [stacks] pin calcom: calcom=calcom/cal.com:v6.2.0, calcom-postgres=postgres:16-alpine
2026/09/28 08:20:29 manager.go:1562: [INFO] [stacks] Stack calcom post-start status:
2026/09/28 08:20:29 manager.go:1565: [INFO] [stacks] calcom calcom/cal.com:v6.2.0 running Up 3 seconds (health: starting)
2026/09/28 08:20:29 manager.go:1565: [INFO] [stacks] calcom-postgres postgres:16-alpine running Up 9 seconds (healthy)
2026/09/28 08:22:29 main.go:3762: [WARN] [deadapp] calcom: crash_loop — 6 in 10m0s; STOPPING it (decision 28)
2026/09/28 08:22:30 manager.go:1305: [INFO] [stacks] Stopping stack: calcom
2026/09/28 08:22:40 manager.go:1314: [INFO] [stacks] Stack calcom stopped successfully (took 10.8s)
2026/09/28 08:22:40 settings.go:1881: [WARN] [settings] restore hold SET for calcom — the app stays stopped until it is cleared
2026/09/28 08:22:40 update_guard.go:827: [WARN] [backup] calcom is STOPPED by the box: crash_loop (trip 1 within 24h0m0s) — Start gives it one more try (decision 28)
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
claper Up 9 minutes (healthy)
claper-postgres Up 10 minutes (healthy)
filebrowser Up 15 minutes (healthy)
privatebin Up 19 minutes (healthy)
paperless-webserver Up 19 minutes (healthy)
paperless-postgres Up 20 minutes (healthy)
paperless-redis Up 20 minutes (healthy)
felhom-controller Up 27 minutes (healthy)
traefik Up 3 days
@@ -0,0 +1,33 @@
cgroup=/sys/fs/cgroup/system.slice/docker-6cc2be9f86e8fda3e63e61be599d7c4cd616a0f4ecbe6f0c5b77b229bf00a1e5.scope
805306368
08:28:27 cur=593969152 oom 0 oom_kill 0 anon 81022976 restarts=5
08:28:31 cur=683855872 oom 0 oom_kill 0 anon 158216192 restarts=5
08:28:35 cur=805126144 oom 0 oom_kill 0 anon 662142976 restarts=5
08:28:39 cur=507691008 oom 0 oom_kill 0 anon 17334272 restarts=6
08:28:43 cur=646803456 oom 0 oom_kill 0 anon 132259840 restarts=6
08:28:48 cur=805306368 oom 0 oom_kill 0 anon 707543040 restarts=6
08:28:52 cur=11423744 oom 1 oom_kill 1 anon 3809280 restarts=6
08:28:56 cur=608956416 oom 0 oom_kill 0 anon 104116224 restarts=7
08:29:00 cur=681771008 oom 0 oom_kill 0 anon 166567936 restarts=7
08:29:04 cur=804466688 oom 0 oom_kill 0 anon 626081792 restarts=7
08:29:08 cur=16332591104 oom 16 oom_kill 14 anon 982835200
08:29:12 cur=16284028928 oom 16 oom_kill 14 anon 954310656
08:29:16 cur=16285847552 oom 16 oom_kill 14 anon 956977152
08:29:20 cur=16286691328 oom 16 oom_kill 14 anon 957399040
08:29:24 cur=16285790208 oom 16 oom_kill 14 anon 957583360
08:29:28 cur=16285872128 oom 16 oom_kill 14 anon 957526016
08:29:33 cur=16301551616 oom 16 oom_kill 14 anon 956796928
08:29:37 cur=16290660352 oom 16 oom_kill 14 anon 956542976
08:29:41 cur=16290639872 oom 16 oom_kill 14 anon 956497920
08:29:45 cur=16291016704 oom 16 oom_kill 14 anon 956428288
08:29:49 cur=16290803712 oom 16 oom_kill 14 anon 957030400
08:29:53 cur=16290611200 oom 16 oom_kill 14 anon 957075456
08:29:57 cur=16291708928 oom 16 oom_kill 14 anon 957132800
08:30:01 cur=16292163584 oom 16 oom_kill 14 anon 955760640
08:30:05 cur=16291311616 oom 16 oom_kill 14 anon 955813888
08:30:09 cur=16289894400 oom 16 oom_kill 14 anon 954003456
08:30:14 cur=16290275328 oom 16 oom_kill 14 anon 955899904
08:30:18 cur=16291815424 oom 16 oom_kill 14 anon 956694528
08:30:22 cur=16291139584 oom 16 oom_kill 14 anon 956817408
08:30:26 cur=16291594240 oom 16 oom_kill 14 anon 956981248
@@ -0,0 +1,4 @@
Error response from daemon: No such container: calcom
Error response from daemon: No such container: calcom
@@ -0,0 +1,6 @@
10:26:10 app never answered on c-cal/api/auth/providers (last rc=0 code=404)
wait False
setup ('404', '404 page not found\n')
setup again (must refuse) ('404', '404 page not found\n')
GET /drilld1401a 404 19
GET /nobody20cc4f 404 19
@@ -0,0 +1,6 @@
10:30:45 [X] stop -> 200 {'ok': True, 'message': 'Stack calcom stop completed'}
10:31:17 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'calcom', 'volumes_removed': ['calcom_calcom_postgres_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': '
10:31:25 [X] after remove: deployed=False leftovers='/opt/docker/stacks/calcom'
0
ls: cannot access '/opt/docker/stacks/calcom/app.yaml': No such file or directory
@@ -0,0 +1,13 @@
claper 2.5.1
[sh -c /app/bin/claper eval Claper.Release.migrate && /app/bin/claper eval Claper.Release.seeds && /app/bin/claper start] []
warning: the VM is running with native name encoding of latin1 which may cause Elixir to malfunction as it expects utf8. Please ensure your locale is set to UTF-8 (which can be verified by running "locale" in your shell) or set the ELIXIR_ERL_OPTIONS="+fnu" environment variable
default admin exists: true
warning: the VM is running with native name encoding of latin1 which may cause Elixir to malfunction as it expects utf8. Please ensure your locale is set to UTF-8 (which can be verified by running "locale" in your shell) or set the ELIXIR_ERL_OPTIONS="+fnu" environment variable
default password claper works: true
warning: the VM is running with native name encoding of latin1 which may cause Elixir to malfunction as it expects utf8. Please ensure your locale is set to UTF-8 (which can be verified by running "locale" in your shell) or set the ELIXIR_ERL_OPTIONS="+fnu" environment variable
control: unknown email: false
08:16:36.932 [info] == Running 20240730123205 Claper.Repo.Migrations.RemoveIsAdminFromUsers.change/0 forward
Created admin role
Created default admin user:
Email: admin@claper.co
@@ -0,0 +1,9 @@
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
image: ghcr.io/claperco/claper:2.5
total used free shared buff/cache available
Mem: 25898 808 18374 39 6754 25089
10:15:41 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:16:52 [1] deployed, controller state=running, pinned={'claper': 'ghcr.io/claperco/claper:2.5', 'claper-postgres': 'postgres:16-alpine'}
deployed True
@@ -0,0 +1,4 @@
registered id=2
probe authenticates: true
control: wrong password: false
@@ -0,0 +1,8 @@
10:32:19 claper: register_user :: DRILL_SEEDED=3
10:32:23 claper: seeded user drill-e157296d@gate.invalid
seed ok True
10:32:30 claper: readback — the seeded account authenticates=True
verify True
10:32:37 claper: readback — the seeded account authenticates=False
10:32:37 claper: :: DRILL_ANSWER=false
verify with a token that was never seeded (must be False): False
@@ -0,0 +1,48 @@
#!/usr/bin/env python3
"""benchq.py <queue-name> <app> <svc=ref>[,<svc=ref>] [<app> <moves> ...]
Runs `upgrade-test.py --move` on bench LXC 9401 for each app IN ORDER, then copies that edge's
evidence OFF the bench into apps/<app>/bench/ before starting the next (R-320: the evidence leaves
the machine at the end of the phase that made it). One line per finished edge in benchq-<queue>.txt.
"""
import json, os, subprocess, sys, time
HERE = os.path.dirname(os.path.abspath(__file__))
BENCH = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/753351c4-7b70-46b6-9338-40951c9512c5/scratchpad/bench.sh"
q = sys.argv[1]
args = sys.argv[2:]
out = open(os.path.join(HERE, "..", os.environ.get("BQ", "B"), f"benchq-{q}.txt"), "a", buffering=1)
for i in range(0, len(args), 2):
app, moves = args[i], args[i + 1].split(",")
label = app
if "@" in app: # romm@engine: a second edge of one app gets its own folder
app, tag = app.split("@", 1)
label = f"{app}-{tag}"
t0 = time.time()
out.write(f"{time.strftime('%H:%M:%S')} START {app} {moves}\n")
pre = f"cd /opt/upg && [ -d evidence/MV-{app} ] && mv evidence/MV-{app} evidence/MV-{app}-$(date +%s); "
subprocess.run([BENCH], input=pre, capture_output=True, text=True, timeout=120)
cmd = (f"cd /opt/upg && timeout 2400 python3 upgrade-test.py --soak 600 --move {app} {' '.join(moves)} "
f"> /opt/upg/MV-{app}.log 2>&1; echo rc=$?")
r = subprocess.run([BENCH], input=cmd, capture_output=True, text=True, timeout=2700)
dest = os.path.join(HERE, "..", os.environ.get("BQ", "B"), "apps", label, "bench")
os.makedirs(dest, exist_ok=True)
pull = (f"cd /opt/upg && tar czf /tmp/ev-{app}.tgz MV-{app}.log evidence/MV-{app} 2>/dev/null; "
f"base64 -w0 /tmp/ev-{app}.tgz")
r2 = subprocess.run([BENCH], input=pull, capture_output=True, text=True, timeout=600)
import base64, io, tarfile
try:
tarfile.open(fileobj=io.BytesIO(base64.b64decode(r2.stdout.strip()))).extractall(dest)
except Exception as e:
out.write(f" evidence pull FAILED for {app}: {e}\n")
v = {}
try:
v = json.load(open(os.path.join(dest, "evidence", f"MV-{app}", "verdict.json")))
except Exception:
pass
mem = v.get("memory") or {}
peaks = {n: c.get("peak_pct") for n, c in (mem.get("containers") or {}).items()}
out.write(f"{time.strftime('%H:%M:%S')} DONE {app} {r.stdout.strip()} verdict={v.get('verdict')} "
f"read_before={v.get('seed_read_before')} read_after={v.get('seed_read_after')} "
f"abort={v.get('abort')} marks={v.get('marks')} peaks={peaks} "
f"detail={str(v.get('abort_detail'))[:160]!r} ({round(time.time()-t0)}s)\n")
out.write(f"{time.strftime('%H:%M:%S')} QUEUE {q} FINISHED\n")
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,507 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = os.environ.get('SC', '/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/a061fcea-2c08-478c-9fd6-991dc8e748f0/scratchpad')
EV = os.environ.get('EV', '/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/pg-calcom-claper-2026-09-28')
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
import secrets as _s
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Serialised across concurrent walks: one git working tree, one lock."""
import fcntl
with open(f"{SC}/drill.lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
return _drill_bump(app, frm, to, service_hint)
def _drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}