docs: R-691 (2) built (07 §6.5), R-701 (a) measured, R-702 claper default admin, R-703 calcom OOM; evidence
gates / gates (push) Successful in 25s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-28 10:37:42 +02:00
parent 03df4c8220
commit a8f790c0c6
39 changed files with 4023 additions and 4 deletions
@@ -562,7 +562,7 @@ anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknow
**What it is.** "Remove the app, keep my data" leaves the app's drive folder (`<drive>/appdata/<app>`, or the folder
its definition binds through `${HDD_PATH}`) in place. A reinstall over it no longer runs silently into the old files
(R-657): the install asks „A megőrzött adataimat használom" / "Use my kept data" (the database from the newest copy of
THIS drive's install — the app's own unit or the second-drive mirror — loaded under the kept files, then the
THIS drive's install — the app's own unit, the second-drive mirror or (since v0.277.0) the off-site copy — loaded under the kept files, then the
template's `after_load:`, e.g. nextcloud's `occ files:scan --all`) or „Tiszta lappal kezdem" / "Start fresh".
**Where it lives.** "Start fresh" RENAMES the folder — same drive, never a copy, never across drives — to
@@ -570,6 +570,15 @@ template's `after_load:`, e.g. nextcloud's `occ files:scan --all`) or „Tiszta
Load has its database). `<drive>/kept` is in `ProtectedHDDPaths`, never under `userdata/`, never bound by a live app.
The page „Megőrzött adatok" / "Kept data" lists every dated folder and every `appdata/<x>` no installed app binds.
**The off-site copy counts too (controller v0.277.0, R-691 (2)).** „Use my kept data" and Load consider the app's
newest OFF-SITE snapshot beside the local copies, and use it when it is NEWER than every local copy or the only one (a tie
goes local). The page names the copy and its date („távoli mentés, 2026-09-28 04:15" / "the off-site copy, …"). The
repository is asked once per page, bounded to 20 s; an unreachable one offers nothing. The load downloads the unit ALONE
(never the files — they are the kept ones), and refuses it before anything moves when it was taken of another drive,
holds no database, or **does not record its data version** (§6.6 — a unit written before v0.275.0 is never loaded blind);
a mixed or mismatched unit is refused by §6.6's own rule. Then a dated folder's files move back and the one unit restore
runs. The downloaded copy is removed on every path. A failed download or a refusal leaves the kept files as they were.
**It is NOT backed up.** No tier captures `<drive>/kept` or a leftover `appdata/<x>` of a removed app; the page says
so („Erről nem készül mentés." / "This is not backed up."). What brings it back into an app is the unit it carries
(or the app's own unit on the drive), listed per row as „Visszatölthető innen" / "Can be loaded from".
@@ -96,3 +96,16 @@ claim page: 200
qemu processes: 0
disk reverted to virgin
DooPlex copy of the passphrase shredded
## 2026-09-28T08:35:05Z teardown layer 3 (hub): delete the scratch customer by the hub's own flow
HTTP/1.1 303 See Other
Location: /configs?flash=deleted
2026/09/28 10:35:05 [INFO] customer DELETE cascade started for drill-g0276 (journal #21, 1 host(s))
2026/09/28 10:35:05 [INFO] delete drill-g0276: host drill-g0276-b861fc deleted (escrow DEMOTED to retained custody)
2026/09/28 10:35:06 [INFO] tenantsync: deprovision ok for drill-g0276 (ns=drill-g0276, existed=false)
2026/09/28 10:35:06 [INFO] reset drill-g0276: PBS tenancy deprovisioned
2026/09/28 10:35:06 [INFO] [claim] reset to unclaimed for drill-g0276 (customer RESET) — next onboarding mints a fresh code
2026/09/28 10:35:06 [INFO] delete drill-g0276: residue purged (reports=1 app_telemetry=1 app_log_tails=0 log_tail_requests=0 notif_prefs=0 selfbind_tokens=0 appliance_registrations=0)
2026/09/28 10:35:06 [INFO] customer DELETE cascade COMPLETE for drill-g0276 (journal #21) — full teardown
-- after:
0
0
@@ -0,0 +1,34 @@
## 2026-09-28T08:04:59Z demo-hp 9201 paperless-ngx — read-only
paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 5 hours (healthy)
paperless-redis redis:7-alpine Up 5 hours (healthy)
paperless-postgres postgres:18-alpine Up 5 hours (healthy)
PG_VERSION in the running datadir: 18
server 18.6|documents 0|tags 0|correspondents 0|users 3
paperless-ngx_paperless_data
paperless-ngx_paperless_postgres_data
paperless-ngx_paperless_postgres_data.pre-update-20260928T022253Z
paperless-ngx_paperless_redis_data
-- controller log, conversion lines, last 2 nights:
2026/09/28 01:30:05 tier2.go:425: [INFO] [backup] Tier 2 copied paperless-ngx → /mnt/sys_drive/felhom-data/backups/secondary/paperless-ngx (399.8 MB, 1 leg(s), 3s) [SSD: state-only]
2026/09/28 02:22:44 unattended.go:320: [INFO] [update-leg] paperless-ngx: step pressed paperless-postgres=postgres:16-alpine, paperless-redis=redis:7-alpine, paperless-webserver=ghcr.io/paperless-ngx/paperless-ngx:2.20.15 → paperless-post
2026/09/28 02:22:44 update.go:786: [INFO] [stacks] update paperless-ngx: ladder — the last step (1 of 1) — the catalog's current definition
2026/09/28 02:22:44 update.go:800: [INFO] [stacks] update paperless-ngx: this step CONVERTS paperless-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume paperless-ngx_paperless_postgres_data
2026/09/28 02:22:44 update.go:814: [INFO] [stacks] update paperless-ngx: precondition met — Tier 2 (second drive) copy from 2026-09-28T00:30:01Z (1h53m0s old, limit 24h0m0s)
2026/09/28 02:22:46 undo.go:287: [INFO] [stacks] update paperless-ngx: the undo copy will hold 3 named volume(s), 398.1 MiB
2026/09/28 02:22:46 pgconvert.go:340: [INFO] [stacks] update paperless-ngx: conversion space — database volume paperless-ngx_paperless_postgres_data is 361 MB; the dump needs up to 451 MB + the 2 GB floor; 28.6 GB free beside the stack di
2026/09/28 02:22:46 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:53 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:53 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:57 update.go:1365: [INFO] [stacks] update paperless-ngx: phase copying
2026/09/28 02:22:57 update.go:1365: [INFO] [stacks] update paperless-ngx: phase converting
2026/09/28 02:22:57 pgconvert.go:423: [INFO] [stacks] update paperless-ngx: CONVERTING paperless-postgres PostgreSQL 16 → 18 (volume paperless-ngx_paperless_postgres_data, kept copy paperless-ngx_paperless_postgres_data.pre-update-2026092
2026/09/28 02:22:59 pgconvert.go:468: [INFO] [stacks] update paperless-ngx: pg_dumpall from 16 done in 591ms — 488 kB, completion line present; before: 2 database(s), 72 table(s), 1537 row(s)
2026/09/28 02:23:00 update.go:1365: [INFO] [stacks] update paperless-ngx: phase converting
2026/09/28 02:23:00 pgconvert.go:484: [INFO] [stacks] update paperless-ngx: database volume paperless-ngx_paperless_postgres_data emptied (its copy paperless-ngx_paperless_postgres_data.pre-update-20260928T022253Z holds the old datadir, mar
2026/09/28 02:23:07 pgconvert.go:513: [INFO] [stacks] update paperless-ngx: loaded into 18 in 1.68s (dropped the entrypoint's empty database(s) [paperless]; skipped CREATE ROLE for [paperless])
2026/09/28 02:23:07 pgconvert.go:529: [INFO] [stacks] update paperless-ngx: CONVERTED 16 → 18 in 9.93s — the check is equal (2 database(s), 72 table(s), 1537 row(s)), PG_VERSION 18
2026/09/28 02:23:49 update.go:1015: [INFO] [stacks] update paperless-ngx: KEEPING the pre-conversion datadir copy paperless-ngx_paperless_postgres_data.pre-update-20260928T022253Z (PostgreSQL 16) until a backup of the converted app is prove
2026/09/28 02:23:54 unattended.go:327: [INFO] [update-leg] paperless-ngx: step ended done after 70.1 s
@@ -0,0 +1 @@
## app list BEFORE Part E (deployed): ['adventurelog', 'bentopdf', 'bookstack', 'calibre-web', 'docmost', 'kimai', 'opengist', 'paperless-ngx', 'privatebin', 'romm']
@@ -0,0 +1,3 @@
2026-09-28T10:15:02 POST /backup/offbox/toggle app=nextcloud enabled=on -> HTTP/2 302 ['location: /backups/remote?flash=flash.offbox.app_setting_updated']
settings app_backup.nextcloud = {'enabled': False, 'offbox': True}
@@ -0,0 +1,10 @@
10:12:58 # kp deploy — 2026-09-28T10:12:55+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:13:01 before: deployed= False | appdata: paperless romm
10:13:04 deploy (no choice) -> 202 {"ok": true, "message": "Telepítés elindítva – az állapot a kártyán követhető"}
10:14:30 deployed after (85.5, 'running')
10:14:33 # kp seed — 2026-09-28T10:14:30+0200; guest 9201; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:14:41 nextcloud: occ user:add :: The account "drill895c04" was created successfully Display name set to "drill895c04"
10:14:45 nextcloud: seeded user drill895c04
10:14:45 seeded (the account): True uid= drill895c04
10:14:50 PUT before-backup.txt -> 201
10:14:50 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
@@ -0,0 +1,10 @@
10:01:57 # kp deploy — 2026-09-28T10:01:53+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:02:00 before: deployed= False | appdata: paperless
10:02:03 deploy (no choice) -> 202 {"ok": true, "message": "Telepítés elindítva – az állapot a kártyán követhető"}
10:03:23 deployed after (80.4, 'running')
10:03:27 # kp seed — 2026-09-28T10:03:23+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:03:35 nextcloud: occ user:add :: The account "drill2bf734" was created successfully Display name set to "drill2bf734"
10:03:38 nextcloud: seeded user drill2bf734
10:03:38 seeded (the account): True uid= drill2bf734
10:03:44 PUT before-backup.txt -> 201
10:03:44 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
@@ -0,0 +1,100 @@
10:03:57 # kp window 10:06 — 2026-09-28T10:03:54+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:03:57 POST /backups/window window_start=10:06 -> 303
10:04:03 2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 11:06 (next run 2026-09-28 11:06 CEST)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 11:51 (next run 2026-09-28 11:51 CEST)
2026/09/28 08:03:57 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run
2026/09/28 08:03:57 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run
2026/09/28 08:03:57 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run
2026/09/28 08:03:57 backup_handlers.go:65: [INFO] [web] backup window set to 10:06 (legs 10:06/11:06/11:51)
10:07:00 # kp logs 15m — 2026-09-28T10:06:57+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:07:04 2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: ring-spill (every 30s)
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job db-dump scheduled for 2026-09-29 02:30 CEST
2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: offsite-credential-retry (every 5m0s)
2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: backup-cache (every 5m0s)
2026/09/28 07:59:29 scheduler.go:102: [INFO] [scheduler] Registered periodic job: conversion-copy-release (every 1h0m0s)
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job tier2-backup scheduled for 2026-09-29 03:30 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offbox-backup scheduled for 2026-09-29 04:15 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-abandon-sweep scheduled for 2026-09-29 05:10 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-29 06:00 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-29 05:30 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-29 04:00 CEST
2026/09/28 07:59:29 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-29 03:30 CEST
2026/09/28 07:59:29 scheduler.go:206: [INFO] [scheduler] Starting scheduler with 17 jobs
2026/09/28 07:59:29 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 07:59:29 handlers.go:3622: [INFO] [web] FileBrowser: drive /mnt/felhom-drives/scratch_hdd/userdata/nextcloud not mounted — skipping userdata skeleton
2026/09/28 07:59:29 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 handlers.go:3622: [INFO] [web] FileBrowser: drive /mnt/felhom-drives/scratch_hdd/userdata/nextcloud not mounted — skipping userdata skeleton
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:00:02 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:01:29 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan
2026/09/28 08:01:29 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 64ms)
2026/09/28 08:02:03 router.go:431: [INFO] [api] Deploy requested for stack: nextcloud
2026/09/28 08:02:03 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:02:03 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:02:03 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:03 deploy.go:442: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH, DOMAIN, SUBDOMAIN, DB_PASSWORD]
2026/09/28 08:02:03 manager.go:1601: [INFO] [stacks] Deploying stack nextcloud — checking 3 images...
2026/09/28 08:02:09 healthprobe.go:190: [WARN] Health probe nextcloud: API GET :80/status.php → Get "http://nextcloud-redis:80/status.php": dial tcp: lookup nextcloud-redis on 127.0.0.11:53: no such host
2026/09/28 08:02:14 deploy.go:532: [INFO] [stacks] Stack nextcloud deployed successfully (took 10.9s)
2026/09/28 08:02:14 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:14 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:14 installed.go:420: [INFO] [stacks] installed-images nextcloud: recorded 3 service(s) (nextcloud=nextcloud:34.0.4-apache (sha256:a5ace30c695a…), nextcloud-db=mariadb:12.3 (sha256:805c8e104bd5…), nextcloud-redis=redis:7-alpine (sha256:858f009f9709…))
2026/09/28 08:02:14 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/28 08:02:14 pin.go:93: [INFO] [stacks] pin nextcloud: nextcloud=nextcloud:34.0.4-apache, nextcloud-db=mariadb:12.3, nextcloud-redis=redis:7-alpine
2026/09/28 08:02:17 manager.go:1562: [INFO] [stacks] Stack nextcloud post-start status:
2026/09/28 08:02:17 manager.go:1565: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
2026/09/28 08:02:17 manager.go:1565: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 13 seconds (healthy)
2026/09/28 08:02:17 manager.go:1565: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 13 seconds (healthy)
2026/09/28 08:03:29 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan
2026/09/28 08:03:29 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 103ms)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 10:06 (next run 2026-09-28 10:06 CEST)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 11:06 (next run 2026-09-28 11:06 CEST)
2026/09/28 08:03:57 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 11:51 (next run 2026-09-28 11:51 CEST)
4b4a913c4600 nextcloud-db nextcloud mariadb:12.3
4221e639f48b nextcloud-redis nextcloud redis:7-alpine
2026/09/28 08:03:57 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:03:57 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/09/28 08:04:29 scheduler.go:346: [INFO] [scheduler] Running job: backup-cache
2026/09/28 08:04:29 scheduler.go:346: [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/28 08:04:29 scheduler.go:363: [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
2026/09/28 08:04:29 scheduler.go:346: [INFO] [scheduler] Running job: system-health
4b4a913c4600 nextcloud-db nextcloud mariadb:12.3
4221e639f48b nextcloud-redis nextcloud redis:7-alpine
2026/09/28 08:04:30 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/28 08:04:30 scheduler.go:363: [INFO] [scheduler] Job system-health completed (took 151ms)
2026/09/28 08:04:30 scheduler.go:363: [INFO] [scheduler] Job backup-cache completed (took 156ms)
2026/09/28 08:05:29 scheduler.go:346: [INFO] [scheduler] Running job: stack-scan
2026/09/28 08:05:29 scheduler.go:363: [INFO] [scheduler] Job stack-scan completed (took 104ms)
2026/09/28 08:06:00 scheduler.go:346: [INFO] [scheduler] Running job: db-dump
4b4a913c4600 nextcloud-db nextcloud mariadb:12.3
4221e639f48b nextcloud-redis nextcloud redis:7-alpine
2026/09/28 08:06:00 backup.go:571: [INFO] [backup] Discovered 2 database(s): nextcloud-db(mariadb), paperless-postgres(postgres)
2026/09/28 08:06:00 dbdump.go:410: [INFO] [backup] DB dump: nextcloud-db → nextcloud-mariadb.sql (852.5 KB, 471ms, 131 tables)
2026/09/28 08:06:00 backup.go:946: [INFO] [backup] Stopping nextcloud for safe volume dump
2026/09/28 08:06:00 manager.go:1305: [INFO] [stacks] Stopping stack: nextcloud
2026/09/28 08:06:02 manager.go:1314: [INFO] [stacks] Stack nextcloud stopped successfully (took 1.9s)
2026/09/28 08:06:03 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_db_data → 164.5 MB
2026/09/28 08:06:09 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 761.6 MB
2026/09/28 08:06:09 backup.go:845: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 187.5 KB
2026/09/28 08:06:09 backup.go:956: [INFO] [backup] Restarting nextcloud after volume dump
2026/09/28 08:06:09 manager.go:1216: [INFO] [stacks] Starting stack: nextcloud
2026/09/28 08:06:21 manager.go:1231: [INFO] [stacks] Stack nextcloud started successfully (took 11.4s)
2026/09/28 08:06:24 manager.go:1562: [INFO] [stacks] Stack nextcloud post-start status:
2026/09/28 08:06:24 manager.go:1565: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
2026/09/28 08:06:24 manager.go:1565: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 14 seconds (healthy)
2026/09/28 08:06:24 manager.go:1565: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 14 seconds (healthy)
2026/09/28 08:06:43 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/09/28 08:06:43 scheduler.go:363: [INFO] [scheduler] Job db-dump completed (took 43.252s)
10:07:08 # kp window 02:30 — 2026-09-28T10:07:04+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:07:08 POST /backups/window window_start=02:30 -> 303
10:07:14 2026/09/28 08:07:08 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 11:06 → 03:30 (next run 2026-09-29 03:30 CEST)
2026/09/28 08:07:08 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 11:51 → 04:15 (next run 2026-09-29 04:15 CEST)
2026/09/28 08:07:08 scheduler.go:67: [DEBUG] [scheduler] daily job tier2-backup: rescheduled — recomputing next run
2026/09/28 08:07:08 scheduler.go:67: [DEBUG] [scheduler] daily job db-dump: rescheduled — recomputing next run
2026/09/28 08:07:08 scheduler.go:67: [DEBUG] [scheduler] daily job offbox-backup: rescheduled — recomputing next run
2026/09/28 08:07:08 backup_handlers.go:65: [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15)
@@ -0,0 +1,42 @@
10:07:26 # kp remove-keep — 2026-09-28T10:07:23+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:07:28 stop 200
10:08:11 remove (keep drive data, remove_backups=False) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": [], "hdd_paths_preserved": ["/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)"], "verified": true}, "message": "Stack nextcloud removed"}
10:08:22 after: deployed= False | appdata: nextcloud paperless
-rw-r--r-- 1 root root 5760 2026-09-28T08:06:43 manifest.json
drwxr-xr-x 2 root root 4096 2026-09-28T08:06:09 volume-dumps
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps:
total 864
drwxr-xr-x 2 root root 4096 2026-09-28T08:06:00 .
drwxr-xr-x 5 root root 4096 2026-09-28T08:06:43 ..
-rw-r--r-- 1 root root 872958 2026-09-28T08:06:00 nextcloud-mariadb.sql
ls: cannot access '/mnt/*/*/backups/secondary/nextcloud': No such file or directory
ls: cannot access '/mnt/*/backups/secondary/nextcloud': No such file or directory
10:08:25 # kp ask — 2026-09-28T10:08:22+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:08:28 deploy, no choice, lang=hu -> 409 use_offered=True use_desc='Az adatbázist innen töltjük vissza: saját mentés, 2026-09-28 10:06. A régi fájlokkal indítjuk az alkalmazást. Ami ezután változott, hiányozhat.' use_off=''
10:08:31 deploy, no choice, lang=en -> 409 use_offered=True use_desc='We load the database from its own backup, 2026-09-28 10:06 and start the app with the old files. Changes made after that time can be missing.' use_off=''
10:08:31 still not installed: False
10:08:34 # kp use — 2026-09-28T10:08:32+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:08:40 deploy kept_data=use -> 202 {"ok": true, "message": "A(z) Nextcloud megőrzött adatainak betöltése elindult."}
10:09:01 deployed after (20.1, 'degraded')
10:09:38 controller lines:
2026/09/28 08:08:40 kept_install.go:97: [INFO] [api] Deploy nextcloud: USE MY KEPT DATA — loading from /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (tier 1, 2026-09-28T08:06:00Z) under the kept files [/mnt/felhom-drives/scratch_hdd/appdata/nextcloud]
2026/09/28 08:08:40 restore_unit.go:538: [WARN] [backup] Restore nextcloud: generated replacement for [NEXTCLOUD_ADMIN_PASSWORD] — the credential was reset (old value unrecoverable); no data-encrypting key was involved, but a regenerated DATABASE password will not match the restored data directory's stored hash (R-12
2026/09/28 08:08:40 restore_unit.go:393: [INFO] [backup] Restoring nextcloud from recovery unit /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud: images=3, secrets recovered=2/3, data_keys=0
2026/09/28 08:08:41 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_db_data for nextcloud
2026/09/28 08:08:42 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_html for nextcloud
2026/09/28 08:08:49 restore.go:152: [INFO] [backup] Restoring Docker volume nextcloud_nextcloud_redis_data for nextcloud
2026/09/28 08:08:49 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for nextcloud
2026/09/28 08:08:49 deploy.go:718: [INFO] [stacks] nextcloud: the restore generated [NEXTCLOUD_ADMIN_PASSWORD] — the app's own login came back with its data; the page will not show the new value as the password
2026/09/28 08:08:50 restore_db.go:87: [INFO] [backup] Restore nextcloud: replaying DB dump into nextcloud-db (mariadb)
2026/09/28 08:08:56 restore_db.go:97: [INFO] [backup] Restore nextcloud: replayed 1 DB dump(s)
2026/09/28 08:09:11 restore_unit.go:478: [INFO] [backup] Restore-from-unit completed: nextcloud — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
2026/09/28 08:09:11 kept_load.go:331: [INFO] [backup] kept load nextcloud from /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud done in 30s (volumes 3/3, dbs 1/1)
2026/09/28 08:09:12 kept.go:521: [INFO] [stacks] after_load nextcloud: nextcloud [php occ files:scan --all] in 656ms (err=<nil>): Starting scan for user 1 out of 2 (admin)
10:09:38 restore status: {"running": false, "op": "restore", "stack": "nextcloud", "started_at": "2026-09-28T08:08:40.940253282Z", "last": {"op": "restore", "stack": "nextcloud", "ok": true, "message": "A(z) Nextcloud a megőrzött adataival fut.", "finished_at": "2026-09-28T08:09:11.382296795Z"}, "last_recent": true}
10:09:41 # kp read — 2026-09-28T10:09:39+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:09:48 occ user:info — present ['drill2bf734'], absent [] (+ a uid that cannot exist, the control): {'drill2bf734': True, 'nobody85b783': False}
10:09:49 PROPFIND as the seeded user -> 207; present ['before-backup.txt']: [True]; absent []: []; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
10:09:52 state: running images: nextcloud=nextcloud:34.0.4-apache nextcloud-redis=redis:7-alpine nextcloud-db=mariadb:12.3
@@ -0,0 +1,39 @@
10:10:08 # kp remove-keep delete-backups — 2026-09-28T10:10:04+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:10:10 stop 200
10:10:53 remove (keep drive data, remove_backups=True) -> 200 {"ok": true, "data": {"removed": "nextcloud", "volumes_removed": ["nextcloud_nextcloud_db_data", "nextcloud_nextcloud_html", "nextcloud_nextcloud_redis_data"], "hdd_paths_removed": [], "hdd_paths_preserved": ["/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)"], "backup_paths_removed": ["/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (928M)"], "verified": true}, "message": "Stack n
10:11:04 after: deployed= False | appdata: nextcloud paperless
ls: cannot access '/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud': No such file or directory
ls: cannot access '/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps': No such file or directory
ls: cannot access '/mnt/*/*/backups/secondary/nextcloud': No such file or directory
ls: cannot access '/mnt/*/backups/secondary/nextcloud': No such file or directory
10:11:07 # kp kept-delete — 2026-09-28T10:11:04+0200; guest 9202; controller gitea.dooplex.hu/admin/felhom-controller:0.277.0
10:11:07 kept nextcloud items: ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud']
10:11:07 POST /kept-data/delete /mnt/felhom-drives/scratch_hdd/appdata/nextcloud -> HTTP/2 302 ['location: /kept-data?flash=A%28z%29+Nextcloud+meg%C5%91rz%C3%B6tt+adatai+t%C3%B6r%C3%B6lve.']
10:11:10 appdata now: ls: cannot access '/mnt/felhom-drives/scratch_hdd/kept': No such file or directory /mnt/felhom-drives/scratch_hdd/appdata: nextcloud paperless
ls: cannot access '/opt/docker/stacks/nextcloud/app.yaml': No such file or directory
0
0
/mnt/felhom-drives/scratch_hdd/backups/primary/:
/mnt/felhom-drives/scratch_hdd/userdata/:
calibre-web
documents
downloads
emby
immich
jellyfin
komga
media
navidrome
nextcloud
paperless-ngx
plex
radarr
romm
roms
sonarr
/mnt/felhom-drives/scratch_hdd/userdata/nextcloud
removed-empty-harness-dir
@@ -0,0 +1,3 @@
before: gitea.dooplex.hu/admin/felhom-controller:0.276.0
gitea.dooplex.hu/admin/felhom-controller:0.277.0
gitea.dooplex.hu/admin/felhom-controller:0.277.0 Up 25 seconds (healthy)
@@ -0,0 +1,9 @@
## 2026-09-28T08:11:50Z floor 0.276.0 -> 0.277.0 (min_agent 0.131.0 from the v0.277.0 header); golden vouched 0.276.0
HTTP/1.1 303 See Other
Location: /configuration?flash=floor_set
name="min_controller_version" value="0.277.0"
## 2026-09-28T08:12:31Z both demo boxes, read from the boxes:
felhom-pve 9201: gitea.dooplex.hu/admin/felhom-controller:0.277.0 Up 28 seconds (healthy)
demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.277.0 Up 24 seconds (healthy)
2026/09/28 10:11:53 [INFO] managed floor SERVED for demo-felhom: floor 0.277.0, agent requirement "0.131.0" from declared (golden 0.276.0)
2026/09/28 10:11:53 [INFO] managed floor SERVED for demo-hp: floor 0.277.0, agent requirement "0.131.0" from declared (golden 0.276.0)
@@ -0,0 +1,7 @@
# after restoring the fix
--- PASS: TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest (0.00s)
--- PASS: TestR691_OffsiteLoadDownloadsJudgesThenRestores (0.00s)
--- PASS: TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone (0.01s)
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.024s
--- PASS: TestR691_InstallChoiceNamesTheOffsiteCopy (0.02s)
ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.027s
@@ -0,0 +1,31 @@
# RP1 — KeptBestCopy returns KeptDBCopy alone (0.276.0's choice: the off-site copy is never looked at)
189,191c189
< local, lok := m.KeptDBCopy(app, drive)
< off, ook := m.KeptOffsiteCopy(ctx, app)
< return KeptNewer(local, lok, off, ook)
---
> return m.KeptDBCopy(app, drive) // RED-PROOF RP1
=== RUN TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest
r691_kept_offsite_test.go:114: only off-site copy: got {UnitDir: Tier:0 Time:0001-01-01 00:00:00 +0000 UTC DriveLabel: SnapshotID: UnitPath:} ok=false, want the off-site snapshot ab12cd34
--- FAIL: TestR691_OffsiteCopyIsOfferedWhenItIsTheOnlyOrNewest (0.00s)
=== RUN TestR691_OffsiteLoadDownloadsJudgesThenRestores
r691_kept_offsite_test.go:162: no off-site copy offered
--- FAIL: TestR691_OffsiteLoadDownloadsJudgesThenRestores (0.00s)
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/data_version_not_recorded
r691_kept_offsite_test.go:210: no off-site copy offered
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/taken_of_another_drive
r691_kept_offsite_test.go:210: no off-site copy offered
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/definition_is_not_the_data's
r691_kept_offsite_test.go:210: no off-site copy offered
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/download_failed
r691_kept_offsite_test.go:210: no off-site copy offered
--- FAIL: TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone (0.00s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.012s
=== RUN TestR691_InstallChoiceNamesTheOffsiteCopy
kept_install_test.go:147: en: use_offered=false use_desc="" use_off="There is no backup of the database, so the app cannot load the old files. You can look at the files under Kept data."
--- FAIL: TestR691_InstallChoiceNamesTheOffsiteCopy (0.02s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.024s
FAIL
@@ -0,0 +1,15 @@
# RP2 — the data-version refusal removed (a unit with no data block would be loaded blind)
240c240
< if man == nil || man.Data == nil {
---
> if man == nil { // RED-PROOF RP2
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/data_version_not_recorded
r691_kept_offsite_test.go:218: the load reported success
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/taken_of_another_drive
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/definition_is_not_the_data's
=== RUN TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone/download_failed
--- FAIL: TestR691_OffsiteLoadRefusalsLeaveTheKeptFilesAlone (0.01s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.019s
FAIL
@@ -0,0 +1,5 @@
# RP3 — observed, not staged: the FIRST draft of LoadKeptOffsite removed the downloaded copy in a `defer`, i.e. AFTER
# EndRestoreOp and after(ok) had already reported the outcome. The tests failed on it before the fix (copied from the run):
r691_kept_offsite_test.go:184: the downloaded copy was left behind at /tmp/TestR691_OffsiteLoadDownloadsJudgesThenRestores2898855198/002/backups/offsite-proof/nextcloud (err=<nil>)
r691_kept_offsite_test.go:233: the downloaded copy was left behind after a refusal (err=<nil>) (x3 sub-cases)
# Fix: cleanup() runs before fail()/EndRestoreOp on every path.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,190 @@
"""kp.py <step> [arg] — kept data + off-site restore, live, one step per call (2026-09-28, controller 0.277.0).
EVIDENCE, NOT PRODUCT. Every act goes through the product's own endpoints (the ones the pages call):
POST /api/stacks/nextcloud/deploy (with and without kept_data) · /stop · /remove · POST /backups/window
POST /kept-data/delete · GET /kept-data · the restore pages' form posts
and the app's OWN front doors for data: `occ user:add` inside nextcloud's container, and WebDAV as the
seeded user (R-156: nothing planted in a volume, no SQL).
Env: GUEST=9202|9201 (demo-hp), SC=<scratchpad> (seed tokens live there, never in the evidence tree),
OUT=<evidence file>. The throwaway app is nextcloud; the sub-domain is SUB (default c-nc).
"""
import json, os, re, subprocess, sys, time
import walk as w, fixtures
step = sys.argv[1]
arg = sys.argv[2] if len(sys.argv) > 2 else ""
APP = "nextcloud"
SUB = os.environ.get("SUB", "c-nc")
HDD = os.environ.get("HDD", {"9202": "/mnt/felhom-drives/scratch_hdd", "9201": "/mnt/felhom-drives/hdd_1"}[w.GUEST])
w.DRIVE = HDD + "/userdata" # the helper hard-codes 9202's drive for the folders a deploy needs
TOK = os.path.join(w.SC, f"seed-tokens-{w.GUEST}.json")
OUT = open(os.environ["OUT"], "a", buffering=1)
def say(*a):
w.say(*a)
OUT.write(time.strftime("%H:%M:%S ") + " ".join(map(str, a)) + "\n")
def g(cmd, timeout=600):
return w.guest(cmd, timeout=timeout)
def toks():
try:
return json.load(open(TOK))
except Exception:
return {}
def save_toks(t):
old = os.umask(0o077)
json.dump(t, open(TOK, "w"), default=str)
os.umask(old)
def deploy(choice=None, lang=""):
vals = w.deploy_values(APP, SUB)
vals["HDD_PATH"] = HDD
body = {"values": vals}
if choice:
body["kept_data"] = choice
return w.ctl("POST", f"/api/stacks/{APP}/deploy" + ("?lang=en" if lang else ""), body)
def wait_deployed(limit=900):
t0 = time.time()
while time.time() - t0 < limit:
st = w.stack(APP)
if st.get("deployed") and (st.get("app_config") or {}).get("pinned_images") and st.get("state") in ("running", "unhealthy", "degraded"):
return round(time.time() - t0, 1), st.get("state")
time.sleep(5)
return None, w.stack(APP).get("state")
def dav(mode, name=""):
"""A file through Nextcloud's OWN WebDAV as the seeded user."""
t = toks()["nextcloud"]
base = f"{w.BASE}/remote.php/dav/files/{t['uid']}/"
a = ["curl", "-sk", "--max-time", "60", "-u", f"{t['uid']}:{t['pw']}", "-H", f"Host: {SUB}.{w.DOMAIN}", "-w", "\n%{http_code}"]
if mode == "put":
r = subprocess.run(a + ["-X", "PUT", "--data-binary", f"kept-offsite {name} {time.strftime('%FT%T')}", base + name], capture_output=True, text=True)
return r.stdout.strip().splitlines()[-1] if r.stdout.strip() else "?"
r = subprocess.run(a + ["-X", "PROPFIND", "-H", "Depth: 1", base], capture_output=True, text=True)
lines = r.stdout.strip().splitlines()
code = lines[-1] if lines else "?"
names = sorted(set(re.findall(r"<d:href>[^<]*/([^/<]+)</d:href>", r.stdout)))
return code, names
def files_report(expect_present, expect_absent):
code, names = dav("ls")
ok = code == "207" and all(n in names for n in expect_present) and not any(n in names for n in expect_absent)
say(f"PROPFIND as the seeded user -> {code}; present {expect_present}: {[n in names for n in expect_present]}; "
f"absent {expect_absent}: {[n not in names for n in expect_absent]}; listing={names}")
return ok
def users_report(want_present, want_absent):
nc = fixtures.FIXTURES["nextcloud"]
res = {}
for u in want_present + want_absent + ["nobody" + os.urandom(3).hex()]:
got, _ = nc._info(w, u)
res[u] = got
ok = all(res[u] is True for u in want_present) and all(res[u] is False for u in want_absent)
say(f"occ user:info — present {want_present}, absent {want_absent} (+ a uid that cannot exist, the control): {res}")
return ok
w.login()
say(f"# kp {step} {arg} — {time.strftime('%FT%T%z')}; guest {w.GUEST}; controller {g('cat /etc/felhom-controller-image').strip()}")
if step == "deploy":
st = w.stack(APP)
say("before: deployed=", st.get("deployed"), "| appdata:", g(f"ls {HDD}/appdata 2>&1 | tr '\\n' ' '"))
code, d = deploy()
say(f"deploy (no choice) -> {code} {json.dumps(d, ensure_ascii=False)[:300]}")
say("deployed after", wait_deployed())
elif step == "seed":
t = toks()
t["nextcloud"] = fixtures.FIXTURES["nextcloud"].seed(w, SUB, say)
save_toks(t)
say("seeded (the account):", t["nextcloud"] is not None, "uid=", (t["nextcloud"] or {}).get("uid"))
say(f"PUT before-backup.txt -> {dav('put', 'before-backup.txt')}")
files_report(["before-backup.txt"], [])
elif step == "marker":
# Written AFTER the snapshot that will be restored: a second account + a file.
nc = fixtures.FIXTURES["nextcloud"]
uid = "marker" + os.urandom(3).hex()
out = g(f"docker exec -u www-data -e OC_PASS=Marker-{os.urandom(8).hex()} nextcloud php occ user:add --password-from-env {uid} 2>&1")
say(f"occ user:add {uid} :: {' '.join(out.split())[:160]}")
t = toks(); t["marker"] = uid; save_toks(t)
say(f"PUT after-snapshot.txt -> {dav('put', 'after-snapshot.txt')}")
users_report([t["nextcloud"]["uid"], uid], [])
files_report(["before-backup.txt", "after-snapshot.txt"], [])
elif step == "read":
t = toks()
present = [t["nextcloud"]["uid"]] + ([] if arg == "no-marker" else ([t["marker"]] if t.get("marker") else []))
absent = [t["marker"]] if (arg == "no-marker" and t.get("marker")) else []
users_report(present, absent)
files_report(["before-backup.txt"], [])
say("state:", w.stack(APP).get("state"), "images:", g(f"docker ps --filter label=com.docker.compose.project={APP} --format '{{{{.Names}}}}={{{{.Image}}}}' | tr '\\n' ' '"))
elif step == "ask":
for lang in ("", "en"):
code, d = deploy(None, lang)
dd = d.get("data") or {}
say(f"deploy, no choice, lang={lang or 'hu'} -> {code} use_offered={dd.get('use_offered')} use_desc={dd.get('use_desc')!r} use_off={dd.get('use_off')!r}")
say("still not installed:", w.stack(APP).get("deployed"))
elif step == "use":
since = g("date -u +%Y-%m-%dT%H:%M:%SZ").strip()
code, d = deploy("use")
say(f"deploy kept_data=use -> {code} {json.dumps(d, ensure_ascii=False)[:240]}")
say("deployed after", wait_deployed(1500))
for i in range(90):
if int((g(f"docker logs --since {since} felhom-controller 2>&1 | grep -c -E 'after_load|kept load .* FAILED'").strip() or "0")) > 0:
break
time.sleep(5)
time.sleep(15)
say("controller lines:\n" + g(f"docker logs --since {since} felhom-controller 2>&1 | grep -iE 'kept|after_load|restor|offbox|snapshot' | grep -v DEBUG | cut -c1-320 | head -50"))
say("restore status:", json.dumps(w.ctl("GET", "/api/backup/restore-status")[1].get("data"), ensure_ascii=False)[:400])
elif step == "remove-keep":
# arg: "keep-backups" (default) or "delete-backups"
code, d = w.ctl("POST", f"/api/stacks/{APP}/stop"); say("stop", code)
time.sleep(15)
rb = arg == "delete-backups"
code, d = w.ctl("POST", f"/api/stacks/{APP}/remove", {"remove_hdd_data": False, "remove_backups": rb})
say(f"remove (keep drive data, remove_backups={rb}) -> {code} {json.dumps(d, ensure_ascii=False)[:400]}")
time.sleep(8)
say("after: deployed=", w.stack(APP).get("deployed"), "| appdata:", g(f"ls {HDD}/appdata | tr '\\n' ' '; echo; ls -la --time-style=+%FT%T {HDD}/backups/primary/{APP} {HDD}/backups/primary/{APP}/db-dumps 2>&1 | tail -8; ls -d /mnt/*/*/backups/secondary/{APP} /mnt/*/backups/secondary/{APP} 2>&1"))
elif step == "window":
sess = open(f"{w.SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{w.SC}/csrf{os.getpid()}.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", f"window_start={arg}", f"{w.BASE}/backups/window"])
say(f"POST /backups/window window_start={arg} -> {r.stdout}")
time.sleep(3)
say(g("docker logs --since 1m felhom-controller 2>&1 | grep -E 'rescheduled|window set' | tail -6"))
elif step == "logs":
say(g(f"docker logs --since {arg} felhom-controller 2>&1 | grep -iE '{APP}|job|offbox|snapshot|kept' | grep -v DEBUG | cut -c1-300 | tail -80"))
elif step == "kept":
for lang in ("", "en"):
h = w.page("/kept-data" + ("?lang=en" if lang else ""))
t = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ", h))
i = t.find("extcloud")
say(f"GET /kept-data{'?lang=en' if lang else ''}: ...{t[max(0, i-200):i+500]}..." if i >= 0 else f"GET /kept-data{'?lang=en' if lang else ''}: no nextcloud row")
elif step == "kept-delete":
h = w.page("/kept-data")
paths = sorted(set(re.findall(r'name="path" value="([^"]*nextcloud[^"]*)"', h)))
say("kept nextcloud items:", paths)
sess = open(f"{w.SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{w.SC}/csrf{os.getpid()}.txt").read().strip()
for p in paths:
r = w.sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", f"path={p}", "--data-urlencode", "confirm=Nextcloud",
f"{w.BASE}/kept-data/delete"])
loc = [l for l in r.stdout.split("\n") if l.lower().startswith("location:")]
say(f"POST /kept-data/delete {p} -> {r.stdout.splitlines()[0] if r.stdout else '?'} {loc[:1]}")
say("appdata now:", g(f"ls {HDD}/appdata {HDD}/kept 2>&1 | tr '\\n' ' '"))
else:
sys.exit(f"unknown step {step}")
@@ -0,0 +1,507 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = os.environ.get('SC', '/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/a061fcea-2c08-478c-9fd6-991dc8e748f0/scratchpad')
EV = os.environ.get('EV', '/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/kept-offsite-2026-09-28')
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
import secrets as _s
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Serialised across concurrent walks: one git working tree, one lock."""
import fcntl
with open(f"{SC}/drill.lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
return _drill_bump(app, frm, to, service_hint)
def _drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
@@ -0,0 +1,9 @@
## 2026-09-28T08:33:16Z bench create — BEFORE
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
local-lvm lvmthin active 56487936 31384697 25103238 55.56%
nvme-scratch dir active 983379700 50609912 882743176 5.15%
total used free shared buff/cache available
Mem: 29 6 4 0 19 23
Configuration file 'nodes/demo-hp/lxc/9401.conf' does not exist
@@ -0,0 +1,17 @@
+ pveam download local debian-13-standard_13.6-1_amd64.tar.zst
+ tail -1
download of 'http://download.proxmox.com/images/system/debian-13-standard_13.6-1_amd64.tar.zst' to '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' finished
+ pct create 9401 local:vztmpl/debian-13-standard_13.6-1_amd64.tar.zst --hostname upgrade-harness --cores 6 --memory 10240 --swap 0 --rootfs nvme-scratch:40 --net0 name=eth0,bridge=vmbr0,ip=dhcp --unprivileged 1 --features nesting=1,keyctl=1 --onboot 0
+ tail -2
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:mDMV8ol8xGZRh95eFLQBtzrp/wlPiAmCskPzerTpe7Y root@upgrade-harness
+ pct start 9401
+ sleep 12
+ pct exec 9401 -- bash -c 'hostname; ip -4 -o addr show eth0 | awk "{print \$4}"; apt-get update -qq >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -qq docker.io docker-compose python3 python3-yaml curl >/dev/null 2>&1; docker --version; docker compose version; python3 --version; docker network create traefik-public >/dev/null && echo traefik-public-net; curl -s -o /dev/null -w net-%{http_code} https://ghcr.io/'
upgrade-harness
192.168.0.157/24
Docker version 26.1.5+dfsg1, build a72d7cd
Docker Compose version 2.26.1-4
Python 3.13.5
traefik-public-net
net-301
@@ -0,0 +1,3 @@
10:34:47 START claper ['claper-postgres=postgres:17-alpine']
10:34:56 DONE claper rc=1 verdict=None read_before=None read_after=None abort=None marks=None peaks={} detail='None' (9s)
10:34:56 QUEUE q1 FINISHED
@@ -0,0 +1 @@
10:35:51 START claper ['claper-postgres=postgres:17-alpine']
@@ -0,0 +1,45 @@
start ('200', {'ok': True, 'data': {'state': 'starting'}, 'message': 'Stack calcom start requested — state now: starting'})
calcom Up 2 seconds (health: starting)
calcom-postgres Up About a minute (healthy)
restarts=4 oom=false exit=0
📲 Updated app: 'zoom'
+ yarn start
• turbo 2.7.1
• Packages in scope: @calcom/web
• Running start in 1 packages
• Remote caching disabled
/calcom/node_modules/turbo/bin/turbo:273
throw e;
^
Error: Command failed: /calcom/node_modules/turbo-linux-64/bin/turbo run start --filter=@calcom/web
at genericNodeError (node:internal/errors:984:15)
at wrappedFn (node:internal/errors:538:14)
at checkExecSyncError (node:child_process:891:11)
at Object.execFileSync (node:child_process:927:15)
at Object.<anonymous> (/calcom/node_modules/turbo/bin/turbo:266:17)
at Module._compile (node:internal/modules/cjs/loader:1521:14)
at Module._extensions..js (node:internal/modules/cjs/loader:1623:10)
at Module.load (node:internal/modules/cjs/loader:1266:32)
at Module._load (node:internal/modules/cjs/loader:1091:12)
at Function.executeUserEntryPoint [as runMain] (node:internal/modules/run_main:164:12) {
status: null,
signal: 'SIGKILL',
output: [ null, null, null ],
pid: 137,
stdout: null,
stderr: null
}
Node.js v20.20.0
+ scripts/replace-placeholder.sh http://localhost:3000 https://c-cal.enkisfelhom.hu
Replacing all statically built instances of http://localhost:3000 with https://c-cal.enkisfelhom.hu.
+ scripts/wait-for-it.sh -- echo database is up
Error: you need to provide a host and port to test.
Usage:
scripts/wait-for-it.sh host:port|url [-t timeout] [-- command args]
-q | --quiet Do not output any status messages
-t TIMEOUT | --timeout=timeout Timeout in seconds, zero for no timeout
-- COMMAND ARGS Execute command with args after the test finishes
+ npx prisma migrate deploy --schema /calcom/packages/prisma/schema.prisma
@@ -0,0 +1,3 @@
10:17:34 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:21:00 [1] deployed, controller state=running, pinned={'calcom': 'calcom/cal.com:v6.2.0', 'calcom-postgres': 'postgres:16-alpine'}
deployed True
@@ -0,0 +1,33 @@
2026/09/28 08:17:34 router.go:431: [INFO] [api] Deploy requested for stack: calcom
2026/09/28 08:17:34 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:17:34 deploy.go:442: [INFO] [stacks] Deploying stack calcom with 5 env vars: [DOMAIN, SUBDOMAIN, NEXTAUTH_SECRET, CALENDSO_ENCRYPTION_KEY, DB_PASSWORD]
2026/09/28 08:17:34 manager.go:1601: [INFO] [stacks] Deploying stack calcom — checking 2 images...
2026/09/28 08:20:26 deploy.go:532: [INFO] [stacks] Stack calcom deployed successfully (took 171.7s)
2026/09/28 08:20:26 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:20:26 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:20:26 installed.go:420: [INFO] [stacks] installed-images calcom: recorded 2 service(s) (calcom=calcom/cal.com:v6.2.0 (sha256:ace3bb1219fb…), calcom-postgres=postgres:16-alpine (sha256:721873c34ceb…))
2026/09/28 08:20:26 [INFO] [stacks] SaveAppConfig: saved config for calcom
2026/09/28 08:20:26 pin.go:93: [INFO] [stacks] pin calcom: calcom=calcom/cal.com:v6.2.0, calcom-postgres=postgres:16-alpine
2026/09/28 08:20:29 manager.go:1562: [INFO] [stacks] Stack calcom post-start status:
2026/09/28 08:20:29 manager.go:1565: [INFO] [stacks] calcom calcom/cal.com:v6.2.0 running Up 3 seconds (health: starting)
2026/09/28 08:20:29 manager.go:1565: [INFO] [stacks] calcom-postgres postgres:16-alpine running Up 9 seconds (healthy)
2026/09/28 08:22:29 main.go:3762: [WARN] [deadapp] calcom: crash_loop — 6 in 10m0s; STOPPING it (decision 28)
2026/09/28 08:22:30 manager.go:1305: [INFO] [stacks] Stopping stack: calcom
2026/09/28 08:22:40 manager.go:1314: [INFO] [stacks] Stack calcom stopped successfully (took 10.8s)
2026/09/28 08:22:40 settings.go:1881: [WARN] [settings] restore hold SET for calcom — the app stays stopped until it is cleared
2026/09/28 08:22:40 update_guard.go:827: [WARN] [backup] calcom is STOPPED by the box: crash_loop (trip 1 within 24h0m0s) — Start gives it one more try (decision 28)
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
2026/09/28 08:24:30 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/calcom/docker-compose.yml
claper Up 9 minutes (healthy)
claper-postgres Up 10 minutes (healthy)
filebrowser Up 15 minutes (healthy)
privatebin Up 19 minutes (healthy)
paperless-webserver Up 19 minutes (healthy)
paperless-postgres Up 20 minutes (healthy)
paperless-redis Up 20 minutes (healthy)
felhom-controller Up 27 minutes (healthy)
traefik Up 3 days
@@ -0,0 +1,33 @@
cgroup=/sys/fs/cgroup/system.slice/docker-6cc2be9f86e8fda3e63e61be599d7c4cd616a0f4ecbe6f0c5b77b229bf00a1e5.scope
805306368
08:28:27 cur=593969152 oom 0 oom_kill 0 anon 81022976 restarts=5
08:28:31 cur=683855872 oom 0 oom_kill 0 anon 158216192 restarts=5
08:28:35 cur=805126144 oom 0 oom_kill 0 anon 662142976 restarts=5
08:28:39 cur=507691008 oom 0 oom_kill 0 anon 17334272 restarts=6
08:28:43 cur=646803456 oom 0 oom_kill 0 anon 132259840 restarts=6
08:28:48 cur=805306368 oom 0 oom_kill 0 anon 707543040 restarts=6
08:28:52 cur=11423744 oom 1 oom_kill 1 anon 3809280 restarts=6
08:28:56 cur=608956416 oom 0 oom_kill 0 anon 104116224 restarts=7
08:29:00 cur=681771008 oom 0 oom_kill 0 anon 166567936 restarts=7
08:29:04 cur=804466688 oom 0 oom_kill 0 anon 626081792 restarts=7
08:29:08 cur=16332591104 oom 16 oom_kill 14 anon 982835200
08:29:12 cur=16284028928 oom 16 oom_kill 14 anon 954310656
08:29:16 cur=16285847552 oom 16 oom_kill 14 anon 956977152
08:29:20 cur=16286691328 oom 16 oom_kill 14 anon 957399040
08:29:24 cur=16285790208 oom 16 oom_kill 14 anon 957583360
08:29:28 cur=16285872128 oom 16 oom_kill 14 anon 957526016
08:29:33 cur=16301551616 oom 16 oom_kill 14 anon 956796928
08:29:37 cur=16290660352 oom 16 oom_kill 14 anon 956542976
08:29:41 cur=16290639872 oom 16 oom_kill 14 anon 956497920
08:29:45 cur=16291016704 oom 16 oom_kill 14 anon 956428288
08:29:49 cur=16290803712 oom 16 oom_kill 14 anon 957030400
08:29:53 cur=16290611200 oom 16 oom_kill 14 anon 957075456
08:29:57 cur=16291708928 oom 16 oom_kill 14 anon 957132800
08:30:01 cur=16292163584 oom 16 oom_kill 14 anon 955760640
08:30:05 cur=16291311616 oom 16 oom_kill 14 anon 955813888
08:30:09 cur=16289894400 oom 16 oom_kill 14 anon 954003456
08:30:14 cur=16290275328 oom 16 oom_kill 14 anon 955899904
08:30:18 cur=16291815424 oom 16 oom_kill 14 anon 956694528
08:30:22 cur=16291139584 oom 16 oom_kill 14 anon 956817408
08:30:26 cur=16291594240 oom 16 oom_kill 14 anon 956981248
@@ -0,0 +1,4 @@
Error response from daemon: No such container: calcom
Error response from daemon: No such container: calcom
@@ -0,0 +1,6 @@
10:26:10 app never answered on c-cal/api/auth/providers (last rc=0 code=404)
wait False
setup ('404', '404 page not found\n')
setup again (must refuse) ('404', '404 page not found\n')
GET /drilld1401a 404 19
GET /nobody20cc4f 404 19
@@ -0,0 +1,6 @@
10:30:45 [X] stop -> 200 {'ok': True, 'message': 'Stack calcom stop completed'}
10:31:17 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'calcom', 'volumes_removed': ['calcom_calcom_postgres_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': '
10:31:25 [X] after remove: deployed=False leftovers='/opt/docker/stacks/calcom'
0
ls: cannot access '/opt/docker/stacks/calcom/app.yaml': No such file or directory
@@ -0,0 +1,13 @@
claper 2.5.1
[sh -c /app/bin/claper eval Claper.Release.migrate && /app/bin/claper eval Claper.Release.seeds && /app/bin/claper start] []
warning: the VM is running with native name encoding of latin1 which may cause Elixir to malfunction as it expects utf8. Please ensure your locale is set to UTF-8 (which can be verified by running "locale" in your shell) or set the ELIXIR_ERL_OPTIONS="+fnu" environment variable
default admin exists: true
warning: the VM is running with native name encoding of latin1 which may cause Elixir to malfunction as it expects utf8. Please ensure your locale is set to UTF-8 (which can be verified by running "locale" in your shell) or set the ELIXIR_ERL_OPTIONS="+fnu" environment variable
default password claper works: true
warning: the VM is running with native name encoding of latin1 which may cause Elixir to malfunction as it expects utf8. Please ensure your locale is set to UTF-8 (which can be verified by running "locale" in your shell) or set the ELIXIR_ERL_OPTIONS="+fnu" environment variable
control: unknown email: false
08:16:36.932 [info] == Running 20240730123205 Claper.Repo.Migrations.RemoveIsAdminFromUsers.change/0 forward
Created admin role
Created default admin user:
Email: admin@claper.co
@@ -0,0 +1,9 @@
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
image: ghcr.io/claperco/claper:2.5
total used free shared buff/cache available
Mem: 25898 808 18374 39 6754 25089
10:15:41 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:16:52 [1] deployed, controller state=running, pinned={'claper': 'ghcr.io/claperco/claper:2.5', 'claper-postgres': 'postgres:16-alpine'}
deployed True
@@ -0,0 +1,4 @@
registered id=2
probe authenticates: true
control: wrong password: false
@@ -0,0 +1,8 @@
10:32:19 claper: register_user :: DRILL_SEEDED=3
10:32:23 claper: seeded user drill-e157296d@gate.invalid
seed ok True
10:32:30 claper: readback — the seeded account authenticates=True
verify True
10:32:37 claper: readback — the seeded account authenticates=False
10:32:37 claper: :: DRILL_ANSWER=false
verify with a token that was never seeded (must be False): False
@@ -0,0 +1,48 @@
#!/usr/bin/env python3
"""benchq.py <queue-name> <app> <svc=ref>[,<svc=ref>] [<app> <moves> ...]
Runs `upgrade-test.py --move` on bench LXC 9401 for each app IN ORDER, then copies that edge's
evidence OFF the bench into apps/<app>/bench/ before starting the next (R-320: the evidence leaves
the machine at the end of the phase that made it). One line per finished edge in benchq-<queue>.txt.
"""
import json, os, subprocess, sys, time
HERE = os.path.dirname(os.path.abspath(__file__))
BENCH = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/753351c4-7b70-46b6-9338-40951c9512c5/scratchpad/bench.sh"
q = sys.argv[1]
args = sys.argv[2:]
out = open(os.path.join(HERE, "..", os.environ.get("BQ", "B"), f"benchq-{q}.txt"), "a", buffering=1)
for i in range(0, len(args), 2):
app, moves = args[i], args[i + 1].split(",")
label = app
if "@" in app: # romm@engine: a second edge of one app gets its own folder
app, tag = app.split("@", 1)
label = f"{app}-{tag}"
t0 = time.time()
out.write(f"{time.strftime('%H:%M:%S')} START {app} {moves}\n")
pre = f"cd /opt/upg && [ -d evidence/MV-{app} ] && mv evidence/MV-{app} evidence/MV-{app}-$(date +%s); "
subprocess.run([BENCH], input=pre, capture_output=True, text=True, timeout=120)
cmd = (f"cd /opt/upg && timeout 2400 python3 upgrade-test.py --soak 600 --move {app} {' '.join(moves)} "
f"> /opt/upg/MV-{app}.log 2>&1; echo rc=$?")
r = subprocess.run([BENCH], input=cmd, capture_output=True, text=True, timeout=2700)
dest = os.path.join(HERE, "..", os.environ.get("BQ", "B"), "apps", label, "bench")
os.makedirs(dest, exist_ok=True)
pull = (f"cd /opt/upg && tar czf /tmp/ev-{app}.tgz MV-{app}.log evidence/MV-{app} 2>/dev/null; "
f"base64 -w0 /tmp/ev-{app}.tgz")
r2 = subprocess.run([BENCH], input=pull, capture_output=True, text=True, timeout=600)
import base64, io, tarfile
try:
tarfile.open(fileobj=io.BytesIO(base64.b64decode(r2.stdout.strip()))).extractall(dest)
except Exception as e:
out.write(f" evidence pull FAILED for {app}: {e}\n")
v = {}
try:
v = json.load(open(os.path.join(dest, "evidence", f"MV-{app}", "verdict.json")))
except Exception:
pass
mem = v.get("memory") or {}
peaks = {n: c.get("peak_pct") for n, c in (mem.get("containers") or {}).items()}
out.write(f"{time.strftime('%H:%M:%S')} DONE {app} {r.stdout.strip()} verdict={v.get('verdict')} "
f"read_before={v.get('seed_read_before')} read_after={v.get('seed_read_after')} "
f"abort={v.get('abort')} marks={v.get('marks')} peaks={peaks} "
f"detail={str(v.get('abort_detail'))[:160]!r} ({round(time.time()-t0)}s)\n")
out.write(f"{time.strftime('%H:%M:%S')} QUEUE {q} FINISHED\n")
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,507 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = os.environ.get('SC', '/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/a061fcea-2c08-478c-9fd6-991dc8e748f0/scratchpad')
EV = os.environ.get('EV', '/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/pg-calcom-claper-2026-09-28')
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
import secrets as _s
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Serialised across concurrent walks: one git working tree, one lock."""
import fcntl
with open(f"{SC}/drill.lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
return _drill_bump(app, frm, to, service_hint)
def _drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
+4 -2
View File
@@ -808,11 +808,13 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). | **OPEN — P3, gaps (1)–(4) only; owner: CC** |
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** |
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** | **OPEN — P3; owner: CC — NARROWED to (2)** |
| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** **-- 2026-09-28 (controller v0.277.0): (2) BUILT** — `KeptBestCopy` offers the off-site copy when it is newer than every local copy or the only one; the page names the copy and its date; `LoadKeptOffsite` downloads the unit alone, refuses a unit of another drive, with no data, or with no recorded data version (`07` §6.5/§6.6), then restores. Red-proofed (`audits/kept-offsite-2026-09-28/redproofs/`). Tier 1 regression live on 9202 (the choice named „saját mentés, 2026-09-28 10:06”, seed + file back). Floor 0.277.0, both demo boxes. **STILL OPEN: the live off-site proof** — a throwaway nextcloud on demo-hp 9201 joined the off-site copy 2026-09-28 10:15; its first snapshot runs the night of 09-28/29 (Part E (b)). Tier 2 not runnable live (9202 has one drive). | **WATCHING — P3; owner: CC (live off-site proof)** |
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. | **OPEN — P3; owner: operator (which option), CC measures** |
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. **-- 2026-09-28 option (a) MEASURED — NOT ENOUGH:** `pct fstrim` 9201 + 9202 from the host (rc 0): `local-lvm` 62.32 % → **50.60 %**, free 20.3 → **26.6 GiB** — still under the 31 GiB the preflight needs. The pool is 53.9 GiB and guest 9201 itself holds ~26 GiB, so no trim can reach 31 GiB. The agent's own verdict after the trim: `tier=felhom-pbs due=true` (archive 2026-09-24T20:06:25Z, not proven). `nvme-scratch` has ~820 GiB free. **Left to the operator: (b) or (c).** `audits/evidence-golden-0276-2026-09-28/phaseD1-reclaim.txt` | **OPEN — P3; owner: operator ((b) or (c)); (a) measured, not enough** |
| **R-702** | **[P1-HIGH] Every claper install creates an admin `admin@claper.co` with the public password `claper`, and the app is published on the household's domain.** Measured 2026-09-28 on scratch guest 9202 (catalog `claper` template, `ghcr.io/claperco/claper:2.5` = 2.5.1): the image's own start command runs `Claper.Release.seeds`, which logs `Created default admin user: Email: admin@claper.co`; asked through claper's own CLI (`bin/claper rpc`), `get_user_by_email_and_password("admin@claper.co", "claper")` answered **true**, an unknown e-mail answered false (control). The template routes `<sub>.<domain>` through traefik and the tunnel, so any claper a household installs can be logged into by anyone who knows claper's README. Not measured: whether any box runs claper today (R-632 lists it as never deployed), whether upstream reads an env var for the seed admin. **Needs:** a decision on the fix shape — (a) the catalog passes a generated admin password (if upstream supports it), (b) the controller changes the seeded admin's password after the first start, (c) pull claper from the catalog until (a)/(b). Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-claper-default-admin.txt`. | **OPEN — P1; owner: operator (fix shape), CC implements** |
| **R-703** | **[P2] calcom v6.2.0 cannot start at its catalog memory limit — a fresh install crash-loops and the box stops it.** Measured 2026-09-28 on 9202: install from the live catalog → `crash_loop — 6 in 10m0s; STOPPING it (decision 28)`; one Start later, the container's own cgroup counted `oom_kill 1` per start at `memory.max` 805306368 (768 MiB) while `anon` reached ~700 MB during `turbo run start` (`signal: 'SIGKILL'`); Docker reported `OOMKilled=false` (R-528's shape). So calcom in the live catalog cannot run on any box. Not measured: the limit it needs. **Needs:** a measured limit (a bench watch at 1.5–2 GiB), then the catalog change. Until then calcom's PostgreSQL move is `inconclusive — the FROM version does not run`. Evidence: `audits/pg-calcom-claper-2026-09-28/box/C0-calcom-crash.txt`, `C0-calcom-memory.txt`. | **OPEN — P2; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
@@ -152,7 +152,10 @@ qemu-system-x86_64 -enable-kvm -cpu host -smp 4 -m 8192 \
1. Revert + boot per §4.0.
2. The debian template is **absent on `virgin`** and **the exact point release rots** — list the
current one (`pveam available --section system | grep debian-13`) and `pveam download local <that>`.
current one (`pveam available --section system | grep 'debian-13-standard_.*_amd64'`) and `pveam download local <that>`.
**Filter on `_amd64`:** the index also lists an `_arm64` build of the same point release, and a version sort
picks it (2026-09-28: the 0.276.0 bake's first attempt did, and was stopped before `pct create` finished; the
2026-09-26 bench LXC did too and would not start).
It was `debian-13-standard_13.6-1_amd64.tar.zst` on 2026-07-31.
3. `scp` in `build-golden.sh` (agent repo `configs/`) + the Gitea token (`~/.gitea-token` on 180,
0600), then run it as a transient unit so it survives a session close, reading the token from the