diff --git a/.claude/rules/unprompted-work.md b/.claude/rules/unprompted-work.md index 49df8ff8..816a2bd7 100644 --- a/.claude/rules/unprompted-work.md +++ b/.claude/rules/unprompted-work.md @@ -6,8 +6,8 @@ unconditional: true > Goal sessions, nightly sessions, "work the register" sessions. **A session that starts from > `/goal` or a standing brief inherits these rules exactly as it inherits the gates.** They are the > part of `PROMPT-TEMPLATE.md` that a task file used to carry and a goal does not. Same wording lives -> in `felhom.eu`, `felhom-controller` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace root's -> unversioned `.claude/rules/`; change all four or none. +> in `felhom.eu`, `felhom-controller`, `felhom-agent` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace +> root's unversioned `.claude/rules/`; change all five or none. ## 1. What you may pick up on your own diff --git a/CONTEXT.md b/CONTEXT.md index 63fcd675..2dcb210b 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -2746,7 +2746,9 @@ their work is carried as backlog rows (**R-165, R-166, R-167**), not as prose he returned ONE string and `BackupCadence()` ONE 24h window, so "local daily AND PBS weekly" was not expressible — which is why the DR tier was `applied` since 07-21 with **one** snapshot on demo-felhom and **zero, ever** on demo-hp. Now: `backup_targets[]` per-tier cadence+retention; - ONE quiesce window for both due tiers (never two app outages for one night); per-tier hub + ONE quiesce window for both due tiers (never two app outages for one night) — **REVERSED 2026-10-06 by `09` §3 + decision 156 (R-518 option A, controller v0.301.0): one short stop per tier, the apps start again at each tier's + `snapshotted`, the next tier waits for a later cycle; the button makes the local copy only** (`07` §6.4); per-tier hub thresholds (host 26h / offsite 8d); fresh-install default; an unprovisioned tier DEFERS. **Operator rulings:** 2-week offsite retention, first backup runs as long as it needs, one backup at a time per guest, drill box dropped from the rollout. diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 15c4599e..f25f3b3a 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -593,6 +593,23 @@ files are neither deleted nor hidden — their visibility is a separate recorded `CaptureRecoveryUnit`'s already-current early return fire, so anything that must happen on every capture — such as bounding the undo copies — has to sit ABOVE that check, not after it. +**[FACT] 2026-10-06 — restoring over a newer schema (R-638, `09` §3 decision 154).** The loader replays a copy ON TOP of +the live database, so it removes only what the copy knows about. **Measured safe on the two main paths** (scratch 9202, +controller 0.299.0, `audits/design-build-2026-10-06/B/`): a copy taken at the OLD version, then an Update that migrated, +then the household's restore of that copy — docmost 0.95.0 → 0.96.0 (42 → 48 tables) came back with exactly the copy's +42 and none of the 6 new ones; romm 5.0.0 → 5.3.0 (27 → 39) came back with exactly 27 (views counted), `Imported DB dump` +logged, the data read back, the old version pinned. They are safe because the unit puts the copy's OWN volumes back +before the replay. **Two side paths fixed by ORDER (controller v0.301.0):** the no-manifest fallback `RestoreApp` now +starts only the database services, replays, then starts the app (it started the whole stack at the CURRENT definition +first, so a newer app could migrate the old data before the replay); and a unit restore whose volume leg failed no +longer replays (the copy's dump over a database volume that was NOT put back). The loader is unchanged; no delete step +was added. **Known limit, not fixed (R-893):** after a failed OFF-SITE replay, the rollback loads the pre-restore copy +(the NEWER state) over the OLDER volume just put back (`offbox_reconstitute.go` rollback), so tables the newer version +removed stay, and when the snapshot's older definition was written, the rollback branch does not put the newer one +back — the older app starts on rolled-back data. An order change cannot fix it: the only undo is a logical dump and the +volume it should land in was replaced. Option B (a loader that rebuilds instead of overlays) or a pre-restore volume +copy would; both are larger than this ruling. + **[DESIGN] 2026-08-22 — the failure ladder of a database restore: replay → rollback → hold.** Recorded here rather than only in a closed register row, because a decision that survives only inside a closed work item is a decision nobody will find. @@ -672,7 +689,19 @@ successes only. After an agent restart the success is read back from the tier's **What a run may do** (R-518, cheap half). A tier the agent reports `storage: absent` is dropped before anything is stopped, logged, and reported once as `backup_tier_skipped`; `unknown` is never skipped. -**Still open:** quiescing per tier, so a slow second tier does not keep every app down. +**One stop per tier (controller v0.301.0, R-518 option A, `09` §3 decision 156 — this REVERSES R-82's „one quiesce +window for both due tiers, never two app outages for one night").** A window runs only the FIRST due tier and starts +the apps again at that tier's `snapshotted`; the upload then finishes with the apps running. Another due tier waits +for a later cycle, in its own short window. A tier that refuses to start (BUSY or an error) inside a window still lets +the next tier try in the same window — no copy has been made yet, so it is still one copy per stop. **The button +(„Mentés most") makes the LOCAL copy only**, with one short stop; the off-site copy follows at the next night run. If +the local tier's storage is absent the button makes no copy and stops nothing (logged at ERROR) — the off-site tier is +never its stand-in. **Measured, read-only, 2026-10-06:** demo-felhom's night off-site job started 06:21:08 and reached +`snapshotted` at 06:21:10, its one app running again at 06:21:18 — the off-site part of a stop is seconds; the rest is +the apps' own stop and start (demo-hp 2026-10-05: stop 21 s, start 46 s). **Not shown live:** a press under the new +rule (the scratch guest has no agent connection; the demo boxes take only deliveries and read-backs) — the first +night run on the demo boxes is the proof owed (R-518). The page states about 1–1.5 minutes (an estimate from the +parts above, not a measured press). **Measured 2026-10-05 on demo-hp (9 apps, controller v0.295.0):** „Mentés most" stopped the apps at 09:19:08Z, the local tier ran 09:19:29–09:24:09, the PBS tier was busy (the controller logged a retry in 15 min; no second stop was seen in the next 55 min), the last app was back at 09:24:55Z — the longest stop **5 min 47 s**, for the local tier alone. The button text and its confirm (v0.296.0) give both diff --git a/documentation/architecture/08-alarm-ladder.md b/documentation/architecture/08-alarm-ladder.md index 45e0c094..833a9c4e 100644 --- a/documentation/architecture/08-alarm-ladder.md +++ b/documentation/architecture/08-alarm-ladder.md @@ -236,6 +236,14 @@ was 4,530 kills in six hours ≈ 375 per 30 min; a hiccup is 1–3. Live on 9202 reached 21 kills 2.5 min after start and sent ONE storm; at 49 kills still one. **Not an app-down state:** the app still reads `running` and `IsDownState` is unchanged (§4). **Limit:** a container whose `OOMKilled` flag stays false in an LXC guest (R-528) is never read, so it never storms. +**Measured again 2026-10-06 on scratch 9202 (Docker 29.8.2; R-528's false flags were on 29.8.0, 2026-09-15): the flag +was TRUE in all four shapes tried** — a child process killed while the container kept running (`oom_kill` 0 → 3, +`OOMKilled=true`, still running) and three runs where the kernel also killed the main process (exit 137, +`OOMKilled=true`); `docker events` carried an `oom` event each time. So R-528's starting point did not reproduce on +today's engine, and its build (read the counter of every container; „probably out of memory" on exit 137) was STOPPED +before any code, by the brief's measure-first rule — `09` §3 decision 155 stands for the operator to confirm or drop. +Cost measured for the record: one `docker exec … cat memory.events` across demo-hp's 21 containers takes 1.6 s wall +(0.47 s user, 0.44 s sys); cloudflared has no `cat` and reads as unknown. `audits/design-build-2026-10-06/C/`. **The box STOPS a crash loop or an out-of-memory storm (decision 28 of `09` §3, R-667, controller v0.269.0 / hub v0.123.0, 2026-09-24).** Alarming was not enough: gokapi crash-looped for hours at 385 → 546 restarts diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 8cf9cc10..baf5092d 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -565,6 +565,16 @@ R-636's louder repeated alarm. is case-insensitive (termix's router ignores case). Final trick run: 113 tries on 11 apps, 0 got in. Same day, operator: CC changes the admin passwords of demo-hp's installed bookstack and calibre-web and stores them in the operator's credentials file (not in any repo). + **2026-10-06 (R-717, controller v0.301.0): the window reopens a switch a COMMAND closed.** `after_setup` gains + `open_command` / `open_success` (same `service`, `user`, `args_env` and argv-safe rules as `command`). The window + marks the lock `opening` on disk first, lifts the env and runs `open_command`; only full success records `lifted`. + A failed open closes it again at once and the app page says so (`app_info.signup_native_open_failed`). The close + (`command`) runs when the window ends, at every controller start (the loop's first pass) and after a successful + update (`markNativeLockForReapply`); a failed close is retried every 2 min and logged at ERROR. A template with + `command` and no `open_command` keeps today's window (address block + env) and logs once that its own switch cannot + reopen. **Wishlist** uses it (node's own sqlite module, `system_config.enableSignup` in group `global`), proven live + on 9202. **Opengist cannot**: its container has no sqlite tool and no script runtime, and its CLI has no settings + command — it keeps the address block alone (R-717 narrowed). 50. **Every newly installed app goes off-site by itself when the customer has off-site** — *operator ruling 2026-09-30 (R-720, option A).* It restores the intent of the 2026-09-16 ruling (`07` §6: Tier 3 ON for every new customer), which a per-app switch starting OFF had undone. If the apps will not fit the customer's quota, the page says so, @@ -905,6 +915,27 @@ its length, and both fixes cost something the household would notice — operato 127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer. *Operator ruling 2026-10-05.* (R-861) +### 2026-10-06 (14:24) — three operator rulings and the reviewer's three design picks (recorded before the work) + +151. **R-469 — closed.** The engine-major gate stays as decision 35's permanent per-app check. *Operator ruling + 2026-10-06 14:24* (option A). +152. **The agent repo gets its copy of the shared rule file** (`.claude/rules/unprompted-work.md`), identical to the other + copies. It adds rules and loosens nothing. *Operator ruling 2026-10-06 14:24* (option A). +153. **The design items are built in the next session** (option A of the reviewer's 2026-10-06 proposal). *Operator + ruling 2026-10-06 14:24.* +154. **R-638 — option A: measure, then fix the order on the three side paths.** If the measurement shows the two main + restore paths are safe, the row closes on A alone; option B (rebuild instead of overlay) stays a note in the row. The + reverse-direction rollback (exposure 3) is fixed in the same session if the order fix covers it; if not, it becomes + a known limit in `07` §6 and a row. *Reviewer's pick, 2026-10-06, standing (the operator saw it and did not change it).* +155. **R-528 — options A then C** (read the kill counter of every running container; the crash-loop alarm says + „probably out of memory" on exit 137). Option B (the host-side agent read) waits until A + C have run a week on the + demo boxes. *Reviewer's pick, 2026-10-06, standing.* +156. **R-518 — option A: one stop per backup tier.** Each window runs only the first due tier and resumes the apps at its + `snapshotted`; the next due tier waits for a later cycle. A manual press makes the LOCAL copy only, with one short + stop; the off-site copy follows at the next night run. **This reverses the recorded R-82 choice „ONE quiesce window + for both due tiers (never two app outages for one night)".** *Reviewer's pick, 2026-10-06, standing — the operator's + chat answer decides it; the brief carries it as A.* + ### 2026-10-06 (13:25) — two operator rulings (recorded before the work) 149. **R-890 — the admin seed is allowed on the test boxes** (scratch 9202 and the disposable Tester 1 box) as well as diff --git a/documentation/audits/design-build-2026-10-06/B/B0-repoint-drill.txt b/documentation/audits/design-build-2026-10-06/B/B0-repoint-drill.txt new file mode 100644 index 00000000..5e311079 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/B0-repoint-drill.txt @@ -0,0 +1,2 @@ +26: repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + diff --git a/documentation/audits/design-build-2026-10-06/B/docmost-run1-installed-live.txt b/documentation/audits/design-build-2026-10-06/B/docmost-run1-installed-live.txt new file mode 100644 index 00000000..27426477 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/docmost-run1-installed-live.txt @@ -0,0 +1,10 @@ +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.299.0 +drill: 8647f6f DRILL docmost: the OLD definition (6d8cd87^) for R-638 slice 0 | images: ['docmost/docmost:0.95.0', 'postgres:16-alpine', 'redis:7-alpine'] + docmost: /api/auth/setup http=200 rc=0 + docmost: login as the seeded user http=200 ok=True +[2] old version installed, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}, live tables 48 +[3] the ladder entry used: [({'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}, {'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'})] +drill: fb47929 DRILL docmost: the NEW definition (6d8cd87) + its one ladder entry, R-638 slice 0 | images: ['docmost/docmost:0.96.0', 'postgres:16-alpine', 'redis:7-alpine'] +[3] update -> None, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}, live tables 48, added by the migration: 0 [] + docmost: login as the seeded user http=200 ok=True +RESULT the update did not migrate (no done, or no new table) — slice 0 cannot measure; stopping before restore diff --git a/documentation/audits/design-build-2026-10-06/B/docmost-run1-tables.json b/documentation/audits/design-build-2026-10-06/B/docmost-run1-tables.json new file mode 100644 index 00000000..e98be00e --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/docmost-run1-tables.json @@ -0,0 +1,117 @@ +{ + "app": "docmost", + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.299.0", + "pinned_old": { + "docmost": "docmost/docmost:0.96.0", + "docmost-postgres": "postgres:18-alpine", + "docmost-redis": "redis:7-alpine" + }, + "tables_old_live": [ + "ai_chat_messages", + "ai_chats", + "api_keys", + "attachments", + "audit", + "auth_accounts", + "auth_providers", + "backlinks", + "base_properties", + "base_rows", + "base_views", + "billing", + "comments", + "favorites", + "file_tasks", + "group_users", + "groups", + "kysely_migration", + "kysely_migration_lock", + "labels", + "notifications", + "oauth_authorization_codes", + "oauth_clients", + "oauth_grants", + "oauth_tokens", + "page_access", + "page_history", + "page_labels", + "page_permissions", + "page_transclusion_references", + "page_transclusions", + "page_verifications", + "page_verifiers", + "pages", + "public_spaces", + "scim_tokens", + "shares", + "siem_destinations", + "space_members", + "spaces", + "templates", + "user_mfa", + "user_sessions", + "user_tokens", + "users", + "watchers", + "workspace_invitations", + "workspaces" + ], + "update_final_phase": null, + "pinned_new": { + "docmost": "docmost/docmost:0.96.0", + "docmost-postgres": "postgres:18-alpine", + "docmost-redis": "redis:7-alpine" + }, + "tables_new_live": [ + "ai_chat_messages", + "ai_chats", + "api_keys", + "attachments", + "audit", + "auth_accounts", + "auth_providers", + "backlinks", + "base_properties", + "base_rows", + "base_views", + "billing", + "comments", + "favorites", + "file_tasks", + "group_users", + "groups", + "kysely_migration", + "kysely_migration_lock", + "labels", + "notifications", + "oauth_authorization_codes", + "oauth_clients", + "oauth_grants", + "oauth_tokens", + "page_access", + "page_history", + "page_labels", + "page_permissions", + "page_transclusion_references", + "page_transclusions", + "page_verifications", + "page_verifiers", + "pages", + "public_spaces", + "scim_tokens", + "shares", + "siem_destinations", + "space_members", + "spaces", + "templates", + "user_mfa", + "user_sessions", + "user_tokens", + "users", + "watchers", + "workspace_invitations", + "workspaces" + ], + "tables_added_by_migration": [], + "seed_after_update": true +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/B/docmost-tables.json b/documentation/audits/design-build-2026-10-06/B/docmost-tables.json new file mode 100644 index 00000000..c1413fb7 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/docmost-tables.json @@ -0,0 +1,247 @@ +{ + "app": "docmost", + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.299.0", + "pinned_old": { + "docmost": "docmost/docmost:0.95.0", + "docmost-postgres": "postgres:16-alpine", + "docmost-redis": "redis:7-alpine" + }, + "tables_old_live": [ + "ai_chat_messages", + "ai_chats", + "api_keys", + "attachments", + "audit", + "auth_accounts", + "auth_providers", + "backlinks", + "base_properties", + "base_rows", + "base_views", + "billing", + "comments", + "favorites", + "file_tasks", + "group_users", + "groups", + "kysely_migration", + "kysely_migration_lock", + "labels", + "notifications", + "page_access", + "page_history", + "page_labels", + "page_permissions", + "page_transclusion_references", + "page_transclusions", + "page_verifications", + "page_verifiers", + "pages", + "scim_tokens", + "shares", + "space_members", + "spaces", + "templates", + "user_mfa", + "user_sessions", + "user_tokens", + "users", + "watchers", + "workspace_invitations", + "workspaces" + ], + "update_final_phase": "done", + "pinned_new": { + "docmost": "docmost/docmost:0.96.0", + "docmost-postgres": "postgres:16-alpine", + "docmost-redis": "redis:7-alpine" + }, + "tables_new_live": [ + "ai_chat_messages", + "ai_chats", + "api_keys", + "attachments", + "audit", + "auth_accounts", + "auth_providers", + "backlinks", + "base_properties", + "base_rows", + "base_views", + "billing", + "comments", + "favorites", + "file_tasks", + "group_users", + "groups", + "kysely_migration", + "kysely_migration_lock", + "labels", + "notifications", + "oauth_authorization_codes", + "oauth_clients", + "oauth_grants", + "oauth_tokens", + "page_access", + "page_history", + "page_labels", + "page_permissions", + "page_transclusion_references", + "page_transclusions", + "page_verifications", + "page_verifiers", + "pages", + "public_spaces", + "scim_tokens", + "shares", + "siem_destinations", + "space_members", + "spaces", + "templates", + "user_mfa", + "user_sessions", + "user_tokens", + "users", + "watchers", + "workspace_invitations", + "workspaces" + ], + "tables_added_by_migration": [ + "oauth_authorization_codes", + "oauth_clients", + "oauth_grants", + "oauth_tokens", + "public_spaces", + "siem_destinations" + ], + "seed_after_update": true, + "restore": { + "snapshot_id": "helyi", + "http": "HTTP/2 302", + "seconds": 32.4, + "state_after": "running", + "hold_after": null + }, + "imported_db_dump_line": [ + "2026/10/06 12:55:31 dbdump.go:820: [INFO] [backup] Imported DB dump docmost-postgres.sql into docmost-postgres (postgres)" + ], + "tables_after_restore": [ + "ai_chat_messages", + "ai_chats", + "api_keys", + "attachments", + "audit", + "auth_accounts", + "auth_providers", + "backlinks", + "base_properties", + "base_rows", + "base_views", + "billing", + "comments", + "favorites", + "file_tasks", + "group_users", + "groups", + "kysely_migration", + "kysely_migration_lock", + "labels", + "notifications", + "page_access", + "page_history", + "page_labels", + "page_permissions", + "page_transclusion_references", + "page_transclusions", + "page_verifications", + "page_verifiers", + "pages", + "scim_tokens", + "shares", + "space_members", + "spaces", + "templates", + "user_mfa", + "user_sessions", + "user_tokens", + "users", + "watchers", + "workspace_invitations", + "workspaces" + ], + "seed_after_restore": true, + "unit_files": [ + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/compose/.felhom.yml", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/compose/docker-compose.yml", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/compose/app.yaml", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/data-stamps.json", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_postgres_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_redis_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_storage.tar", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/manifest.json", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/.felhom-tier2-layout" + ], + "dump_files_seen": [ + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_postgres_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_redis_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_storage.tar" + ], + "tables_in_copy_dump": [ + "ai_chat_messages", + "ai_chats", + "api_keys", + "attachments", + "audit", + "auth_accounts", + "auth_providers", + "backlinks", + "base_properties", + "base_rows", + "base_views", + "billing", + "comments", + "favorites", + "file_tasks", + "group_users", + "groups", + "kysely_migration", + "kysely_migration_lock", + "labels", + "notifications", + "page_access", + "page_history", + "page_labels", + "page_permissions", + "page_transclusion_references", + "page_transclusions", + "page_verifications", + "page_verifiers", + "pages", + "scim_tokens", + "shares", + "space_members", + "spaces", + "templates", + "user_mfa", + "user_sessions", + "user_tokens", + "users", + "watchers", + "workspace_invitations", + "workspaces" + ], + "dump_used_for_control": [ + "/mnt/felhom-drives/scratch_hdd/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql" + ], + "migration_tables_left_after_restore": [], + "live_equals_copy": true, + "pinned_after_restore": { + "docmost": "docmost/docmost:0.95.0", + "docmost-postgres": "postgres:16-alpine", + "docmost-redis": "redis:7-alpine" + }, + "state_after_restore": "running", + "verdict": "SAFE" +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/B/docmost.txt b/documentation/audits/design-build-2026-10-06/B/docmost.txt new file mode 100644 index 00000000..dc5b2dfb --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/docmost.txt @@ -0,0 +1,30 @@ +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.299.0 +docmost already installed on 9202 — removing it first (scratch box) +drill: 686eafd DRILL docmost: the OLD definition (6d8cd87^) for R-638 slice 0 | images: ['docmost/docmost:0.95.0', 'postgres:16-alpine', 'redis:7-alpine'] + docmost: /api/auth/setup http=200 rc=0 + docmost: login as the seeded user http=200 ok=True +[2] old version installed, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}, live tables 42 +[3] the ladder entry used: [({'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}, {'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'})] +drill: 4dee823 DRILL docmost: the NEW definition (6d8cd87) + its one ladder entry, R-638 slice 0 | images: ['docmost/docmost:0.96.0', 'postgres:16-alpine', 'redis:7-alpine'] +[3] update -> done, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}, live tables 48, added by the migration: 6 ['oauth_authorization_codes', 'oauth_clients', 'oauth_grants', 'oauth_tokens', 'public_spaces', 'siem_destinations'] + docmost: login as the seeded user http=200 ok=True +[4] restorable copies offered: [{"time": "2026-10-06T12:53:06Z", "short_id": "helyi", "tier": 1, "drive_label": "Bels\u0151 SSD (rendszer)"}] +2026/10/06 12:55:25 handlers.go:1831: [WARN] [web] Restore requested (async): stack=docmost, snapshot=helyi from 172.18.0.5:40656 +2026/10/06 12:55:25 restore_unit.go:381: [INFO] [backup] Restore docmost: the data belongs to [docmost/docmost:0.95.0 postgres:16-alpine redis:7-alpine] (written 2026-10-06T12:53:06Z); the app ran [docmost/docmost:0.96.0@sha256:b56947fcfd08aab8fae12a377e1792784786adbf8b96e4281f14ef4fc072685a postgre +2026/10/06 12:55:25 restore_unit.go:393: [INFO] [backup] Restoring docmost from recovery unit /mnt/sys_drive/felhom-data/backups/primary/docmost: images=3, secrets recovered=2/2, data_keys=0 +2026/10/06 12:55:27 restore.go:152: [INFO] [backup] Restoring Docker volume docmost_docmost_postgres_data for docmost +2026/10/06 12:55:28 restore.go:152: [INFO] [backup] Restoring Docker volume docmost_docmost_redis_data for docmost +2026/10/06 12:55:28 restore.go:152: [INFO] [backup] Restoring Docker volume docmost_docmost_storage for docmost +2026/10/06 12:55:29 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for docmost +2026/10/06 12:55:29 restore_db.go:87: [INFO] [backup] Restore docmost: replaying DB dump into docmost-postgres (postgres) +2026/10/06 12:55:31 dbdump.go:820: [INFO] [backup] Imported DB dump docmost-postgres.sql into docmost-postgres (postgres) +2026/10/06 12:55:31 restore_db.go:97: [INFO] [backup] Restore docmost: replayed 1 DB dump(s) +2026/10/06 12:55:56 restore_unit.go:478: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed +2026/10/06 12:55:56 handlers.go:1843: [INFO] [web] Restore completed (async): stack=docmost in 30.474627183s (volumes 3/3, dbs 1/1) + + docmost: login as the seeded user http=200 ok=True +[5] Imported DB dump lines: ['2026/10/06 12:55:31 dbdump.go:820: [INFO] [backup] Imported DB dump docmost-postgres.sql into docmost-postgres (postgres)'] +[5] seed read back after restore: True +[5] live tables after restore 42; the copy's own dump lists 42; equal=True; migration tables still there: [] +[5] pinned after restore {'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}, state running +RESULT docmost: SAFE diff --git a/documentation/audits/design-build-2026-10-06/B/go-test-green.txt b/documentation/audits/design-build-2026-10-06/B/go-test-green.txt new file mode 100644 index 00000000..0bb5a444 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/go-test-green.txt @@ -0,0 +1,45 @@ +# R-638 option A — full controller suite after slices 1+2, 2026-10-06T13:10:21Z +$ cd felhom-controller/controller && go build ./... && go vet ./... && go test ./... -count=1 +build rc=0 vet rc=0 test rc=0 +--- FAIL lines (none expected): +--- tail: +ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.036s +ok gitea.dooplex.hu/admin/felhom-controller/internal/appexport 0.627s +? gitea.dooplex.hu/admin/felhom-controller/internal/assets [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 1.713s +ok gitea.dooplex.hu/admin/felhom-controller/internal/backupwindow 0.011s +ok gitea.dooplex.hu/admin/felhom-controller/internal/bootrecon 0.011s +ok gitea.dooplex.hu/admin/felhom-controller/internal/bootstrap 14.027s +ok gitea.dooplex.hu/admin/felhom-controller/internal/channelhealth 0.009s +? gitea.dooplex.hu/admin/felhom-controller/internal/cloudflare [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/config 0.009s +ok gitea.dooplex.hu/admin/felhom-controller/internal/crashboot 0.011s +? gitea.dooplex.hu/admin/felhom-controller/internal/crypto [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/dockerexec 0.192s +ok gitea.dooplex.hu/admin/felhom-controller/internal/family 0.025s +ok gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.010s +ok gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.376s +ok gitea.dooplex.hu/admin/felhom-controller/internal/infra 0.014s +? gitea.dooplex.hu/admin/felhom-controller/internal/integrations [no test files] +? gitea.dooplex.hu/admin/felhom-controller/internal/logx [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/mailrelay 0.016s +ok gitea.dooplex.hu/admin/felhom-controller/internal/metrics 0.009s +ok gitea.dooplex.hu/admin/felhom-controller/internal/monitor 0.008s +ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.009s +ok gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.942s +ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply 0.014s +ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.087s +? gitea.dooplex.hu/admin/felhom-controller/internal/recovery [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/report 2.599s +ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.456s +? gitea.dooplex.hu/admin/felhom-controller/internal/selftest [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/selfupdate 0.045s +ok gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.542s +ok gitea.dooplex.hu/admin/felhom-controller/internal/setup 0.011s +ok gitea.dooplex.hu/admin/felhom-controller/internal/sockheal 0.007s +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 397.006s +ok gitea.dooplex.hu/admin/felhom-controller/internal/sync 0.406s +ok gitea.dooplex.hu/admin/felhom-controller/internal/system 0.063s +ok gitea.dooplex.hu/admin/felhom-controller/internal/util 0.040s +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 63.736s +? gitea.dooplex.hu/admin/felhom-controller/scripts [no test files] diff --git a/documentation/audits/design-build-2026-10-06/B/green-r638-new-tests.txt b/documentation/audits/design-build-2026-10-06/B/green-r638-new-tests.txt new file mode 100644 index 00000000..45d4f613 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/green-r638-new-tests.txt @@ -0,0 +1,19 @@ +# R-638 slices 1+2 GREEN after the fix (felhom-controller 3124278 + fix, uncommitted at run time) — 2026-10-06T13:03:17Z +$ cd felhom-controller/controller && go test ./internal/backup/ -run '^TestR638_' -v -count=1 +=== RUN TestR638_FallbackReplaysBeforeTheAppStarts +--- PASS: TestR638_FallbackReplaysBeforeTheAppStarts (0.00s) +=== RUN TestR638_FallbackThroughUnitRestoreKeepsTheOrder +--- PASS: TestR638_FallbackThroughUnitRestoreKeepsTheOrder (0.00s) +=== RUN TestR638_FallbackNoDumpTakesOneFullStart +--- PASS: TestR638_FallbackNoDumpTakesOneFullStart (0.00s) +=== RUN TestR638_FallbackRefusesWhenNoDBServiceIdentifiable +--- PASS: TestR638_FallbackRefusesWhenNoDBServiceIdentifiable (0.01s) +=== RUN TestR638_FallbackVolumeFailureSkipsTheReplay +--- PASS: TestR638_FallbackVolumeFailureSkipsTheReplay (0.00s) +=== RUN TestR638_UnitVolumeFailureNeverCallsTheImporter +--- PASS: TestR638_UnitVolumeFailureNeverCallsTheImporter (0.00s) +=== RUN TestR638_UnitVolumeSuccessStillReplays +--- PASS: TestR638_UnitVolumeSuccessStillReplays (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.020s +rc=0 diff --git a/documentation/audits/design-build-2026-10-06/B/red-slice1-fallback-order.txt b/documentation/audits/design-build-2026-10-06/B/red-slice1-fallback-order.txt new file mode 100644 index 00000000..0ea46040 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/red-slice1-fallback-order.txt @@ -0,0 +1,20 @@ +# R-638 slice 1 RED-PROOF on UNCHANGED code (felhom-controller 3124278 + the new test file only) — 2026-10-06T13:01:56Z +$ cd felhom-controller/controller && go test ./internal/backup/ -run '^TestR638_Fallback' -v -count=1 +=== RUN TestR638_FallbackReplaysBeforeTheAppStarts + r638_restore_order_test.go:89: the FULL stack was already up when the replay fired (calls before the replay: "stop,start") — a newer app can migrate the restored data before the copy is loaded (R-638) +--- FAIL: TestR638_FallbackReplaysBeforeTheAppStarts (0.00s) +=== RUN TestR638_FallbackThroughUnitRestoreKeepsTheOrder + r638_restore_order_test.go:118: sequence = "stop,start", want stop → db-only start → replay → full start +--- FAIL: TestR638_FallbackThroughUnitRestoreKeepsTheOrder (0.00s) +=== RUN TestR638_FallbackNoDumpTakesOneFullStart +--- PASS: TestR638_FallbackNoDumpTakesOneFullStart (0.00s) +=== RUN TestR638_FallbackRefusesWhenNoDBServiceIdentifiable + r638_restore_order_test.go:151: expected a refusal: a dump exists but no database service can be started for it +--- FAIL: TestR638_FallbackRefusesWhenNoDBServiceIdentifiable (0.00s) +=== RUN TestR638_FallbackVolumeFailureSkipsTheReplay + r638_restore_order_test.go:178: the volume failure must surface as the data-error outcome naming the volume, got: +--- FAIL: TestR638_FallbackVolumeFailureSkipsTheReplay (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.009s +FAIL +rc=1 diff --git a/documentation/audits/design-build-2026-10-06/B/red-slice2-no-replay-after-volume-failure.txt b/documentation/audits/design-build-2026-10-06/B/red-slice2-no-replay-after-volume-failure.txt new file mode 100644 index 00000000..fb5284a7 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/red-slice2-no-replay-after-volume-failure.txt @@ -0,0 +1,11 @@ +# R-638 slice 2 RED-PROOF on UNCHANGED code (felhom-controller 3124278 + the new test file only) — 2026-10-06T13:02:03Z +$ cd felhom-controller/controller && go test ./internal/backup/ -run '^TestR638_Unit' -v -count=1 +=== RUN TestR638_UnitVolumeFailureNeverCallsTheImporter + r638_restore_order_test.go:211: the importer ran after a failed volume leg: [/tmp/TestR638_UnitVolumeFailureNeverCallsTheImporter760559713/001/drive/backups/primary/app/db-dumps/app-postgres.sql] +--- FAIL: TestR638_UnitVolumeFailureNeverCallsTheImporter (0.00s) +=== RUN TestR638_UnitVolumeSuccessStillReplays +--- PASS: TestR638_UnitVolumeSuccessStillReplays (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.009s +FAIL +rc=1 diff --git a/documentation/audits/design-build-2026-10-06/B/romm-run1-hddpath-refused.txt b/documentation/audits/design-build-2026-10-06/B/romm-run1-hddpath-refused.txt new file mode 100644 index 00000000..0e0a7e1a --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/romm-run1-hddpath-refused.txt @@ -0,0 +1,3 @@ +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.299.0 +drill: 9a724b6 DRILL romm: the OLD definition (15f9ebf^) for R-638 slice 0 | images: ['rommapp/romm:5.0.0', 'mariadb:11.4', 'redis:7-alpine'] +RESULT the install did not complete diff --git a/documentation/audits/design-build-2026-10-06/B/romm-tables.json b/documentation/audits/design-build-2026-10-06/B/romm-tables.json new file mode 100644 index 00000000..63231c1d --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/romm-tables.json @@ -0,0 +1,199 @@ +{ + "app": "romm", + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.299.0", + "pinned_old": { + "romm": "rommapp/romm:5.0.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "tables_old_live": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_files", + "rom_notes", + "rom_user", + "roms", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collections" + ], + "update_final_phase": "done", + "pinned_new": { + "romm": "rommapp/romm:5.3.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "tables_new_live": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "memory_card_versions", + "memory_cards", + "music_favorite_tracks", + "music_playlist_tracks", + "music_playlists", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_file_doc_meta", + "rom_file_user", + "rom_files", + "rom_identity_keys", + "rom_notes", + "rom_similarity", + "rom_user", + "roms", + "roms_facets", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "streaming_container_adoptions", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collection_roms", + "virtual_collections" + ], + "tables_added_by_migration": [ + "memory_card_versions", + "memory_cards", + "music_favorite_tracks", + "music_playlist_tracks", + "music_playlists", + "rom_file_doc_meta", + "rom_file_user", + "rom_identity_keys", + "rom_similarity", + "roms_facets", + "streaming_container_adoptions", + "virtual_collection_roms" + ], + "seed_after_update": false, + "restore": { + "snapshot_id": "helyi", + "http": "HTTP/2 302", + "seconds": 56.7, + "state_after": "running", + "hold_after": null + }, + "imported_db_dump_line": [ + "2026/10/06 13:13:07 dbdump.go:820: [INFO] [backup] Imported DB dump romm-mariadb.sql into romm-db (mariadb)" + ], + "tables_after_restore": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_files", + "rom_notes", + "rom_user", + "roms", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collections" + ], + "seed_after_restore": true, + "unit_files": [ + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/pre-restore-20261006T131058Z-romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/compose/.felhom.yml", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/compose/docker-compose.yml", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/compose/app.yaml", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/data-stamps.json", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_config.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_redis_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_db_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/manifest.json" + ], + "dump_files_seen": [ + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_config.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_redis_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_db_data.tar" + ], + "tables_in_copy_dump": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_files", + "rom_notes", + "rom_user", + "roms", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collections" + ], + "dump_used_for_control": [ + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/romm-mariadb.sql" + ], + "migration_tables_left_after_restore": [], + "live_equals_copy": true, + "pinned_after_restore": { + "romm": "rommapp/romm:5.0.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "state_after_restore": "running", + "verdict": "SAFE" +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/B/romm.txt b/documentation/audits/design-build-2026-10-06/B/romm.txt new file mode 100644 index 00000000..b1ca7aa3 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/romm.txt @@ -0,0 +1,38 @@ +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.299.0 +drill: 903cc03 DRILL romm: the OLD definition (15f9ebf^) for R-638 slice 0 | images: ['rommapp/romm:5.0.0', 'mariadb:11.4', 'redis:7-alpine'] + romm: POST /api/users http=201 + romm: login as the seeded user http=200 ok=True +[2] old version installed, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, live tables 27 +[3] the ladder entry used: [({'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, {'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'})] +drill: 04a94d3 DRILL romm: the NEW definition (15f9ebf) + its one ladder entry, R-638 slice 0 | images: ['rommapp/romm:5.3.0', 'mariadb:11.4', 'redis:7-alpine'] +[3] update -> done, pinned={'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, live tables 39, added by the migration: 12 ['memory_card_versions', 'memory_cards', 'music_favorite_tracks', 'music_playlist_tracks', 'music_playlists', 'rom_file_doc_meta', 'rom_file_user', 'rom_identity_keys', 'rom_similarity', 'roms_facets', 'streaming_container_adoptions', 'virtual_collection_roms'] + romm: login as the seeded user http=502 ok=False + romm: body +502 Bad Gateway + +

502 Bad Gateway

+
nginx/1.29.8
+ + + +[4] restorable copies offered: [{"time": "2026-10-06T13:10:41Z", "short_id": "helyi", "tier": 1, "drive_label": "SCRATCH HDD"}] +2026/10/06 13:12:51 handlers.go:1831: [WARN] [web] Restore requested (async): stack=romm, snapshot=helyi from 172.18.0.5:40656 +2026/10/06 13:12:51 restore_unit.go:381: [INFO] [backup] Restore romm: the data belongs to [rommapp/romm:5.0.0 mariadb:11.4 redis:7-alpine] (written 2026-10-06T13:10:41Z); the app ran [rommapp/romm:5.3.0@sha256:dc586cb3a2c7316fcffb3dc273171b1964523f0f409e989295d26d2199df1d4e mariadb:11.4@sha256:70cc +2026/10/06 13:12:51 restore_unit.go:393: [INFO] [backup] Restoring romm from recovery unit /mnt/felhom-drives/scratch_hdd/backups/primary/romm: images=3, secrets recovered=3/3, data_keys=0 +2026/10/06 13:13:01 restore.go:152: [INFO] [backup] Restoring Docker volume romm_romm_config for romm +2026/10/06 13:13:01 restore.go:152: [INFO] [backup] Restoring Docker volume romm_romm_db_data for romm +2026/10/06 13:13:02 restore.go:152: [INFO] [backup] Restoring Docker volume romm_romm_redis_data for romm +2026/10/06 13:13:03 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for romm +2026/10/06 13:13:03 restore_db.go:87: [INFO] [backup] Restore romm: replaying DB dump into romm-db (mariadb) +2026/10/06 13:13:07 dbdump.go:820: [INFO] [backup] Imported DB dump romm-mariadb.sql into romm-db (mariadb) +2026/10/06 13:13:07 restore_db.go:97: [INFO] [backup] Restore romm: replayed 1 DB dump(s) +2026/10/06 13:13:47 restore_unit.go:478: [INFO] [backup] Restore-from-unit completed: romm — 3 volume(s) of 3 listed, 1 database(s) of 1 listed +2026/10/06 13:13:47 handlers.go:1843: [INFO] [web] Restore completed (async): stack=romm in 55.871437852s (volumes 3/3, dbs 1/1) + + romm: login as the seeded user http=200 ok=True +[5] Imported DB dump lines: ['2026/10/06 13:13:07 dbdump.go:820: [INFO] [backup] Imported DB dump romm-mariadb.sql into romm-db (mariadb)'] +[5] seed read back after restore: True +[5] live tables after restore 27; the copy's own dump lists 27; equal=True; migration tables still there: [] +[5] pinned after restore {'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, state running +RESULT romm: SAFE +remove -> 200 diff --git a/documentation/audits/design-build-2026-10-06/B/run3-romm-tables.json b/documentation/audits/design-build-2026-10-06/B/run3-romm-tables.json new file mode 100644 index 00000000..0d7e1c2a --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/run3-romm-tables.json @@ -0,0 +1,197 @@ +{ + "app": "romm", + "controller": "gitea.dooplex.hu/admin/felhom-controller:0.299.0", + "pinned_old": { + "romm": "rommapp/romm:5.0.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "tables_old_live": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_files", + "rom_notes", + "rom_user", + "roms", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collections" + ], + "update_final_phase": "done", + "pinned_new": { + "romm": "rommapp/romm:5.3.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "tables_new_live": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "memory_card_versions", + "memory_cards", + "music_favorite_tracks", + "music_playlist_tracks", + "music_playlists", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_file_doc_meta", + "rom_file_user", + "rom_files", + "rom_identity_keys", + "rom_notes", + "rom_similarity", + "rom_user", + "roms", + "roms_facets", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "streaming_container_adoptions", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collection_roms", + "virtual_collections" + ], + "tables_added_by_migration": [ + "memory_card_versions", + "memory_cards", + "music_favorite_tracks", + "music_playlist_tracks", + "music_playlists", + "rom_file_doc_meta", + "rom_file_user", + "rom_identity_keys", + "rom_similarity", + "roms_facets", + "streaming_container_adoptions", + "virtual_collection_roms" + ], + "seed_after_update": false, + "restore": { + "snapshot_id": "helyi", + "http": "HTTP/2 302", + "seconds": 56.7, + "state_after": "running", + "hold_after": null + }, + "imported_db_dump_line": [ + "2026/10/06 13:04:43 dbdump.go:820: [INFO] [backup] Imported DB dump romm-mariadb.sql into romm-db (mariadb)" + ], + "tables_after_restore": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_files", + "rom_notes", + "rom_user", + "roms", + "roms_metadata", + "saves", + "screenshots", + "sibling_roms", + "smart_collections", + "states", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users", + "virtual_collections" + ], + "seed_after_restore": true, + "unit_files": [ + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/pre-restore-20261006T130227Z-romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/compose/.felhom.yml", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/compose/docker-compose.yml", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/compose/app.yaml", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/data-stamps.json", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_config.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_redis_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_db_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/manifest.json" + ], + "dump_files_seen": [ + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/pre-restore-20261006T130227Z-romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/romm-mariadb.sql", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_config.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_redis_data.tar", + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/volume-dumps/romm_romm_db_data.tar" + ], + "tables_in_copy_dump": [ + "alembic_version", + "client_tokens", + "collections", + "collections_roms", + "device_save_sync", + "devices", + "firmware", + "hidden_entities", + "permission_group_grants", + "permission_groups", + "platforms", + "play_sessions", + "rom_files", + "rom_notes", + "rom_user", + "roms", + "saves", + "screenshots", + "smart_collections", + "states", + "sync_sessions", + "track_meta", + "user_permission_overrides", + "users" + ], + "dump_used_for_control": [ + "/mnt/felhom-drives/scratch_hdd/backups/primary/romm/db-dumps/pre-restore-20261006T130227Z-romm-mariadb.sql" + ], + "migration_tables_left_after_restore": [], + "live_equals_copy": false, + "pinned_after_restore": { + "romm": "rommapp/romm:5.0.0", + "romm-db": "mariadb:11.4", + "romm-redis": "redis:7-alpine" + }, + "state_after_restore": "running", + "verdict": "NOT SAFE / NOT SHOWN" +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/B/run3-romm.txt b/documentation/audits/design-build-2026-10-06/B/run3-romm.txt new file mode 100644 index 00000000..5679853f --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/B/run3-romm.txt @@ -0,0 +1,39 @@ +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.299.0 +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.299.0 +drill: 04b0daf DRILL romm: the OLD definition (15f9ebf^) for R-638 slice 0 | images: ['rommapp/romm:5.0.0', 'mariadb:11.4', 'redis:7-alpine'] + romm: POST /api/users http=201 + romm: login as the seeded user http=200 ok=True +[2] old version installed, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, live tables 27 +[3] the ladder entry used: [({'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, {'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'})] +drill: 7b95d47 DRILL romm: the NEW definition (15f9ebf) + its one ladder entry, R-638 slice 0 | images: ['rommapp/romm:5.3.0', 'mariadb:11.4', 'redis:7-alpine'] +[3] update -> done, pinned={'romm': 'rommapp/romm:5.3.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, live tables 39, added by the migration: 12 ['memory_card_versions', 'memory_cards', 'music_favorite_tracks', 'music_playlist_tracks', 'music_playlists', 'rom_file_doc_meta', 'rom_file_user', 'rom_identity_keys', 'rom_similarity', 'roms_facets', 'streaming_container_adoptions', 'virtual_collection_roms'] + romm: login as the seeded user http=502 ok=False + romm: body +502 Bad Gateway + +

502 Bad Gateway

+
nginx/1.29.8
+ + + +[4] restorable copies offered: [{"time": "2026-10-06T13:02:12Z", "short_id": "helyi", "tier": 1, "drive_label": "SCRATCH HDD"}] +2026/10/06 13:04:28 handlers.go:1831: [WARN] [web] Restore requested (async): stack=romm, snapshot=helyi from 172.18.0.5:40656 +2026/10/06 13:04:28 restore_unit.go:381: [INFO] [backup] Restore romm: the data belongs to [rommapp/romm:5.0.0 mariadb:11.4 redis:7-alpine] (written 2026-10-06T13:02:12Z); the app ran [rommapp/romm:5.3.0@sha256:dc586cb3a2c7316fcffb3dc273171b1964523f0f409e989295d26d2199df1d4e mariadb:11.4@sha256:70cc +2026/10/06 13:04:28 restore_unit.go:393: [INFO] [backup] Restoring romm from recovery unit /mnt/felhom-drives/scratch_hdd/backups/primary/romm: images=3, secrets recovered=3/3, data_keys=0 +2026/10/06 13:04:38 restore.go:152: [INFO] [backup] Restoring Docker volume romm_romm_config for romm +2026/10/06 13:04:38 restore.go:152: [INFO] [backup] Restoring Docker volume romm_romm_db_data for romm +2026/10/06 13:04:39 restore.go:152: [INFO] [backup] Restoring Docker volume romm_romm_redis_data for romm +2026/10/06 13:04:39 restore.go:195: [INFO] [backup] Restored 3 Docker volume(s) for romm +2026/10/06 13:04:40 restore_db.go:87: [INFO] [backup] Restore romm: replaying DB dump into romm-db (mariadb) +2026/10/06 13:04:43 dbdump.go:820: [INFO] [backup] Imported DB dump romm-mariadb.sql into romm-db (mariadb) +2026/10/06 13:04:43 restore_db.go:97: [INFO] [backup] Restore romm: replayed 1 DB dump(s) +2026/10/06 13:05:23 restore_unit.go:478: [INFO] [backup] Restore-from-unit completed: romm — 3 volume(s) of 3 listed, 1 database(s) of 1 listed +2026/10/06 13:05:23 handlers.go:1843: [INFO] [web] Restore completed (async): stack=romm in 55.385810686s (volumes 3/3, dbs 1/1) + + romm: login as the seeded user http=200 ok=True +[5] Imported DB dump lines: ['2026/10/06 13:04:43 dbdump.go:820: [INFO] [backup] Imported DB dump romm-mariadb.sql into romm-db (mariadb)'] +[5] seed read back after restore: True +[5] live tables after restore 27; the copy's own dump lists 24; equal=False; migration tables still there: [] +[5] pinned after restore {'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}, state running +RESULT romm: NOT SAFE / NOT SHOWN +remove -> 200 diff --git a/documentation/audits/design-build-2026-10-06/C/C0-oom-signal-measure-9202.txt b/documentation/audits/design-build-2026-10-06/C/C0-oom-signal-measure-9202.txt new file mode 100644 index 00000000..31a404da --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/C/C0-oom-signal-measure-9202.txt @@ -0,0 +1,76 @@ +Tue Oct 6 12:41:22 UTC 2026 +Unable to find image 'alpine:3.20' locally +3.20: Pulling from library/alpine +25f1d6b1951a: Already exists +Digest: sha256:d9e853e87e55526f6b2917df91a2115c36dd7c696a35be12163d44e6e2a4b6bc +Status: Downloaded newer image for alpine:3.20 +started r528-hog (alpine:3.20, --memory 64m, swap 64m) +== before +low 0 +high 0 +max 0 +oom 0 +oom_kill 0 +oom_group_kill 0 +sock_throttled 0 +OOMKilled=false running=true exit=0 restarts=0 +hog 1 exec rc=137 +Error response from daemon: container 573bd9b193fea38ab8301c6d9e4752211aaffb1a47009ed60d5b3cbb433ff5e5 is not running +hog 2 exec rc=1 +Error response from daemon: container 573bd9b193fea38ab8301c6d9e4752211aaffb1a47009ed60d5b3cbb433ff5e5 is not running +hog 3 exec rc=1 +== after three hogs (child processes of the running container) +Error response from daemon: container 573bd9b193fea38ab8301c6d9e4752211aaffb1a47009ed60d5b3cbb433ff5e5 is not running +OOMKilled=true running=false exit=137 restarts=0 +== docker events (oom) in the last 2 minutes +1791290484 oom r528-hog + +=== run 2: the hog raises its own oom_score_adj to 1000, so the kernel picks it, not the container's main process +Tue Oct 6 12:41:40 UTC 2026 +started r528-hog (main: a sleep loop) +memory.oom.group=0 +== before +oom_kill 0 +OOMKilled=false running=true exit=0 +hog 1 exec rc=137 +Error response from daemon: container 4152ea857b6ad23ccc31b95774ff83bc7843c8532fd0dce64c21486cf88a7b56 is not running +hog 2 exec rc=1 +Error response from daemon: container 4152ea857b6ad23ccc31b95774ff83bc7843c8532fd0dce64c21486cf88a7b56 is not running +hog 3 exec rc=1 +== after three hogs +Error response from daemon: container 4152ea857b6ad23ccc31b95774ff83bc7843c8532fd0dce64c21486cf88a7b56 is not running +OOMKilled=true running=false exit=137 restarts=0 +== docker oom events since this run started +1791290484 oom r528-hog + +=== run 3: main process protected (--oom-score-adj -1000), hog at oom_score_adj 1000 +Tue Oct 6 12:41:59 UTC 2026 +started +pid1 oom_score_adj=0 +hog 1 exec rc=137 +Error response from daemon: container eba82647cdb95aebcda6410686d6b59f2178cdf645d84d852b136a00be5bac74 is not running +hog 2 exec rc=1 +Error response from daemon: container eba82647cdb95aebcda6410686d6b59f2178cdf645d84d852b136a00be5bac74 is not running +hog 3 exec rc=1 +== after three hogs +Error response from daemon: container eba82647cdb95aebcda6410686d6b59f2178cdf645d84d852b136a00be5bac74 is not running +OOMKilled=true running=false exit=137 restarts=0 +== docker oom events since this run started +1791290484 oom r528-hog +1791290501 oom r528-hog + +=== run 4: --memory 256m, main = sleep infinity (no forks), hog = one 400 MB block (dd bs=400M) +Tue Oct 6 12:42:18 UTC 2026 +started +hog 1 exec rc=137 +hog 2 exec rc=137 +hog 3 exec rc=137 +== after three hogs +oom 5 +oom_kill 3 +oom_group_kill 0 +OOMKilled=true running=true exit=0 restarts=0 +== docker oom events since this run started +1791290538 oom r528-hog +1791290541 oom r528-hog +1791290543 oom r528-hog diff --git a/documentation/audits/design-build-2026-10-06/C/C1-exec-cost-demo-hp-9201.txt b/documentation/audits/design-build-2026-10-06/C/C1-exec-cost-demo-hp-9201.txt new file mode 100644 index 00000000..c85b1d59 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/C/C1-exec-cost-demo-hp-9201.txt @@ -0,0 +1,11 @@ +Tue Oct 6 12:42:54 UTC 2026 +running containers: 21 +engine 29.8.2 + unreadable: cloudflared +bash: line 11: bc: command not found +one full read: s wall; readable 20, unreadable 1 +bash: line 12: /usr/bin/time: No such file or directory +=== again, timed with bash's own clock +run 1: one full read of 21 containers took 1637 ms wall +run 2: one full read of 21 containers took 1642 ms wall +bash time: 1.661 s wall, 0.474 s user, 0.438 s sys diff --git a/documentation/audits/design-build-2026-10-06/C/C2-teardown.txt b/documentation/audits/design-build-2026-10-06/C/C2-teardown.txt new file mode 100644 index 00000000..c6bc7460 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/C/C2-teardown.txt @@ -0,0 +1,5 @@ +r528-hog +Untagged: alpine:3.20 +Untagged: alpine@sha256:d9e853e87e55526f6b2917df91a2115c36dd7c696a35be12163d44e6e2a4b6bc +Deleted: sha256:bf8527eb54c3680e728d5b4b383a8ba730d72dae7236fbc8dff97ed6b224a731 +0 diff --git a/documentation/audits/design-build-2026-10-06/D/go-test-green.txt b/documentation/audits/design-build-2026-10-06/D/go-test-green.txt new file mode 100644 index 00000000..4674a8be --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/D/go-test-green.txt @@ -0,0 +1,107 @@ +R-518 option A (decision 156) — controller suite after the fix (local clone, branch r518-one-stop-per-tier, +working tree = fix + tests + copy; base felhom-controller 3124278). + +$ cd controller && go build ./... && go vet ./... && echo BUILD_VET_OK +BUILD_VET_OK + +$ cd controller && go test ./... (full output — 43 lines, no FAIL line) +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 10.666s +ok gitea.dooplex.hu/admin/felhom-controller/internal/agentapi 0.182s +ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.145s +ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.028s +ok gitea.dooplex.hu/admin/felhom-controller/internal/appexport 0.615s +? gitea.dooplex.hu/admin/felhom-controller/internal/assets [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 1.777s +ok gitea.dooplex.hu/admin/felhom-controller/internal/backupwindow 0.006s +ok gitea.dooplex.hu/admin/felhom-controller/internal/bootrecon 0.008s +ok gitea.dooplex.hu/admin/felhom-controller/internal/bootstrap 14.029s +ok gitea.dooplex.hu/admin/felhom-controller/internal/channelhealth 0.005s +? gitea.dooplex.hu/admin/felhom-controller/internal/cloudflare [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/config 0.006s +ok gitea.dooplex.hu/admin/felhom-controller/internal/crashboot 0.007s +? gitea.dooplex.hu/admin/felhom-controller/internal/crypto [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/dockerexec 0.189s +ok gitea.dooplex.hu/admin/felhom-controller/internal/family 0.026s +ok gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.011s +ok gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.395s +ok gitea.dooplex.hu/admin/felhom-controller/internal/infra 0.012s +? gitea.dooplex.hu/admin/felhom-controller/internal/integrations [no test files] +? gitea.dooplex.hu/admin/felhom-controller/internal/logx [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/mailrelay 0.018s +ok gitea.dooplex.hu/admin/felhom-controller/internal/metrics 0.010s +ok gitea.dooplex.hu/admin/felhom-controller/internal/monitor 0.007s +ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.011s +ok gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.910s +ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply 0.014s +ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.089s +? gitea.dooplex.hu/admin/felhom-controller/internal/recovery [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/report 2.587s +ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.457s +? gitea.dooplex.hu/admin/felhom-controller/internal/selftest [no test files] +ok gitea.dooplex.hu/admin/felhom-controller/internal/selfupdate 0.043s +ok gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.555s +ok gitea.dooplex.hu/admin/felhom-controller/internal/setup 0.010s +ok gitea.dooplex.hu/admin/felhom-controller/internal/sockheal 0.008s +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 397.094s +ok gitea.dooplex.hu/admin/felhom-controller/internal/sync 0.385s +ok gitea.dooplex.hu/admin/felhom-controller/internal/system 0.064s +ok gitea.dooplex.hu/admin/felhom-controller/internal/util 0.041s +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 63.621s +? gitea.dooplex.hu/admin/felhom-controller/scripts [no test files] + +Filter proof (the new/renamed tests ran, after the fix): +$ go test ./internal/quiesce/ -run 'TestFirstTierSnapshot_ResumesAppWhileUploadContinues|TestManualPress_RunsOnlyLocalTier|TestBothTiersDue_FirstCycleRunsOnlyFirstTier|TestLeftoverTier_RunsInNextCycleInItsOwnWindow|TestNotify_BothFailingTiersAreReported|TestManualRun_' -v -count=1 +=== RUN TestNotify_BothFailingTiersAreReported +--- PASS: TestNotify_BothFailingTiersAreReported (0.00s) +=== RUN TestManualRun_AbsentTierSkipped +--- PASS: TestManualRun_AbsentTierSkipped (0.00s) +=== RUN TestManualRun_UnknownStorageNotSkipped +--- PASS: TestManualRun_UnknownStorageNotSkipped (0.00s) +=== RUN TestBothTiersDue_FirstCycleRunsOnlyFirstTier +--- PASS: TestBothTiersDue_FirstCycleRunsOnlyFirstTier (0.00s) +=== RUN TestFirstTierSnapshot_ResumesAppWhileUploadContinues +--- PASS: TestFirstTierSnapshot_ResumesAppWhileUploadContinues (0.00s) +=== RUN TestLeftoverTier_RunsInNextCycleInItsOwnWindow +--- PASS: TestLeftoverTier_RunsInNextCycleInItsOwnWindow (0.00s) +=== RUN TestManualPress_RunsOnlyLocalTier +--- PASS: TestManualPress_RunsOnlyLocalTier (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.022s + +$ go test ./internal/quiesce/ -run TestManualRunTiers_ -v -count=1 +=== RUN TestManualRunTiers_NoPrimaryAvailable_BacksUpFirstAvailable +--- PASS: TestManualRunTiers_NoPrimaryAvailable_BacksUpFirstAvailable (0.00s) +PASS + +$ go test ./internal/web/ -run 'TestR518_|TestI18n' -v -count=1 (top-level RUN lines) +=== RUN TestI18nParity +=== RUN TestI18nEnglishPages +=== RUN TestI18nParityCoversEveryMarker +=== RUN TestI18nJSContextValuesAreSafe +=== RUN TestI18nDirectRenderPagesFollowLanguage +=== RUN TestI18nDirectRenderPagesHaveNoAdminChrome +=== RUN TestR518_BackupButtonStatesTheShortLocalOnlyStop +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 7.899s + +$ python3 controller/scripts/controller_gates.py --fast (felhom.eu sibling present; summary) + template-id OK (exit 0) + emoji OK (exit 0) + native-confirm OK (exit 0) + offbox-rename OK (exit 0) + app-row-dedup OK (exit 0) + mojibake OK (exit 0) + docker-v OK (exit 0) + secret-markup OK (exit 0) + retrieval-promise OK (exit 0) + debug-routes OK (exit 0) + reuse-refs OK (exit 0) + instructions OK (exit 0) + observations OK (exit 0) + minagent-header OK (exit 0) + i18n OK (exit 0) + go-parity OK (exit 0) + gofmt OK (exit 0) + golden-notice ADVISORY (exit 0, advisory) + +all controller gates OK diff --git a/documentation/audits/design-build-2026-10-06/D/red-first-tier-resume.txt b/documentation/audits/design-build-2026-10-06/D/red-first-tier-resume.txt new file mode 100644 index 00000000..42c8f78e --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/D/red-first-tier-resume.txt @@ -0,0 +1,13 @@ +$ cd controller && go test ./internal/quiesce/ -run 'TestFirstTierSnapshot_ResumesAppWhileUploadContinues' -v -count=1 +RESULT: FAILED (exit code 1) — run on the UNCHANGED quiesce.go (felhom-controller 3124278), new test only. + +=== RUN TestFirstTierSnapshot_ResumesAppWhileUploadContinues + tiers_test.go:223: the apps were still STOPPED when the second `snapshotted` poll answered (restarts per poll = [0 0 0 0]) — they must resume at the first tier's snapshot (decision 156) +--- FAIL: TestFirstTierSnapshot_ResumesAppWhileUploadContinues (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.010s +FAIL + +Reading: two tiers due; local answered snapshotted, snapshotted, snapshotted, done. The restart count +sampled at each local poll was 0 at every poll — the apps stayed stopped through the whole local +upload (R-82: early resume only on the LAST tier). diff --git a/documentation/audits/design-build-2026-10-06/D/red-manual-local-only.txt b/documentation/audits/design-build-2026-10-06/D/red-manual-local-only.txt new file mode 100644 index 00000000..299f7c6f --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/D/red-manual-local-only.txt @@ -0,0 +1,12 @@ +$ cd controller && go test ./internal/quiesce/ -run 'TestManualPress_RunsOnlyLocalTier' -v -count=1 +RESULT: FAILED (exit code 1) — run on the UNCHANGED quiesce.go (felhom-controller 3124278), new test only. + +=== RUN TestManualPress_RunsOnlyLocalTier + tiers_test.go:298: a manual press must back up ONLY the local (primary) tier; started=[felhom-pbs local] +--- FAIL: TestManualPress_RunsOnlyLocalTier (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.006s +FAIL + +Reading: the old manual press (allTiersForManualRun) requested EVERY advertised tier, in agent +order — here the off-site tier felhom-pbs was started too (listed first in this test on purpose). diff --git a/documentation/audits/design-build-2026-10-06/D/red-press-no-local-backs-up-nothing.txt b/documentation/audits/design-build-2026-10-06/D/red-press-no-local-backs-up-nothing.txt new file mode 100644 index 00000000..7a4e195c --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/D/red-press-no-local-backs-up-nothing.txt @@ -0,0 +1,7 @@ +$ go test ./internal/quiesce/ -run TestManualRunTiers_NoPrimaryAvailable_BacksUpNothing -v (manualRunTiers as the helper wrote it: the off-site tier stands in) +=== RUN TestManualRunTiers_NoPrimaryAvailable_BacksUpNothing + tier_skip_test.go:94: want no tier (the local copy's storage is absent), got [{target:felhom-pbs ageSecs: state: primary:false}] +--- FAIL: TestManualRunTiers_NoPrimaryAvailable_BacksUpNothing (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.005s +FAIL diff --git a/documentation/audits/design-build-2026-10-06/E/E0-9202-test-image.txt b/documentation/audits/design-build-2026-10-06/E/E0-9202-test-image.txt new file mode 100644 index 00000000..41d4580e --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/E0-9202-test-image.txt @@ -0,0 +1,2 @@ +Loaded image: gitea.dooplex.hu/admin/felhom-controller:0.301.0 +gitea.dooplex.hu/admin/felhom-controller:0.301.0 Up 20 seconds (healthy) diff --git a/documentation/audits/design-build-2026-10-06/E/E1-wishlist-config-probe.txt b/documentation/audits/design-build-2026-10-06/E/E1-wishlist-config-probe.txt new file mode 100644 index 00000000..375828ea --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/E1-wishlist-config-probe.txt @@ -0,0 +1,9 @@ +15:27:05 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +15:27:25 [1] deployed, controller state=running, pinned={'wishlist': 'ghcr.io/cmintey/wishlist:v0.67.1'} +deploy: True +15:27:25 gate: wishlist is gated — passed as the household (cookie set) +15:27:26 wishlist: /signup http=200 type=success +seed ok: True +v24.20.0 +[] + diff --git a/documentation/audits/design-build-2026-10-06/E/go-test-green.txt b/documentation/audits/design-build-2026-10-06/E/go-test-green.txt new file mode 100644 index 00000000..35c5d9f0 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/go-test-green.txt @@ -0,0 +1,63 @@ +$ cd controller && go build ./... && go vet ./... && go test ./... -count=1 +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 10.753s +ok gitea.dooplex.hu/admin/felhom-controller/internal/agentapi 0.202s +ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.179s +ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.037s +ok gitea.dooplex.hu/admin/felhom-controller/internal/appexport 0.621s +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 1.771s +ok gitea.dooplex.hu/admin/felhom-controller/internal/backupwindow 0.008s +ok gitea.dooplex.hu/admin/felhom-controller/internal/bootrecon 0.012s +ok gitea.dooplex.hu/admin/felhom-controller/internal/bootstrap 14.026s +ok gitea.dooplex.hu/admin/felhom-controller/internal/channelhealth 0.010s +ok gitea.dooplex.hu/admin/felhom-controller/internal/config 0.007s +ok gitea.dooplex.hu/admin/felhom-controller/internal/crashboot 0.007s +ok gitea.dooplex.hu/admin/felhom-controller/internal/dockerexec 0.200s +ok gitea.dooplex.hu/admin/felhom-controller/internal/family 0.023s +ok gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.009s +ok gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.395s +ok gitea.dooplex.hu/admin/felhom-controller/internal/infra 0.013s +ok gitea.dooplex.hu/admin/felhom-controller/internal/mailrelay 0.017s +ok gitea.dooplex.hu/admin/felhom-controller/internal/metrics 0.010s +ok gitea.dooplex.hu/admin/felhom-controller/internal/monitor 0.007s +ok gitea.dooplex.hu/admin/felhom-controller/internal/nightchain 0.010s +ok gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.958s +ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply 0.015s +ok gitea.dooplex.hu/admin/felhom-controller/internal/quiesce 0.086s +ok gitea.dooplex.hu/admin/felhom-controller/internal/report 2.590s +ok gitea.dooplex.hu/admin/felhom-controller/internal/scheduler 0.455s +ok gitea.dooplex.hu/admin/felhom-controller/internal/selfupdate 0.048s +ok gitea.dooplex.hu/admin/felhom-controller/internal/settings 0.580s +ok gitea.dooplex.hu/admin/felhom-controller/internal/setup 0.009s +ok gitea.dooplex.hu/admin/felhom-controller/internal/sockheal 0.006s +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 397.076s +ok gitea.dooplex.hu/admin/felhom-controller/internal/sync 0.393s +ok gitea.dooplex.hu/admin/felhom-controller/internal/system 0.065s +ok gitea.dooplex.hu/admin/felhom-controller/internal/util 0.030s +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 64.818s +rc=0 + +$ cd controller && python3 scripts/controller_gates.py --fast + +============================================================================== +== summary +============================================================================== + template-id OK (exit 0) + emoji OK (exit 0) + native-confirm OK (exit 0) + offbox-rename OK (exit 0) + app-row-dedup OK (exit 0) + mojibake OK (exit 0) + docker-v OK (exit 0) + secret-markup OK (exit 0) + retrieval-promise OK (exit 0) + debug-routes OK (exit 0) + reuse-refs OK (exit 0) + instructions OK (exit 0) + observations OK (exit 0) + minagent-header OK (exit 0) + i18n OK (exit 0) + go-parity OK (exit 0) + gofmt OK (exit 0) + golden-notice ADVISORY (exit 0, advisory) + +all controller gates OK diff --git a/documentation/audits/design-build-2026-10-06/E/live.txt b/documentation/audits/design-build-2026-10-06/E/live.txt new file mode 100644 index 00000000..20db8d0f --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/live.txt @@ -0,0 +1,29 @@ +wishlist installed from the earlier probe — removing it first (scratch box) +9202 controller: gitea.dooplex.hu/admin/felhom-controller:0.301.0 + wishlist: /signup http=200 type=success +[1] the household's first account: True | DB: ['DB enableSignup= users=1'] +[2] the household opens the app (setup done) -> 200 {'data': {'opened': True}, 'error': '', 'ok': True} +[2] DB after the close command: ['DB enableSignup=false users=1'] +2026/10/06 13:31:53 after_setup.go:156: [INFO] [stacks] wishlist: the app's own sign-up switch CLOSED (the gate opened (household); env [], command true) + +[3] stranger, straight at the app: {"type":"error","error":{"message":"This instance is invite only"}} HTTP 401 +[3] DB: ['DB enableSignup=false users=1'] +[4] the household opens the 15-minute window -> 200 {'data': {'open_until': '2026-10-06T13:47:34Z'}, 'error': '', 'ok': True} +[4] DB inside the window: ['DB enableSignup=true users=1'] +[4] a family member, straight at the app: {"type":"success","status":200,"data":"[{\"success\":1},true]"} HTTP 200 +[4] DB: ['DB enableSignup=true users=2'] +2026/10/06 13:32:34 signup_block.go:142: [INFO] [stacks] wishlist: the household opened sign-up until 2026-10-06T13:47:34Z — the loop closes it again +2026/10/06 13:32:34 after_setup.go:205: [INFO] [stacks] wishlist: the app's own sign-up switch opened for the household's window (the household's window; env [], open_command true) + +[5] the switch reads CLOSED again 906 s after the check began: ['DB enableSignup=false users=2'] +[5] stranger after the window, straight at the app: {"type":"error","error":{"message":"This instance is invite only"}} HTTP 401 +[5] DB: ['DB enableSignup=false users=2'] +2026/10/06 13:32:34 signup_block.go:142: [INFO] [stacks] wishlist: the household opened sign-up until 2026-10-06T13:47:34Z — the loop closes it again +2026/10/06 13:32:34 after_setup.go:205: [INFO] [stacks] wishlist: the app's own sign-up switch opened for the household's window (the household's window; env [], open_command true) +2026/10/06 13:47:44 after_setup.go:156: [INFO] [stacks] wishlist: the app's own sign-up switch CLOSED (the loop (window ended or retry); env [], command true) + +[5] app page record: {"after_setup": {"at": "2026-10-06T13:47:44Z", "ok": true}, "setup_gate": {"state": "open", "since": "2026-10-06T13:30:49Z", "hosts": ["wishlist.enkisfelhom.hu"], "opened_at": "2026-10-06T13:31:53Z", "opened_by": "household", "signup_open_until": "2026-10-06T13:47:34Z", "native_lock": "applied"}} +15:48:49 [X] stop -> 200 {'ok': True, 'message': 'Stack wishlist stop completed'} +15:49:21 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wishlist', 'volumes_removed': ['wishlist_wishlist_data', 'wishlist_wishlist_uploads'], 'hdd_paths_removed': [], 'hdd_paths_pre +15:49:29 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wishlist' +remove -> 200 diff --git a/documentation/audits/design-build-2026-10-06/E/red-close-after-update.txt b/documentation/audits/design-build-2026-10-06/E/red-close-after-update.txt new file mode 100644 index 00000000..9dfb69a7 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/red-close-after-update.txt @@ -0,0 +1,14 @@ +# R-717 red-proof: TestAfterSetupR717_TheCloseRunsAgainAfterAnUpdate +# mutation in controller/internal/stacks/update.go: +# --- removed +# m.markNativeLockForReapply(name, dir) // R-717: the loop closes the app's own sign-up switch again after an update +# +++ put instead +# // RED-PROOF: no re-close after an update +$ cd controller && go test ./internal/stacks/ -run ^TestAfterSetupR717_TheCloseRunsAgainAfterAnUpdate$ -count=1 -v +=== RUN TestAfterSetupR717_TheCloseRunsAgainAfterAnUpdate + after_setup_r717_test.go:232: the update did not re-run the close: closes=1 open=true +--- FAIL: TestAfterSetupR717_TheCloseRunsAgainAfterAnUpdate (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.017s +FAIL +# exit code: 1 diff --git a/documentation/audits/design-build-2026-10-06/E/red-failed-close-retry.txt b/documentation/audits/design-build-2026-10-06/E/red-failed-close-retry.txt new file mode 100644 index 00000000..ea171cc1 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/red-failed-close-retry.txt @@ -0,0 +1,14 @@ +# R-717 red-proof: TestAfterSetupR717_AFailedCloseIsRetriedNotTrusted +# mutation in controller/internal/stacks/signup_block.go: +# --- removed +# gap = nativeLockOpenRetry +# +++ put instead +# // RED-PROOF: the 30-minute retry also after a window +$ cd controller && go test ./internal/stacks/ -run ^TestAfterSetupR717_AFailedCloseIsRetriedNotTrusted$ -count=1 -v +=== RUN TestAfterSetupR717_AFailedCloseIsRetriedNotTrusted + after_setup_r717_test.go:210: the loop did not retry the close +--- FAIL: TestAfterSetupR717_AFailedCloseIsRetriedNotTrusted (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.012s +FAIL +# exit code: 1 diff --git a/documentation/audits/design-build-2026-10-06/E/red-failed-open.txt b/documentation/audits/design-build-2026-10-06/E/red-failed-open.txt new file mode 100644 index 00000000..54f6f9ec --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/red-failed-open.txt @@ -0,0 +1,19 @@ +# R-717 red-proof: TestAfterSetupR717_AFailedOpenStaysClosedAndSaysSo +# mutation in controller/internal/stacks/after_setup.go: +# --- removed +# if cerr := m.applyNativeLock(name, true, "an opening that failed"); cerr != nil { +# +++ put instead +# if cerr := error(nil); cerr != nil { // RED-PROOF: no re-close after a failed open +$ cd controller && go test ./internal/stacks/ -run ^TestAfterSetupR717_AFailedOpenStaysClosedAndSaysSo$ -count=1 -v +=== RUN TestAfterSetupR717_AFailedOpenStaysClosedAndSaysSo + after_setup_r717_test.go:170: a failed open left the switch OPEN + [INFO] [stacks] gapp: setup gate CLOSED before the first start — only the household reaches [gapp.example.hu] until the first setup is done + [INFO] [stacks] gapp: setup gate OPENED by household — the app is reached as without a gate + [INFO] [stacks] gapp: the app's own sign-up switch CLOSED (the gate opened (household); env [], command true) + [ERROR] [stacks] gapp: the app's own sign-up switch could NOT be opened (the household's window): after_setup open_command did not report "SIGNUP-OPENED" (err exit 1) — closing it again; a new family member cannot sign up in the app itself + [INFO] [stacks] gapp: the household opened sign-up until 2026-10-06T13:27:46Z — the loop closes it again +--- FAIL: TestAfterSetupR717_AFailedOpenStaysClosedAndSaysSo (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.010s +FAIL +# exit code: 1 diff --git a/documentation/audits/design-build-2026-10-06/E/red-no-open-command-logs-once.txt b/documentation/audits/design-build-2026-10-06/E/red-no-open-command-logs-once.txt new file mode 100644 index 00000000..d8032479 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/red-no-open-command-logs-once.txt @@ -0,0 +1,21 @@ +# R-717 red-proof: TestAfterSetupR717_NoOpenCommandLogsOnceAndKeepsTheWindow +# mutation in controller/internal/stacks/after_setup.go: +# --- removed +# if _, seen := m.nativeOpenMissingLogged.LoadOrStore(name, true); !seen { +# +++ put instead +# if seen := false; !seen { // RED-PROOF: no once-guard +$ cd controller && go test ./internal/stacks/ -run ^TestAfterSetupR717_NoOpenCommandLogsOnceAndKeepsTheWindow$ -count=1 -v +=== RUN TestAfterSetupR717_NoOpenCommandLogsOnceAndKeepsTheWindow + after_setup_r717_test.go:253: the missing open_command was logged 2 times, want once: + [INFO] [stacks] gapp: setup gate CLOSED before the first start — only the household reaches [gapp.example.hu] until the first setup is done + [INFO] [stacks] gapp: setup gate OPENED by household — the app is reached as without a gate + [INFO] [stacks] gapp: the app's own sign-up switch CLOSED (the gate opened (household); env [], command true) + [WARN] [stacks] gapp: the template closes the app's own sign-up switch with a command but has no open_command — the household's window cannot reopen it (only the address block opens) + [INFO] [stacks] gapp: the household opened sign-up until 2026-10-06T13:27:53Z — the loop closes it again + [WARN] [stacks] gapp: the template closes the app's own sign-up switch with a command but has no open_command — the household's window cannot reopen it (only the address block opens) + [INFO] [stacks] gapp: the household opened sign-up until 2026-10-06T13:27:53Z — the loop closes it again +--- FAIL: TestAfterSetupR717_NoOpenCommandLogsOnceAndKeepsTheWindow (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.011s +FAIL +# exit code: 1 diff --git a/documentation/audits/design-build-2026-10-06/E/red-open-then-close.txt b/documentation/audits/design-build-2026-10-06/E/red-open-then-close.txt new file mode 100644 index 00000000..f5564ff9 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/red-open-then-close.txt @@ -0,0 +1,16 @@ +# R-717 red-proof: TestAfterSetupR717_OpenThenCloseEachRunOnce +# mutation in controller/internal/stacks/after_setup.go: +# --- removed +# if err == nil && len(spec.OpenCommand) > 0 { +# err = m.runAfterSetupCommand(name, dir, spec, spec.OpenCommand, spec.OpenSuccess, "open_command") +# } +# +++ put instead +# // RED-PROOF: the open_command is never run (the code before R-717) +$ cd controller && go test ./internal/stacks/ -run ^TestAfterSetupR717_OpenThenCloseEachRunOnce$ -count=1 -v +=== RUN TestAfterSetupR717_OpenThenCloseEachRunOnce + after_setup_r717_test.go:97: the window did not open the app's own switch: opens=0 closes=1 open=false +--- FAIL: TestAfterSetupR717_OpenThenCloseEachRunOnce (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.012s +FAIL +# exit code: 1 diff --git a/documentation/audits/design-build-2026-10-06/E/red-restart-mid-open.txt b/documentation/audits/design-build-2026-10-06/E/red-restart-mid-open.txt new file mode 100644 index 00000000..ccc95a99 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/E/red-restart-mid-open.txt @@ -0,0 +1,22 @@ +# R-717 red-proof: TestAfterSetupR717_ARestartInsideTheWindowEndsClosed +# mutation in controller/internal/stacks/after_setup.go: +# --- removed +# m.mutateAppConfig(name, dir, "after_setup_opening", func(cfg *AppConfig) bool { +# if cfg.SetupGate == nil { +# return false +# } +# cfg.SetupGate.NativeLock = NativeLockOpening +# return true +# }) +# if c := LoadAppConfig(dir); c == nil || c.SetupGate == nil || c.SetupGate.NativeLock != NativeLockOpening { +# +++ put instead +# // RED-PROOF: no "opening" mark before anything opens +# if c := LoadAppConfig(dir); c == nil || c.SetupGate == nil { +$ cd controller && go test ./internal/stacks/ -run ^TestAfterSetupR717_ARestartInsideTheWindowEndsClosed$ -count=1 -v +=== RUN TestAfterSetupR717_ARestartInsideTheWindowEndsClosed + after_setup_r717_test.go:149: a restart left sign-up OPEN in the app past its window (record at the crash: native="applied") +--- FAIL: TestAfterSetupR717_ARestartInsideTheWindowEndsClosed (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s +FAIL +# exit code: 1 diff --git a/documentation/audits/design-build-2026-10-06/F/bench-static-check.txt b/documentation/audits/design-build-2026-10-06/F/bench-static-check.txt new file mode 100644 index 00000000..b21aa91d --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench-static-check.txt @@ -0,0 +1,14 @@ +Tue Oct 6 12:57:05 UTC 2026 +wger-files Up 4 minutes (healthy) +wger Up 4 minutes (healthy) +wger-files /static/css/workout-manager.css: 200 2442B text/css +wger-files /static/bootstrap-compiled.css: 200 277029B +control - the same file straight at wger (Django, DEBUG=False): 404 +login page links: /static/css/workout-manager.7007d84ce531.css /static/bootstrap-compiled.80a6279921f8.css /static/css/bootstrap-custom.400ad578123c.css +static volume: 283M; files: 22725 +wger-files 9.938MiB / 32MiB +wger 246MiB / 384MiB +the page's own link /static/css/workout-manager.7007d84ce531.css: 200 2481B +the page's own link /static/bootstrap-compiled.80a6279921f8.css: 200 277042B +the page's own link /static/css/bootstrap-custom.400ad578123c.css: 200 1006B +a directory listing is refused: 403 diff --git a/documentation/audits/design-build-2026-10-06/F/bench/R762-wger.log b/documentation/audits/design-build-2026-10-06/F/bench/R762-wger.log new file mode 100644 index 00000000..26bf41f9 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/R762-wger.log @@ -0,0 +1,126 @@ +[12:50:17] scratch drive folders cleared before FROM (R-656): none existed +[12:50:17] MV-wger: deploying wger at FROM {'wger': 'wger/server:2.7'} +[12:52:15] FROM settled=True in 92.6s :: {"wger": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[12:52:15] fixture: the BOX walk's own (Wger), through upgrade_boxport +[12:52:16] wger: the generated admin password does not log in (POST /en/user/login -> 200) — running the template's own after_install command (the app's CLI, as the product does after an install) +[12:52:18] wger: after_install :: version 8.3.1, blocking by username FELHOM_AFTER_INSTALL_OK +[12:52:19] wger: POST /api/v2/weightentry/ http=201 +[12:52:19] wger: readback of the seeded weight entry http=200 found=True +[12:52:19] C1 (seed reads back BEFORE): True +[12:52:20] MV-wger: swapping to TO {'wger': 'wger/server:2.7', 'wger-files': 'nginx:1.30.5-alpine'} +[12:52:36] TO up -d rc=0 +[12:53:07] TO settled=True in 31.1s :: {"wger": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "wger-files": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[12:53:07] migration lines observed: 6 +[12:53:08] wger: readback of the seeded weight entry http=200 found=True +[12:53:08] RESULT (seed reads back AFTER): True +[12:53:08] memory watch: 600s, 4 callers on 1 path(s) at 172.18.0.2:8000 +[12:53:23] + 15s wger=284M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=284 +[12:53:38] + 30s wger=284M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=574 +[12:53:53] + 46s wger=281M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=863 +[12:54:09] + 61s wger=262M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=1152 +[12:54:24] + 76s wger=259M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=1441 +[12:54:39] + 91s wger=258M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=1731 +[12:54:54] + 106s wger=257M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=2021 +[12:55:09] + 121s wger=258M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=2309 +[12:55:24] + 136s wger=258M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=2600 +[12:55:40] + 152s wger=257M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=2887 +[12:55:55] + 167s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=3175 +[12:56:10] + 182s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=3461 +[12:56:25] + 197s wger=257M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=3747 +[12:56:40] + 212s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=4037 +[12:56:55] + 227s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=4325 +[12:57:11] + 242s wger=259M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=15M kills=0 rs=0 reqs=4613 +[12:57:26] + 258s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=4897 +[12:57:41] + 273s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=5184 +[12:57:56] + 288s wger=258M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=5471 +[12:58:11] + 303s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=5763 +[12:58:26] + 318s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6051 +[12:58:41] + 334s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6335 +[12:58:57] + 349s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6623 +[12:59:12] + 364s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6912 +[12:59:27] + 379s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=7197 +[12:59:42] + 394s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=7483 +[12:59:57] + 409s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=7769 +[13:00:12] + 424s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8058 +[13:00:28] + 440s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8347 +[13:00:43] + 455s wger=260M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8635 +[13:00:58] + 470s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8920 +[13:01:13] + 485s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=9208 +[13:01:28] + 500s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=9497 +[13:01:43] + 515s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=9787 +[13:01:59] + 530s wger=257M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10074 +[13:02:14] + 546s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10364 +[13:02:29] + 561s wger=257M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10653 +[13:02:44] + 576s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10942 +[13:02:59] + 591s wger=257M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=11230 +[13:03:14] + 606s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=11518 +[13:03:15] memory watch: killed=False tight=[] requests=11519 codes={'302': 11519} +[13:03:15] MV-wger: ABORT — putting the FROM images back +[13:03:58] wger: readback of the seeded weight entry http=200 found=True +[13:03:58] ABORT: app came back in 31.3s; data present=True +{ + "harness_version": 5, + "edge": "MV-wger", + "app": "wger", + "note": "definition step to wger@files", + "from": { + "wger": "wger/server:2.7" + }, + "to": { + "wger": "wger/server:2.7", + "wger-files": "nginx:1.30.5-alpine" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "\u001b[2Kwger | Performing database migrations", + "abort": "starts-and-serves", + "abort_detail": null, + "engine_state_after": null, + "memory": { + "soak_s": 606.7, + "requested_s": 600, + "requests": 11519, + "codes": { + "302": 11519 + }, + "first_kill": null, + "containers": { + "wger": { + "limit": 402653184, + "peak": 402653184, + "peak_pct": 1.0, + "anon_peak_sampled": 202338304, + "anon_peak_pct": 0.503, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "wger-files": { + "limit": 33554432, + "peak": 16572416, + "peak_pct": 0.494, + "anon_peak_sampled": 5337088, + "anon_peak_pct": 0.159, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + } + }, + "unmeasured": [], + "load": "reached" + }, + "marks": [], + "bench_overrides": null, + "duration_s": 31.1, + "measured_at": "2026-10-06T13:03:58Z", + "evidence": "evidence/MV-wger", + "scratch_cleared": [], + "files_changed": [], + "files_changed_detail": [], + "files_ignored": [], + "total_s": 821.0 +} diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/abort-states.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/abort-states.json new file mode 100644 index 00000000..5d1f0247 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/abort-states.json @@ -0,0 +1,14 @@ +{ + "wger": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "wger-files": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + } +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/compose-final.log b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/compose-final.log new file mode 100644 index 00000000..cb5e10b7 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/compose-final.log @@ -0,0 +1,245 @@ +wger | *** Using settings from env: settings.main +wger | level=INFO ts=2026-10-06 15:03:27,479 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | Performing database migrations +wger | level=INFO ts=2026-10-06 15:03:29,033 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | System check identified some issues: +wger | +wger | WARNINGS: +wger | ?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. +wger | HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +wger | Operations to perform: +wger | Apply all migrations: account, actstream, allauth_idp_oidc, auth, authtoken, axes, config, contenttypes, core, easy_thumbnails, exercises, gallery, gym, mailer, manager, measurements, mfa, nutrition, sessions, sites, socialaccount, token_blacklist, trophies, weight +wger | Running migrations: +wger | No migrations to apply. +wger | Your models in app(s): 'exercises', 'gallery' have changes that are not yet reflected in a migration, and so won't be applied. +wger | Run 'manage.py makemigrations' to make new migrations, and then re-run 'manage.py migrate' to apply them. +wger | level=INFO ts=2026-10-06 15:03:32,185 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | System check identified some issues: +wger | +wger | WARNINGS: +wger | ?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. +wger | HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +wger | Set site URL to fitness.gate.invalid +wger | Using django's development server on port 8000... +wger | level=INFO ts=2026-10-06 15:03:34,510 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | level=INFO ts=2026-10-06 15:03:35,797 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | level=INFO ts=2026-10-06 15:03:35,810 module=autoreload path=/home/wger/.local/lib/python3.12/site-packages/django/utils/autoreload.py line=681 message=Watching for file changes with StatReloader +wger | Performing system checks... +wger | +wger | System check identified some issues: +wger | +wger | WARNINGS: +wger | ?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. +wger | HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +wger | +wger | System check identified 1 issue (0 silenced). +wger | October 06, 2026 - 15:03:36 +wger | Django version 6.0.8, using settings 'settings.main' +wger | Starting development server at http://0.0.0.0:8000/ +wger | Quit the server with CONTROL-C. +wger | +wger | WARNING: This is a development server. Do not use it in a production setting. Use a production WSGI or ASGI server instead. +wger | For more information on production servers see: https://docs.djangoproject.com/en/6.0/howto/deployment/ +wger | Can't find name for static file reference: images/logos/logo-social.png. This is expected in some node packages. +wger | Can't find name for static file reference: fontawesomefree/css/all.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/workout-manager.css. This is expected in some node packages. +wger | Can't find name for static file reference: bootstrap-compiled.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/bootstrap-custom.css. This is expected in some node packages. +wger | Can't find name for static file reference: images/favicon.png. This is expected in some node packages. +wger | [06/Oct/2026 15:03:56] "HEAD / HTTP/1.1" 302 0 +wger | [06/Oct/2026 15:03:56] "HEAD /en/ HTTP/1.1" 302 0 +wger | Can't find name for static file reference: images/logos/logo-social.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/favicon.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/logo-font.svg. This is expected in some node packages. +wger | Can't find name for static file reference: bootstrap-compiled.css. This is expected in some node packages. +wger | Can't find name for static file reference: fontawesomefree/css/all.min.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/landing_page.css. This is expected in some node packages. +wger | Can't find name for static file reference: node/bootstrap/dist/js/bootstrap.bundle.min.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/language.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/forms.js. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/logo-font.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/hero.avif. This is expected in some node packages. +wger | Can't find name for static file reference: images/hero.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/screens-1.avif. This is expected in some node packages. +wger | Can't find name for static file reference: images/screens-1.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/screens-2.avif. This is expected in some node packages. +wger | Can't find name for static file reference: images/screens-2.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/screens-3.avif. This is expected in some node packages. +wger | Can't find name for static file reference: images/screens-3.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/community.avif. This is expected in some node packages. +wger | Can't find name for static file reference: images/community.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/dumbbell.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/code.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/fork.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/logo-font-inverse.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/bg.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ca.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/cs.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/de.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/el.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en-au.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en-gb.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ar.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-co.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-mx.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ni.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ve.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/fi.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/fr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/he.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/hr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/it.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ko.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/mk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/nl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/nb.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pt.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pt-br.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ro.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ru.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sv.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ta.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/th.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/tr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/uk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/zh-hans.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/zh-hant.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/play-store/badge.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/app-store/black.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/flathub/black.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/fdroid/get-it-on.svg. This is expected in some node packages. +wger | [06/Oct/2026 15:03:56] "HEAD /en/software/features HTTP/1.1" 200 0 +wger | Can't find name for static file reference: images/logos/logo-social.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/favicon.png. This is expected in some node packages. +wger | Can't find name for static file reference: css/workout-manager.css. This is expected in some node packages. +wger | Can't find name for static file reference: bootstrap-compiled.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/bootstrap-custom.css. This is expected in some node packages. +wger | Can't find name for static file reference: fontawesomefree/css/all.min.css. This is expected in some node packages. +wger | Can't find name for static file reference: node/@wger-project/react-components/build/assets/index.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/language-menu.css. This is expected in some node packages. +wger | Can't find name for static file reference: node/bootstrap/dist/js/bootstrap.bundle.min.js. This is expected in some node packages. +wger | Can't find name for static file reference: node/htmx.org/dist/htmx.min.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/nutrition.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/language.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/forms.js. This is expected in some node packages. +wger | Can't find name for static file reference: node/@wger-project/react-components/build/main.js. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/logo-bg-white.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/play-store/badge.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/app-store/black.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/flathub/black.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/bg.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ca.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/cs.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/de.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/el.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en-au.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en-gb.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ar.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-co.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-mx.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ni.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ve.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/fi.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/fr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/he.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/hr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/it.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ko.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/mk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/nl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/nb.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pt.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pt-br.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ro.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ru.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sv.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ta.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/th.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/tr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/uk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/zh-hans.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/zh-hant.svg. This is expected in some node packages. +wger | Can't find name for static file reference: mfa/js/webauthn-json.js. This is expected in some node packages. +wger | Can't find name for static file reference: mfa/js/webauthn.js. This is expected in some node packages. +wger | Can't find name for static file reference: account/js/onload.js. This is expected in some node packages. +wger | [06/Oct/2026 15:03:57] "GET /en/user/login HTTP/1.1" 200 39022 +wger | level=WARNING ts=2026-10-06 15:03:57,689 module=log path=/home/wger/.local/lib/python3.12/site-packages/django/utils/log.py line=249 message=Forbidden: /api/v2/weightentry/ +wger | [06/Oct/2026 15:03:57] "GET /api/v2/weightentry/?weight=60.37 HTTP/1.1" 403 58 +wger | Can't find name for static file reference: images/logos/logo-social.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/favicon.png. This is expected in some node packages. +wger | Can't find name for static file reference: css/workout-manager.css. This is expected in some node packages. +wger | Can't find name for static file reference: bootstrap-compiled.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/bootstrap-custom.css. This is expected in some node packages. +wger | Can't find name for static file reference: fontawesomefree/css/all.min.css. This is expected in some node packages. +wger | Can't find name for static file reference: node/@wger-project/react-components/build/assets/index.css. This is expected in some node packages. +wger | Can't find name for static file reference: css/language-menu.css. This is expected in some node packages. +wger | Can't find name for static file reference: node/bootstrap/dist/js/bootstrap.bundle.min.js. This is expected in some node packages. +wger | Can't find name for static file reference: node/htmx.org/dist/htmx.min.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/nutrition.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/language.js. This is expected in some node packages. +wger | Can't find name for static file reference: js/forms.js. This is expected in some node packages. +wger | Can't find name for static file reference: node/@wger-project/react-components/build/main.js. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/logo-bg-white.png. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/play-store/badge.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/app-store/black.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/logos/flathub/black.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/bg.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ca.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/cs.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/de.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/el.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en-au.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/en-gb.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ar.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-co.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-mx.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ni.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/es-ve.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/fi.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/fr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/he.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/hr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/it.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ko.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/mk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/nl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/nb.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pt.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/pt-br.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ro.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ru.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sl.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/sv.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/ta.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/th.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/tr.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/uk.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/zh-hans.svg. This is expected in some node packages. +wger | Can't find name for static file reference: images/icons/flags/zh-hant.svg. This is expected in some node packages. +wger | Can't find name for static file reference: mfa/js/webauthn-json.js. This is expected in some node packages. +wger | Can't find name for static file reference: mfa/js/webauthn.js. This is expected in some node packages. +wger | Can't find name for static file reference: account/js/onload.js. This is expected in some node packages. +wger | [06/Oct/2026 15:03:57] "GET /en/user/login HTTP/1.1" 200 39022 +wger | level=INFO ts=2026-10-06 15:03:58,072 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=295 message=AXES: Successful login by {username: "********************", ip_address: "********************", user_agent: "curl/8.14.1", path_info: "/en/user/login"}. +wger | level=INFO ts=2026-10-06 15:03:58,086 module=database path=/home/wger/.local/lib/python3.12/site-packages/axes/handlers/database.py line=425 message=AXES: Cleaned up 1 expired access attempts from database that were older than 2026-10-06 12:58:57.858736+00:00 +wger | Can't find name for static file reference: images/logos/logo-social.png. This is expected in some node packages. +wger | [06/Oct/2026 15:03:58] "POST /en/user/login HTTP/1.1" 302 0 +wger | [06/Oct/2026 15:03:58] "GET /api/v2/weightentry/?weight=199.99 HTTP/1.1" 200 52 +wger | [06/Oct/2026 15:03:58] "GET /api/v2/weightentry/?weight=60.37 HTTP/1.1" 200 158 diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/engine-state.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/engine-state.json new file mode 100644 index 00000000..ec747fa4 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/engine-state.json @@ -0,0 +1 @@ +null \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-after-detail.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-after-detail.json new file mode 100644 index 00000000..9e26dfee --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-after-detail.json @@ -0,0 +1 @@ +{} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-after.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-after.json new file mode 100644 index 00000000..9e26dfee --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-after.json @@ -0,0 +1 @@ +{} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-before-detail.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-before-detail.json new file mode 100644 index 00000000..9e26dfee --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-before-detail.json @@ -0,0 +1 @@ +{} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-before.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-before.json new file mode 100644 index 00000000..9e26dfee --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/files-before.json @@ -0,0 +1 @@ +{} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/memory-samples.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/memory-samples.json new file mode 100644 index 00000000..2cb352b0 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/memory-samples.json @@ -0,0 +1,1122 @@ +[ + { + "t": 15.2, + "containers": { + "wger": { + "limit": 402653184, + "current": 298065920, + "peak": 402653184, + "anon": 200736768, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12689408, + "peak": 14786560, + "anon": 5144576, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 284 + }, + { + "t": 30.3, + "containers": { + "wger": { + "limit": 402653184, + "current": 298090496, + "peak": 402653184, + "anon": 200884224, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12603392, + "peak": 14786560, + "anon": 5181440, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 574 + }, + { + "t": 45.5, + "containers": { + "wger": { + "limit": 402653184, + "current": 295055360, + "peak": 402653184, + "anon": 200904704, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12435456, + "peak": 14786560, + "anon": 5181440, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 863 + }, + { + "t": 60.7, + "containers": { + "wger": { + "limit": 402653184, + "current": 274833408, + "peak": 402653184, + "anon": 200884224, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12472320, + "peak": 14786560, + "anon": 5218304, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1152 + }, + { + "t": 75.8, + "containers": { + "wger": { + "limit": 402653184, + "current": 272379904, + "peak": 402653184, + "anon": 200884224, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12472320, + "peak": 14786560, + "anon": 5218304, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1441 + }, + { + "t": 91.0, + "containers": { + "wger": { + "limit": 402653184, + "current": 270979072, + "peak": 402653184, + "anon": 200888320, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12509184, + "peak": 14786560, + "anon": 5255168, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 1731 + }, + { + "t": 106.1, + "containers": { + "wger": { + "limit": 402653184, + "current": 270352384, + "peak": 402653184, + "anon": 200888320, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12509184, + "peak": 14786560, + "anon": 5255168, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2021 + }, + { + "t": 121.3, + "containers": { + "wger": { + "limit": 402653184, + "current": 270733312, + "peak": 402653184, + "anon": 200888320, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12550144, + "peak": 14868480, + "anon": 5292032, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2309 + }, + { + "t": 136.5, + "containers": { + "wger": { + "limit": 402653184, + "current": 271216640, + "peak": 402653184, + "anon": 200888320, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12550144, + "peak": 14868480, + "anon": 5292032, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2600 + }, + { + "t": 151.6, + "containers": { + "wger": { + "limit": 402653184, + "current": 270180352, + "peak": 402653184, + "anon": 200888320, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12595200, + "peak": 14909440, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 2887 + }, + { + "t": 166.8, + "containers": { + "wger": { + "limit": 402653184, + "current": 270659584, + "peak": 402653184, + "anon": 200888320, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12595200, + "peak": 14909440, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3175 + }, + { + "t": 181.9, + "containers": { + "wger": { + "limit": 402653184, + "current": 270950400, + "peak": 402653184, + "anon": 200892416, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12595200, + "peak": 14909440, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3461 + }, + { + "t": 197.1, + "containers": { + "wger": { + "limit": 402653184, + "current": 269983744, + "peak": 402653184, + "anon": 200908800, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12595200, + "peak": 14909440, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 3747 + }, + { + "t": 212.2, + "containers": { + "wger": { + "limit": 402653184, + "current": 270884864, + "peak": 402653184, + "anon": 200908800, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12595200, + "peak": 14909440, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4037 + }, + { + "t": 227.4, + "containers": { + "wger": { + "limit": 402653184, + "current": 270663680, + "peak": 402653184, + "anon": 200908800, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 12595200, + "peak": 14909440, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4325 + }, + { + "t": 242.5, + "containers": { + "wger": { + "limit": 402653184, + "current": 271679488, + "peak": 402653184, + "anon": 200921088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13312000, + "peak": 16195584, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4613 + }, + { + "t": 257.7, + "containers": { + "wger": { + "limit": 402653184, + "current": 271978496, + "peak": 402653184, + "anon": 200921088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16195584, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 4897 + }, + { + "t": 272.8, + "containers": { + "wger": { + "limit": 402653184, + "current": 271896576, + "peak": 402653184, + "anon": 201084928, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5184 + }, + { + "t": 288.0, + "containers": { + "wger": { + "limit": 402653184, + "current": 271486976, + "peak": 402653184, + "anon": 201084928, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5471 + }, + { + "t": 303.2, + "containers": { + "wger": { + "limit": 402653184, + "current": 271978496, + "peak": 402653184, + "anon": 201084928, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 5763 + }, + { + "t": 318.3, + "containers": { + "wger": { + "limit": 402653184, + "current": 271724544, + "peak": 402653184, + "anon": 201084928, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6051 + }, + { + "t": 333.5, + "containers": { + "wger": { + "limit": 402653184, + "current": 272207872, + "peak": 402653184, + "anon": 201097216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6335 + }, + { + "t": 348.6, + "containers": { + "wger": { + "limit": 402653184, + "current": 272142336, + "peak": 402653184, + "anon": 201097216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6623 + }, + { + "t": 363.8, + "containers": { + "wger": { + "limit": 402653184, + "current": 272478208, + "peak": 402653184, + "anon": 201097216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 6912 + }, + { + "t": 378.9, + "containers": { + "wger": { + "limit": 402653184, + "current": 272125952, + "peak": 402653184, + "anon": 201097216, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7197 + }, + { + "t": 394.1, + "containers": { + "wger": { + "limit": 402653184, + "current": 271986688, + "peak": 402653184, + "anon": 201146368, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7483 + }, + { + "t": 409.2, + "containers": { + "wger": { + "limit": 402653184, + "current": 271773696, + "peak": 402653184, + "anon": 201146368, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 7769 + }, + { + "t": 424.4, + "containers": { + "wger": { + "limit": 402653184, + "current": 271884288, + "peak": 402653184, + "anon": 201146368, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8058 + }, + { + "t": 439.5, + "containers": { + "wger": { + "limit": 402653184, + "current": 271945728, + "peak": 402653184, + "anon": 201146368, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16310272, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8347 + }, + { + "t": 454.7, + "containers": { + "wger": { + "limit": 402653184, + "current": 273219584, + "peak": 402653184, + "anon": 202248192, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8635 + }, + { + "t": 469.8, + "containers": { + "wger": { + "limit": 402653184, + "current": 269266944, + "peak": 402653184, + "anon": 202268672, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 8920 + }, + { + "t": 485.0, + "containers": { + "wger": { + "limit": 402653184, + "current": 269324288, + "peak": 402653184, + "anon": 202268672, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9208 + }, + { + "t": 500.2, + "containers": { + "wger": { + "limit": 402653184, + "current": 269180928, + "peak": 402653184, + "anon": 202268672, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9497 + }, + { + "t": 515.3, + "containers": { + "wger": { + "limit": 402653184, + "current": 269258752, + "peak": 402653184, + "anon": 202268672, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 9787 + }, + { + "t": 530.5, + "containers": { + "wger": { + "limit": 402653184, + "current": 269692928, + "peak": 402653184, + "anon": 202268672, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10074 + }, + { + "t": 545.6, + "containers": { + "wger": { + "limit": 402653184, + "current": 269316096, + "peak": 402653184, + "anon": 202338304, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10364 + }, + { + "t": 560.9, + "containers": { + "wger": { + "limit": 402653184, + "current": 269864960, + "peak": 402653184, + "anon": 202297344, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10653 + }, + { + "t": 576.0, + "containers": { + "wger": { + "limit": 402653184, + "current": 268996608, + "peak": 402653184, + "anon": 202321920, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13950976, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 10942 + }, + { + "t": 591.2, + "containers": { + "wger": { + "limit": 402653184, + "current": 269762560, + "peak": 402653184, + "anon": 202330112, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13910016, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 11230 + }, + { + "t": 606.3, + "containers": { + "wger": { + "limit": 402653184, + "current": 268980224, + "peak": 402653184, + "anon": 202330112, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + }, + "wger-files": { + "limit": 33554432, + "current": 13910016, + "peak": 16572416, + "anon": 5337088, + "oom_kill": 0, + "restarts": 0, + "oomkilled_flag": false, + "status": "running", + "cgroup": true + } + }, + "requests": 11518 + } +] \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/migration-lines.txt b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/migration-lines.txt new file mode 100644 index 00000000..8d423c31 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/migration-lines.txt @@ -0,0 +1,6 @@ +wger | Performing database migrations +wger | Apply all migrations: account, actstream, allauth_idp_oidc, auth, authtoken, axes, config, contenttypes, core, easy_thumbnails, exercises, gallery, gym, mailer, manager, measurements, mfa, nutrition, sessions, sites, socialaccount, token_blacklist, trophies, weight +wger | Running migrations: +wger | No migrations to apply. +wger | Your models in app(s): 'exercises', 'gallery' have changes that are not yet reflected in a migration, and so won't be applied. +wger | Run 'manage.py makemigrations' to make new migrations, and then re-run 'manage.py migrate' to apply them. \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/run.log b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/run.log new file mode 100644 index 00000000..19ec984f --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/run.log @@ -0,0 +1,60 @@ +[12:50:17] scratch drive folders cleared before FROM (R-656): none existed +[12:50:17] MV-wger: deploying wger at FROM {'wger': 'wger/server:2.7'} +[12:52:15] FROM settled=True in 92.6s :: {"wger": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[12:52:15] fixture: the BOX walk's own (Wger), through upgrade_boxport +[12:52:16] wger: the generated admin password does not log in (POST /en/user/login -> 200) — running the template's own after_install command (the app's CLI, as the product does after an install) +[12:52:18] wger: after_install :: version 8.3.1, blocking by username FELHOM_AFTER_INSTALL_OK +[12:52:19] wger: POST /api/v2/weightentry/ http=201 +[12:52:19] wger: readback of the seeded weight entry http=200 found=True +[12:52:19] C1 (seed reads back BEFORE): True +[12:52:20] MV-wger: swapping to TO {'wger': 'wger/server:2.7', 'wger-files': 'nginx:1.30.5-alpine'} +[12:52:36] TO up -d rc=0 +[12:53:07] TO settled=True in 31.1s :: {"wger": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}, "wger-files": {"status": "running", "health": "healthy", "restarts": 0, "exit": 0}} +[12:53:07] migration lines observed: 6 +[12:53:08] wger: readback of the seeded weight entry http=200 found=True +[12:53:08] RESULT (seed reads back AFTER): True +[12:53:08] memory watch: 600s, 4 callers on 1 path(s) at 172.18.0.2:8000 +[12:53:23] + 15s wger=284M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=284 +[12:53:38] + 30s wger=284M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=574 +[12:53:53] + 46s wger=281M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=863 +[12:54:09] + 61s wger=262M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=1152 +[12:54:24] + 76s wger=259M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=1441 +[12:54:39] + 91s wger=258M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=1731 +[12:54:54] + 106s wger=257M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=2021 +[12:55:09] + 121s wger=258M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=2309 +[12:55:24] + 136s wger=258M/384M peak=384M kills=0 rs=0 wger-files=11M/32M peak=14M kills=0 rs=0 reqs=2600 +[12:55:40] + 152s wger=257M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=2887 +[12:55:55] + 167s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=3175 +[12:56:10] + 182s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=3461 +[12:56:25] + 197s wger=257M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=3747 +[12:56:40] + 212s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=4037 +[12:56:55] + 227s wger=258M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=14M kills=0 rs=0 reqs=4325 +[12:57:11] + 242s wger=259M/384M peak=384M kills=0 rs=0 wger-files=12M/32M peak=15M kills=0 rs=0 reqs=4613 +[12:57:26] + 258s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=4897 +[12:57:41] + 273s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=5184 +[12:57:56] + 288s wger=258M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=5471 +[12:58:11] + 303s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=5763 +[12:58:26] + 318s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6051 +[12:58:41] + 334s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6335 +[12:58:57] + 349s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6623 +[12:59:12] + 364s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=6912 +[12:59:27] + 379s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=7197 +[12:59:42] + 394s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=7483 +[12:59:57] + 409s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=7769 +[13:00:12] + 424s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8058 +[13:00:28] + 440s wger=259M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8347 +[13:00:43] + 455s wger=260M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8635 +[13:00:58] + 470s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=8920 +[13:01:13] + 485s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=9208 +[13:01:28] + 500s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=9497 +[13:01:43] + 515s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=9787 +[13:01:59] + 530s wger=257M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10074 +[13:02:14] + 546s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10364 +[13:02:29] + 561s wger=257M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10653 +[13:02:44] + 576s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=10942 +[13:02:59] + 591s wger=257M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=11230 +[13:03:14] + 606s wger=256M/384M peak=384M kills=0 rs=0 wger-files=13M/32M peak=15M kills=0 rs=0 reqs=11518 +[13:03:15] memory watch: killed=False tight=[] requests=11519 codes={'302': 11519} +[13:03:15] MV-wger: ABORT — putting the FROM images back +[13:03:58] wger: readback of the seeded weight entry http=200 found=True +[13:03:58] ABORT: app came back in 31.3s; data present=True \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/to-full.log b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/to-full.log new file mode 100644 index 00000000..2d8a2ca0 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/to-full.log @@ -0,0 +1,71 @@ +wger-files | /docker-entrypoint.sh: /docker-entrypoint.d/ is not empty, will attempt to perform configuration +wger | *** Using settings from env: settings.main +wger-files | /docker-entrypoint.sh: Looking for shell scripts in /docker-entrypoint.d/ +wger-files | /docker-entrypoint.sh: Launching /docker-entrypoint.d/10-listen-on-ipv6-by-default.sh +wger-files | 10-listen-on-ipv6-by-default.sh: info: Getting the checksum of /etc/nginx/conf.d/default.conf +wger-files | 10-listen-on-ipv6-by-default.sh: info: Enabled listen on IPv6 in /etc/nginx/conf.d/default.conf +wger-files | /docker-entrypoint.sh: Sourcing /docker-entrypoint.d/15-local-resolvers.envsh +wger | level=INFO ts=2026-10-06 14:52:37,325 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | Running in production mode, running collectstatic now +wger | level=INFO ts=2026-10-06 14:52:39,095 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | +wger | 11362 static files copied to '/home/wger/static', 11362 post-processed. +wger | Performing database migrations +wger | level=INFO ts=2026-10-06 14:52:57,357 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | System check identified some issues: +wger | +wger | WARNINGS: +wger | ?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. +wger | HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +wger | Operations to perform: +wger | Apply all migrations: account, actstream, allauth_idp_oidc, auth, authtoken, axes, config, contenttypes, core, easy_thumbnails, exercises, gallery, gym, mailer, manager, measurements, mfa, nutrition, sessions, sites, socialaccount, token_blacklist, trophies, weight +wger | Running migrations: +wger | No migrations to apply. +wger | Your models in app(s): 'exercises', 'gallery' have changes that are not yet reflected in a migration, and so won't be applied. +wger | Run 'manage.py makemigrations' to make new migrations, and then re-run 'manage.py migrate' to apply them. +wger | level=INFO ts=2026-10-06 14:53:00,351 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | System check identified some issues: +wger | +wger | WARNINGS: +wger | ?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. +wger-files | /docker-entrypoint.sh: Launching /docker-entrypoint.d/20-envsubst-on-templates.sh +wger | HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +wger | Set site URL to fitness.gate.invalid +wger | Using django's development server on port 8000... +wger | level=INFO ts=2026-10-06 14:53:02,636 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger | level=INFO ts=2026-10-06 14:53:03,920 module=apps path=/home/wger/.local/lib/python3.12/site-packages/axes/apps.py line=53 message=AXES: BEGIN version 8.3.1, blocking by username +wger-files | /docker-entrypoint.sh: Launching /docker-entrypoint.d/30-tune-worker-processes.sh +wger-files | /docker-entrypoint.sh: Configuration complete; ready for start up +wger-files | 2026/10/06 12:52:36 [notice] 1#1: using the "epoll" event method +wger-files | 2026/10/06 12:52:36 [notice] 1#1: nginx/1.30.5 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: built by gcc 15.2.0 (Alpine 15.2.0) +wger-files | 2026/10/06 12:52:36 [notice] 1#1: OS: Linux 7.0.14-20-pve +wger-files | 2026/10/06 12:52:36 [notice] 1#1: getrlimit(RLIMIT_NOFILE): 524288:524288 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker processes +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker process 30 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker process 31 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker process 32 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker process 33 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker process 34 +wger-files | 2026/10/06 12:52:36 [notice] 1#1: start worker process 35 +wger-files | 127.0.0.1 - - [06/Oct/2026:12:53:06 +0000] "GET / HTTP/1.1" 200 896 "-" "Wget" "-" +wger | level=INFO ts=2026-10-06 14:53:03,931 module=autoreload path=/home/wger/.local/lib/python3.12/site-packages/django/utils/autoreload.py line=681 message=Watching for file changes with StatReloader +wger | Performing system checks... +wger | +wger | System check identified some issues: +wger | +wger | WARNINGS: +wger | ?: (axes.W006) AXES_LOCKOUT_PARAMETERS does not contain 'ip_address'. This configuration allows attackers to bypass rate limits by rotating User-Agents or Cookies. +wger | HINT: Add 'ip_address' to AXES_LOCKOUT_PARAMETERS. +wger | +wger | System check identified 1 issue (0 silenced). +wger | October 06, 2026 - 14:53:04 +wger | Django version 6.0.8, using settings 'settings.main' +wger | Starting development server at http://0.0.0.0:8000/ +wger | Quit the server with CONTROL-C. +wger | +wger | WARNING: This is a development server. Do not use it in a production setting. Use a production WSGI or ASGI server instead. +wger | For more information on production servers see: https://docs.djangoproject.com/en/6.0/howto/deployment/ +wger | [06/Oct/2026 14:53:05] "HEAD / HTTP/1.1" 302 0 +wger | [06/Oct/2026 14:53:05] "HEAD /en/ HTTP/1.1" 302 0 +wger | [06/Oct/2026 14:53:06] "HEAD /en/software/features HTTP/1.1" 200 0 diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/to-states.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/to-states.json new file mode 100644 index 00000000..5d1f0247 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/to-states.json @@ -0,0 +1,14 @@ +{ + "wger": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + }, + "wger-files": { + "status": "running", + "health": "healthy", + "restarts": 0, + "exit": 0 + } +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/verdict.json b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/verdict.json new file mode 100644 index 00000000..5f584bcd --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/bench/evidence/MV-wger/verdict.json @@ -0,0 +1,66 @@ +{ + "harness_version": 5, + "edge": "MV-wger", + "app": "wger", + "note": "definition step to wger@files", + "from": { + "wger": "wger/server:2.7" + }, + "to": { + "wger": "wger/server:2.7", + "wger-files": "nginx:1.30.5-alpine" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "\u001b[2Kwger | Performing database migrations", + "abort": "starts-and-serves", + "abort_detail": null, + "engine_state_after": null, + "memory": { + "soak_s": 606.7, + "requested_s": 600, + "requests": 11519, + "codes": { + "302": 11519 + }, + "first_kill": null, + "containers": { + "wger": { + "limit": 402653184, + "peak": 402653184, + "peak_pct": 1.0, + "anon_peak_sampled": 202338304, + "anon_peak_pct": 0.503, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + }, + "wger-files": { + "limit": 33554432, + "peak": 16572416, + "peak_pct": 0.494, + "anon_peak_sampled": 5337088, + "anon_peak_pct": 0.159, + "oom_kills": 0, + "restarts": 0, + "oomkilled_flag": false, + "measured": true + } + }, + "unmeasured": [], + "load": "reached" + }, + "marks": [], + "bench_overrides": null, + "duration_s": 31.1, + "measured_at": "2026-10-06T13:03:58Z", + "evidence": "evidence/MV-wger", + "scratch_cleared": [], + "files_changed": [], + "files_changed_detail": [], + "files_ignored": [], + "total_s": 821.0 +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/box/box-verdict-wger.json b/documentation/audits/design-build-2026-10-06/F/box/box-verdict-wger.json new file mode 100644 index 00000000..ca351194 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/box/box-verdict-wger.json @@ -0,0 +1,32 @@ +{ + "app": "wger", + "venue": "box 9202 (drill catalog, controller 0.301.0 test image), the product's guarded Update (R-762 definition step)", + "from": { + "wger": "wger/server:2.7" + }, + "to": { + "wger": "wger/server:2.7", + "wger-files": "nginx:1.30.5-alpine" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "duration_s": 59.4, + "final_phase": "done", + "measured_at": "2026-10-06T13:35:32Z", + "note": "written by hand from step.txt: wgerstep.py crashed AFTER the step when it decoded the PNG body as text; the after-checks were re-read binary-safe (step.txt '[5] (finished by hand\u2026)')", + "r762": { + "photo": "/media/gallery/1/36151595-cdde-4493-849d-72a4ff32d493.png", + "photo_before": "404", + "photo_after": "200 178 image/png", + "css_before": "404 x3", + "css_after": "200 x3 (the login page's own hashed links)", + "control_unknown_static": "404", + "memory": [ + "wger-files 7.613MiB / 32MiB", + "wger 241MiB / 384MiB" + ] + }, + "evidence": "felhom.eu/documentation/audits/design-build-2026-10-06/F/box/step.txt" +} \ No newline at end of file diff --git a/documentation/audits/design-build-2026-10-06/F/box/step.txt b/documentation/audits/design-build-2026-10-06/F/box/step.txt new file mode 100644 index 00000000..de3b452b --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/box/step.txt @@ -0,0 +1,27 @@ +drill: c182492 DRILL wger: un-hidden for the R-762 box step (DRILL only) +[1] installed, pinned={'wger': 'wger/server:2.7'} + wger: the generated admin password does not log in (POST /en/user/login -> 403) — running the template's own after_install command (the app's CLI, as the product does after an install) + wger: after_install :: version 8.3.1, blocking by username FELHOM_AFTER_INSTALL_OK + wger: POST /api/v2/weightentry/ http=201 + wger: readback of the seeded weight entry http=200 found=True +[2] POST /api/v2/gallery/ (a 178 B PNG) -> 201 +[2] the photo /media/gallery/1/36151595-cdde-4493-849d-72a4ff32d493.png read BEFORE the step -> 404; CSS before {'/static/css/workout-manager.css': '404', '/static/bootstrap-compiled.css': '404', '/static/css/bootstrap-custom.css': '404'} +drill: 095a292 DRILL wger: the R-762 definition (wger-files) + its ladder entry (box proof) | to: {'wger': 'wger/server:2.7', 'wger-files': 'nginx:1.30.5-alpine'} + phase +0.0s backing-up | err=None + phase +12.3s safety-dump | err=None + phase +13.3s pulling | err=None + phase +16.4s copying | err=None + phase +27.7s starting | err=None + phase +28.7s verifying | err=None + phase +59.4s done | err=None + wger: readback of the seeded weight entry http=200 found=True +[5] (finished by hand after the script's text decode of the PNG body failed) the photo /media/gallery/1/36151595-cdde-4493-849d-72a4ff32d493.png -> 200 178 image/png (posted: 178 B) +[5] the login page's own CSS links: {'/static/css/workout-manager.7007d84ce531.css': '200 2481 text/css', '/static/bootstrap-compiled.80a6279921f8.css': '200 277042 text/css', '/static/css/bootstrap-custom.400ad578123c.css': '200 1006 text/css'} +[5] control, an unknown file under /static/ -> 404 153 text/html +[5] memory: ['wger-files 7.613MiB / 32MiB', 'wger 241MiB / 384MiB'] +[5] pinned: {'wger': 'wger/server:2.7', 'wger-files': 'nginx:1.30.5-alpine'} state: running +[5] badges after: {'hu': [{'title': 'Ez az alkalmazás a legfrissebb elérhető változatot futtatja.', 'text': 'Naprakész'}], 'en': [{'title': 'This app is running the newest version available.', 'text': 'Up to date'}]} +15:37:45 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'} +15:38:17 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media', 'wger_wger_static'], 'hdd_paths_removed': [], 'hdd_paths_prese +15:38:25 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger' +remove -> 200 diff --git a/documentation/audits/design-build-2026-10-06/F/red-definition-edge.txt b/documentation/audits/design-build-2026-10-06/F/red-definition-edge.txt new file mode 100644 index 00000000..1481ce3a --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/F/red-definition-edge.txt @@ -0,0 +1,65 @@ +test_all_three_conditions_allow (__main__.BenchAdminSeedGuard.test_all_three_conditions_allow) ... ok +test_each_condition_alone_refuses (__main__.BenchAdminSeedGuard.test_each_condition_alone_refuses) ... ok +test_the_bench_seeds_through_the_admin_invite_and_never_shows_the_token (__main__.BenchAdminSeedGuard.test_the_bench_seeds_through_the_admin_invite_and_never_shows_the_token) ... /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:177: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/.felhom-h-mxcpia_3' mode='r' encoding='utf-8'> + headers.append(open(extra[i + 1][1:]).read().strip()) +ResourceWarning: Enable tracemalloc to get the object allocation traceback +ok +test_the_bench_venue_names_itself (__main__.BenchAdminSeedGuard.test_the_bench_venue_names_itself) ... ok +test_the_box_walk_never_signs_in_as_admin (__main__.BenchAdminSeedGuard.test_the_box_walk_never_signs_in_as_admin) ... ok +test_a_household_shaped_box_is_never_seeded (__main__.BoxAdminSeedGuard.test_a_household_shaped_box_is_never_seeded) ... ok +test_a_refused_invite_is_inconclusive (__main__.BoxAdminSeedGuard.test_a_refused_invite_is_inconclusive) ... ok +test_all_conditions_allow (__main__.BoxAdminSeedGuard.test_all_conditions_allow) ... ok +test_each_condition_alone_refuses (__main__.BoxAdminSeedGuard.test_each_condition_alone_refuses) ... ok +test_only_a_drill_address_is_invited (__main__.BoxAdminSeedGuard.test_only_a_drill_address_is_invited) ... ok +test_the_test_box_seeds_through_the_invite_inside_the_box (__main__.BoxAdminSeedGuard.test_the_test_box_seeds_through_the_invite_inside_the_box) ... ok +test_names_changed_added_removed_and_nothing_else (__main__.ChangedFiles.test_names_changed_added_removed_and_nothing_else) ... ok +test_never_a_bare_root_never_outside (__main__.ClearScratch.test_never_a_bare_root_never_outside) ... /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:52: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/bench-scratch-cy0e4bk1/hdd/appdata/x/config/config.php' mode='w' encoding='utf-8'> + open(p, "w").write("last run") +ResourceWarning: Enable tracemalloc to get the object allocation traceback +ok +test_the_apps_own_folders_are_cleared_and_said (__main__.ClearScratch.test_the_apps_own_folders_are_cleared_and_said) ... /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:52: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/bench-scratch-eox1rlxi/hdd/appdata/nextcloud/config/config.php' mode='w' encoding='utf-8'> + open(p, "w").write("last run") +ResourceWarning: Enable tracemalloc to get the object allocation traceback +/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py:52: ResourceWarning: unclosed file <_io.TextIOWrapper name='/tmp/bench-scratch-eox1rlxi/hdd/appdata/immich/config/config.php' mode='w' encoding='utf-8'> + open(p, "w").write("last run") +ResourceWarning: Enable tracemalloc to get the object allocation traceback +ok +test_nothing_new_is_refused (__main__.DefinitionEdge.test_nothing_new_is_refused) ... ok +test_the_edge_names_the_new_service (__main__.DefinitionEdge.test_the_edge_names_the_new_service) ... ERROR +test_the_to_render_reads_the_new_definition (__main__.DefinitionEdge.test_the_to_render_reads_the_new_definition) ... FAIL +test_answers_of_any_code_are_the_app_answering (__main__.LoadVerdict.test_answers_of_any_code_are_the_app_answering) ... ok +test_every_request_errored_is_inconclusive (__main__.LoadVerdict.test_every_request_errored_is_inconclusive) ... ok +test_under_half_is_inconclusive_and_none_is_inconclusive (__main__.LoadVerdict.test_under_half_is_inconclusive_and_none_is_inconclusive) ... ok +test_a_listed_name_that_grew_or_vanished_still_marks (__main__.MarkerIgnore.test_a_listed_name_that_grew_or_vanished_still_marks) ... ok +test_a_moved_tree_the_file_walk_cannot_name_keeps_the_mark (__main__.MarkerIgnore.test_a_moved_tree_the_file_walk_cannot_name_keeps_the_mark) ... ok +test_an_unlisted_changed_file_still_marks (__main__.MarkerIgnore.test_an_unlisted_changed_file_still_marks) ... ok +test_every_entry_has_a_reason (__main__.MarkerIgnore.test_every_entry_has_a_reason) ... ok +test_the_list_belongs_to_its_app (__main__.MarkerIgnore.test_the_list_belongs_to_its_app) ... ok +test_the_six_measured_markers_do_not_mark_and_each_carries_a_reason (__main__.MarkerIgnore.test_the_six_measured_markers_do_not_mark_and_each_carries_a_reason) ... ok +test_env_is_0600_and_shredded (__main__.SecretHygiene.test_env_is_0600_and_shredded) ... ok +test_evidence_files_are_redacted (__main__.SecretHygiene.test_evidence_files_are_redacted) ... ok +test_redact_longest_first (__main__.SecretHygiene.test_redact_longest_first) ... ok +test_secret_values_are_the_generated_fields_and_the_seed_password (__main__.SecretHygiene.test_secret_values_are_the_generated_fields_and_the_seed_password) ... ok + +====================================================================== +ERROR: test_the_edge_names_the_new_service (__main__.DefinitionEdge.test_the_edge_names_the_new_service) +---------------------------------------------------------------------- +Traceback (most recent call last): + File "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py", line 376, in test_the_edge_names_the_new_service + self.assertEqual(e["to_template"], "app1@new") + ~^^^^^^^^^^^^^^^ +KeyError: 'to_template' + +====================================================================== +FAIL: test_the_to_render_reads_the_new_definition (__main__.DefinitionEdge.test_the_to_render_reads_the_new_definition) +---------------------------------------------------------------------- +Traceback (most recent call last): + File "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts/test_upgrade_bench.py", line 384, in test_the_to_render_reads_the_new_definition + self.assertIn("nginx:1.30.5-alpine", (wd / "docker-compose.yml").read_text()) + ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +AssertionError: 'nginx:1.30.5-alpine' not found in 'services:\n app1:\n image: app/one:1.0\n' + +---------------------------------------------------------------------- +Ran 30 tests in 0.028s + +FAILED (failures=1, errors=1) diff --git a/documentation/audits/design-build-2026-10-06/delivery/vouch-golden-floors.txt b/documentation/audits/design-build-2026-10-06/delivery/vouch-golden-floors.txt new file mode 100644 index 00000000..ef0eb073 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/delivery/vouch-golden-floors.txt @@ -0,0 +1,21 @@ +== vouch 2026-10-06T14:04:05Z: agent 0.149.0, golden 0.301.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +== floors 2026-10-06T14:04:35Z: 0.301.0 +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/06 16:04:35 [INFO] Artifact manifest set: agent=0.149.0 golden=0.301.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66" +2026/10/06 16:04:35 [INFO] Customer demo-hp controller-version floor override set to "0.301.0" (declared MinAgent "") +2026/10/06 16:04:35 [INFO] Customer demo-felhom controller-version floor override set to "0.301.0" (declared MinAgent "") +2026/10/06 16:04:36 [INFO] Customer tester-1 controller-version floor override set to "0.301.0" (declared MinAgent "") +== floors again 2026-10-06T14:04:56Z: 0.301.0 with min_agent 0.131.0 (the first save declared none; allowed because the floor equals the golden, re-saved to match the 0.300.0 floors) +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/06 16:04:35 [INFO] Customer demo-hp controller-version floor override set to "0.301.0" (declared MinAgent "") +2026/10/06 16:04:35 [INFO] Customer demo-felhom controller-version floor override set to "0.301.0" (declared MinAgent "") +2026/10/06 16:04:36 [INFO] Customer tester-1 controller-version floor override set to "0.301.0" (declared MinAgent "") +2026/10/06 16:04:56 [INFO] Customer demo-hp controller-version floor override set to "0.301.0" (declared MinAgent "0.131.0") +2026/10/06 16:04:57 [INFO] Customer demo-felhom controller-version floor override set to "0.301.0" (declared MinAgent "0.131.0") +2026/10/06 16:04:57 [INFO] Customer tester-1 controller-version floor override set to "0.301.0" (declared MinAgent "0.131.0") diff --git a/documentation/audits/design-build-2026-10-06/tools/r638slice0.py b/documentation/audits/design-build-2026-10-06/tools/r638slice0.py new file mode 100644 index 00000000..7d4d2482 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/tools/r638slice0.py @@ -0,0 +1,158 @@ +"""r638slice0.py — R-638 slice 0, MEASUREMENT ONLY, on scratch 9202 (drill catalog). `09` §3 decision 154. + + 1 the drill's template is put at the OLD definition (the commit before the R-462 move), ladder emptied; + 2 install fresh through the product; seed through the app's own door; read back (C1); the live table list; + 3 a DRILL commit moves it to the NEW definition with the one proven ladder entry; sync; the product's guarded + Update (its backing-up phase makes the Tier-2 copy at the OLD version); the table list after the migration; + 4 the named unit restore of that pre-update copy through the household's button (POST /backup/restore); + 5 positive observable: the controller's `Imported DB dump` line + the seed read back; + control from a different channel: the live table list against the table list READ FROM THE COPY'S OWN DUMP FILE + on the Tier-2 drive — every table the migration added must be GONE; + 6 removed through the product (KEEP=1 keeps it). +Evidence: ../B/.txt + ../B/-tables.json.""" +import json, os, re, subprocess, sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +import upgrade_fixtures_box as fixtures + +APP = sys.argv[1] +CFG = { + "docmost": {"sub": "docmost", "old": "6d8cd87^", "new": "6d8cd87", "db": "docmost-postgres", + "tables": "psql -U \"$POSTGRES_USER\" -d \"$POSTGRES_DB\" -Atc \"select tablename from pg_tables where schemaname='public' order by 1\""}, + "romm": {"sub": "romm", "extra": {"HDD_PATH": "/mnt/felhom-drives/scratch_hdd"}, "old": "15f9ebf^", "new": "15f9ebf", "db": "romm-db", + "tables": "MYSQL_PWD=\"$MYSQL_PASSWORD\" mariadb -u\"$MYSQL_USER\" \"$MYSQL_DATABASE\" -Nse 'show tables'"}, +}[APP] +LIVE = "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu" +D = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +HERE = os.path.dirname(os.path.abspath(__file__)) +EV = os.path.join(HERE, "..", "B") +log = open(os.path.join(EV, f"{APP}.txt"), "a", buffering=1) +REC = {"app": APP} + + +def say(*a): + w.say(*a); log.write(" ".join(map(str, a)) + "\n") + + +def git(*a, cwd=D): + return subprocess.run(["git", "-C", cwd, *a], check=True, capture_output=True, text=True).stdout + + +def live_tables(): + out = w.guest(f"docker exec -i {CFG['db']} sh -s 2>&1 <<'SQLX'\n{CFG['tables']}\nSQLX\n") + return sorted(t for t in out.split() if re.fullmatch(r"[A-Za-z0-9_]+", t)) + + +def set_template(rev, ladder_entries, msg): + git("pull", "-q", "--rebase", "origin", "main") + comp = git("show", f"{rev}:templates/{APP}/docker-compose.yml", cwd=LIVE) + open(f"{D}/templates/{APP}/docker-compose.yml", "w").write(comp) + sys.path.insert(0, f"{LIVE}/scripts") + import ladder + fy = open(f"{LIVE}/templates/{APP}/.felhom.yml").read() + entries, _, _ = ladder.parse(fy) + head = fy.split("\nupdate_ladder:")[0].rstrip("\n") + "\n" + if ladder_entries: + keep = [e for e in entries if e in ladder_entries] + for e in keep: + head = ladder.append_entry(head, e) + open(f"{D}/templates/{APP}/.felhom.yml", "w").write(head) + subprocess.run(["rm", "-rf", f"{D}/templates/{APP}/steps"], check=True) + subprocess.run(["git", "-C", D, "add", "-A", f"templates/{APP}"], check=True) + git("commit", "-q", "--allow-empty", "-m", msg) + git("push", "-q", "origin", "main") + say("drill:", git("log", "--oneline", "-1").strip(), "| images:", re.findall(r"image:\s*(\S+)", comp)) + return entries + + +fx = fixtures.FIXTURES[APP] +w.login() +REC["controller"] = w.guest("docker inspect -f '{{.Config.Image}}' felhom-controller").strip() +say("9202 controller:", REC["controller"]) +if w.stack(APP).get("deployed"): + say(f"{APP} already installed on 9202 — removing it first (scratch box)") + w.remove(APP) +sys.path.insert(0, f"{LIVE}/scripts") +import ladder +all_entries = ladder.parse(open(f"{LIVE}/templates/{APP}/.felhom.yml").read())[0] +set_template(CFG["old"], [], f"DRILL {APP}: the OLD definition ({CFG['old']}) for R-638 slice 0") +old_imgs = re.findall(r"image:\s*(\S+)", git("show", f"{CFG['old']}:templates/{APP}/docker-compose.yml", cwd=LIVE)) +w.sync_rescan(APP, old_imgs[0]) # wait until the box's catalog carries the OLD definition (run 1 installed the live one) +if not w.deploy(APP, CFG["sub"], CFG.get("extra")): + sys.exit(say("RESULT the install did not complete") or 1) +tok = fx.seed(w, CFG["sub"], say) +if tok is None or not fx.verify(w, CFG["sub"], tok, say): + sys.exit(say("RESULT C1 failed — the seed did not read back before") or 1) +REC["pinned_old"] = (w.stack(APP).get("app_config") or {}).get("pinned_images") +if old_imgs[0] not in (REC["pinned_old"] or {}).values(): + sys.exit(say(f"RESULT the box installed {REC['pinned_old']}, not the OLD definition — stopping") or 1) +REC["tables_old_live"] = live_tables() +say(f"[2] old version installed, pinned={REC['pinned_old']}, live tables {len(REC['tables_old_live'])}") + +new_comp = git("show", f"{CFG['new']}:templates/{APP}/docker-compose.yml", cwd=LIVE) +new_imgs = dict(re.findall(r"^ ([a-z0-9-]+):\s*\n(?:.*\n)*?\s+image:\s*(\S+)", new_comp, re.M)) +step = all_entries[:1] # docmost 0.95.0 -> 0.96.0 (pg16), romm 5.0.0 -> 5.3.0 (mariadb 11.4) +say(f"[3] the ladder entry used: {[ (e['from'], e['to']) for e in step ]}") +set_template(CFG["new"], step, f"DRILL {APP}: the NEW definition ({CFG['new']}) + its one ladder entry, R-638 slice 0") +w.sync_rescan(APP, next(iter(step[0]["to"].values())) if step else None) +since = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip() +res = w.press_update(APP, poll=1, cap_s=1800) +REC["update_final_phase"] = res.get("final_phase") +REC["pinned_new"] = (w.stack(APP).get("app_config") or {}).get("pinned_images") +time.sleep(10) +REC["tables_new_live"] = live_tables() +added = sorted(set(REC["tables_new_live"]) - set(REC["tables_old_live"])) +REC["tables_added_by_migration"] = added +say(f"[3] update -> {res.get('final_phase')}, pinned={REC['pinned_new']}, live tables {len(REC['tables_new_live'])}, " + f"added by the migration: {len(added)} {added}") +REC["seed_after_update"] = fx.verify(w, CFG["sub"], tok, say) +if res.get("final_phase") != "done" or not added: + say("RESULT the update did not migrate (no done, or no new table) — slice 0 cannot measure; stopping before restore") + json.dump(REC, open(os.path.join(EV, f"{APP}-tables.json"), "w"), indent=2) + sys.exit(1) + +snaps = w.snapshots(APP) +say(f"[4] restorable copies offered: {json.dumps(snaps)[:600]}") +pick = snaps[0] if snaps else None +sid = pick and (pick.get("id") or pick.get("snapshot_id") or pick.get("short_id")) +since_r = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip() +r = w.restore(APP, sid) +REC["restore"] = {k: r.get(k) for k in ("snapshot_id", "http", "seconds", "state_after", "hold_after")} +time.sleep(15) +lines = w.guest(f"docker logs --since {since_r} felhom-controller 2>&1 | grep -v DEBUG | grep -i -E 'restor|Imported DB dump|replay|volume' | cut -c1-300") +log.write(lines + "\n") +REC["imported_db_dump_line"] = [l for l in lines.splitlines() if "Imported DB dump" in l] +REC["tables_after_restore"] = live_tables() +REC["seed_after_restore"] = fx.verify(w, CFG["sub"], tok, say) +# control from a different channel: the table list in the copy's own dump file on the Tier-2 drive +REC["unit_files"] = w.guest(f"find /mnt/felhom-drives/scratch_hdd/backups -path '*{APP}*' -type f 2>/dev/null | head -40").split() +# the RESTORED copy's dump — never a `pre-restore-*` safety dump the restore itself wrote (run 3 of romm read one) +dumps = "\n".join(f for f in REC["unit_files"] if re.search(r"(\.sql(\.gz|\.zst)?$|dump)", f) and "pre-restore" not in f) +REC["dump_files_seen"] = dumps.split() +copy_tables = [] +if dumps.split(): + f0 = dumps.split()[0] + cat = "zcat" if f0.endswith(".gz") else ("zstdcat" if f0.endswith(".zst") else "cat") + # tables AND views: `SHOW TABLES` / pg_tables list what the live side lists, so the copy side must name the same kinds + # (romm has three views; run 3 counted only CREATE TABLE and read 24 of 27). pg_tables lists tables only. + kinds = "TABLE|VIEW" if APP == "romm" else "TABLE" + out = w.guest(f"{cat} '{f0}' | grep -o -E 'CREATE (ALGORITHM=[A-Z]+ )?(DEFINER=[^ ]+ )?(SQL SECURITY [A-Z]+ )?({kinds}) (IF NOT EXISTS )?[`\"]?([a-z]+\\.)?[`\"]?[A-Za-z0-9_]+' | sed -E 's/.*[ .`\"]//' | sort -u") + copy_tables = sorted(set(t for t in out.split() if re.fullmatch(r"[A-Za-z0-9_]+", t))) +REC["tables_in_copy_dump"] = copy_tables +REC["dump_used_for_control"] = dumps.split()[:1] +left = sorted(set(REC["tables_after_restore"]) & set(added)) +REC["migration_tables_left_after_restore"] = left +REC["live_equals_copy"] = REC["tables_after_restore"] == copy_tables +say(f"[5] Imported DB dump lines: {REC['imported_db_dump_line']}") +say(f"[5] seed read back after restore: {REC['seed_after_restore']}") +say(f"[5] live tables after restore {len(REC['tables_after_restore'])}; the copy's own dump lists {len(copy_tables)}; " + f"equal={REC['live_equals_copy']}; migration tables still there: {left}") +REC["pinned_after_restore"] = (w.stack(APP).get("app_config") or {}).get("pinned_images") +REC["state_after_restore"] = w.stack(APP).get("state") +say(f"[5] pinned after restore {REC['pinned_after_restore']}, state {REC['state_after_restore']}") +ok = bool(REC["imported_db_dump_line"]) and REC["seed_after_restore"] and not left and REC["live_equals_copy"] +REC["verdict"] = "SAFE" if ok else "NOT SAFE / NOT SHOWN" +say(f"RESULT {APP}: {REC['verdict']}") +json.dump(REC, open(os.path.join(EV, f"{APP}-tables.json"), "w"), indent=2) +if not os.environ.get("KEEP"): + say(f"remove -> {w.remove(APP)}") diff --git a/documentation/audits/design-build-2026-10-06/tools/r717live.py b/documentation/audits/design-build-2026-10-06/tools/r717live.py new file mode 100644 index 00000000..f250058e --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/tools/r717live.py @@ -0,0 +1,93 @@ +"""r717live.py — R-717 live proof on scratch 9202 (test controller 0.301.0, drill catalog with wishlist's after_setup). + + 1 wishlist installed fresh; the household's first account (the fixture, through the setup gate); + 2 the household opens the app (POST /apps//setup-gate/open) → after_setup's close command runs; + the DB value read (enableSignup=false) + the controller's own lines; + 3 a STRANGER straight at the app (the container's address — past the address block, the strictest case): + sign-up refused; the user count before/after; + 4 the household's 15-minute window (POST /apps//signup-window) → the open command; DB = true; a family member + signs up straight at the app → success; the user count +1; + 5 after the window ends (polled; ~15 min): the close command ran again; DB = false; a stranger refused; count unchanged. +Evidence: ../E/live.txt.""" +import json, os, secrets, sys, time +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +import upgrade_fixtures_box as fixtures + +APP, SUB = "wishlist", "wishlist" +HERE = os.path.dirname(os.path.abspath(__file__)) +log = open(os.path.join(HERE, "..", "E", "live.txt"), "a", buffering=1) + + +def say(*a): + w.say(*a); log.write(" ".join(map(str, a)) + "\n") + + +DBJS = r'''const {DatabaseSync}=require('node:sqlite');const db=new DatabaseSync('/usr/src/app/data/prod.db',{readOnly:true}); +const v=db.prepare("select value from system_config where key='enableSignup' and groupId='global'").get(); +const n=db.prepare("select count(*) as n from user").get();console.log('DB enableSignup='+(v?v.value:'')+' users='+n.n);''' + + +def db(): + return w.guest("docker exec -i wishlist node - <<'JS'\n" + DBJS + "\nJS\n").strip().splitlines()[-1:] + + +def stranger(tag): + """A sign-up straight at the app's own address inside the box — past the box's address block.""" + u = f"{tag}{secrets.token_hex(3)}" + body = f"name={tag}&username={u}&email={u}%40example.invalid&password=Drill-{secrets.token_hex(8)}&tokenId=" + out = w.guest(f"""ip=$(docker inspect -f '{{{{range .NetworkSettings.Networks}}}}{{{{.IPAddress}}}} {{{{end}}}}' wishlist | awk '{{print $1}}') +curl -s -m 30 -H 'Host: {SUB}.{w.DOMAIN}' -H 'Origin: https://{SUB}.{w.DOMAIN}' -H 'x-sveltekit-action: true' \\ + -H 'Content-Type: application/x-www-form-urlencoded' -w '\\nHTTP %{{http_code}}' --data '{body}' http://$ip:3000/signup""") + return " ".join(out.split())[-220:] + + +def ctl_lines(since): + return w.guest(f"docker logs --since {since} felhom-controller 2>&1 | grep -v DEBUG | grep -i -E 'after_setup|native|signup|sign-up|window' | cut -c1-300") + + +w.login() +if w.stack(APP).get("deployed"): + say("wishlist installed from the earlier probe — removing it first (scratch box)") + w.remove(APP) +w.sync_rescan() +time.sleep(5) +w.sync_rescan() +say("9202 controller:", w.guest("docker inspect -f '{{.Config.Image}}' felhom-controller").strip()) +if not w.deploy(APP, SUB): + sys.exit(say("RESULT install failed") or 1) +fx = fixtures.FIXTURES[APP] +tok = fx.seed(w, SUB, say) +say("[1] the household's first account:", tok is not None, "| DB:", db()) +t0 = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip() +code, d = w.ctl("POST", f"/apps/{APP}/setup-gate/open", {}) +say("[2] the household opens the app (setup done) ->", code, str(d)[:160]) +time.sleep(25) +say("[2] DB after the close command:", db()) +log.write(ctl_lines(t0) + "\n") +say("[3] stranger, straight at the app:", stranger("stranger")) +say("[3] DB:", db()) +t1 = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip() +code, d = w.ctl("POST", f"/apps/{APP}/signup-window", {}) +say("[4] the household opens the 15-minute window ->", code, str(d)[:160]) +time.sleep(25) +say("[4] DB inside the window:", db()) +say("[4] a family member, straight at the app:", stranger("family")) +say("[4] DB:", db()) +log.write(ctl_lines(t1) + "\n") +t2 = time.time() +while time.time() - t2 < 1500: + time.sleep(30) + v = db() + if v and "enableSignup=false" in v[0]: + say(f"[5] the switch reads CLOSED again {round(time.time() - t2)} s after the check began: {v}") + break +else: + say("[5] the switch did NOT close within 25 minutes:", db()) +say("[5] stranger after the window, straight at the app:", stranger("late")) +say("[5] DB:", db()) +log.write(ctl_lines(t1) + "\n") +st = w.stack(APP) +say("[5] app page record:", json.dumps({k: (st.get("app_config") or {}).get(k) for k in ("after_setup", "setup_gate")})[:600]) +if not os.environ.get("KEEP"): + say("remove ->", w.remove(APP)) diff --git a/documentation/audits/design-build-2026-10-06/tools/repoint.py b/documentation/audits/design-build-2026-10-06/tools/repoint.py new file mode 100644 index 00000000..6bd3fed8 --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/tools/repoint.py @@ -0,0 +1,37 @@ +#!/usr/bin/env python3 +"""Point 9202 at the drill catalog, or put the saved controller.yaml back. `09` §6.5.""" +import io, os, re, sys +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +SAVE = f"{VOL}/controller.yaml.pre-designbuild" +DRILL = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" +def creds(): + for l in io.open(os.path.expanduser("~/.git-credentials")).read().split("\n"): + m = re.match(r"https://(admin):([^@]+)@gitea\.dooplex\.hu", l) + if m: return m.group(1), m.group(2) + sys.exit("no admin credential") +if sys.argv[1] == "drill": + u, t = creds() + out = w.guest(f"""set -e +test -f {SAVE} || cp -p {VOL}/controller.yaml {SAVE} +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml"; s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml +""") + print(out.replace(t, "")) +elif sys.argv[1] == "restore": + print(w.guest(f"""set -e +cp -p {SAVE} {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +grep -n 'repo_url' {VOL}/controller.yaml; grep -c 'token: ""' {VOL}/controller.yaml || true +cmp {SAVE} {VOL}/controller.yaml && echo RESTORED-IDENTICAL""")) diff --git a/documentation/audits/design-build-2026-10-06/tools/wgerstep.py b/documentation/audits/design-build-2026-10-06/tools/wgerstep.py new file mode 100644 index 00000000..f3fefafa --- /dev/null +++ b/documentation/audits/design-build-2026-10-06/tools/wgerstep.py @@ -0,0 +1,140 @@ +"""wgerstep.py — R-762 on scratch 9202 (drill catalog), ONE process, through the product. + + 1 DRILL: wger un-hidden (DRILL only); install fresh at the live definition; seed (the fixture) + read back (C1); + 2 a progress photo posted to the gallery (the household's API with its session); its /media URL read: 404 expected + (the R-762 fault — the control that the later 200 is the fix, not the test); + 3 the login page's own CSS links read: 404 expected; + 4 a DRILL commit replaces the compose with the new definition + one ladder entry; sync; the product's guarded Update; + 5 after: the seed, the same photo URL (200 + the same byte count), the CSS links (200), wger-files' memory; + 6 box verdict JSON; removed through the product (KEEP=1 keeps it). +Evidence: ../F/box/step.txt + box-verdict-wger.json.""" +import json, os, re, struct, subprocess, sys, tempfile, time, zlib +sys.path.insert(0, "/mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu/scripts") +import box_walk as w +import upgrade_fixtures_box as fixtures + +APP, SUB = "wger", "fitness" +NEWDEF = sys.argv[1] +HERE = os.path.dirname(os.path.abspath(__file__)) +EVD = os.path.join(HERE, "..", "F", "box"); os.makedirs(EVD, exist_ok=True) +log = open(f"{EVD}/step.txt", "a", buffering=1) +D = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +REC = {"app": APP} + + +def say(*a): + w.say(*a); log.write(" ".join(map(str, a)) + "\n") + + +def git(*a): + return subprocess.run(["git", "-C", D, *a], check=True, capture_output=True, text=True).stdout + + +def png(path, n=64): + """A real PNG (n x n, random-ish colour) — the gallery validates the image.""" + import secrets + rgb = secrets.token_bytes(3) + raw = b"".join(b"\x00" + rgb * n for _ in range(n)) + def chunk(t, d): + return struct.pack(">I", len(d)) + t + d + struct.pack(">I", zlib.crc32(t + d) & 0xffffffff) + data = (b"\x89PNG\r\n\x1a\n" + chunk(b"IHDR", struct.pack(">IIBBBBB", n, n, 8, 2, 0, 0, 0)) + + chunk(b"IDAT", zlib.compress(raw)) + chunk(b"IEND", b"")) + open(path, "wb").write(data) + return len(data) + + +def css_links(): + rc, code, page = w.app_curl(SUB, "/en/user/login") + links = re.findall(r'(/static/[^"]+\.css)', page or "")[:3] + return {l: w.app_curl(SUB, l)[1] for l in links} + + +fx = fixtures.FIXTURES[APP] +w.login() +if w.stack(APP).get("deployed"): + say("wger already installed — removing it first (scratch box)") + w.remove(APP) +git("pull", "-q", "--rebase", "origin", "main") +fy = f"{D}/templates/{APP}/.felhom.yml" +s = open(fy).read() +if "lifecycle: hidden" in s: + open(fy, "w").write(s.replace("lifecycle: hidden", "lifecycle: available", 1)) + git("commit", "-q", "-am", "DRILL wger: un-hidden for the R-762 box step (DRILL only)") + git("push", "-q", "origin", "main") +say("drill:", git("log", "--oneline", "-1").strip()) +w.sync_rescan() +time.sleep(5) +w.sync_rescan() +if not w.deploy(APP, SUB): + sys.exit(say("RESULT the install did not complete") or 1) +before = (w.stack(APP).get("app_config") or {}).get("pinned_images") +say(f"[1] installed, pinned={before}") +tok = fx.seed(w, SUB, say) +if tok is None or not fx.verify(w, SUB, tok, say): + sys.exit(say(f"RESULT C1 failed: {getattr(fx, 'tried', '')}") or 1) + +jar = tempfile.mktemp(prefix="wger-jar-") +hdr, why = fx._login(w, SUB, tok["pw"], jar) +pic = tempfile.mktemp(prefix="wger-photo-", suffix=".png") +size = png(pic) +rc, code, out = w.app_curl(SUB, "/api/v2/gallery/", *hdr, "-F", f"image=@{pic};type=image/png", "-F", "date=2026-10-06", + "-F", "description=r762", method="POST") +say(f"[2] POST /api/v2/gallery/ (a {size} B PNG) -> {code}") +img = None +try: + img = json.loads(out).get("image") +except Exception: + say(f" body {out[:200]}") +path = re.sub(r"^https?://[^/]+", "", img or "") +REC["photo_path"], REC["photo_bytes"] = path, size +code_b = w.app_curl(SUB, path)[1] if path else None +REC["photo_before"] = code_b +REC["css_before"] = css_links() +say(f"[2] the photo {path} read BEFORE the step -> {code_b}; CSS before {REC['css_before']}") + +git("pull", "-q", "--rebase", "origin", "main") +newc = open(os.path.join(NEWDEF, "docker-compose.yml")).read() +open(f"{D}/templates/{APP}/docker-compose.yml", "w").write(newc) +to = dict(re.findall(r"^ ([a-z0-9-]+):\s*\n(?:(?! [a-z0-9-]+:\s*\n).*\n)*?\s+image:\s*(\S+)", newc, re.M)) +s = open(fy).read() +entry = {"from": before, "to": to, "verdict": "proven", "tested_at": "DRILL", "harness_version": 5, + "evidence": "DRILL (box proof in progress)", "marks": {"files_may_change": False, "needs_person": None, "memory_tight": False}} +s = s.rstrip("\n") + "\n - " + json.dumps(entry) + "\n" +s = re.sub(r'^ mem_limit: .*$', ' mem_limit: "416M"', s, count=1, flags=re.M) +open(fy, "w").write(s) +git("commit", "-q", "-am", f"DRILL wger: the R-762 definition (wger-files) + its ladder entry (box proof)") +git("push", "-q", "origin", "main") +say("drill:", git("log", "--oneline", "-1").strip(), "| to:", to) +w.sync_rescan(APP, to.get("wger-files")) +REC["badge_before"] = w.badges(APP) +since = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip() +res = w.press_update(APP, poll=1, cap_s=1800) +for p in res.get("phases", []): + log.write(f" phase +{p['t']}s {p['phase']} | err={p['error']}\n") +time.sleep(15) +read = fx.verify(w, SUB, tok, say) +code_a = w.app_curl(SUB, path)[1] if path else None +rc, _, body = w.app_curl(SUB, path, "-o", "/dev/null", "-w", "%{size_download}") if path else (1, None, "") +REC["photo_after"], REC["photo_after_bytes"] = code_a, body.strip()[-12:] +REC["css_after"] = css_links() +REC["memory"] = w.guest("docker stats --no-stream --format '{{.Name}} {{.MemUsage}}' | grep -i wger").strip().splitlines() +lines = w.guest(f"docker logs --since {since} felhom-controller 2>&1 | grep -E 'update {APP}' | grep -v DEBUG | cut -c1-400") +log.write(lines + "\n") +st = w.stack(APP); after = (st.get("app_config") or {}).get("pinned_images") +REC["badge_after"] = w.badges(APP) +say(f"[5] after: phase={res.get('final_phase')} pinned={after} seed={read} photo={code_a} ({REC['photo_after_bytes']}) " + f"css={REC['css_after']} mem={REC['memory']}") +ok = (res.get("final_phase") == "done" and read and after == to and code_a == "200" + and REC["css_after"] and all(c == "200" for c in REC["css_after"].values())) +verdict = {"app": APP, "venue": "box 9202 (drill catalog), the product's guarded Update (R-762 definition step)", + "from": before, "to": after, "verdict": "proven" if ok else "failed", + "seed_read_before": True, "seed_read_after": read, "healthy_after": st.get("state") == "running", + "duration_s": res.get("duration_s"), "final_phase": res.get("final_phase"), "measured_at": since, + "evidence": "felhom.eu/documentation/audits/design-build-2026-10-06/F/box/step.txt", "r762": REC} +json.dump(verdict, open(f"{EVD}/box-verdict-{APP}.json", "w"), indent=2) +say(f"RESULT {verdict['verdict']}") +for f in (jar, pic): + if os.path.exists(f): + os.unlink(f) +if not os.environ.get("KEEP"): + say(f"remove -> {w.remove(APP)}") diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index aaff0a5e..8e9c4190 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -129,7 +129,7 @@ stopping line that lies. | **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-612** | Apps & catalog | P3 | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** **— FIXED 2026-09-23 night (catalog `a5a729a`): 512M.** Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (`The seed command has been executed`), and wishlist then moved v0.66.0 → v0.67.1 with its test record. **Still open:** a failed first-boot seed is invisible to the box — nothing reads the seed's exit. | **READY — P3, narrowed to making a failed seed visible; owner: CC (catalog)** **Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | -| **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** | — | — | CC | +| **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. **-- 2026-10-06 (afternoon), measured again on 9202 (live template, wger 2.7):** `/static/css/workout-manager.css` 404 straight at the app, `/home/wger/static` 4 KB, settings `DEBUG False`; the image has gunicorn but no whitenoise and runs Django's `runserver` (no `WGER_USE_GUNICORN`, R-755). So `DJANGO_DEBUG=False` alone would collect the files and still serve none: the fix needs a server for `/static` + `/media` (a second container) — medium, not taken. `audits/r890-instructions-2026-10-06/C/wger.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** **2026-10-06: NARROWED — the files are served.** Catalog `cf1ed43`: `DJANGO_DEBUG=False` + a `wger-files` nginx serving `/static` and `/media`, a definition step proven on the bench (harness v5) and on 9202 (photo 404 → 200, CSS 404 → 200). **What remains is R-755's merged half:** wger still runs Django's `runserver`; upstream's gunicorn runs 3 workers, which do not fit wger's 384 MB — a memory decision and a new proof. Known cost of the fix: the collected static files are 283 MB in a named volume, in every backup of wger. Ready to show? Its files and sign-up are fixed and proven; the server question is open — `audits/design-build-2026-10-06/`F/. | — | Decide the gunicorn worker count against the memory limit, prove it on both venues; then the operator decides whether wger is shown. If nothing: wger stays hidden | CC | | **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | | **R-577** | Apps & catalog | P4 | **[P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them.** FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (`launcher_shared`, `launcher_share_password`) deliberately do NOT, and `TestGuestSharePagesHaveNoGlobe` pins that so it stays a decision rather than an oversight. **Why it is the operator's and not CC's:** a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The `felhom_lang` cookie already built would fit them exactly (display-only, their own browser, never the household's setting). **Fix shape, if the operator says yes:** add `{{template "lang_globe" .}}` to both shells with the anonymous form, and one render case per page per language. | **READY - rank P3-LOW; owner: operator (the decision), CC (the change)** **Re-ranked 2026-10-03: P3->P4: feature decision for the operator; Hungarian default works today.** | — | — | operator | | **R-707** | Apps & catalog | P4 | **[P2] 37 apps still start with a login a stranger can take (`09` §3 decision 45).** Audit of all 53 apps: `app-catalog-felhom.eu/FIRST-ADMIN.md` (class, fix route, status, measured or read). Open: **3 hard-coded defaults** — calibre-web (`admin / admin123`, measured working on demo-hp and 9202; its own `cps.py -s` route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (`changeme@example.com / MyPassword`), wger (`admin / adminadmin`); **34 open first-run screens** (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). **Stale notes:** romm's `default_creds` `admin / admin` answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. **Measured on demo-hp 2026-09-28 (read-only):** bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via `after_install:`, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). **2026-09-29 (controller v0.280.0, catalog `d0e7e2e`):** every class-3 app fixed — mealie, wger, calibre-web by `after_install` (calibre-web with the new `password:24:special`), proven on 9202 fresh installs (`audits/login-gate-2026-09-29/D/`); the setup gate (decision 46, spike PASSED) built and live on immich, n8n, audiobookshelf (probes measured) and uptime-kuma (button) (`…/C/`); romm's and zipline's stale notes removed. **Left: 30 class-4 apps** — gate each (probe measured on 9202 where one exists — 11 upstream candidates listed in `…/B/B-VERDICT.md` §3; the button otherwise). **2026-09-29 afternoon (controller v0.281.0, catalog `6faf432`):** 28 more class-4 apps gated — 32 of 34 — each proven on 9202 (`audits/gate-rollout-2026-09-29/`B): stranger → gate page / 401, household reached the first-setup screen, the gate opened (9 by a measured probe, the rest by the press), the app answered after. seerr, outline, rallly: gated, their opening needs a media server / e-mail (not proven). **Left:** wanderer (R-714); plant-it is not installable. | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — seerr, outline and rallly are gated, but the gate OPENING is not proven (needs a media server / e-mail)) — **CLOSED — 2026-09-29 (the rest → R-714)** | — | — | CC | @@ -156,8 +156,9 @@ stopping line that lies. | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | -| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. | — | — | CC | +| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. | — | Read back the first night runs on demo-hp and demo-felhom under v0.301.0 (the agent's `snapshotted` line, the apps' StartedAt, the controller's window lines); close with that evidence. If nothing: the row stays open | CC | | **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-638.md`** — for the operator. | — | — | CC | +| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | | **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** **2026-10-06: leg (a) PUSHED** to the live catalog (`c265b37`). **2026-10-06 (burn-down night, later): leg (b) NEEDS A DESIGN.** `07` §7.4 sets no direction; refusing the restore without the DB password, or `ALTER USER` after it, each change restore behaviour on customer data. | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | @@ -208,7 +209,7 @@ stopping line that lies. | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`28a5203`; the catalog clone stores no credentials; the token is supplied per fetch; R-615's repo comparison ignores credentials on both sides, so a token never re-clones (`TestR616_TokenSetSameRepoNoRecloneOriginClean`) and a credentialed origin is cleaned at the next pull. The operator's Gitea admin token rotation (the row's second half) is still owed). Ships with the next controller release; close after delivery. **2026-10-06: DELIVERED** in controller v0.298.0 (the clone stores no credentials). Left: the operator's Gitea admin token rotation. | — | — | CC + operator | -| **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). **-- 2026-10-06 (afternoon): NEEDS A DESIGN.** The controller's `after_setup` command form runs only when the lock is SET (`internal/stacks/after_setup.go` `applyNativeLock`, `lock && len(spec.Command) > 0`); the household's 15-minute window lifts only the env form. A database switch closed by a command would stay closed through the window, so the household could not let a family member sign up. Needs a lift command (an `open` twin) in the controller first. | **OPEN — P3; owner: CC** | — | — | CC | +| **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). **-- 2026-10-06 (afternoon): NEEDS A DESIGN.** The controller's `after_setup` command form runs only when the lock is SET (`internal/stacks/after_setup.go` `applyNativeLock`, `lock && len(spec.Command) > 0`); the household's 15-minute window lifts only the env form. A database switch closed by a command would stay closed through the window, so the household could not let a family member sign up. Needs a lift command (an `open` twin) in the controller first. | **OPEN — P3; owner: CC** **2026-10-06: NARROWED to opengist.** Controller v0.301.0 gives `after_setup` an `open_command`; **wishlist uses it, proven live on 9202** (after the setup the switch read closed and a stranger got 401 "invite only"; inside the window open and a family member signed up, users 1 → 2; at the window's end the close ran again 10 s later, a stranger 401, users stayed 2) — `audits/design-build-2026-10-06/`E/. **Opengist cannot:** its container has no sqlite tool and no script runtime, and its CLI has no settings command; it keeps the address block alone. | — | Opengist only: needs a way to write its setting (a helper container in the template, or an upstream CLI/env switch) — a design. If nothing: opengist stays closed by the address block alone | CC | | **R-775** | Security & access | P3 | **[P2-MEDIUM] Grimmory: a stranger's 5 wrong sign-ins lock EVERY visitor out of the web login for 15 minutes — so Grimmory was not published.** MEASURED 2026-10-01 on 9202 (drill catalog, v3.4.1, through traefik): after 5 wrong tries for `admin` every further sign-in answered 429 — the household's right password AND a different name — and stayed 429 for 10+ minutes of retries. Read in the jar: `AuthRateLimitService` — Caffeine `expireAfterWrite(ofMinutes(15))`, `MAX_ATTEMPTS 5`, keys `login:ip:` and `login:user:`; Spring `forward-headers-strategy: native` takes the address from X-Forwarded-For, and behind the tunnel every visitor is the tunnel container's address (R-753) — the wger shape (R-752), with no setting to change it. Everything else in the checklist passed (bench + box step v3.4.1 → v3.5.0, gate by its own probe, OPDS through traefik); two smaller findings for the publishing session: on a reinstall over the first install's kept books, a new upload was saved to the drive but not added to the library (`box/grimmory/reinstall-c1.txt`, not investigated); and the remove + restore round trip (2.5) cannot be shown on 9202 for a drive app — its backup lives on the scratch drive, which is not a registered drive (R-756). The template waits in `audits/new-apps-2026-10-01/wip/grimmory/`. **Needs (operator):** (A) publish with a sentence on the page that wrong guesses by others can lock the login for 15 minutes (MEASURED: during the lock an e-reader's OPDS feed still answered 200 with its own login, wrong 401 — `box/grimmory/opds-under-lock.txt`), or (B) wait until the box passes each visitor's real address (R-753). `audits/new-apps-2026-10-01/box/grimmory/throttle.txt` **-- 2026-10-01 (evening):** option B's precondition SHIPPED (controller v0.286.1, R-753): Grimmory's Tomcat RemoteIpValve walks from the right and counts `172.16.0.0/12` as a proxy (READ in source, Spring Boot 4.1.1 — not yet measured with Grimmory's own lock), so `login:ip:` becomes per visitor; `login:user:` still lets a stranger lock the public name `admin` 15 min. A third route was spiked and passed: Grimmory behind the permanent family gate with its e-reader paths excepted (R-780). Recommendation: publish behind the family gate if R-780 is built; otherwise B with a measured 3.6. **UPDATE 2026-10-02 — NARROWED, Grimmory PUBLISHED behind the family gate:** a stranger cannot reach Grimmory's web sign-in at all (6 tries through the simulated tunnel: the gate's 401, then the household signs in 200 — `audits/family-gate-2026-10-02/A/items.txt`), and the e-reader exceptions keep Grimmory's own login. The 2.5 round trip now WORKS on 9202 (`audits/family-gate-2026-10-02/B/box/life.txt`). **What is left:** a family member past the gate can still lock a NAME (the admin's) for 15 minutes with 5 wrong tries — hard-coded in Grimmory; and the reinstall-over-kept-books finding (`new-apps-2026-10-01/box/grimmory/reinstall-c1.txt`) is still not investigated. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P2→P3: Grimmory is now behind the family gate; only a family member can still lock a name.** | — | — | CC | | **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-831** | Security & access | P3 | **The Hetzner storage API token (`HETZNER_TOKEN`, the storage project's token in `Secret/storagebox`) was printed into the 2026-10-03 session transcript** — CC read the gitignored `manifests/storagebox.secret.yaml` and its redaction pattern missed the quoted value. It can create, reset and delete Storage Box sub-accounts. Not rotated by the operator's choice (decision 73). **Rotation, whenever chosen (3 steps):** create a new token in the storage project in the Hetzner console → patch `Secret/storagebox` key `HETZNER_TOKEN` in `felhom-system` and `kubectl rollout restart deployment/hub` → delete the old token in the console. Rule for sessions: never print a file that holds secrets — read the one field needed. | **WAITING-ON-OPERATOR — rotation is his call** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | @@ -233,7 +234,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-243** | Monitoring & notifications | P2 | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | — | — | operator | -| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. | — | — | CC | +| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. **2026-10-06: STOPPED BEFORE ANY CODE — the measurement contradicted the design** (decision 155 said A then C). On scratch 9202, Docker 29.8.2, the `OOMKilled` flag was TRUE in all four shapes tried: a child process killed while the container kept running (`oom_kill` 0 → 3, flag true) and three main-process kills (exit 137, flag true); an `oom` event each time. The false flags of 2026-09-15 were on Docker 29.8.0. Cost read for the record: one exec read of 21 containers on demo-hp = 1.6 s. `audits/design-build-2026-10-06/`C/. | — | Operator: (A) drop A + C — today's engine reports the flag; keep the row as a watch for a box that reports a false flag; (B) build A + C anyway as insurance (≈ 1.6 s per 5-min scan on demo-hp). If nothing: nothing is built; the alarm keeps reading the flag | operator | | **R-79** | Monitoring & notifications | P3 | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the issues and warnings travel to the hub as sentences; changing them is the two-repo spike the row itself names. | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC | | **R-211** | Monitoring & notifications | P3 | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC | | **R-333** | Monitoring & notifications | P3 | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | @@ -277,12 +278,13 @@ stopping line that lies. | **R-89** | Business & legal | P4 | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 24 rows (P3 2, P4 22) +## Process & tooling — 25 rows (P3 2, P4 23) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC | | **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | +| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config`, no entry in `operations/nodes.md` and no Proxmox host known to this workspace — releases reach it only through the hub's signed jobs and floors. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** | the Tester 1 box's host and an SSH route | Operator: name where the Tester 1 box runs and whether CC may reach it by SSH; then `box_walk.py` gains a target table (host, guest, base URL) and `BOX_ADMIN_SEED_GUESTS` its row. If nothing: the walk stays on scratch 9202 | CC | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | diff --git a/documentation/tests/golden-0.301.0-2026-10-06/02-round-trip.txt b/documentation/tests/golden-0.301.0-2026-10-06/02-round-trip.txt new file mode 100644 index 00000000..520dbfd9 --- /dev/null +++ b/documentation/tests/golden-0.301.0-2026-10-06/02-round-trip.txt @@ -0,0 +1,2 @@ +registry file sha256: 96e94fed70ca3066486acf66902d18994be076e0a877d1e0ca78f4aa06c355b9 +bake GOLDEN_SHA256=96e94fed70ca3066486acf66902d18994be076e0a877d1e0ca78f4aa06c355b9 diff --git a/documentation/tests/golden-0.301.0-2026-10-06/03-teardown.txt b/documentation/tests/golden-0.301.0-2026-10-06/03-teardown.txt new file mode 100644 index 00000000..a7689ab2 --- /dev/null +++ b/documentation/tests/golden-0.301.0-2026-10-06/03-teardown.txt @@ -0,0 +1,4 @@ +purging CT 9100 from related configurations.. +0 +3 +VM off (no qemu process); disk on virgin diff --git a/documentation/tests/golden-0.301.0-2026-10-06/README.md b/documentation/tests/golden-0.301.0-2026-10-06/README.md new file mode 100644 index 00000000..bb59e9d1 --- /dev/null +++ b/documentation/tests/golden-0.301.0-2026-10-06/README.md @@ -0,0 +1,37 @@ +# Golden 0.301.0 — bake + publish, 2026-10-06 (evening; R-518 + R-638 + R-717 release) + +Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–4, in the drill VM on DooPlex. + +| | Previous (`../golden-0.300.0-2026-10-06/`) | This bake | +|---|---|---| +| `build-golden.sh` | v3.2.0 | same file, unchanged (agent repo `configs/`) | +| Controller | `felhom-controller:0.300.0` | **`felhom-controller:0.301.0`** (MinAgent 0.131.0, unchanged) | +| Docker engine | approved set (the six pairs of the 0.300.0 `PINNED` line) | same six pairs, `GOLDEN_DOCKER_PKGS` set in the runner from the start | +| Template | `debian-13-standard_13.6-1_amd64.tar.zst` | same (after `pveam update`) | + +## Pass markers (from `bake.log`, this folder) + +``` +[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie … docker-ce=5:29.8.2-1~debian.13~trixie … + docker OK (overlay2; data-root /var/lib/docker) +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.301.0 +GOLDEN_SHA256=96e94fed70ca3066486acf66902d18994be076e0a877d1e0ca78f4aa06c355b9 +``` + +No `excluding`, no `FATAL`, no `GOLDEN_DOCKER_PKGS not set`. Round trip: the registry's file hashed = the bake's sha +(`02-round-trip.txt`). Token: the unit's properties grep = 0; the committed log grep = 0, with a working control (= 1 on +a copy with the token appended, the copy shredded). + +## Vouch and floors + +Hub, 2026-10-06 14:04Z: agent 0.149.0, golden 0.301.0, min_agent 0.131.0 (bundle sha kept); floors 0.301.0 with +MinAgent 0.131.0 for demo-hp, demo-felhom, tester-1 only. Global floor not touched; Tester 2 not touched +(`../../audits/design-build-2026-10-06/delivery/vouch-golden-floors.txt`). + +## Teardown + +Build guest 9100 destroyed (`pct list` empty); token, runner and log shredded in the VM; VM off (no qemu process); disk +on `virgin` (`03-teardown.txt`). diff --git a/documentation/tests/golden-0.301.0-2026-10-06/bake.log b/documentation/tests/golden-0.301.0-2026-10-06/bake.log new file mode 100644 index 00000000..3d4e56f1 --- /dev/null +++ b/documentation/tests/golden-0.301.0-2026-10-06/bake.log @@ -0,0 +1,340 @@ +[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.301.0 +[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) … + Logical volume "vm-9100-disk-0" created. + Logical volume pve/vm-9100-disk-0 changed. +Creating filesystem with 8388608 4k blocks and 2097152 inodes +Filesystem UUID: 4df958ac-fa81-4c95-b290-8060e35dfc3d +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, + 4096000, 7962624 + Logical volume "vm-9100-disk-1" created. + Logical volume pve/vm-9100-disk-1 changed. +Creating filesystem with 6291456 4k blocks and 1572864 inodes +Filesystem UUID: adf92f46-9c87-464d-8044-b1f50dc01894 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, +extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' +Total bytes read: 553512960 (528MiB, 89MiB/s) +Detected container architecture: amd64 +Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ... +done: SHA256:WCKhTWOVatFQP7VDdlJs8G8r+AyysGzGK8YGQYcBV5Q root@felhom-golden +Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ... +done: SHA256:8/fMKUzxOSnr2yNvqIrFypYSelIaXpqvdajfuZdAjtY root@felhom-golden +Creating SSH host key 'ssh_host_rsa_key' - this may take some time ... +done: SHA256:6aSXDFpyACIEGYqeRJEED/rjJafSuQQtpYs7rbaxwMk root@felhom-golden +[golden] starting + installing Docker (official repo, trixie channel) … +[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory + installed: containerd.io 2.3.6-1~debian.13~trixie + installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie + installed: docker-ce 5:29.8.2-1~debian.13~trixie + installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie + installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie + installed: docker-compose-plugin 5.6.0-1~debian.13~trixie +[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a +[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49 +[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation … +[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds … +[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) … +Unable to find image 'hello-world:latest' locally +latest: Pulling from library/hello-world +4f55086f7dd0: Pulling fs layer +4f55086f7dd0: Verifying Checksum +4f55086f7dd0: Download complete +4f55086f7dd0: Pull complete +Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8 +Status: Downloaded newer image for hello-world:latest + docker OK (overlay2; data-root /var/lib/docker) + live-restore: on + /var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4 + /mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4 + both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576 +[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.301.0 (no registry cred at deploy) … + +WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'. +Configure a credential helper to remove this warning. See +https://docs.docker.com/go/credential-store/ + +0.301.0: Pulling from admin/felhom-controller +774043ccc8cc: Pulling fs layer +ab6b448d4be9: Pulling fs layer +23a5bfa58353: Pulling fs layer +862a57157567: Pulling fs layer +da380b34a313: Pulling fs layer +50372d43d046: Pulling fs layer +862a57157567: Waiting +da380b34a313: Waiting +50372d43d046: Waiting +23a5bfa58353: Verifying Checksum +23a5bfa58353: Download complete +774043ccc8cc: Verifying Checksum +774043ccc8cc: Download complete +862a57157567: Verifying Checksum +862a57157567: Download complete +50372d43d046: Verifying Checksum +50372d43d046: Download complete +ab6b448d4be9: Verifying Checksum +ab6b448d4be9: Download complete +da380b34a313: Verifying Checksum +da380b34a313: Download complete +774043ccc8cc: Pull complete +ab6b448d4be9: Pull complete +23a5bfa58353: Pull complete +862a57157567: Pull complete +da380b34a313: Pull complete +50372d43d046: Pull complete +Digest: sha256:0e80f6af59490dafdb0adf49d314f4aa5253882739a4c6bc09742a564655fb0f +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.301.0 +gitea.dooplex.hu/admin/felhom-controller:0.301.0 +[golden] asking the controller which infra images it manages … +[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 … +v3.7.13: Pulling from library/traefik +e2de96513ba9: Pulling fs layer +b686a4f73445: Pulling fs layer +78cb21c375ca: Pulling fs layer +acb2f33459b1: Pulling fs layer +acb2f33459b1: Waiting +e2de96513ba9: Verifying Checksum +e2de96513ba9: Download complete +b686a4f73445: Verifying Checksum +b686a4f73445: Download complete +acb2f33459b1: Verifying Checksum +acb2f33459b1: Download complete +e2de96513ba9: Pull complete +78cb21c375ca: Verifying Checksum +78cb21c375ca: Download complete +b686a4f73445: Pull complete +78cb21c375ca: Pull complete +acb2f33459b1: Pull complete +Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0 +Status: Downloaded newer image for traefik:v3.7.13 +docker.io/library/traefik:v3.7.13 +2026.9.3: Pulling from cloudflare/cloudflared +2cc7ee286bf3: Pulling fs layer +c172f21841df: Pulling fs layer +218cf840d0d9: Pulling fs layer +f6069939f718: Pulling fs layer +d6b1b89eccac: Pulling fs layer +2780920e5dbf: Pulling fs layer +7c12895b777b: Pulling fs layer +3214acf345c0: Pulling fs layer +52630fc75a18: Pulling fs layer +dd64bf2dd177: Pulling fs layer +b839dfae01f6: Pulling fs layer +ebddc55facdc: Pulling fs layer +c4bc6f35ff5e: Pulling fs layer +b96fe2995f90: Pulling fs layer +58c0c263dc73: Pulling fs layer +bd8962e29291: Pulling fs layer +cac2ae0193cb: Pulling fs layer +f0383d5ebc47: Pulling fs layer +7c12895b777b: Waiting +3214acf345c0: Waiting +52630fc75a18: Waiting +dd64bf2dd177: Waiting +b839dfae01f6: Waiting +ebddc55facdc: Waiting +c4bc6f35ff5e: Waiting +b96fe2995f90: Waiting +58c0c263dc73: Waiting +bd8962e29291: Waiting +cac2ae0193cb: Waiting +f0383d5ebc47: Waiting +f6069939f718: Waiting +d6b1b89eccac: Waiting +2780920e5dbf: Waiting +2cc7ee286bf3: Download complete +c172f21841df: Download complete +218cf840d0d9: Download complete +f6069939f718: Verifying Checksum +f6069939f718: Download complete +d6b1b89eccac: Verifying Checksum +d6b1b89eccac: Download complete +2780920e5dbf: Verifying Checksum +2780920e5dbf: Download complete +7c12895b777b: Verifying Checksum +7c12895b777b: Download complete +3214acf345c0: Verifying Checksum +3214acf345c0: Download complete +52630fc75a18: Verifying Checksum +52630fc75a18: Download complete +dd64bf2dd177: Verifying Checksum +dd64bf2dd177: Download complete +2cc7ee286bf3: Pull complete +b839dfae01f6: Download complete +ebddc55facdc: Verifying Checksum +ebddc55facdc: Download complete +c4bc6f35ff5e: Verifying Checksum +c4bc6f35ff5e: Download complete +c172f21841df: Pull complete +b96fe2995f90: Verifying Checksum +b96fe2995f90: Download complete +58c0c263dc73: Verifying Checksum +58c0c263dc73: Download complete +bd8962e29291: Verifying Checksum +bd8962e29291: Download complete +cac2ae0193cb: Verifying Checksum +cac2ae0193cb: Download complete +f0383d5ebc47: Verifying Checksum +f0383d5ebc47: Download complete +218cf840d0d9: Pull complete +f6069939f718: Pull complete +d6b1b89eccac: Pull complete +2780920e5dbf: Pull complete +7c12895b777b: Pull complete +3214acf345c0: Pull complete +52630fc75a18: Pull complete +dd64bf2dd177: Pull complete +b839dfae01f6: Pull complete +ebddc55facdc: Pull complete +c4bc6f35ff5e: Pull complete +b96fe2995f90: Pull complete +58c0c263dc73: Pull complete +bd8962e29291: Pull complete +cac2ae0193cb: Pull complete +f0383d5ebc47: Pull complete +Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c +Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3 +docker.io/cloudflare/cloudflared:2026.9.3 +1.5.6-stable: Pulling from gtstef/filebrowser +55afa1ecc21d: Pulling fs layer +8ed8f35f8d4f: Pulling fs layer +989b226a579c: Pulling fs layer +660aeead31d5: Pulling fs layer +4f4fb700ef54: Pulling fs layer +adce24567e4c: Pulling fs layer +f17ea56b313b: Pulling fs layer +6b6f3b3efe88: Pulling fs layer +4ed1ca4f3fce: Pulling fs layer +e6fc9c6a5757: Pulling fs layer +d47782d1182a: Pulling fs layer +f17ea56b313b: Waiting +6b6f3b3efe88: Waiting +4ed1ca4f3fce: Waiting +e6fc9c6a5757: Waiting +d47782d1182a: Waiting +660aeead31d5: Waiting +4f4fb700ef54: Waiting +adce24567e4c: Waiting +55afa1ecc21d: Verifying Checksum +55afa1ecc21d: Download complete +660aeead31d5: Verifying Checksum +660aeead31d5: Download complete +4f4fb700ef54: Verifying Checksum +4f4fb700ef54: Download complete +8ed8f35f8d4f: Verifying Checksum +8ed8f35f8d4f: Download complete +989b226a579c: Verifying Checksum +989b226a579c: Download complete +f17ea56b313b: Verifying Checksum +f17ea56b313b: Download complete +55afa1ecc21d: Pull complete +6b6f3b3efe88: Verifying Checksum +6b6f3b3efe88: Download complete +4ed1ca4f3fce: Verifying Checksum +4ed1ca4f3fce: Download complete +adce24567e4c: Verifying Checksum +adce24567e4c: Download complete +d47782d1182a: Verifying Checksum +d47782d1182a: Download complete +e6fc9c6a5757: Verifying Checksum +e6fc9c6a5757: Download complete +8ed8f35f8d4f: Pull complete +989b226a579c: Pull complete +660aeead31d5: Pull complete +4f4fb700ef54: Pull complete +adce24567e4c: Pull complete +f17ea56b313b: Pull complete +6b6f3b3efe88: Pull complete +4ed1ca4f3fce: Pull complete +e6fc9c6a5757: Pull complete +d47782d1182a: Pull complete +Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8 +Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable +docker.io/gtstef/filebrowser:1.5.6-stable +1.1.0: Pulling from admin/felhom-samba +897d797d2723: Pulling fs layer +3051591aa250: Pulling fs layer +ce57a3f93416: Pulling fs layer +fb94eeec2fe1: Pulling fs layer +fb94eeec2fe1: Waiting +ce57a3f93416: Verifying Checksum +ce57a3f93416: Download complete +fb94eeec2fe1: Verifying Checksum +fb94eeec2fe1: Download complete +897d797d2723: Verifying Checksum +897d797d2723: Download complete +897d797d2723: Pull complete +3051591aa250: Verifying Checksum +3051591aa250: Download complete +3051591aa250: Pull complete +ce57a3f93416: Pull complete +fb94eeec2fe1: Pull complete +Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0 +gitea.dooplex.hu/admin/felhom-samba:1.1.0 +[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'. +[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'. +[golden] baking the first-boot SSH host-key regeneration unit (F3) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'. +[golden] identity-clean + minimize … +[golden] stop + archive … +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +INFO: archive file size: 618MB +INFO: Finished Backup of VM 9100 (00:00:29) +[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_06-15_56_42.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive) +[golden] publishing golden (648310469 bytes, sha256 96e94fed70ca3066…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.301.0/golden.tar.zst +[golden] pre-delete existing: HTTP 404 (404/204 expected) +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.301.0 +GOLDEN_SHA256=96e94fed70ca3066486acf66902d18994be076e0a877d1e0ca78f4aa06c355b9 +[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.301.0 / 96e94fed70ca3066486acf66902d18994be076e0a877d1e0ca78f4aa06c355b9 +[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)