From 5a349d9884ddeedfd5b0b7fd858e2dee6076bcb0 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 23 Sep 2026 10:28:56 +0200 Subject: [PATCH] =?UTF-8?q?Undo=20bake-off:=20copy=20the=20folder=20wins?= =?UTF-8?q?=20(09=20=C2=A73=20decisions=2019-20,=20=C2=A76.1a)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the full-system backup waits for the update leg, inside its window; built later). - Bake-off on 9202, docmost / romm / vikunja: both methods pass every case; the folder copy wins because an app with no database server gets no dump, so dump-and-load would need the folder copy anyway. 1-5 s extra downtime, ~420 MB/s, disk = the volumes. - R-645 filed: lifting an update hold by hand lets the recovery unit be re-captured with the failed definition within seconds. Documents and evidence only; product code follows in the controller. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../architecture/09-update-architecture.md | 28 + .../01-repoint-to-drill.txt | 10 + .../02-three-controls.txt | 10 + .../70-rate-nextcloud-deploy.txt | 8 + .../71-rate-nextcloud-fill.txt | 5 + .../72-rate-measure.txt | 13 + .../73-rate-teardown.txt | 16 + .../audits/undo-bakeoff-2026-09-23/README.md | 59 + .../audits/undo-bakeoff-2026-09-23/bakeoff.py | 266 +++++ .../docmost-10-prep.txt | 13 + .../docmost-20-fcopy.txt | 8 + .../docmost-30-break-update.txt | 9 + .../docmost-40-undoF.txt | 15 + .../docmost-41-undoF-redo.txt | 17 + .../docmost-50-update2-undoD.txt | 26 + .../docmost-60-cutoff.txt | 16 + .../undo-bakeoff-2026-09-23/fixtures.py | 1008 +++++++++++++++++ .../audits/undo-bakeoff-2026-09-23/repoint.py | 60 + .../undo-bakeoff-2026-09-23/romm-1-prep.txt | 15 + .../undo-bakeoff-2026-09-23/romm-2-fcopy.txt | 10 + .../undo-bakeoff-2026-09-23/romm-3-break.txt | 2 + .../undo-bakeoff-2026-09-23/romm-4-update.txt | 7 + .../undo-bakeoff-2026-09-23/romm-5-undoF.txt | 15 + .../undo-bakeoff-2026-09-23/romm-6-update.txt | 7 + .../undo-bakeoff-2026-09-23/romm-7-undoD.txt | 17 + .../undo-bakeoff-2026-09-23/romm-8-cutoff.txt | 6 + .../audits/undo-bakeoff-2026-09-23/spike.py | 160 +++ .../vikunja-1-prep.txt | 14 + .../vikunja-2-fcopy.txt | 9 + .../vikunja-3-break.txt | 2 + .../vikunja-4-update.txt | 7 + .../vikunja-5-undoF.txt | 13 + .../vikunja-6-update.txt | 6 + .../vikunja-7-undoD.txt | 8 + .../vikunja-8-cutoff.txt | 11 + .../audits/undo-bakeoff-2026-09-23/walk.py | 493 ++++++++ documentation/backlog/OPEN-ITEMS.md | 1 + 37 files changed, 2390 insertions(+) create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/01-repoint-to-drill.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/02-three-controls.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/70-rate-nextcloud-deploy.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/71-rate-nextcloud-fill.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/72-rate-measure.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/73-rate-teardown.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/README.md create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/bakeoff.py create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-10-prep.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-20-fcopy.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-30-break-update.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-41-undoF-redo.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-50-update2-undoD.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/docmost-60-cutoff.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/fixtures.py create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/repoint.py create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-1-prep.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-2-fcopy.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-3-break.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-4-update.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-5-undoF.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-6-update.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-7-undoD.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/romm-8-cutoff.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/spike.py create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-1-prep.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-2-fcopy.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-3-break.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-4-update.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-5-undoF.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-6-update.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-7-undoD.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/vikunja-8-cutoff.txt create mode 100644 documentation/audits/undo-bakeoff-2026-09-23/walk.py diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index fb7e4995..ae8883ed 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -265,6 +265,21 @@ builds them (`audits/update-rulings-2026-09-23/`). 18. **Fleet view** (Q7). The report carries, per compose service (so the database too), the installed reference, the catalog reference and the badge state. **Built later**, when the fleet grows. +### 2026-09-23 (afternoon) — two more operator rulings + +19. **The undo's copy method is chosen by a bake-off, not by assumption.** Two methods are measured on + the same apps — *dump and load* (the database as a text file, loaded back) and *copy the folder* + (the app's named volumes copied with its containers stopped, put back on failure). The simpler one + that passes every case is built; if both pass, the folder copy wins on simplicity unless its + downtime or disk cost fails the bake-off's limits. **Why:** the morning spike found two traps in the + dump route (a cut-off file loads as success; a failed MariaDB load is half old, half new) and the + folder route had not been measured at all. Result: §6.1a. + +20. **On update nights, the full-system backup waits for the update leg, inside its own window** + (answers R-643 / §6.4 part 7's open point). The leg stops starting new steps at **W+5h**, so the + full-system backup keeps at least one hour of its [W+2h, W+6h) window. **Built with §6.4 part 7, + not before.** + **RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update (`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4); R-636's louder repeated alarm. @@ -798,6 +813,19 @@ product path that loads a safety dump back exists (`rollbackSafetyDump`), but on restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices them. +**THE COPY METHOD, chosen by the bake-off of 2026-09-23 afternoon (decision 19): +COPY THE FOLDER.** `audits/undo-bakeoff-2026-09-23/README.md`. Both methods passed every case on all +three apps (seeds before and after the backup, ledger equal, a cut-off copy caught before anything is +swapped or loaded). The folder copy wins because an app with no database server gets no dump at all, +so the dump route would have needed the folder copy anyway. **Shape:** after the pull and just before +`up` — where the app is stopped anyway to be recreated — each NAMED volume the app owns is copied with +`cp -a` into a sibling volume `.pre-update-` by a helper container, which writes a +finished-marker last; a bind-mounted user folder is never copied and never touched. On a failed health +check the copy is validated (marker present) and put back, the old definition and pin are restored +from the journal's own copies, and the old version is checked with the OLD `.felhom.yml` probe. +Measured cost: ≈ 1–5 s extra downtime for the three apps, ≈ 420 MB/s, disk = the volumes' size (the +update refuses before moving anything when the copy would breach the 2 GB floor). + **The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0 were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached diff --git a/documentation/audits/undo-bakeoff-2026-09-23/01-repoint-to-drill.txt b/documentation/audits/undo-bakeoff-2026-09-23/01-repoint-to-drill.txt new file mode 100644 index 00000000..f2c89786 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/01-repoint-to-drill.txt @@ -0,0 +1,10 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/02-three-controls.txt b/documentation/audits/undo-bakeoff-2026-09-23/02-three-controls.txt new file mode 100644 index 00000000..18271f9e --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/02-three-controls.txt @@ -0,0 +1,10 @@ +sync 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'message': 'Sablonok na +control 1 (9202 follows drill): a837c3a DRILL: FROM states for the undo bake-off + image: docmost/docmost:0.95.0 + image: vikunja/vikunja:2.3.0 + image: rommapp/romm:5.0.0 + +control 2 (9201 on live): cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462) + +control 3 (live main): cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/70-rate-nextcloud-deploy.txt b/documentation/audits/undo-bakeoff-2026-09-23/70-rate-nextcloud-deploy.txt new file mode 100644 index 00000000..b811e032 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/70-rate-nextcloud-deploy.txt @@ -0,0 +1,8 @@ +10:22:59 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud'] +10:22:59 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH'] +10:22:59 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +10:23:44 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.1-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'} +10:23:44 deployed: True +10:23:48 nextcloud: occ user:add :: The account "drilladbeb4" was created successfully Display name set to "drilladbeb4" +10:23:51 nextcloud: seeded user drilladbeb4 +10:23:51 seeded: True diff --git a/documentation/audits/undo-bakeoff-2026-09-23/71-rate-nextcloud-fill.txt b/documentation/audits/undo-bakeoff-2026-09-23/71-rate-nextcloud-fill.txt new file mode 100644 index 00000000..fc8ff399 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/71-rate-nextcloud-fill.txt @@ -0,0 +1,5 @@ +10:26:02 300 WebDAV uploads through the front door: {'201': 300} +10:26:03 PROPFIND lists 300 drill files (readback through the app); http 207 +10:26:06 449 +185374751 /v + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/72-rate-measure.txt b/documentation/audits/undo-bakeoff-2026-09-23/72-rate-measure.txt new file mode 100644 index 00000000..41689b8f --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/72-rate-measure.txt @@ -0,0 +1,13 @@ +10:26:49 == D's cost on the same database: mariadb-dump while running +dump 0.44s 1062572 B +== F on the same database: stop, copy, start +stop 1.58s +cp -a 185 MB: 0.81s +tar 185 MB: 0.53s +start 0.63s +== a 2 GiB synthetic volume (random bytes, not app data) for the extrapolation +page cache could NOT be dropped (LXC) - warm-ish numbers +cp -a 2 GiB: 4.87s +tar 2 GiB: 4.55s +/dev/loop1 69G 48G 18G 74% /var/lib/docker + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/73-rate-teardown.txt b/documentation/audits/undo-bakeoff-2026-09-23/73-rate-teardown.txt new file mode 100644 index 00000000..44ee7094 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/73-rate-teardown.txt @@ -0,0 +1,16 @@ +10:27:13 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'} +10:27:18 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó +10:27:18 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +10:27:46 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], +10:27:53 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud' +10:27:56 removed docmost_docmost_postgres_data.pre-undo +removed docmost_docmost_redis_data.pre-undo +removed docmost_docmost_storage.pre-undo +removed romm_romm_config.pre-undo +removed romm_romm_db_data.pre-undo +removed romm_romm_redis_data.pre-undo +removed vikunja_vikunja_data.pre-undo +removed vikunja_vikunja_db.pre-undo +nextcloud drive folder removed +0 + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/README.md b/documentation/audits/undo-bakeoff-2026-09-23/README.md new file mode 100644 index 00000000..987dd817 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/README.md @@ -0,0 +1,59 @@ +# The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D) + +`09` §3 decision 19. Venue: scratch guest **9202**, controller v0.262.1, drill catalog (reset to live +`cfcfe5278428` first; `a837c3a7d8bd` FROM states), `update.health_timeout: 90s`. Same three apps and +edges as the morning spike (`../update-rulings-2026-09-23/`), each a real migrating edge made to fail +a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it, +both through the app's own front door. Driver: `bakeoff.py` (stages per app: prep → fcopy → break → +update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min. + +**Method, stated so it is not over-read.** Both undos were done BY HAND in the order the product +would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that +exists) and the product's own Start supplied the app's secrets. + +## The table + +| criterion | limit | **F** docmost / romm / vikunja | **D** docmost / romm / vikunja | +|---|---|---|---| +| seed A and seed B read back after the undo | both, every app | **yes·yes / yes·yes / yes·yes** | yes·yes / yes·yes / — (no safety dump exists: R-641) | +| tables / ledger equal to before the update | equal | **42 · 48 = / alembic 0095 = ¹ / 36 · 117 =** | 42 · 48 = / **24 base tables** · 0095 = / — | +| a cut-off copy detected before anything is swapped or loaded | detected | **docmost 4 of 4 cuts** (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); **romm** (105.8 of 160.2 MB, no marker); vikunja **could not be cut** — 2.9 MB completes before a kill lands ² | docmost and romm: the half file lacks `-- PostgreSQL database dump complete` / `-- Dump completed` → refused; vikunja — | +| extra downtime before the update | ≤ 30 s | stop + copy: **≈ 2.4 s / ≈ 5.3 s / ≈ 1 s** (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) | 0 — the dump runs against the running database | +| rate | stated | cp -a **185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s**; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. **A 5 GB database ≈ 12 s at this rate**; on a cold cache or a spinning system disk expect 25–50 s | `mariadb-dump` of the same Nextcloud: 0.44 s, 1.06 MB | +| disk needed | stated; refuse near the 2 GB floor | the app's named volumes: **51 MB / 175 MB / 2.5 MB** (Nextcloud 185 MB) | the dump: 136 KB / 63 KB / — (Nextcloud 1 MB) | +| product code the build needs | estimated | **≈ 300 lines**, one mechanism for every app class | ≈ 450: the marker, two engine-specific empty-then-load paths, **and F's own machinery anyway** for apps with no database server | + +¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed +before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source +bytes, file counts equal, `romm-2-fcopy.txt`), and the migration ledger read back equal. +² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut. + +## The choice — F, copy the folder (decision 19) + +**Both methods pass every row.** Decision 19 then says the folder copy wins on simplicity unless its +downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk +cost is refused near the 2 GB floor by construction. **The decisive fact is the third column:** an +app with no database server gets no safety dump, so D would have to carry F's volume copy anyway. +F is one mechanism; D is F plus two engine loaders. + +**Copy tool: `cp -a` into a sibling volume** (`.pre-update-`), not `tar`: the two ran at +the same speed, and a sibling volume puts the undo back with the same `cp -a` and needs no space on +the app's own drive. + +## Three things the bake-off found that the build must respect + +1. **The undo must keep its own copy of the old definition.** Lifting the hold made the box re-capture + docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the + unit started the new version on the restored data, which migrated again (`docmost-40-undoF.txt`, + invalid run, kept as evidence; redo `docmost-41`). R-639, seen live. +2. **Killing `docker run` does not stop the copy** — the container runs on and finishes + (`docmost-60-cutoff.txt`, first block). The product must judge a copy by the helper container's own + exit status and the finished-marker it writes last, never by the client. +3. **Where the copy is taken decides the downtime.** Taken after the pull and just before `up` — where + the update stops the app anyway to recreate it — the extra downtime is the copy alone. + +## Evidence + +`docmost-*`, `romm-*`, `vikunja-*` (per stage), `70`–`73` (the rate test on Nextcloud, removed after), +`01`/`02` (repoint + three controls). Nextcloud and every `.pre-undo`/`.cut`/`.rate` volume were +removed by name at the end of the phase (`73-rate-teardown.txt`). diff --git a/documentation/audits/undo-bakeoff-2026-09-23/bakeoff.py b/documentation/audits/undo-bakeoff-2026-09-23/bakeoff.py new file mode 100644 index 00000000..7bd759f9 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/bakeoff.py @@ -0,0 +1,266 @@ +#!/usr/bin/env python3 +"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on +the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202, +drill catalog. + + F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure. + D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and + recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads. + +EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync, +rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is +lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own +Start supplies the app's secrets (they are never decrypted here). + +Usage: python3 bakeoff.py stages: prep fcopy break update undoF undoD cutoff state +""" +import json, os, sys, time +sys.path.insert(0, ".") +import walk as w +from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \ + romm_verify_b, vik_seed_b, vik_verify_b + +EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999), + "romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999), + "vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)} +UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost", + "vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja", + "romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"} +SUFFIX = ".pre-undo" + + +def seed_b(app, sub, A): + return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A) + + +def verify_b(app, sub, A, B): + r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B) + return all(r.values()) if isinstance(r, dict) else r + + +def db_state(app): + """The table count and the migration ledger, asked of the engine (or the SQLite file, read-only).""" + if app == "docmost": + return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; " + "docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip() + if app == "romm": + q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '" + base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45") + alembic = q("select version_num from alembic_version") + return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip() + return w.guest("""python3 - <<'PY' +import sqlite3 +c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True) +t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0] +m=c.execute("select count(*), max(id) from migration").fetchone() +print(f"tables={t} ledger={m[0]} newest={m[1]}") +PY""").strip() + + +def volumes(app): + return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split() + if not v.endswith(SUFFIX)] + + +def containers(app): + return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split() + + +def stage_prep(app): + s = {"app": app, "sub": SUB[app]} + w.say(f"=== prep {app}: deploy at the drill FROM pin") + w.say("deployed:", w.deploy(app, s["sub"])) + s["seedA"] = FX[app].seed(w, s["sub"], w.say) + w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say)) + w.backup_now(app) + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60) + s["seedB"] = seed_b(app, s["sub"], s["seedA"]) + w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"])) + s["volumes"] = volumes(app) + w.say("named volumes:", s["volumes"]) + save(app, s) + + +def stage_fcopy(app): + """F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and + separately a tar, for the rate), start. Measures the EXTRA downtime and the disk.""" + s = load(app) + s["state_before_update"] = db_state(app) + w.say("db state before the update:", s["state_before_update"]) + vols = s["volumes"] + w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml")) + script = f""" +set -u +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}') +t0=$(date +%s.%N) +docker stop $cs >/dev/null +t1=$(date +%s.%N) +for v in {' '.join(vols)}; do + docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1 + docker volume create "$v{SUFFIX}" >/dev/null + a=$(date +%s.%N) + docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v" + b=$(date +%s.%N) + bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1) + files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l') + cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1) + cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l') + mkdir -p /root/bk; c=$(date +%s.%N) + docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N) + tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar + python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))" +done +t2=$(date +%s.%N) +docker start $cs >/dev/null +t3=$(date +%s.%N) +python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))" +df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}' +""" + out = w.guest(script, timeout=1800) + w.say(out) + t0 = time.time() + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + w.say(f"front door answering again {round(time.time()-t0,1)}s after the start") + s["fcopy"] = out + s["fcopy_at"] = ts() + save(app, s) + + +def stage_break(app): + s = load(app) + frm, to, p0, p1 = EDGE[app] + s["break_commit"] = break_edge(app, frm, to, p0, p1) + w.say("badge caught up after", w.sync_rescan(app, to), "s") + save(app, s) + + +def stage_update(app): + s = load(app) + r = w.press_update(app, poll=1.0) + w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False)) + s.setdefault("updates", []).append(r) + save(app, s) + + +LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold' +docker restart felhom-controller >/dev/null; sleep 15""" + +# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre--compose.yml, taken +# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s +# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new +# version again. That is R-639 seen live, and why the product undo keeps its own copies. +PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml +grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }} +cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml +sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4""" + + +def lift_and_pinback(app): + frm, to, _, _ = EDGE[app] + t = time.time() + w.say(w.guest(LIFT.format(app=app))) + w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to))) + w.login() + return round(time.time() - t, 1) + + +def stage_undoF(app): + s = load(app) + w.say(f"=== undo by F: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + vols = s["volumes"] + script = f""" +t0=$(date +%s.%N) +cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null +for v in {' '.join(vols)}; do + docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v" +done +python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))" +""" + w.say(w.guest(script, timeout=1800)) + t0 = time.time() + c, d = w.ctl("POST", f"/api/stacks/{app}/start") + w.say("product start ->", c) + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_undoD(app): + """D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy.""" + s = load(app) + w.say(f"=== undo by D: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + if app == "vikunja": + w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it " + "is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.") + return + c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse + time.sleep(8) + if app == "docmost": + load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1""" + marker = "-- PostgreSQL database dump complete" + else: + load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1""" + marker = "-- Dump completed" + script = f""" +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B" +docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null +t0=$(date +%s.%N) +if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi +{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out +python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))" +docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null +""" + w.say(w.guest(script, timeout=900)) + t0 = time.time() + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_cutoff(app): + """The cut-off copy, both methods, DETECTED before anything is loaded or swapped. + F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker + the build would write only after cp exits 0. D: the safety dump cut in half — the marker check.""" + s = load(app) + big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0)) + script = f""" +v={big} +docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null +# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured +# on docmost 2026-09-23, rc 137 with a complete copy and a marker). +docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null +sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1 +echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)" +echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap" +docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1) +if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql + echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load" + rm -f /root/cut.sql +else echo "D: no safety dump exists for this app (R-641)"; fi +""" + w.say(f"=== cut-off copy: {app} (largest volume {big})") + out = w.guest(script, timeout=600) + w.say(out) + s["cutoff"] = out + save(app, s) + + +if __name__ == "__main__": + w.login() + {"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update, + "undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff, + "state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-10-prep.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-10-prep.txt new file mode 100644 index 00000000..af669cfe --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-10-prep.txt @@ -0,0 +1,13 @@ +09:53:18 === prep docmost: deploy at the drill FROM pin +09:53:18 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +09:54:58 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +09:54:58 deployed: True +09:54:59 docmost: /api/auth/setup http=200 rc=0 +09:54:59 docmost: login as the seeded user http=200 ok=True +09:54:59 C1 A: True +09:54:59 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +09:55:40 [4] backup idle; last=None +09:55:40 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cd43-623a-7ddf-9b62-213653bc5558","name":"drillB3aab00","description":"","slug":"drillb3aab00","logo":null,"visibility":"private","defaultRol +09:55:40 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +09:55:40 B reads back: True +09:55:43 named volumes: ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'] diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-20-fcopy.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-20-fcopy.txt new file mode 100644 index 00000000..6a9d31c3 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-20-fcopy.txt @@ -0,0 +1,8 @@ +09:55:53 db state before the update: 42 48 ledger, newest 20260620T010047-personal-spaces +09:56:05 VOL docmost_docmost_postgres_data bytes=51156449 files=1561 | copy bytes=51156449 files=1561 | cp -a 0.98s | tar 0.54s (51838464 B) +VOL docmost_docmost_redis_data bytes=39819 files=6 | copy bytes=39819 files=6 | cp -a 0.46s | tar 0.45s (37376 B) +VOL docmost_docmost_storage bytes=4096 files=1 | copy bytes=4096 files=1 | cp -a 0.46s | tar 0.39s (1536 B) +stop 0.52s copy(all, cp -a + the tar measurement) 9.00s start 0.75s +free on the docker root: 21.04 GiB + +09:56:17 front door answering again 12.2s after the start diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-30-break-update.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-30-break-update.txt new file mode 100644 index 00000000..d6a62e34 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-30-break-update.txt @@ -0,0 +1,9 @@ +09:56:25 [5] drill commit 386a34da911a: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0) +09:56:29 badge caught up after 4.5 s +09:56:30 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +09:56:30 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +09:56:31 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None +09:57:44 + 74.9s phase=starting label=Indítás az új verzióval… err=None hold=None +09:57:46 + 76.9s phase=verifying label=Működés ellenőrzése… err=None hold=None +09:59:18 + 168.1s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +09:59:18 {"final_phase": "failed", "hold_reason": "A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 168.1} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt new file mode 100644 index 00000000..54f79bd4 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt @@ -0,0 +1,15 @@ +09:59:23 === undo by F: docmost +09:59:41 [INFO] [settings] restore hold CLEARED for docmost + +09:59:43 pinned_images: + docmost: docmost/docmost:0.95.0 + docmost-postgres: postgres:16-alpine + docmost-redis: redis:7-alpine + +09:59:43 hold lift + pin back took 20.2 s (not part of a product undo) +09:59:48 volumes put back in 2.20s + +09:59:59 product start -> 200 +10:00:12 docmost: login as the seeded user http=200 ok=True +10:00:12 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +10:00:15 F RESULT: healthy=True after 23.6s A=True B=True db now: 48 52 ledger, newest 20260904T171920-public-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-41-undoF-redo.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-41-undoF-redo.txt new file mode 100644 index 00000000..6d9467bf --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-41-undoF-redo.txt @@ -0,0 +1,17 @@ +10:00:57 === REDO of the F undo with the bake-off's OWN pre-update definition (the first try read the re-captured unit — see docmost-40) +10:01:00 image: docmost/docmost:0.95.0 + +10:01:02 pinned_images: + docmost: docmost/docmost:0.95.0 + docmost-postgres: postgres:16-alpine + docmost-redis: redis:7-alpine + +10:01:07 stop + volumes put back in 2.91s + +10:01:18 product start -> 200 +10:01:33 docmost/docmost:0.95.0 Up 15 seconds (healthy) +successfully started + +10:01:34 docmost: login as the seeded user http=200 ok=True +10:01:34 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +10:01:36 F RESULT (redo): healthy=True after 23.6s A=True B=True db now: 42 48 ledger, newest 20260620T010047-personal-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-50-update2-undoD.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-50-update2-undoD.txt new file mode 100644 index 00000000..245d4a92 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-50-update2-undoD.txt @@ -0,0 +1,26 @@ +10:01:43 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +10:01:43 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +10:01:44 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +10:01:45 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None +10:01:46 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +10:03:18 + 95.4s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +10:03:18 {"final_phase": "failed", "hold_reason": "A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4} +10:03:18 === undo by D: docmost +10:03:36 [INFO] [settings] restore hold CLEARED for docmost + +10:03:38 pinned_images: + docmost: docmost/docmost:0.95.0 + docmost-postgres: postgres:16-alpine + docmost-redis: redis:7-alpine + +10:03:38 hold lift + pin back took 20.2 s (not part of a product undo) +10:04:01 undo copy: pre-restore-20260923T080143Z-docmost-postgres.sql 136043 B +marker present +load rc=0 +ALTER TABLE +ALTER TABLE +marker check + load 1.87s + +10:04:13 docmost: login as the seeded user http=200 ok=True +10:04:14 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +10:04:16 D RESULT: healthy=True after 12.2s A=True B=True db now: 42 48 ledger, newest 20260620T010047-personal-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/docmost-60-cutoff.txt b/documentation/audits/undo-bakeoff-2026-09-23/docmost-60-cutoff.txt new file mode 100644 index 00000000..06cb8a7c --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/docmost-60-cutoff.txt @@ -0,0 +1,16 @@ +10:04:32 === cut-off copy: docmost (largest volume docmost_docmost_postgres_data) +10:04:39 copy rc=137 (137 = killed) +source bytes=69334709 cut copy bytes=69334709 +finished-marker in the cut copy: 1 -> F refuses to swap +D: whole copy marker: 1 cut copy marker: 0 -> D refuses to load + +10:04:55 NOTE: the attempt above killed the docker CLIENT (rc 137) while the copy CONTAINER ran on and finished — so nothing was cut and the marker was written. Redo: kill the copy CONTAINER itself mid-copy. +10:05:07 kill after 0.15s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1 +kill after 0.25s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1 +kill after 0.35s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1 + +10:05:31 kill after 0s: container exit=137 source bytes=69172397 cut copy bytes=42921516 finished-marker=0 +kill after 0.02s: container exit=137 source bytes=69172397 cut copy bytes=48586287 finished-marker=0 +kill after 0.05s: container exit=137 source bytes=69172397 cut copy bytes=59159114 finished-marker=0 +kill after 0.08s: container exit=137 source bytes=69172397 cut copy bytes=68445276 finished-marker=0 + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/fixtures.py b/documentation/audits/undo-bakeoff-2026-09-23/fixtures.py new file mode 100644 index 00000000..3a681e96 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/fixtures.py @@ -0,0 +1,1008 @@ +#!/usr/bin/env python3 +"""Box-side seed/verify fixtures for walk.py, guest 9202. + +THE ONE RULE (R-156), carried verbatim from `app-catalog-felhom.eu/scripts/upgrade_fixtures.py`: +*nothing is ever seeded into a volume by hand.* Every seed here goes in through the app's OWN +interface — its HTTP API through the household's real front door (traefik, `Host: .`), +or its own CLI running inside its own container. A raw SQL INSERT or a planted file is never used. + +If an app has no non-browser route, its fixture returns None and the edge is recorded +`inconclusive — no non-browser seed route`, WITH WHAT WAS TRIED. That is a result, not a gap. + +Each fixture: + seed(w, sub, say) -> an opaque token, or None + verify(w, sub, tok, say) -> True / False +verify() must ask the APP, never the filesystem: a migration is supposed to rewrite files. +Where a fixture can prove itself (a negative control that must read as absent) it does so on EVERY +call, so a readback that has broken into always saying "found" fails instead of passing everything. +""" +import base64, json, re, secrets, time + + +def _gx(w, container, *cmd, timeout=240): + """Run a command inside the app's OWN container on 9202 (its own CLI, not our SQL).""" + import shlex + line = " ".join(shlex.quote(c) for c in cmd) + return w.guest(f"docker exec {container} {line} 2>&1", timeout=timeout) + + +# ============================================================================================= +class PrivateBin: + """PrivateBin's own JSON API. A paste is a POST and reading it back is a GET — an + application-level round trip. File-backed, no database: this single seed IS the file half.""" + sub = "paste" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200",)): + return None + marker = "upg-" + secrets.token_hex(8) + ct = base64.b64encode(marker.encode()).decode() + body = json.dumps({ + "v": 2, + "adata": [[base64.b64encode(secrets.token_bytes(16)).decode(), + base64.b64encode(secrets.token_bytes(8)).decode(), + 100000, 256, 128, "aes", "gcm", "none"], "plaintext", 0, 0], + "ct": ct, "meta": {"expire": "never"}}) + rc, code, out = w.app_curl(sub, "/", "-H", "X-Requested-With: JSONHttpRequest", + "-H", "Content-Type: application/json", + data=body, method="POST") + try: + j = json.loads(out) + except Exception: + say(f" privatebin: POST returned non-JSON (http {code}): {out[:200]}") + return None + if j.get("status") != 0 or not j.get("id"): + say(f" privatebin: POST refused: {out[:250]}") + return None + say(f" privatebin: seeded paste id={j['id']}") + return {"id": j["id"], "marker": ct} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200",), tries=36): + return False + # negative control, every call: a paste id that cannot exist must NOT read back + rc, code, out = w.app_curl(sub, "/?pasteid=" + secrets.token_hex(8), + "-H", "X-Requested-With: JSONHttpRequest") + if t["marker"] in out: + say(" privatebin: READBACK UNUSABLE — a paste id that cannot exist returned the marker") + return False + rc, code, out = w.app_curl(sub, "/?pasteid=" + t["id"], + "-H", "X-Requested-With: JSONHttpRequest") + got = code == "200" and t["marker"] in out + say(f" privatebin: readback http={code} marker_present={got}") + return got + + +# ============================================================================================= +class Docmost: + """Docmost's own REST API: create the first workspace+user, then prove the account survives by + asking the app to AUTHENTICATE it. Login is version-stable across the API churn.""" + sub = "docs" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302", "404")): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + body = json.dumps({"workspaceName": "drill", "name": "drill", "email": email, "password": pw}) + rc, code, out = w.app_curl(sub, "/api/auth/setup", "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" docmost: /api/auth/setup http={code} rc={rc}") + if code not in ("200", "201"): + say(f" docmost: setup refused: {out[:250]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "404"), tries=36): + return False + # negative control: a password that was never set must NOT authenticate + bad = json.dumps({"email": t["email"], "password": "definitely-" + secrets.token_hex(8)}) + rc, code, _ = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" docmost: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"email": t["email"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" docmost: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" docmost: login body {out[:200]}") + return ok + + +# ============================================================================================= +class BookStack: + """BookStack mints no API token without a browser, so BOTH halves go through `php artisan` — + BookStack's OWN CLI, inside its own container, against its own User model. + + The exit code carries no information here (`bookstack:reset-mfa` exits 1 for a user it FOUND + and for one it did not), so the discriminator is the OUTPUT: the positive sentence required and + the not-found sentence required absent. The negative control runs on every verify. + + LIMITATION (R-460): this seeds the DATABASE half only. The FILE half needs the API token the + app cannot mint headlessly — so a bookstack edge is at best HALF-proven here. + """ + sub = "wiki" + + def _artisan(self, w, *args): + for path in ("/app/www/artisan", "/var/www/html/artisan"): + out = _gx(w, "bookstack", "php", path, *args) + if "Could not open input file" not in out: + return " ".join(out.split()) + return " ".join(out.split()) + + def _lookup(self, w, email): + out = self._artisan(w, "bookstack:reset-mfa", f"--email={email}") + found = f"Email: {email}" in out + missing = "could not be found" in out + if found == missing: + return None, out + return found, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + out = self._artisan(w, "bookstack:create-admin", f"--email={email}", + f"--name=drill-{secrets.token_hex(3)}", f"--password={pw}") + say(f" bookstack: artisan create-admin :: {out[:140]}") + if "successfully created" not in out: + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" bookstack: the app never served /login") + return False + absent, _ = self._lookup(w, f"nobody-{secrets.token_hex(6)}@gate.invalid") + if absent is not False: + say(f" bookstack: READBACK UNUSABLE — an email that cannot exist did not read absent ({absent})") + return False + found, out = self._lookup(w, t["email"]) + say(f" bookstack: readback of the seeded account found={found} :: {out[:140]}") + return found is True + + +# ============================================================================================= +class Gitea: + """Gitea's own admin CLI creates the first user; its own REST API (basic auth) then creates a + repository and reads it back. Both are the app's own interfaces.""" + sub = "git" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = _gx(w, "gitea", "su", "git", "-c", + f"gitea admin user create --username {user} --password {pw} " + f"--email {user}@gate.invalid --admin --must-change-password=false") + say(f" gitea: admin user create :: {' '.join(out.split())[:140]}") + if "has been successfully created" not in out and "successfully created" not in out: + return None + repo = "drillrepo" + secrets.token_hex(3) + rc, code, body = w.app_curl(sub, "/api/v1/user/repos", "-u", f"{user}:{pw}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": repo, "private": True}), method="POST") + say(f" gitea: create repo http={code}") + if code not in ("201", "200"): + say(f" gitea: repo refused {body[:200]}") + return None + return {"user": user, "pw": pw, "repo": repo} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + rc, code, _ = w.app_curl(sub, f"/api/v1/repos/{t['user']}/nope{secrets.token_hex(4)}", + "-u", f"{t['user']}:{t['pw']}") + if code == "200": + say(" gitea: READBACK UNUSABLE — a repo that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/repos/{t['user']}/{t['repo']}", + "-u", f"{t['user']}:{t['pw']}") + ok = code == "200" and t["repo"] in body + say(f" gitea: readback of the seeded repo http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Navidrome: + """Navidrome's own REST API: create the first admin through /auth/createAdmin, then prove the + account survives by logging in through the same door.""" + sub = "music" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, out = w.app_curl(sub, "/auth/createAdmin", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), method="POST") + say(f" navidrome: createAdmin http={code}") + if code not in ("200", "201"): + say(f" navidrome: refused {out[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" navidrome: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"username": t["user"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" navidrome: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vaultwarden: + """Vaultwarden's own account API: register an account, then prove it survives by asking the app + to issue a token for it (its own login endpoint, the household's own route).""" + sub = "vault" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/alive", want=("200",)): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + # Vaultwarden stores an already-hashed master key; the value is opaque to the server. + key = base64.b64encode(secrets.token_bytes(32)).decode() + body = json.dumps({"email": email, "name": "drill", "masterPasswordHash": key, + "key": "0." + base64.b64encode(secrets.token_bytes(48)).decode(), + "kdf": 0, "kdfIterations": 600000}) + rc, code, out = w.app_curl(sub, "/api/accounts/register", + "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" vaultwarden: register http={code}") + if code not in ("200", "204"): + say(f" vaultwarden: refused {out[:250]}") + return None + return {"email": email, "key": key} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/alive", want=("200",), tries=36): + return False + def login(pwhash): + return w.app_curl(sub, "/identity/connect/token", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=("grant_type=password&scope=api%20offline_access" + f"&client_id=web&deviceType=9&deviceIdentifier=drill" + f"&deviceName=drill&username={t['email']}&password={pwhash}"), + method="POST") + rc, code, _ = login(base64.b64encode(secrets.token_bytes(32)).decode()) + if code == "200": + say(" vaultwarden: READBACK UNUSABLE — a wrong master key authenticated") + return False + rc, code, out = login(t["key"].replace("+", "%2B").replace("=", "%3D").replace("/", "%2F")) + ok = code == "200" and "access_token" in out + say(f" vaultwarden: token for the seeded account http={code} ok={ok}") + if not ok: + say(f" vaultwarden: body {out[:200]}") + return ok + + +# ============================================================================================= +class Django: + """A Django app's OWN management CLI, inside its own container, against its own User model. + + Same category as BookStack's `php artisan`: the app's own code and its own ORM, never a raw SQL + INSERT and never a planted file (R-156). `createsuperuser --noinput` is Django's own documented + non-interactive route, and the readback asks the SAME ORM whether the account exists. + + THE FIXTURE PROVES ITSELF ON EVERY CALL: each verify() also asks for a username that cannot + exist and requires the answer False. A readback that has broken into always saying True + therefore fails instead of passing everything. + + LIMITATION, recorded rather than papered over: this seeds the DATABASE half only. An app whose + data is also FILES (adventurelog's images) has a file half this fixture does not touch. + """ + + def __init__(self, container, sub, ready_path="/", ready=("200", "302", "301", "404"), + python="python", workdir=None): + # `python` and `workdir` are per-app because the image decides them: adventurelog's + # interpreter is on PATH, tandoor ships a VENV and the bare `python` cannot import Django + # at all ("Couldn't import Django. Are you sure it's installed…"). Measured, not guessed. + self.container = container + self.sub = sub + self.ready_path = ready_path + self.ready = ready + self.python = python + self.workdir = workdir + + def _wd(self): + return f"-w {self.workdir} " if self.workdir else "" + + def _manage(self, w, code): + # -c is passed to `manage.py shell`; the app's own shell, its own ORM. + return w.guest( + f"docker exec {self._wd()}{self.container} {self.python} manage.py shell " + f"-c {json.dumps(code)} 2>&1", timeout=300) + + def _exists(self, w, username): + # ONE LINE, semicolon-separated. A `\n` inside a double-quoted shell argument reaches + # python as a literal backslash-n and is a SyntaxError — which is exactly how the first + # adventurelog run read as `inconclusive`. The fixture refused to guess, which is right, + # but the instrument was the thing that was broken. + out = self._manage(w, ( + "from django.contrib.auth import get_user_model; " + f"print('DRILL_ANSWER=' + str(get_user_model().objects.filter(username={username!r}).exists()))" + )) + m = re.search(r"DRILL_ANSWER=(True|False)", out) + return (m.group(1) == "True") if m else None, " ".join(out.split())[-300:] + + def seed(self, w, sub, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -e DJANGO_SUPERUSER_PASSWORD={pw} {self._wd()}{self.container} " + f"{self.python} manage.py createsuperuser --noinput " + f"--username {user} --email {user}@gate.invalid 2>&1", timeout=300) + say(f" {self.container}: createsuperuser :: {' '.join(out.split())[:160]}") + got, detail = self._exists(w, user) + if got is not True: + say(f" {self.container}: the account did not appear in the app's own ORM :: {detail[:200]}") + return None + say(f" {self.container}: seeded superuser {user}") + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + say(f" {self.container}: the app never served {self.ready_path}") + return False + absent, detail = self._exists(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" {self.container}: READBACK UNUSABLE — a username that cannot exist did not " + f"read as absent ({absent}) :: {detail[:200]}") + return False + found, detail = self._exists(w, t["user"]) + say(f" {self.container}: readback of the seeded account found={found}") + if found is not True: + say(f" {self.container}: :: {detail[:250]}") + return found is True + + +# ============================================================================================= +class Nextcloud: + """Nextcloud's OWN admin CLI, `occ`, inside its own container: its own code, its own user + backend. Not a SQL INSERT and not a planted file (R-156). + + `occ user:info` is the readback, and it PROVES ITSELF on every call: a uid that cannot exist + must answer "user not found". A readback that has broken into always succeeding therefore + fails instead of passing everything. + + This is the app chosen for the MariaDB engine-major edge (`09` §3 decision 5, R-469 lifted): + the app image does NOT move, only the `mariadb:` sidecar, so the edge carries exactly one + migration and a failure is readable. + """ + sub = "cloud" + + def _occ(self, w, *args, timeout=420): + import shlex + line = " ".join(shlex.quote(a) for a in args) + return w.guest(f"docker exec -u www-data nextcloud php occ {line} 2>&1", timeout=timeout) + + def _info(self, w, uid): + out = self._occ(w, "user:info", uid) + flat = " ".join(out.split()) + if "user not found" in flat.lower() or "could not be found" in flat.lower(): + return False, flat + if f"user_id: {uid}" in flat or f"- user_id: {uid}" in flat or f"user_id: {uid}" in out: + return True, flat + return None, flat + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + return None + uid = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -u www-data -e OC_PASS={pw} nextcloud php occ user:add " + f"--password-from-env --display-name={uid} {uid} 2>&1", timeout=420) + say(f" nextcloud: occ user:add :: {' '.join(out.split())[:160]}") + got, flat = self._info(w, uid) + if got is not True: + say(f" nextcloud: the account did not appear via occ user:info :: {flat[:220]}") + return None + say(f" nextcloud: seeded user {uid}") + return {"uid": uid, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + say(" nextcloud: the app never served /status.php") + return False + absent, flat = self._info(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" nextcloud: READBACK UNUSABLE — a uid that cannot exist did not read absent " + f"({absent}) :: {flat[:200]}") + return False + found, flat = self._info(w, t["uid"]) + say(f" nextcloud: readback of the seeded user found={found}") + if found is not True: + say(f" nextcloud: :: {flat[:250]}") + return found is True + + +# ============================================================================================= +class Grafana: + """Grafana's own HTTP API as the admin the DEPLOY created. The password is the one the + controller showed the household — read from the app's own `app.yaml`, not invented — and the + data (a folder) goes in and comes back through the app's own REST API.""" + sub = "grafana" + + def _auth(self, w, name="grafana"): + # app.yaml stores this ENCRYPTED (`ENC:…`), so it cannot be read back off the box — which + # is correct, and is why the harness uses the value IT generated for the deploy. + pw = (w.GENERATED.get(name) or {}).get("GF_SECURITY_ADMIN_PASSWORD") or "admin" + return f"admin:{pw}" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return None + au = self._auth(w) + title = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/folders", "-u", au, + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="POST") + say(f" grafana: create folder http={code}") + if code not in ("200", "201"): + say(f" grafana: refused {body[:220]}") + return None + try: + uid = json.loads(body)["uid"] + except Exception: + say(f" grafana: no uid in {body[:200]}") + return None + return {"uid": uid, "title": title} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return False + au = self._auth(w) + rc, code, _ = w.app_curl(sub, "/api/folders/nope" + secrets.token_hex(5), "-u", au) + if code == "200": + say(" grafana: READBACK UNUSABLE — a folder uid that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/folders/{t['uid']}", "-u", au) + ok = code == "200" and t["title"] in body + say(f" grafana: readback of the seeded folder http={code} ok={ok}") + return ok + + +# ============================================================================================= +class AudiobookShelf: + """audiobookshelf's own /init endpoint creates the first root account; its own /login proves + the account survived. Both are the app's own API.""" + sub = "audiobooks" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/init", "-H", "Content-Type: application/json", + data=json.dumps({"newRoot": {"username": user, "password": pw}}), + method="POST") + say(f" audiobookshelf: /init http={code}") + if code not in ("200", "204"): + say(f" audiobookshelf: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code == "200": + say(" audiobookshelf: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": t["pw"]}), + method="POST") + ok = code == "200" and t["user"] in body + say(f" audiobookshelf: login as the seeded root http={code} ok={ok}") + return ok + + +# ============================================================================================= +class ActualBudget: + """Actual's own bootstrap API sets the server password; its own login proves it survived.""" + sub = "budget" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/account/bootstrap", + "-H", "Content-Type: application/json", + data=json.dumps({"password": pw}), method="POST") + say(f" actualbudget: /account/bootstrap http={code} :: {body[:140]}") + if code not in ("200", "201") or '"status":"ok"' not in body: + return None + return {"pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def login(p): + return w.app_curl(sub, "/account/login", "-H", "Content-Type: application/json", + data=json.dumps({"loginMethod": "password", "password": p}), + method="POST") + rc, code, body = login("wrong-" + secrets.token_hex(6)) + if '"status":"ok"' in body: + say(" actualbudget: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = '"status":"ok"' in body + say(f" actualbudget: login with the seeded password http={code} ok={ok}") + if not ok: + say(f" actualbudget: body {body[:200]}") + return ok + + +# ============================================================================================= +class Mealie: + """Mealie ships a documented first-run admin. We log in as it through the app's own OAuth-style + token endpoint, create a recipe through the app's own API, and read the recipe back.""" + sub = "recipes" + + def _token(self, w, sub, pw="MyPassword"): + rc, code, body = w.app_curl( + sub, "/api/auth/token", "-H", "Content-Type: application/x-www-form-urlencoded", + data=f"username=changeme%40example.com&password={pw}", method="POST") + if code != "200": + return None, f"http={code} {body[:200]}" + try: + return json.loads(body)["access_token"], "" + except Exception: + return None, body[:200] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return None + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate as the first-run admin :: {why}") + return None + name = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/recipes", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": name}), method="POST") + say(f" mealie: create recipe http={code}") + if code not in ("200", "201"): + say(f" mealie: refused {body[:220]}") + return None + slug = body.strip().strip('"') + return {"slug": slug, "name": name} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return False + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate after the update :: {why}") + return False + rc, code, _ = w.app_curl(sub, "/api/recipes/nope" + secrets.token_hex(5), + "-H", f"Authorization: Bearer {tok}") + if code == "200": + say(" mealie: READBACK UNUSABLE — a slug that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/recipes/{t['slug']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["name"] in body + say(f" mealie: readback of the seeded recipe http={code} ok={ok}") + return ok + + +# ============================================================================================= +class N8n: + """n8n's own owner-setup API creates the first account; its own login proves it survived.""" + sub = "auto" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill" + secrets.token_hex(8) + "1" + rc, code, body = w.app_curl(sub, "/rest/owner/setup", "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "firstName": "drill", + "lastName": "drill", "password": pw}), + method="POST") + say(f" n8n: /rest/owner/setup http={code}") + if code not in ("200", "201"): + say(f" n8n: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return False + def login(p): + return w.app_curl(sub, "/rest/login", "-H", "Content-Type: application/json", + data=json.dumps({"emailOrLdapLoginId": t["email"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" n8n: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" and t["email"] in body + say(f" n8n: login as the seeded owner http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Zipline: + """Zipline's own setup/login API. Zipline 4 creates the first user through its own endpoint.""" + sub = "img" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/healthcheck", want=("200",), tries=90): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=30): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + for path in ("/api/auth/register", "/api/auth/setup"): + rc, code, body = w.app_curl(sub, path, "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), + method="POST") + say(f" zipline: {path} http={code} :: {body[:160]}") + if code in ("200", "201"): + return {"user": user, "pw": pw} + say(" zipline: neither register nor setup accepted a first user") + return None + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=60): + return False + def login(p): + return w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" zipline: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" + say(f" zipline: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vikunja: + """Vikunja's own REST API: register a user, log in, create a project, read the project back. + Four calls, all the app's own front door.""" + sub = "tasks" + + def _token(self, w, sub, t, pw=None): + rc, code, body = w.app_curl(sub, "/api/v1/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], + "password": pw or t["pw"]}), method="POST") + if code != "200": + return None, f"http={code} {body[:160]}" + try: + return json.loads(body)["token"], "" + except Exception: + return None, body[:160] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/v1/register", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw, + "email": f"{user}@gate.invalid"}), + method="POST") + say(f" vikunja: register http={code}") + if code not in ("200", "201"): + say(f" vikunja: refused {body[:220]}") + return None + t = {"user": user, "pw": pw} + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: could not log in after registering :: {why}") + return None + title = "drill-" + secrets.token_hex(5) + # Vikunja CREATES with PUT, not POST — a POST answers `405 Method Not Allowed`, which + # reads like a broken fixture and is really the wrong verb. Measured 2026-09-21. + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="PUT") + say(f" vikunja: create project http={code}") + if code not in ("200", "201"): + say(f" vikunja: project refused {body[:220]}") + return None + t["title"] = title + t["pid"] = json.loads(body).get("id") + return t + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return False + bad, why = self._token(w, sub, t, pw="wrong-" + secrets.token_hex(6)) + if bad: + say(" vikunja: READBACK UNUSABLE — a wrong password authenticated") + return False + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: the seeded account no longer authenticates :: {why}") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{t['pid']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["title"] in body + say(f" vikunja: readback of the seeded project http={code} ok={ok}") + return ok + + +# ============================================================================================= +class OpenGist: + """Opengist's own sign-up and sign-in FORMS. + + Two things had to be measured. Its sign-up is CSRF-protected: a bare POST answers 500 with an + HTML page, which reads like a broken app and is really a missing token — fetch the form, keep + its cookie, send its `_csrf` back. And its REST API refuses the account's own password + (`401 {"message":"Bad crendentials"}`) because it wants a token the app will not mint without a + browser. So the SEEDED DATA is the account itself and the READBACK is a real sign-in, which is + the same shape the docmost and navidrome fixtures use. + + LIMITATION, recorded rather than papered over: this is the DATABASE half. A gist's CONTENT is + not seeded, because that needs the API token above. + """ + sub = "gist" + + def _form(self, w, sub, path, jar, fields): + rc, code, html = w.app_curl(sub, path, "-b", jar, "-c", jar) + m = re.search(r'name="_csrf"[^>]*value="([^"]+)"', html or "") + if not m: + return None, f"no _csrf on {path} (http={code})" + body = "&".join([f"_csrf={m.group(1)}"] + [f"{k}={v}" for k, v in fields.items()]) + rc, code, out = w.app_curl(sub, path, "-b", jar, "-c", jar, + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + return code, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, out = self._form(w, sub, "/register", jar, {"username": user, "password": pw}) + say(f" opengist: /register (with its own _csrf) http={code}") + if code not in ("200", "302", "303"): + say(f" opengist: refused {str(out)[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + # Wait for the LOGIN FORM, not for the root page. Measured 2026-09-21: immediately after a + # successful update the root answers while /login does not yet carry its `_csrf`, so the + # sign-in silently fails and the app looks like it lost the account. It had not. + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" opengist: /login never came back after the update") + return False + for _ in range(24): + rc, code, html = w.app_curl(sub, "/login") + if code == "200" and '_csrf' in (html or ""): + break + time.sleep(5) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar, + {"username": t["user"], "password": "wrong-" + secrets.token_hex(5)}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar) + if t["user"] in (home or ""): + say(" opengist: READBACK UNUSABLE — a wrong password signed in") + return False + jar2 = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar2, {"username": t["user"], "password": t["pw"]}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar2) + ok = t["user"] in (home or "") + say(f" opengist: sign-in as the seeded account http={code} name_on_page={ok}") + return ok + + +# ============================================================================================= +class Papra: + """Papra's own e-mail sign-up and sign-in endpoints.""" + sub = "papra" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/auth/sign-up/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "password": pw, + "name": "drill"}), method="POST") + say(f" papra: sign-up http={code}") + if code not in ("200", "201"): + say(f" papra: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def signin(p): + return w.app_curl(sub, "/api/auth/sign-in/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": t["email"], "password": p}), method="POST") + rc, code, _ = signin("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" papra: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = signin(t["pw"]) + ok = code == "200" + say(f" papra: sign-in as the seeded account http={code} ok={ok}") + return ok + + +# ============================================================================================= +class HomeAssistant: + """Home Assistant's own onboarding API creates the owner account and hands back a code the + same API exchanges for a token. Both are the app's own documented non-browser route.""" + sub = "ha" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/onboarding/users", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "name": "drill", "username": user, + "password": pw, "language": "en"}), + method="POST") + say(f" home-assistant: /api/onboarding/users http={code}") + if code not in ("200", "201"): + say(f" home-assistant: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def _login(self, w, sub, user, pw): + """The app's own login flow: start it, then answer it. A 200 with a step_id of + `mfa`/`init` means the credentials were REFUSED; only `create_entry` is a pass.""" + rc, code, body = w.app_curl(sub, "/auth/login_flow", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "handler": ["homeassistant", None], + "redirect_uri": f"https://{sub}.felhom.invalid/", + "type": "authorize"}), method="POST") + if code not in ("200", "201"): + return None, f"flow start http={code} {body[:160]}" + try: + fid = json.loads(body)["flow_id"] + except Exception: + return None, body[:160] + rc, code, body = w.app_curl(sub, f"/auth/login_flow/{fid}", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "username": user, "password": pw}), + method="POST") + try: + j = json.loads(body) + except Exception: + return None, body[:160] + return (j.get("result") if j.get("type") == "create_entry" else None), body[:200] + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return False + bad, why = self._login(w, sub, t["user"], "wrong-" + secrets.token_hex(6)) + if bad: + say(" home-assistant: READBACK UNUSABLE — a wrong password authenticated") + return False + good, why = self._login(w, sub, t["user"], t["pw"]) + ok = bool(good) + say(f" home-assistant: login as the seeded owner ok={ok}") + if not ok: + say(f" home-assistant: {why}") + return ok + + +# ============================================================================================= +class Romm: + """RomM's own user API, driven the way RomM's own front end drives it. + + Three things had to be measured rather than guessed, and each one answered a 403 or a 422 that + looked like a different fault: RomM sets a **`romm_csrftoken` cookie** on any GET and requires + it back in an **`x-csrftoken` header** (a bare POST is `403 CSRF token verification failed`, + which reads like an auth problem); the fields go in the **JSON body**, not the query string (a + query-string POST is `422 Field required` for every field it was just given); and `email` is + required alongside username, password and role. + + On a fresh install with no admin the first `POST /api/users` is accepted unauthenticated; + afterwards it is not — which is what makes the readback (`POST /api/login` as that user) a real + authentication rather than a repeat of the seed. + + LIMITATION: this is the DATABASE half. RomM's other half is the ROM library on the drive, which + this does not populate. + """ + sub = "arcade" + + def _csrf(self, w, sub): + jar = f"/tmp/romm-{secrets.token_hex(4)}.jar" + w.app_curl(sub, "/api/heartbeat", "-c", jar) + out = w.sh(["bash", "-lc", f"grep -i csrf {jar} | awk '{{print $7}}'"]).stdout or "" + return jar, out.strip() + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + jar, tok = self._csrf(w, sub) + if not tok: + say(" romm: no romm_csrftoken cookie was set on /api/heartbeat") + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl( + sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": pw, "role": "admin"}), method="POST") + say(f" romm: POST /api/users http={code}") + if code not in ("200", "201"): + say(f" romm: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + return False + jar, tok = self._csrf(w, sub) + rc, code, _ = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:wrong-{secrets.token_hex(5)}", method="POST") + if code == "200": + say(" romm: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:{t['pw']}", method="POST") + ok = code == "200" + say(f" romm: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" romm: body {body[:200]}") + return ok + + +FIXTURES = { + "home-assistant": HomeAssistant(), + "romm": Romm(), + "vikunja": Vikunja(), + "opengist": OpenGist(), + "papra": Papra(), + "mealie": Mealie(), + "n8n": N8n(), + "zipline": Zipline(), + "grafana": Grafana(), + "audiobookshelf": AudiobookShelf(), + "actualbudget": ActualBudget(), + "nextcloud": Nextcloud(), + "adventurelog": Django("adventurelog", "travel", "/admin/login/"), + "tandoor": Django("tandoor", "recipes", "/accounts/login/", + python="/opt/recipes/venv/bin/python", workdir="/opt/recipes"), + "privatebin": PrivateBin(), + "docmost": Docmost(), + "bookstack": BookStack(), + "gitea": Gitea(), + "navidrome": Navidrome(), + "vaultwarden": Vaultwarden(), +} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/repoint.py b/documentation/audits/undo-bakeoff-2026-09-23/repoint.py new file mode 100644 index 00000000..ff32bbaa --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/repoint.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config. + +`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is +`controller.yaml.pre-bakeoff` (NOT the older `.pre-28`, which a restore must never pick up). +""" +import re, sys, io +sys.path.insert(0, '.') +import walk as w + +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" + + +def creds(): + for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"): + m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l) + if m: + return m.group(1), m.group(2) + raise SystemExit("no admin credential") + + +def to_drill(): + u, t = creds() + print(w.guest(f""" +set -e +test -f {VOL}/controller.yaml.pre-bakeoff || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-bakeoff +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml" +s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +if not re.search(r'^update:', s, re.M): + s += "update:\\n health_timeout: 90s\\n" +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -A2 '^update:' {VOL}/controller.yaml +""")) + + +def restore(): + print(w.guest(f""" +set -e +cp -p {VOL}/controller.yaml.pre-bakeoff {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -c '^update:' {VOL}/controller.yaml || true +""")) + + +if __name__ == "__main__": + to_drill() if sys.argv[1] == "drill" else restore() diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-1-prep.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-1-prep.txt new file mode 100644 index 00000000..287d9fec --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-1-prep.txt @@ -0,0 +1,15 @@ +10:05:44 === prep romm: deploy at the drill FROM pin +10:05:47 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm'] +10:05:47 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH'] +10:05:47 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +10:07:02 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} +10:07:02 deployed: True +10:07:12 romm: POST /api/users http=201 +10:07:13 romm: login as the seeded user http=200 ok=True +10:07:13 C1 A: True +10:07:13 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +10:08:24 [4] backup idle; last=None +10:08:54 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillb12740b","email":"drillb12740b@gate.invalid","enabled":true,"role":"user","permission_group_id" +10:08:54 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +10:08:54 B reads back: True +10:08:57 named volumes: ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'] diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-2-fcopy.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-2-fcopy.txt new file mode 100644 index 00000000..ea4757f9 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-2-fcopy.txt @@ -0,0 +1,10 @@ +10:08:59 db state before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source +10:09:01 image: rommapp/romm:5.0.0 + +10:09:15 VOL romm_romm_config bytes=4158 files=2 | copy bytes=4158 files=2 | cp -a 0.40s | tar 0.38s (2560 B) +VOL romm_romm_db_data bytes=160302192 files=273 | copy bytes=160302192 files=273 | cp -a 0.68s | tar 0.50s (160458240 B) +VOL romm_romm_redis_data bytes=14512195 files=6 | copy bytes=14512195 files=6 | cp -a 0.42s | tar 0.40s (14509056 B) +stop 3.79s copy(all, cp -a + the tar measurement) 7.75s start 0.68s +free on the docker root: 19.04 GiB + +10:09:49 front door answering again 33.8s after the start diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-3-break.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-3-break.txt new file mode 100644 index 00000000..f88e38c5 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-3-break.txt @@ -0,0 +1,2 @@ +10:09:50 [5] drill commit de819653d3d0: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0) +10:09:55 badge caught up after 4.4 s diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-4-update.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-4-update.txt new file mode 100644 index 00000000..e5d9009a --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-4-update.txt @@ -0,0 +1,7 @@ +10:09:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +10:09:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +10:09:56 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None +10:10:08 + 13.3s phase=starting label=Indítás az új verzióval… err=None hold=None +10:10:14 + 19.5s phase=verifying label=Működés ellenőrzése… err=None hold=None +10:11:54 + 118.9s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. +10:11:54 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 119.0} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-5-undoF.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-5-undoF.txt new file mode 100644 index 00000000..64a5fb5d --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-5-undoF.txt @@ -0,0 +1,15 @@ +10:11:54 === undo by F: romm +10:12:12 [INFO] [settings] restore hold CLEARED for romm + +10:12:14 pinned_images: + romm: rommapp/romm:5.0.0 + romm-db: mariadb:11.4 + romm-redis: redis:7-alpine + +10:12:14 hold lift + pin back took 20.2 s (not part of a product undo) +10:12:18 volumes put back in 1.64s + +10:12:30 product start -> 200 +10:13:05 romm: login as the seeded user http=200 ok=True +10:13:05 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +10:13:08 F RESULT: healthy=True after 45.3s A=True B=True db now: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source (before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-6-update.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-6-update.txt new file mode 100644 index 00000000..5dc91492 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-6-update.txt @@ -0,0 +1,7 @@ +10:13:09 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +10:13:09 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +10:13:10 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +10:13:11 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None +10:13:14 + 5.2s phase=verifying label=Működés ellenőrzése… err=None hold=None +10:14:48 + 99.5s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. +10:14:48 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 99.5} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-7-undoD.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-7-undoD.txt new file mode 100644 index 00000000..e747c7a0 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-7-undoD.txt @@ -0,0 +1,17 @@ +10:14:48 === undo by D: romm +10:15:06 [INFO] [settings] restore hold CLEARED for romm + +10:15:08 pinned_images: + romm: rommapp/romm:5.0.0 + romm-db: mariadb:11.4 + romm-redis: redis:7-alpine + +10:15:08 hold lift + pin back took 20.1 s (not part of a product undo) +10:15:40 undo copy: pre-restore-20260923T081309Z-romm-mariadb.sql 62943 B +marker present +load rc=0 +marker check + load 1.41s + +10:16:15 romm: login as the seeded user http=200 ok=True +10:16:16 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +10:16:18 D RESULT: healthy=True after 33.6s A=True B=True db now: tables=24 alembic=0095_virtual_collections_source (before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/romm-8-cutoff.txt b/documentation/audits/undo-bakeoff-2026-09-23/romm-8-cutoff.txt new file mode 100644 index 00000000..79640568 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/romm-8-cutoff.txt @@ -0,0 +1,6 @@ +10:16:26 === cut-off copy: romm (largest volume romm_romm_db_data) +10:16:35 copy container exit=137 (137 = killed) +source bytes=160215912 cut copy bytes=105844788 +finished-marker in the cut copy: 0 -> F refuses to swap +D: whole copy marker: 1 cut copy marker: 0 -> D refuses to load + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/spike.py b/documentation/audits/undo-bakeoff-2026-09-23/spike.py new file mode 100644 index 00000000..f0de8d13 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/spike.py @@ -0,0 +1,160 @@ +#!/usr/bin/env python3 +"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202. + +EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup, +sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being +spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the +steps the product would take, each one timed. + +State between stages lives in state-.json so each stage can be run, read, and only then +followed by the next (the hand undo needs a person looking at what the previous step left). +""" +import json, os, sys, time, re +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +HERE = os.path.dirname(os.path.abspath(__file__)) +FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()} +SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"} + + +def st_path(app): + return os.path.join(HERE, f"state-{app}.json") + + +def load(app): + return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {} + + +def save(app, s): + json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False) + + +def ts(): + return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + + +# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only +# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it +# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier. +def docmost_seed_b(sub, A): + jar = "/tmp/dm.jar" + w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + name = "drillB" + os.urandom(3).hex() + rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json", + data=json.dumps({"name": name, "slug": name.lower()}), method="POST") + w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}") + return {"space": name} if code in ("200", "201") else None + + +def docmost_verify_b(sub, A, B): + jar = "/tmp/dm.jar" + rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + if code not in ("200", "201"): + w.say(f" docmost B: cannot log in (http={code})"); return False + rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json", + data="{}", method="POST") + ok = code in ("200", "201") and B["space"] in out + neg = "drillBnever" in out + w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def break_edge(app, frm, to, port_from, port_to): + """The failing edge: a REAL migrating image move, plus — in the DRILL template only — the + health probe pointed at a port the app does not answer. Both in one drill commit.""" + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read() + m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f) + assert m, "probe port not found" + f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():] + open(fy, "w").write(f) + h = w.drill_bump(app, frm, to) + return h + + +def pg_state(container, db, user): + return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1") + + +def romm_seed_b(sub, A): + """A SECOND RomM user, created by the first (admin) one — written after the backup.""" + jar, tok = FX["romm"]._csrf(w, sub) + user = "drillb" + os.urandom(3).hex() + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}), + method="POST") + w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}") + return {"user": user} if code in ("200", "201") else None + + +def romm_verify_b(sub, A, B): + jar, tok = FX["romm"]._csrf(w, sub) + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}") + ok = code == "200" and B["user"] in body + neg = "drillbnever" in body + w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def tree(paths): + cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths) + return w.guest(cmd) + + +def my_state(container="romm-db", db="romm"): + # The root password is used INSIDE the container from its own env — it never leaves it. + q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '" + return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) " +echo -n "alembic=$({q("select version_num from alembic_version")}) " +echo -n "users=$({q("select count(*) from users")})" +""") + + +def vik_seed_b(sub, A): + """A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the + app writes into its files volume), uploaded through the app's own attachment API.""" + tok, why = FX["vikunja"]._token(w, sub, A) + title = "drillB-" + os.urandom(4).hex() + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT") + if code not in ("200", "201"): + w.say(f" vikunja B: project refused {code} {body[:120]}"); return None + pid = json.loads(body)["id"] + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT") + tid = json.loads(body)["id"] if code in ("200", "201") else None + content = "drill attachment " + os.urandom(8).hex() + fn = "/tmp/vik-att.txt"; open(fn, "w").write(content) + rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}", + "-F", f"files=@{fn}", method="PUT") + w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}") + return {"title": title, "pid": pid, "tid": tid, "att": content} + + +def vik_verify_b(sub, A, B): + tok, why = FX["vikunja"]._token(w, sub, A) + if not tok: + w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False} + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}") + proj = code == "200" and B["title"] in body + rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}") + att_ok = False + try: + atts = json.loads(body) + if atts: + aid = atts[0]["id"] + rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}") + att_ok = c3 == "200" and B["att"] in b3 + except Exception as e: + w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}") + w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}") + return {"project": proj, "attachment": att_ok} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-1-prep.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-1-prep.txt new file mode 100644 index 00000000..030b126e --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-1-prep.txt @@ -0,0 +1,14 @@ +10:16:50 === prep vikunja: deploy at the drill FROM pin +10:16:50 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +10:16:55 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'} +10:16:55 deployed: True +10:16:56 vikunja: register http=200 +10:16:56 vikunja: create project http=201 +10:16:56 vikunja: readback of the seeded project http=200 ok=True +10:16:56 C1 A: True +10:16:56 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +10:17:56 [4] backup idle; last=None +10:17:57 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill6d84b2","created":"2026-09 +10:17:57 vikunja B: project readback=True attachment content readback=True +10:17:57 B reads back: True +10:18:00 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db'] diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-2-fcopy.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-2-fcopy.txt new file mode 100644 index 00000000..27173aff --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-2-fcopy.txt @@ -0,0 +1,9 @@ +10:18:03 db state before the update: tables=36 ledger=117 newest=SCHEMA_INIT +10:18:05 image: vikunja/vikunja:2.3.0 + +10:18:12 VOL vikunja_vikunja_data bytes=4129 files=2 | copy bytes=4129 files=2 | cp -a 0.42s | tar 0.41s (2560 B) +VOL vikunja_vikunja_db bytes=2471792 files=4 | copy bytes=2471792 files=4 | cp -a 0.42s | tar 0.40s (2470912 B) +stop 0.18s copy(all, cp -a + the tar measurement) 5.03s start 0.23s +free on the docker root: 18.18 GiB + +10:18:14 front door answering again 2.1s after the start diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-3-break.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-3-break.txt new file mode 100644 index 00000000..416f492e --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-3-break.txt @@ -0,0 +1,2 @@ +10:18:15 [5] drill commit 9772e7c0b9f5: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0) +10:18:20 badge caught up after 4.5 s diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-4-update.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-4-update.txt new file mode 100644 index 00000000..7e2b3bdb --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-4-update.txt @@ -0,0 +1,7 @@ +10:18:20 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +10:18:20 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +10:18:21 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None +10:18:24 + 4.1s phase=starting label=Indítás az új verzióval… err=None hold=None +10:18:25 + 5.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +10:19:56 + 95.4s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +10:19:56 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-5-undoF.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-5-undoF.txt new file mode 100644 index 00000000..f2484ad5 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-5-undoF.txt @@ -0,0 +1,13 @@ +10:19:56 === undo by F: vikunja +10:20:14 [INFO] [settings] restore hold CLEARED for vikunja + +10:20:16 pinned_images: + vikunja: vikunja/vikunja:2.3.0 + +10:20:16 hold lift + pin back took 20.2 s (not part of a product undo) +10:20:19 volumes put back in 0.90s + +10:20:19 product start -> 200 +10:20:22 vikunja: readback of the seeded project http=200 ok=True +10:20:22 vikunja B: project readback=True attachment content readback=True +10:20:24 F RESULT: healthy=True after 2.7s A=True B=True db now: tables=36 ledger=117 newest=SCHEMA_INIT (before the update: tables=36 ledger=117 newest=SCHEMA_INIT) diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-6-update.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-6-update.txt new file mode 100644 index 00000000..08e31289 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-6-update.txt @@ -0,0 +1,6 @@ +10:20:25 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +10:20:25 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +10:20:26 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None +10:20:27 + 2.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +10:21:58 + 93.3s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +10:21:58 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 93.3} diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-7-undoD.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-7-undoD.txt new file mode 100644 index 00000000..6b3a4325 --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-7-undoD.txt @@ -0,0 +1,8 @@ +10:21:58 === undo by D: vikunja +10:22:16 [INFO] [settings] restore hold CLEARED for vikunja + +10:22:18 pinned_images: + vikunja: vikunja/vikunja:2.3.0 + +10:22:18 hold lift + pin back took 20.1 s (not part of a product undo) +10:22:18 vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run. diff --git a/documentation/audits/undo-bakeoff-2026-09-23/vikunja-8-cutoff.txt b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-8-cutoff.txt new file mode 100644 index 00000000..404ab31f --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/vikunja-8-cutoff.txt @@ -0,0 +1,11 @@ +10:22:23 === cut-off copy: vikunja (largest volume vikunja_vikunja_db) +10:22:27 copy container exit=0 (137 = killed) +source bytes=2920872 cut copy bytes=2920872 +finished-marker in the cut copy: 1 -> F refuses to swap +D: no safety dump exists for this app (R-641) + +10:22:40 RETRY — 2.9 MB copied inside 0.02 s, so the first cut missed; kill with no delay: +10:22:48 try 1: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1 +try 2: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1 +try 3: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1 + diff --git a/documentation/audits/undo-bakeoff-2026-09-23/walk.py b/documentation/audits/undo-bakeoff-2026-09-23/walk.py new file mode 100644 index 00000000..ba30e5bb --- /dev/null +++ b/documentation/audits/undo-bakeoff-2026-09-23/walk.py @@ -0,0 +1,493 @@ +#!/usr/bin/env python3 +"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints. + +EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses: + POST /api/stacks//deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan + POST /api/stacks//update · POST /api/stacks//remove +and reads GET /api/stacks/. No controller code exists for it. + +The walk, per `09` §6.4 and the update-night brief §4: + 1 deploy from the DRILL catalog at the LIVE pin + 2 seed through the app's OWN front door (R-156: never a volume, never SQL) + 3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing + 4 „Mentés most" + 5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages + 6 press the guarded Update, record every phase with timestamps + 7 read the seed back through the front door + 8 the four version observables side by side + 9 write the verdict record in `09`'s JSON shape + +`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`. +""" +import argparse, json, os, re, subprocess, sys, time +from datetime import datetime, timezone + +SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad" +EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-bakeoff-2026-09-23" +DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +BASE = "https://192.168.0.114" +HOSTHDR = "Host: felhom.enkisfelhom.hu" +DOMAIN = "enkisfelhom.hu" +HP = "demo-hp" + +LOG = [] + + +def say(*a): + line = " ".join(str(x) for x in a) + ts = datetime.now().strftime("%H:%M:%S") + print(f"{ts} {line}", flush=True) + LOG.append(f"{ts} {line}") + + +def sh(args, timeout=300, inp=None): + try: + return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp) + except (subprocess.TimeoutExpired, OSError) as e: + return subprocess.CompletedProcess(args, 124, "", f"{e}") + + +def guest(script, timeout=600): + """Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting).""" + r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP, + "cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; " + "pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"], + timeout=timeout, inp=script) + return r.stdout or "" + + +def login(): + pw = open(f"{SC}/.ctlpw").read().strip() + sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR, + "-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"]) + h = open(f"{SC}/hdr.txt").read() + m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I) + if not m: + sys.exit("login failed: no session cookie") + open(f"{SC}/sess.txt", "w").write(m.group(0)) + r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"]) + c = re.search(r'/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN. + + Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because + a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password + (grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only + reason this was safe to discover by running it (live-probes rule). + + A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on + the scratch drive first — the same act the drive browser performs for a household. + """ + code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields") + fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or [] + values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub} + made = [] + for f in fields: + ev, ty = f.get("env_var"), f.get("type") + if ev in values: + continue + # `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses + # when the caller sends none, deliberately ("the user needs to know their password"), + # while `.felhom.yml` declares `required: false` and the API serves that verbatim. A + # caller that trusts the contract gets a 400. Measured tonight on grafana; filed. + if not f.get("required") and ty != "password": + continue # the controller generates the optional secrets itself + if ty == "path": + p = f"{DRIVE}/{name}" + values[ev] = p + made.append(p) + elif ty in ("secret", "password"): + import secrets as _s + values[ev] = "Drill-" + _s.token_hex(12) + GENERATED.setdefault(name, {})[ev] = values[ev] + elif f.get("default"): + values[ev] = f["default"] + else: + values[ev] = f"drill-{name}" + if made: + guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made)) + say(f" [1] made the drive paths this app requires: {made}") + extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")] + if extra: + say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}") + return values + + +def deploy(name, sub, extra_values=None): + st = stack(name) + if st.get("deployed"): + say(f" [1] {name} already deployed — reusing") + return True + values = deploy_values(name, sub) + if extra_values: + values.update(extra_values) + code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values}) + say(f" [1] deploy -> {code} {str(d)[:120]}") + if code != "202": + return False + # WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the + # container `healthy` while the controller's own state read `unhealthy` — a gate on `running` + # alone therefore times out on an app that is up. The state is RECORDED rather than required; + # the real gate is the fixture's own `wait_app`, which asks whether the APP answers. + seen = None + for _ in range(90): + time.sleep(5) + st = stack(name) + seen = st.get("state") + # `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21: + # tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm + # read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the + # deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark + # (`runComposeDeploy` writes it), so that is what to wait for. + pins = (st.get("app_config") or {}).get("pinned_images") + if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"): + say(f" [1] deployed, controller state={seen}, " + f"pinned={(st.get('app_config') or {}).get('pinned_images')}") + if seen != "running": + say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, " + f"not treated as a failure; the fixture's front-door wait is the real gate") + return True + say(f" [1] never became deployed (last controller state={seen!r})") + return False + + +def backup_now(name): + code, d = ctl("POST", "/api/backup/run") + say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}") + for _ in range(90): + time.sleep(5) + c2, s = ctl("GET", "/api/backup/status") + dd = s.get("data") or {} + if not dd.get("running", False): + say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}") + return True + say(" [4] backup still running after 7.5 min — carrying on") + return False + + +def drill_bump(app, frm, to, service_hint=None): + """Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates). + + `frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in + two images (adventurelog's backend and frontend) moves both in one edge, while its engine + sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move + every service that carries the app's own version and no others. + """ + comp = f"{DRILL}/templates/{app}/docker-compose.yml" + fy = f"{DRILL}/templates/{app}/.felhom.yml" + s = open(comp).read() + froms = [x.strip() for x in frm.split(",") if x.strip()] + tos = [x.strip() for x in to.split(",") if x.strip()] + if len(froms) != len(tos): + say(f" [5] from/to lists differ in length: {froms} vs {tos}") + return None + for f1, t1 in zip(froms, tos): + if f"image: {f1}" not in s: + say(f" [5] FROM ref not found in compose: {f1}") + return None + s = s.replace(f"image: {f1}", f"image: {t1}") + open(comp, "w").write(s) + f = open(fy).read() + today = datetime.now().strftime("%Y-%m-%d") + f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M) + open(fy, "w").write(f) + sh(["git", "-C", DRILL, "add", "-A"]) + sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"]) + r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})") + return h + + +def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5): + """Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP. + + R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and + `catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not + enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the + Update that followed moved nothing and still reported "Frissitve". So when the caller knows + which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the + NUMBER R-607 asks for and has never had. + """ + t0 = time.time() + ctl("POST", "/api/sync") + time.sleep(2) + ctl("POST", "/api/stacks/rescan") + time.sleep(2) + if not expect_app or not expect_ref: + return None + for i in range(tries): + cat = stack(expect_app).get("catalog_images") or {} + if expect_ref in cat.values(): + waited = round(time.time() - t0, 1) + if i: + say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up " + f"to {expect_ref} — R-607's window, measured") + return waited + time.sleep(delay) + ctl("POST", "/api/sync") + time.sleep(1) + ctl("POST", "/api/stacks/rescan") + say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — " + f"catalog_images = {stack(expect_app).get('catalog_images')}") + return None + + +def badges(name): + out = {} + for lang, suffix in (("hu", ""), ("en", "?lang=en")): + h = page(f"/apps/{name}{suffix}") + m = re.findall(r']*title="([^"]*)"[^>]*>([^<]*)<', h) + out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3] + return out + + +def press_update(name, poll=1.0, cap_s=1800): + code, d = ctl("POST", f"/api/stacks/{name}/update") + say(f" [6] Update -> {code} {str(d)[:220]}") + if code not in ("202", "200"): + return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0} + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < cap_s: + st = stack(name) + ph = st.get("update_phase") + if ph != seen: + seen = ph + rec = {"t": round(time.time() - t0, 1), "phase": ph, + "label": st.get("update_phase_label"), "updating": st.get("updating"), + "error": st.get("update_error"), "hold": st.get("hold_reason")} + phases.append(rec) + say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} " + f"err={rec['error']} hold={rec['hold']}") + if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3: + break + time.sleep(poll) + st = stack(name) + return {"accepted": True, "http": code, "phases": phases, + "duration_s": round(time.time() - t0, 1), + "final_phase": st.get("update_phase"), "update_error": st.get("update_error"), + "hold_reason": st.get("hold_reason"), "state": st.get("state")} + + +def observables(name): + st = stack(name) + ac = st.get("app_config") or {} + live = guest(f""" +grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//' +echo '---inspect---' +for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do + echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}' +done +""") + a, _, b = live.partition("---inspect---") + return { + "pinned_images": ac.get("pinned_images"), + "installed_images": {k: (v.get("ref") if isinstance(v, dict) else v) + for k, v in (ac.get("installed_images") or {}).items()}, + "catalog_images": st.get("catalog_images"), + "live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()], + "docker_inspect": [x for x in b.strip().splitlines() if x.strip()], + } + + +def app_logs(name, lines=400): + """The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is + one string with escaped newlines — a scan over the envelope sees a single enormous line and + finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96 + rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines.""" + code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}") + if isinstance(d, dict): + data = d.get("data") + if isinstance(data, dict) and isinstance(data.get("logs"), str): + return data["logs"] + if isinstance(d.get("_raw"), str): + return d["_raw"] + return str(d) + + +def write_verdict(rec, appdir): + os.makedirs(appdir, exist_ok=True) + p = os.path.join(appdir, "verdict.json") + json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False) + say(f" [9] verdict {rec['verdict']} -> {p}") + + +def remove(name): + """Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint + refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up.""" + c1, d1 = ctl("POST", f"/api/stacks/{name}/stop") + say(f" [X] stop -> {c1} {str(d1)[:100]}") + for _ in range(24): + time.sleep(5) + if stack(name).get("state") != "running": + break + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": True, "remove_backups": True}) + say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}") + if code == "409": + # R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive + # path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202 + # `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits + # this. The household's other choice — remove the app, KEEP the data — is accepted, and the + # harness takes it, then tidies its own directory by name at teardown. + say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)" + " — removing the app and KEEPING the drive data instead") + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": False, "remove_backups": True}) + say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}") + time.sleep(5) + st = stack(name) + left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; " + f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'") + say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}") + return code + + +def app_env(name, key): + """Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the + app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller + shows them the same value. Data still goes in through the app's own front door.""" + out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1") + if ":" in out: + return out.split(":", 1)[1].strip().strip('"').strip("'") + return "" + + +def snapshots(name): + """The restorable copies the backups page offers for this app.""" + code, d = ctl("GET", f"/api/backup/snapshots?stack={name}") + data = d.get("data") if isinstance(d, dict) else None + if isinstance(data, dict): + for k in ("snapshots", "items", "restore_points"): + if isinstance(data.get(k), list): + return data[k] + return data if isinstance(data, list) else [] + + +def restore(name, snapshot_id=None, wait_s=1200): + """The household's own way out: the „Visszaállítás a mentésből" button on the backups page. + + A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`, + `snapshot_id` — because that is the button the sentence tells them to press. + """ + snaps = snapshots(name) + if snapshot_id is None: + if not snaps: + say(f" [R] no restorable copy offered for {name}") + return {"ok": False, "why": "no snapshot offered", "snapshots": snaps} + first = snaps[0] + snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id") + say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)") + sess = open(f"{SC}/sess.txt").read().strip() + csrf = open(f"{SC}/csrf.txt").read().strip() + r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}", + "-X", "POST", + "--data-urlencode", f"_csrf={csrf}", + "--data-urlencode", f"stack_name={name}", + "--data-urlencode", f"snapshot_id={snapshot_id}", + f"{BASE}/backup/restore"], timeout=180) + head = (r.stdout or "").split("\n")[0].strip() + loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")] + say(f" [R] POST /backup/restore -> {head} {loc[:1]}") + t0 = time.time() + last = None + while time.time() - t0 < wait_s: + code, d = ctl("GET", "/api/backup/restore-status") + dd = d.get("data") or {} + cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message")) + if cur != last: + say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}") + last = cur + if not dd.get("running", False) and time.time() - t0 > 5: + break + time.sleep(2) + st = stack(name) + say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} " + f"phase={st.get('update_phase')}") + return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps, + "http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1), + "state_after": st.get("state"), "hold_after": st.get("hold_reason"), + "observables_after": observables(name)} diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index beb68f6a..69129937 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -814,6 +814,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-642** | **[P3-LOW] `POST /api/stacks/{name}/start` answers 200 *Stack … start completed* while the app is crash-looping.** MEASURED 2026-09-23 on 9202 twice: docmost 0.95.0 and romm 5.0.0 started on data their newer versions had migrated — both refused and restarted in a loop (`Restarting (1)`, front door 404) behind a 200. The same false-green class as R-443 (closed for the Update) and R-635, on the Start. It matters now because an undo (R-637) must never read the start's return as success. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-43`, `romm-43`. | **OPEN — P3; owner: CC** | | **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. | **WAITING-ON-OPERATOR — `09` §6.4's one open point; owner: CC once answered** | | **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** | +| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". | **OPEN — P3; owner: CC** |