The undo, built and proven live: controller v0.263.2 (09 decision 15)
gates / gates (push) Successful in 25s

- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
  live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
  the backup, after it and seconds before the press read back; cut-off copy
  held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
  ruled; R-646 opened. STATUS asks the floor question.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 12:25:13 +02:00
parent 5a349d9884
commit 05ea21e918
36 changed files with 2727 additions and 124 deletions
File diff suppressed because one or more lines are too long
@@ -744,9 +744,12 @@ closed by construction: nothing reports an update complete on the compose exit c
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
| 8 | `done` | installed images recorded, journal cleared | — |
| 5a | `copying` **(v0.263.0)** | `compose stop`, then each NAMED volume `cp -a` into `<volume>.pre-update-<stamp>` by a helper that writes a finished-marker last (bind folders never) | copies removed, **pin put back, the old version started again** — nothing new ran |
| 6 | `starting` | `compose up -d --remove-orphans` | **UNDO** (below), HOLD only if the undo fails |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **UNDO** (below), HOLD only if the undo fails. Before v0.263.0: stop + HOLD, the pin stays (Scenario F) |
| 7a | `undoing` **(v0.263.0)** | every copy validated (marker) BEFORE anything is poured back; volumes emptied and refilled; definition, pin and the pinned version's `.felhom.yml` record put back from the job's own copies; `up`; health with the OLD version's probe | HOLD, the sentence prefixed *„A frissítés nem sikerült, és az automatikus visszaállítás sem."* + the data state (`untouched` / `half` / `not_started`); copies kept |
| 7b | `undone` **(v0.263.0)** | installed images recorded, copies removed, `app.yaml` `last_update_undone`, journal cleared | — |
| 8 | `done` | installed images recorded, copies removed, `last_update_undone` cleared, journal cleared | — |
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
`update.health_timeout` (default `5m`).
@@ -792,7 +795,21 @@ the harness has proven it, is slice 6's.
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
### 6.1a The undo (decision 15) — SPIKED BY HAND 2026-09-23, not built
### 6.1a The undo (decision 15) — SPIKED BY HAND, then BUILT: controller v0.263.2 (2026-09-23)
**SHIPPED AND PROVEN LIVE on 9202** (`audits/undo-live-2026-09-23/README.md`): docmost, romm and
vikunja each made a real migrating update fail its (deliberately wrong) probe; the product undid all
three in 30–52 s, with data written before the backup, after it, and seconds before the press all read
back through each app's front door and the ledgers equal; the page carries one line in the request's
language. A cut-off copy → HOLD saying *the data is as the new version left it*; a power cut during
the undo → resumed after boot and completed; a person's press after an undo → `done`. **Two defects
only the live box could show, fixed the same day:** the undo's probe was never asked while the
current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` taken at update time
was already the new one, because `.felhom.yml` flows in on every catalog sync (v0.263.2: the pinned
version's file is now recorded in `applied-meta/` whenever a version is pinned). **Residual (R-646):**
an app pinned before v0.263.2 has no such record until its next pin.
The spike, as it was run by hand before any build:
Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made
to fail a deliberately wrong probe, each held by today's product, each then undone by hand.
@@ -1025,7 +1042,7 @@ what the part can do to a household's data if it is wrong, not how likely that i
| # | part | rulings / rows | cost | depends on | risk to customer data |
|---|---|---|---|---|---|
| **1** | **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
| **1** | **SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23** (`audits/undo-live-2026-09-23/`). **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none |
| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
| **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
@@ -0,0 +1,2 @@
gitea.dooplex.hu/admin/felhom-controller:0.263.0
gitea.dooplex.hu/admin/felhom-controller:0.263.0 Up 25 seconds (healthy)
@@ -0,0 +1,23 @@
11:15:35 === docmost: live undo by the product
11:15:36 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cd8c-8fb0-741d-88b3-097dc97c4ec6","name":"drillBf2da7a","description":"","slug":"drillbf2da7a","logo":null,"visibility":"private","defaultRol
11:15:36 seed C written right before the Update: True
11:15:38 db before: 42 48 ledger, newest 20260620T010047-personal-spaces
11:15:39 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
11:15:39 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
11:15:41 + 2.1s phase=pulling label=Új verzió letöltése…
11:15:42 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
11:15:44 + 5.8s phase=starting label=Indítás az új verzióval…
11:15:55 + 16.8s phase=verifying label=Működés ellenőrzése…
11:17:26 + 107.1s phase=undoing label=Visszaállítás az előző változatra…
11:19:13 + 214.3s phase=failed label=A frissítés nem sikerült
11:19:13 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
11:19:15 observables: {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": []}
11:21:17 app never answered on docs/ (last rc=0 code=404)
11:21:17 docmost: login as the seeded user http=404 ok=False
11:21:17 docmost: login body 404 page not found
11:21:17 docmost B: cannot log in (http=404)
11:21:17 docmost B: cannot log in (http=404)
11:21:19 READBACK A=False B=False C=False db after: Error response from daemon: No such container: docmost-postgres Error response from daemon: No such container: docmost-postgres (before: 42 48 ledger, newest 20260620T010047-personal-spaces)
11:21:19 PAGE: {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}}
11:21:21 leftover copies: 3
@@ -0,0 +1 @@
gitea.dooplex.hu/admin/felhom-controller:0.263.1 Up 25 seconds (healthy)
@@ -0,0 +1 @@
gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 25 seconds (healthy)
@@ -0,0 +1,22 @@
11:37:09 === romm: live undo by the product
11:37:10 romm seed B: POST /api/users (as the admin) http=201 {"id":3,"username":"drillb5fbc5e","email":"drillb5fbc5e@gate.invalid","enabled":true,"role":"user","permission_group_id"
11:37:10 seed C written right before the Update: True
11:37:13 db before: tables=24 alembic=0095_virtual_collections_source
11:37:13 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
11:37:13 + 0.0s phase=checking label=Ellenőrzés…
11:37:14 + 0.6s phase=safety-dump label=Adatbázis pillanatkép…
11:37:15 + 1.6s phase=pulling label=Új verzió letöltése…
11:37:17 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
11:37:21 + 7.9s phase=starting label=Indítás az új verzióval…
11:37:32 + 19.0s phase=verifying label=Működés ellenőrzése…
11:39:03 + 109.5s phase=undoing label=Visszaállítás az előző változatra…
11:40:58 + 225.0s phase=failed label=A frissítés nem sikerült
11:40:58 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.'
11:41:01 observables: {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": []}
11:43:02 app never answered on arcade/api/heartbeat (last rc=0 code=404)
11:53:05 app never answered on arcade/api/heartbeat (last rc=0 code=404)
11:53:05 romm B: GET /api/users http=404 seeded-user-listed=False (negative control listed=False)
11:53:05 romm B: GET /api/users http=404 seeded-user-listed=False (negative control listed=False)
11:53:08 READBACK A=False B=False C=False db after: tables=Error response from daemon: No such container: romm-db alembic=Error response from daemon: No such container: romm-db (before: tables=24 alembic=0095_virtual_collections_source)
11:53:08 PAGE: {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."}}
11:53:10 leftover copies: 3
@@ -0,0 +1,17 @@
11:55:43 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
11:56:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
11:56:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
11:56:22 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'}
11:56:27 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
11:56:27 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
11:56:54 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat
11:57:01 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm'
11:57:02 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
11:57:33 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
11:57:41 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
0
romm-drive-folder-removed
sync 200
1a7a491 DRILL: FROM states again (correct probes) for the live undo proof
@@ -0,0 +1,14 @@
11:57:46 === prep docmost: deploy at the drill FROM pin
11:57:46 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
11:58:16 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
11:58:16 deployed: True
11:58:16 docmost: /api/auth/setup http=200 rc=0
11:58:17 docmost: login as the seeded user http=200 ok=True
11:58:17 C1 A: True
11:58:17 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
11:58:57 [4] backup idle; last=None
11:58:58 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cdb4-4376-7727-a390-a5b9e7f6c581","name":"drillB059474","description":"","slug":"drillb059474","logo":null,"visibility":"private","defaultRol
11:58:58 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
11:58:58 B reads back: True
11:59:01 named volumes: ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage']
port: 3000
@@ -0,0 +1,16 @@
11:59:08 === prep romm: deploy at the drill FROM pin
11:59:11 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm']
11:59:11 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH']
11:59:11 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
12:00:16 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}
12:00:16 deployed: True
12:00:18 romm: POST /api/users http=201
12:00:19 romm: login as the seeded user http=200 ok=True
12:00:19 C1 A: True
12:00:19 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
12:01:19 [4] backup idle; last=None
12:01:51 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillbd5863e","email":"drillbd5863e@gate.invalid","enabled":true,"role":"user","permission_group_id"
12:01:51 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
12:01:51 B reads back: True
12:01:54 named volumes: ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data']
port: 8080
@@ -0,0 +1,15 @@
12:02:01 === prep vikunja: deploy at the drill FROM pin
12:02:01 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
12:02:06 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
12:02:06 deployed: True
12:02:07 vikunja: register http=200
12:02:07 vikunja: create project http=201
12:02:07 vikunja: readback of the seeded project http=200 ok=True
12:02:07 C1 A: True
12:02:07 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
12:03:08 [4] backup idle; last=None
12:03:08 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drilld918b0","created":"2026-09
12:03:08 vikunja B: project readback=True attachment content readback=True
12:03:08 B reads back: True
12:03:11 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
port: 3456
@@ -0,0 +1,2 @@
11:59:04 [5] drill commit 58a76b596330: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0)
11:59:08 badge caught up after 4.5 s
@@ -0,0 +1,2 @@
12:01:57 [5] drill commit 9f9283d26b22: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0)
12:02:01 badge caught up after 4.4 s
@@ -0,0 +1,2 @@
12:03:14 [5] drill commit 6075526e6b27: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
12:03:19 badge caught up after 4.5 s
@@ -0,0 +1,20 @@
12:03:41 === docmost: live undo by the product
12:03:42 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cdb8-9866-7a10-bde0-42e824b55295","name":"drillBbdf31d","description":"","slug":"drillbbdf31d","logo":null,"visibility":"private","defaultRol
12:03:42 seed C written right before the Update: True
12:03:44 db before: 42 48 ledger, newest 20260620T010047-personal-spaces
12:03:44 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
12:03:44 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
12:03:46 + 1.6s phase=pulling label=Új verzió letöltése…
12:03:48 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
12:03:50 + 5.3s phase=starting label=Indítás az új verzióval…
12:04:01 + 16.3s phase=verifying label=Működés ellenőrzése…
12:05:32 + 107.2s phase=undoing label=Visszaállítás az előző változatra…
12:06:02 + 137.6s phase=undone label=Visszaállítva az előző változatra
12:06:02 END phase=undone err=None hold=None
12:06:05 observables: {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": ["docmost docmost/docmost:0.95.0 running=true restarts=0", "docmost-postgres postgres:16-alpine running=true restarts=0", "docmost-redis redis:7-alpine running=true restarts=0"]}
12:06:05 docmost: login as the seeded user http=200 ok=True
12:06:06 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
12:06:06 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
12:06:08 READBACK A=True B=True C=True db after: 42 48 ledger, newest 20260620T010047-personal-spaces (before: 42 48 ledger, newest 20260620T010047-personal-spaces)
12:06:08 PAGE: {"hu": {"undone_line": "A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
12:06:10 leftover copies: 0
@@ -0,0 +1,20 @@
12:06:10 === romm: live undo by the product
12:06:12 romm seed B: POST /api/users (as the admin) http=201 {"id":3,"username":"drillb2d189e","email":"drillb2d189e@gate.invalid","enabled":true,"role":"user","permission_group_id"
12:06:12 seed C written right before the Update: True
12:06:14 db before: tables=24 alembic=0095_virtual_collections_source
12:06:14 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
12:06:14 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
12:06:16 + 1.6s phase=pulling label=Új verzió letöltése…
12:06:18 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
12:06:24 + 10.0s phase=starting label=Indítás az új verzióval…
12:06:36 + 21.0s phase=verifying label=Működés ellenőrzése…
12:08:06 + 111.9s phase=undoing label=Visszaállítás az előző változatra…
12:08:58 + 163.9s phase=undone label=Visszaállítva az előző változatra
12:08:58 END phase=undone err=None hold=None
12:09:01 observables: {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": ["romm rommapp/romm:5.0.0 running=true restarts=0", "romm-db mariadb:11.4 running=true restarts=0", "romm-redis redis:7-alpine running=true restarts=0"]}
12:09:04 romm: login as the seeded user http=200 ok=True
12:09:04 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
12:09:05 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
12:09:07 READBACK A=True B=True C=True db after: tables=24 alembic=0095_virtual_collections_source (before: tables=24 alembic=0095_virtual_collections_source)
12:09:07 PAGE: {"hu": {"undone_line": "A(z) romm frissítése 2026-09-23 12:08-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of romm at 2026-09-23 12:08 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
12:09:09 leftover copies: 0
@@ -0,0 +1,19 @@
12:09:09 === vikunja: live undo by the product
12:09:09 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drilld918b0","created":"2026-09
12:09:09 seed C written right before the Update: True
12:09:11 db before: tables=36 ledger=117 newest=SCHEMA_INIT
12:09:12 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
12:09:12 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
12:09:13 + 1.1s phase=pulling label=Új verzió letöltése…
12:09:14 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt…
12:09:15 + 3.7s phase=verifying label=Működés ellenőrzése…
12:10:46 + 94.5s phase=undoing label=Visszaállítás az előző változatra…
12:10:49 + 97.1s phase=undone label=Visszaállítva az előző változatra
12:10:49 END phase=undone err=None hold=None
12:10:51 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]}
12:10:52 vikunja: readback of the seeded project http=200 ok=True
12:10:52 vikunja B: project readback=True attachment content readback=True
12:10:52 vikunja B: project readback=True attachment content readback=True
12:10:54 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT)
12:10:54 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 12:10-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 12:10 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
12:10:56 leftover copies: 0
@@ -0,0 +1,20 @@
12:11:18 === romm: power cut DURING the undo
12:11:18 romm seed B: POST /api/users (as the admin) http=201 {"id":4,"username":"drillb48b886","email":"drillb48b886@gate.invalid","enabled":true,"role":"user","permission_group_id"
12:11:18 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
12:11:18 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
12:11:20 + 1.6s phase=pulling label=Új verzió letöltése…
12:11:21 + 2.9s phase=copying label=Az adatok másolása a frissítés előtt…
12:11:28 + 9.5s phase=starting label=Indítás az új verzióval…
12:11:39 + 20.5s phase=verifying label=Működés ellenőrzése…
12:13:09 + 111.0s phase=undoing label=Visszaállítás az előző változatra…
12:13:15 >>> POWER CUT (pct stop 9202) in phase undoing: rc=0 in 5.2s
12:13:18 guest started again: rc=0
12:13:31 2026/09/23 10:13:25 update.go:1171: [WARN] [stacks] update recovery: romm was interrupted while UNDOING (started 2026-09-23T10:11:18Z) — marking it Updating and RESUMING the undo
2026/09/23 10:13:25 update.go:1224: [INFO] [stacks] update romm: resuming the UNDO after a controller restart
12:14:22 after the restart: phase=undone hold=None
12:14:27 romm: login as the seeded user http=200 ok=True
12:14:28 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
12:14:28 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
12:14:31 READBACK A=True B=True C=True db: tables=24 alembic=0095_virtual_collections_source
12:14:33 leftover copies: 0
@@ -0,0 +1,26 @@
12:14:33 === vikunja: a cut-off undo copy
12:14:33 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
12:14:33 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
12:14:34 + 1.1s phase=pulling label=Új verzió letöltése…
12:14:35 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt…
12:14:37 + 3.7s phase=starting label=Indítás az új verzióval…
12:14:38 + 4.2s phase=verifying label=Működés ellenőrzése…
12:14:40 >>> finished-marker removed from one copy:
copy: vikunja_vikunja_data.pre-update-20260923T101436Z
total 12
drwxr-xr-x 3 root root 4096 Sep 23 10:14 .
drwxr-xr-x 1 root root 4096 Sep 23 10:14 ..
drwxr-xr-x 2 root root 4096 Sep 23 10:13 data
-rw-r--r-- 1 root root 0 Sep 23 10:14 felhom-undo-complete
data
12:16:08 + 94.8s phase=undoing label=Visszaállítás az előző változatra…
12:16:09 + 95.4s phase=failed label=A frissítés nem sikerült
12:16:09 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
12:16:11 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}}
12:16:11 box language -> en: http 302
12:16:11 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}}
12:16:11 box language -> hu: http 302
12:16:13 copies kept: vikunja_vikunja_data.pre-update-20260923T101436Z
vikunja_vikunja_db.pre-update-20260923T101436Z
@@ -0,0 +1,17 @@
12:16:42 === docmost: a person presses Update again after the undo (probe fixed in the catalog)
12:16:42 before, the page: {"hu": {"undone_line": "A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
12:16:42 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
12:16:42 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
12:16:44 + 1.6s phase=pulling label=Új verzió letöltése…
12:16:45 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
12:16:48 + 5.3s phase=starting label=Indítás az új verzióval…
12:16:59 + 16.3s phase=verifying label=Működés ellenőrzése…
12:17:19 + 36.8s phase=done label=Frissítve
12:17:19 END phase=done err=None hold=None
12:17:20 docmost: login as the seeded user http=200 ok=True
12:17:20 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
12:17:20 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
12:17:23 READBACK A=True B=True C=True installed={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
12:17:23 after, the page: {"hu": {"undone_line": null, "hold": null}, "en": {"undone_line": null, "hold": null}}
12:17:25 app.yaml last_update_undone: 0
@@ -0,0 +1 @@
12:17:57 POST /api/stacks/romm/start -> 200 {'ok': True, 'data': {'state': 'running'}, 'message': 'Stack romm start requested — state now: running'}
@@ -0,0 +1,50 @@
2026/09/23 10:13:25 update.go:1171: [WARN] [stacks] update recovery: romm was interrupted while UNDOING (started 2026-09-23T10:11:18Z) — marking it Updating and RESUMING the undo
2026/09/23 10:13:25 update.go:1224: [INFO] [stacks] update romm: resuming the UNDO after a controller restart
2026/09/23 10:13:25 update.go:826: [ERROR] [stacks] update romm FAILED after the new version was started: resumed after a restart during the undo
2026/09/23 10:13:25 update.go:818: [INFO] [stacks] update romm: kept 24813 bytes of the app's own log at /opt/docker/stacks/romm/hold-logs/20260923T101325Z/compose-logs.txt before stopping it (R-621)
2026/09/23 10:13:25 update.go:1106: [INFO] [stacks] update romm: phase undoing
2026/09/23 10:13:25 undo.go:357: [WARN] [stacks] update romm: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: resumed after a restart during the undo)
2026/09/23 10:14:20 undo.go:404: [INFO] [stacks] update romm: UNDONE in 55s — the previous version is running on the data from before the update (the app's health check passed)
2026/09/23 10:14:33 update.go:487: [INFO] [stacks] update vikunja: accepted — guarded update started
2026/09/23 10:14:33 update.go:1106: [INFO] [stacks] update vikunja: phase checking
2026/09/23 10:14:33 update.go:632: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T10:04:52Z (10m0s old, limit 24h0m0s)
2026/09/23 10:14:33 update.go:1106: [INFO] [stacks] update vikunja: phase safety-dump
2026/09/23 10:14:33 update.go:664: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) []
2026/09/23 10:14:34 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB
2026/09/23 10:14:34 update.go:1106: [INFO] [stacks] update vikunja: phase pinning
2026/09/23 10:14:34 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0)
2026/09/23 10:14:34 update.go:1106: [INFO] [stacks] update vikunja: phase pulling
2026/09/23 10:14:35 update.go:1106: [INFO] [stacks] update vikunja: phase copying
2026/09/23 10:14:36 update.go:1106: [INFO] [stacks] update vikunja: phase copying
2026/09/23 10:14:36 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T101436Z in 469ms
2026/09/23 10:14:36 update.go:1106: [INFO] [stacks] update vikunja: phase copying
2026/09/23 10:14:37 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T101436Z in 444ms
2026/09/23 10:14:37 update.go:1106: [INFO] [stacks] update vikunja: phase starting
2026/09/23 10:14:37 update.go:1106: [INFO] [stacks] update vikunja: phase verifying
2026/09/23 10:16:07 update.go:826: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy)
2026/09/23 10:16:08 update.go:818: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T101607Z/compose-logs.txt before stopping it (R-621)
2026/09/23 10:16:08 update.go:1106: [INFO] [stacks] update vikunja: phase undoing
2026/09/23 10:16:08 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy))
2026/09/23 10:16:08 undo.go:363: [ERROR] [stacks] update vikunja: the undo copy vikunja_vikunja_data.pre-update-20260923T101436Z has no finished-marker — it is cut off or missing; NOTHING is put back
2026/09/23 10:16:08 update.go:838: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T101436Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T101436Z}]
2026/09/23 10:16:08 update_guard.go:536: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T10:04:52Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched")
2026/09/23 10:16:42 update.go:487: [INFO] [stacks] update docmost: accepted — guarded update started
2026/09/23 10:16:42 update.go:1106: [INFO] [stacks] update docmost: phase checking
2026/09/23 10:16:42 update.go:632: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T10:03:45Z (13m0s old, limit 24h0m0s)
2026/09/23 10:16:42 update.go:1106: [INFO] [stacks] update docmost: phase safety-dump
2026/09/23 10:16:43 update.go:664: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260923T101642Z-docmost-postgres.sql]
2026/09/23 10:16:44 undo.go:252: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 64.8 MiB
2026/09/23 10:16:44 update.go:1106: [INFO] [stacks] update docmost: phase pinning
2026/09/23 10:16:44 pin.go:365: [INFO] [stacks] update docmost: pin advanced to the catalog's current definition (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:16-alpine, docmost-redis=redis:7-alpine)
2026/09/23 10:16:44 update.go:1106: [INFO] [stacks] update docmost: phase pulling
2026/09/23 10:16:45 update.go:1106: [INFO] [stacks] update docmost: phase copying
2026/09/23 10:16:46 update.go:1106: [INFO] [stacks] update docmost: phase copying
2026/09/23 10:16:46 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260923T101646Z in 673ms
2026/09/23 10:16:46 update.go:1106: [INFO] [stacks] update docmost: phase copying
2026/09/23 10:16:47 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260923T101646Z in 409ms
2026/09/23 10:16:47 update.go:1106: [INFO] [stacks] update docmost: phase copying
2026/09/23 10:16:47 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260923T101646Z in 415ms
2026/09/23 10:16:47 update.go:1106: [INFO] [stacks] update docmost: phase starting
2026/09/23 10:16:58 update.go:1106: [INFO] [stacks] update docmost: phase verifying
2026/09/23 10:17:19 update.go:785: [INFO] [stacks] update docmost: healthy after 20s (the app's health check passed)
2026/09/23 10:17:19 update.go:793: [INFO] [stacks] update docmost: DONE in 37s
@@ -0,0 +1,59 @@
12:18:27 undo copies before removal: vikunja_vikunja_data.pre-update-20260923T101436Z vikunja_vikunja_db.pre-update-20260923T101436Z
12:18:28 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
12:19:00 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
12:19:07 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
12:19:10 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'}
12:19:15 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
12:19:15 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
12:19:42 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat
12:19:49 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm'
12:19:49 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
12:20:21 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
12:20:28 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
12:20:30 undo copies after removal (must be none): 0
12:20:45 docmost: containers=0 volumes=0
romm: containers=0 volumes=0
vikunja: containers=0 volumes=0
romm drive folder removed
removed docmost/docmost:0.95.0
removed docmost/docmost:0.96.0
removed rommapp/romm:5.0.0
removed rommapp/romm:5.3.0
removed vikunja/vikunja:2.3.0
removed vikunja/vikunja:2.6.0
removed nextcloud:34.0.1-apache
removed gitea.dooplex.hu/admin/felhom-controller:0.263.0
removed gitea.dooplex.hu/admin/felhom-controller:0.263.1
controller.yaml.pre-lock-probe
dload.out
pre-docmost-compose.yml
pre-romm-compose.yml
pre-vikunja-compose.yml
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: <redacted>
username: ""
hub:
0
sync 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'me
cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
controller.yaml == saved pre-bakeoff copy
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 48 seconds (healthy)
filebrowser gtstef/filebrowser:1.3.3-stable Up 8 minutes (healthy)
gokapi f0rc3/gokapi:v1.9.6 Restarting (1) 25 seconds ago
paperless-postgres postgres:16-alpine Up 8 minutes (healthy)
paperless-redis redis:7-alpine Up 8 minutes (healthy)
paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 8 minutes (healthy)
privatebin privatebin/pdo:2.0.6 Up 8 minutes (healthy)
traefik traefik:v3.6.7 Up 8 minutes
cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main
controller.yaml.pre-lock-probe
image lines identical
@@ -0,0 +1,52 @@
# The undo, live on 9202 — controller v0.263.0 → v0.263.2, 2026-09-23
Venue: scratch guest **9202** (demo-hp), drill catalog (`admin/app-catalog-drill`, reset to live
`cfcfe5278428` before and after), `update.health_timeout: 90s`. **Method: endpoint-level** — every
act is the endpoint the UI invokes (deploy, backup, sync, rescan, update, start, remove, the language
switch) and every page is fetched as HTML (`/apps/<app>?lang=hu|en`); no browser. The product performs
the undo; nothing here does. Each edge is a real migrating image move with, in the DRILL template only,
a health probe on a port the app does not answer. **Seed A** before the backup, **seed B** after it,
**seed C** seconds before the press — C exists only in the undo's own last-second copy.
## Not done, or changed
- **Three releases, not one.** v0.263.0 was built, deployed to 9202 only, and its first two live
proofs FAILED honestly (HELD, data put back, old version judged "did not start"). Two defects the unit
tests could not see; each fixed, red-proofed and released the same session: **v0.263.1** (the undo's
probe was gated on `running` while the current probe held the app `unhealthy`) and **v0.263.2** (the
"old" `.felhom.yml` saved at update time was already the new one — it flows in on every catalog sync).
Only 9202 ever ran 0.263.0/0.263.1; both images were removed from it by name. **No floor was raised.**
- **The mail and the operator event** of decision 15 are not built (`09` §6.4 part 2).
- **The hold sentence body stays Hungarian on an English box** (R-606); only the new prefix follows
the box language.
## Results (all on v0.263.2 unless stated)
| proof | result | evidence |
|---|---|---|
| docmost (PostgreSQL) undone by the product | failed probe at +107 s → `undoing` → **`undone` at +137.6 s**; A, B, C read back; `42 tables, 48 ledger` before and after; 0 copies left | `40-undo-docmost.txt` |
| romm (MariaDB) | `undone` at +163.9 s (undo 52 s); A, B, C; `24 tables, alembic 0095` before and after | `40-undo-romm.txt` |
| vikunja (SQLite in a volume) | `undone` at +97.1 s (undo 3 s); A, B, C; `36 tables, ledger 117` before and after | `40-undo-vikunja.txt` |
| the page line, both languages | hu: *„A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en: *"The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost."* (same for romm, vikunja) | `40-undo-*.txt` |
| a cut-off copy (vikunja: the finished-marker taken from one copy during `verifying`) | `undoing` → **`failed` in 0.6 s, nothing poured back**; hold (hu box): *„A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése …"*; en box: *"The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja …"*; both copies kept | `60-cutoff-vikunja.txt` |
| a power cut during the undo (romm: `pct stop 9202` the moment the phase read `undoing`) | after boot: *update recovery: romm was interrupted while UNDOING … RESUMING the undo* → **`undone`**; A, B, C; ledger equal; 0 copies left | `50-powercut-romm.txt` |
| a person presses Update after an undo (docmost; the drill catalog then fixed the probe) | `done` at +36.8 s on 0.96.0; A, B, C; the undone line gone from both pages; `last_update_undone` gone from `app.yaml` | `70-manual-press-docmost.txt` |
| a failed undo holds honestly (v0.263.0, docmost and romm — the defect above) | HOLD *„… Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el."* — true: the data had been put back; the old version was never probed | `10-docmost-undo.txt`, `20-romm-undo.txt` |
| removal deletes kept copies (vikunja, held with 2 copies) | 2 → **0** after `POST /api/stacks/vikunja/remove` | `95-teardown.txt` |
| R-642, the Start answer | `Stack romm start requested — state now: running` | `80-r642-start-answer.txt` |
Extra downtime of the copy (the app is stopped there anyway to be recreated): the `copying` phase
lasted 2.1 s (docmost), 6.8 s (romm, incl. a 3.8 s stop), 1.6 s (vikunja).
## Teardown — three layers
- **machine (9202):** docmost, romm, vikunja removed through the product (no containers, no volumes,
no undo copies); romm's drive folder removed by name (the product kept it, R-442 as in every drill);
test images removed by name (docmost 0.95.0/0.96.0, romm 5.0.0/5.3.0, vikunja 2.3.0/2.6.0, nextcloud
34.0.1-apache, controller 0.263.0/0.263.1); `controller.yaml` restored from the saved copy and read
back identical; catalog cache re-cloned from the live repo (`cfcfe52`). **9202 stays on controller
0.263.2** (self-update off; the fleet floor is untouched at 0.262.1). Apps afterwards: the same three
as at the start of the day (gokapi still crash-looping, R-644).
- **host (demo-hp):** nothing provisioned; the guest was stopped and started once for the power cut.
- **hub:** nothing touched.
- **drill repo:** reset to live `main`, image lines identical, `has_actions: false`.
@@ -0,0 +1,266 @@
#!/usr/bin/env python3
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
drill catalog.
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
Start supplies the app's secrets (they are never decrypted here).
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
"""
import json, os, sys, time
sys.path.insert(0, ".")
import walk as w
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
romm_verify_b, vik_seed_b, vik_verify_b
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
SUFFIX = ".pre-undo"
def seed_b(app, sub, A):
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
def verify_b(app, sub, A, B):
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
return all(r.values()) if isinstance(r, dict) else r
def db_state(app):
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
if app == "docmost":
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
if app == "romm":
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
alembic = q("select version_num from alembic_version")
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
return w.guest("""python3 - <<'PY'
import sqlite3
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
m=c.execute("select count(*), max(id) from migration").fetchone()
print(f"tables={t} ledger={m[0]} newest={m[1]}")
PY""").strip()
def volumes(app):
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
if not v.endswith(SUFFIX)]
def containers(app):
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
def stage_prep(app):
s = {"app": app, "sub": SUB[app]}
w.say(f"=== prep {app}: deploy at the drill FROM pin")
w.say("deployed:", w.deploy(app, s["sub"]))
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
w.backup_now(app)
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
s["volumes"] = volumes(app)
w.say("named volumes:", s["volumes"])
save(app, s)
def stage_fcopy(app):
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
s = load(app)
s["state_before_update"] = db_state(app)
w.say("db state before the update:", s["state_before_update"])
vols = s["volumes"]
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
script = f"""
set -u
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
t0=$(date +%s.%N)
docker stop $cs >/dev/null
t1=$(date +%s.%N)
for v in {' '.join(vols)}; do
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
docker volume create "$v{SUFFIX}" >/dev/null
a=$(date +%s.%N)
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
b=$(date +%s.%N)
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
mkdir -p /root/bk; c=$(date +%s.%N)
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
done
t2=$(date +%s.%N)
docker start $cs >/dev/null
t3=$(date +%s.%N)
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
"""
out = w.guest(script, timeout=1800)
w.say(out)
t0 = time.time()
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
s["fcopy"] = out
s["fcopy_at"] = ts()
save(app, s)
def stage_break(app):
s = load(app)
frm, to, p0, p1 = EDGE[app]
s["break_commit"] = break_edge(app, frm, to, p0, p1)
w.say("badge caught up after", w.sync_rescan(app, to), "s")
save(app, s)
def stage_update(app):
s = load(app)
r = w.press_update(app, poll=1.0)
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
s.setdefault("updates", []).append(r)
save(app, s)
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
docker restart felhom-controller >/dev/null; sleep 15"""
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
def lift_and_pinback(app):
frm, to, _, _ = EDGE[app]
t = time.time()
w.say(w.guest(LIFT.format(app=app)))
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
w.login()
return round(time.time() - t, 1)
def stage_undoF(app):
s = load(app)
w.say(f"=== undo by F: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
vols = s["volumes"]
script = f"""
t0=$(date +%s.%N)
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
for v in {' '.join(vols)}; do
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
done
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
"""
w.say(w.guest(script, timeout=1800))
t0 = time.time()
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
w.say("product start ->", c)
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_undoD(app):
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
s = load(app)
w.say(f"=== undo by D: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
if app == "vikunja":
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
return
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
time.sleep(8)
if app == "docmost":
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
marker = "-- PostgreSQL database dump complete"
else:
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
marker = "-- Dump completed"
script = f"""
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
t0=$(date +%s.%N)
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
"""
w.say(w.guest(script, timeout=900))
t0 = time.time()
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_cutoff(app):
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
s = load(app)
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
script = f"""
v={big}
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
rm -f /root/cut.sql
else echo "D: no safety dump exists for this app (R-641)"; fi
"""
w.say(f"=== cut-off copy: {app} (largest volume {big})")
out = w.guest(script, timeout=600)
w.say(out)
s["cutoff"] = out
save(app, s)
if __name__ == "__main__":
w.login()
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,194 @@
#!/usr/bin/env python3
"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes.
Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/<n>,
fetches the app page in both languages, and reads the seeds back through each app's own front door.
Seed C is written immediately before each Update, after every backup — so only the undo's own
last-second copy can bring it back.
"""
import json, re, sys, time, html as H
sys.path.insert(0, ".")
import walk as w
from spike import load, save, ts
import bakeoff as b
HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}
def seed_c(app, s):
return b.seed_b(app, s["sub"], s["seedA"])
def page_lines(app):
out = {}
for lang in ("hu", "en"):
h = H.unescape(w.page(f"/apps/{app}?lang={lang}"))
m = re.search(r'data-update-undone="true">([^<]*)<', h)
hold = re.search(r'data-held="true">([^<]*)<', h)
out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None}
return out
def press(app, poll=0.5, on_phase=None):
code, d = w.ctl("POST", f"/api/stacks/{app}/update")
w.say(f" Update -> {code} {str(d)[:140]}")
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < 1500:
try:
st = w.stack(app)
except Exception as e:
st = {}
ph = st.get("update_phase")
if ph != seen and ph is not None:
seen = ph
phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label")))
w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}")
if on_phase and on_phase(ph):
return phases, "interrupted"
if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2:
break
time.sleep(poll)
st = w.stack(app)
w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}")
return phases, st
def readback(app, s, with_c=True):
sub = s["sub"]
w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2)
A = b.FX[app].verify(w, sub, s["seedA"], w.say)
B = b.verify_b(app, sub, s["seedA"], s["seedB"])
C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None
return {"A": A, "B": B, "C": C}
def stage_undo(app):
s = load(app)
w.say(f"=== {app}: live undo by the product")
s["seedC"] = seed_c(app, s); s["seedC_at"] = ts()
w.say(" seed C written right before the Update:", bool(s["seedC"]))
before = b.db_state(app); w.say(" db before:", before)
phases, st = press(app)
obs = w.observables(app)
w.say(" observables:", json.dumps(obs))
rb = readback(app, s)
after = b.db_state(app)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})")
pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False))
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")},
"readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs}
save(app, s)
def set_box_language(lang):
"""POST /settings/language — the household's language switch (form + the dashboard's CSRF)."""
sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}",
"-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}",
f"{w.BASE}/settings/language"])
w.say(f" box language -> {lang}: http {r.stdout.strip()}")
def stage_powercut(app):
"""Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard
stop); boot it again and let the controller resume the undo."""
import subprocess
s = load(app)
w.say(f"=== {app}: power cut DURING the undo")
s["seedC"] = seed_c(app, s)
cut = {}
def on_phase(ph):
if ph == "undoing":
t = time.time()
r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120)
cut["at"] = ts(); cut["rc"] = r.returncode
w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s")
return True
return False
phases, _ = press(app, poll=0.3, on_phase=on_phase)
if not cut:
w.say(" the cut never landed in `undoing` — recorded as a MISS"); return
r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180)
w.say(f" guest started again: rc={r.returncode}")
for i in range(60):
time.sleep(5)
try:
w.login(); st = w.stack(app)
if st:
break
except SystemExit:
continue
w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6"))
t0 = time.time()
while time.time() - t0 < 600:
st = w.stack(app)
if not st.get("updating") and st.get("update_phase") in ("undone", "failed"):
break
time.sleep(3)
w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}")
rb = readback(app, s)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}")
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb}
save(app, s)
def stage_cutoff(app):
"""Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of
the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it
back and HOLD with the new prefix."""
s = load(app)
w.say(f"=== {app}: a cut-off undo copy")
done = {}
def on_phase(ph):
if ph == "verifying" and not done:
out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c"
docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""")
done["out"] = out
w.say(" >>> finished-marker removed from one copy:\n" + out)
return False
phases, st = press(app, poll=0.5, on_phase=on_phase)
before = b.db_state(app)
pl = page_lines(app)
w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False))
set_box_language("en")
pl_en = page_lines(app)
w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False))
set_box_language("hu")
w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}"))
s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done}
save(app, s)
def stage_fixprobe_and_press(app):
"""After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work
and end the undone note."""
s = load(app)
frm, to, p0, p1 = b.EDGE[app]
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f)
w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"])
w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120)
w.sync_rescan(app, to)
time.sleep(20)
w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan")
w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)")
w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False))
phases, st = press(app)
rb = readback(app, s)
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}")
pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False))
w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml"))
s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl}
save(app, s)
if __name__ == "__main__":
w.login()
{"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff,
"fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2])
@@ -0,0 +1,60 @@
#!/usr/bin/env python3
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
`controller.yaml.pre-bakeoff` (NOT the older `.pre-28`, which a restore must never pick up).
"""
import re, sys, io
sys.path.insert(0, '.')
import walk as w
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
def creds():
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
if m:
return m.group(1), m.group(2)
raise SystemExit("no admin credential")
def to_drill():
u, t = creds()
print(w.guest(f"""
set -e
test -f {VOL}/controller.yaml.pre-bakeoff || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-bakeoff
python3 - <<'PY'
import re
p = "{VOL}/controller.yaml"
s = open(p).read()
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
if not re.search(r'^update:', s, re.M):
s += "update:\\n health_timeout: 90s\\n"
open(p, "w").write(s)
PY
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -A2 '^update:' {VOL}/controller.yaml
"""))
def restore():
print(w.guest(f"""
set -e
cp -p {VOL}/controller.yaml.pre-bakeoff {VOL}/controller.yaml
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -c '^update:' {VOL}/controller.yaml || true
"""))
if __name__ == "__main__":
to_drill() if sys.argv[1] == "drill" else restore()
@@ -0,0 +1,20 @@
set -u
python3 - <<'PY' > 30-reset-apps.txt 2>&1
import walk as w, time
w.login()
for app in ("docmost","romm","vikunja"):
w.remove(app)
print(w.guest("docker volume ls -q --filter label=felhom.undo-copy-of | wc -l; for a in docmost romm vikunja; do docker volume ls -q --filter label=felhom.undo-copy-of=$a; done; rm -rf /mnt/felhom-drives/scratch_hdd/userdata/romm && echo romm-drive-folder-removed"))
for i in range(10):
c,d=w.ctl("POST","/api/sync")
if c=="200": print("sync",c); break
time.sleep(15)
w.ctl("POST","/api/stacks/rescan")
print(w.guest("C=/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache; git -C $C log --oneline -1"))
PY
for app in docmost romm vikunja; do
timeout 1200 python3 bakeoff.py prep $app > 31-prep-$app.txt 2>&1; echo "prep $app rc=$?"
echo "cat /opt/docker/stacks/$app/applied-meta/.felhom.yml 2>&1 | grep -A3 healthcheck | grep port" | /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad/g.sh >> 31-prep-$app.txt 2>&1
timeout 600 python3 bakeoff.py break $app > 32-break-$app.txt 2>&1; echo "break $app rc=$?"
done
echo SETUP-DONE
@@ -0,0 +1,160 @@
#!/usr/bin/env python3
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
steps the product would take, each one timed.
State between stages lives in state-<app>.json so each stage can be run, read, and only then
followed by the next (the hand undo needs a person looking at what the previous step left).
"""
import json, os, sys, time, re
sys.path.insert(0, ".")
import walk as w
import fixtures as fx
HERE = os.path.dirname(os.path.abspath(__file__))
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
def st_path(app):
return os.path.join(HERE, f"state-{app}.json")
def load(app):
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
def save(app, s):
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
def ts():
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
def docmost_seed_b(sub, A):
jar = "/tmp/dm.jar"
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
name = "drillB" + os.urandom(3).hex()
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
return {"space": name} if code in ("200", "201") else None
def docmost_verify_b(sub, A, B):
jar = "/tmp/dm.jar"
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
if code not in ("200", "201"):
w.say(f" docmost B: cannot log in (http={code})"); return False
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
data="{}", method="POST")
ok = code in ("200", "201") and B["space"] in out
neg = "drillBnever" in out
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
return ok and not neg
def break_edge(app, frm, to, port_from, port_to):
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
health probe pointed at a port the app does not answer. Both in one drill commit."""
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read()
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
assert m, "probe port not found"
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
open(fy, "w").write(f)
h = w.drill_bump(app, frm, to)
return h
def pg_state(container, db, user):
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
def romm_seed_b(sub, A):
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
jar, tok = FX["romm"]._csrf(w, sub)
user = "drillb" + os.urandom(3).hex()
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
method="POST")
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
return {"user": user} if code in ("200", "201") else None
def romm_verify_b(sub, A, B):
jar, tok = FX["romm"]._csrf(w, sub)
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}")
ok = code == "200" and B["user"] in body
neg = "drillbnever" in body
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
return ok and not neg
def tree(paths):
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
return w.guest(cmd)
def my_state(container="romm-db", db="romm"):
# The root password is used INSIDE the container from its own env — it never leaves it.
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
echo -n "alembic=$({q("select version_num from alembic_version")}) "
echo -n "users=$({q("select count(*) from users")})"
""")
def vik_seed_b(sub, A):
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
app writes into its files volume), uploaded through the app's own attachment API."""
tok, why = FX["vikunja"]._token(w, sub, A)
title = "drillB-" + os.urandom(4).hex()
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
if code not in ("200", "201"):
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
pid = json.loads(body)["id"]
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
tid = json.loads(body)["id"] if code in ("200", "201") else None
content = "drill attachment " + os.urandom(8).hex()
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
"-F", f"files=@{fn}", method="PUT")
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
return {"title": title, "pid": pid, "tid": tid, "att": content}
def vik_verify_b(sub, A, B):
tok, why = FX["vikunja"]._token(w, sub, A)
if not tok:
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
proj = code == "200" and B["title"] in body
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
att_ok = False
try:
atts = json.loads(body)
if atts:
aid = atts[0]["id"]
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
att_ok = c3 == "200" and B["att"] in b3
except Exception as e:
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
return {"project": proj, "attachment": att_ok}
@@ -0,0 +1,493 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-live-2026-09-23"
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
BASE = "https://192.168.0.114"
HOSTHDR = "Host: felhom.enkisfelhom.hu"
DOMAIN = "enkisfelhom.hu"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
"cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; "
"pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
code, d = ctl("POST", "/api/backup/run")
say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}")
for _ in range(90):
time.sleep(5)
c2, s = ctl("GET", "/api/backup/status")
dd = s.get("data") or {}
if not dd.get("running", False):
say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}")
return True
say(" [4] backup still running after 7.5 min — carrying on")
return False
def drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
+4
View File
@@ -335,3 +335,7 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
| **R-539** | **The restart brake caught a FAST crash loop and was blind to a SLOW one (P3).** Closed by operator ruling 3 of 2026-09-16 in agent **v0.132.0** (`18d03bd`, tag `v0.132.0`, sha256 `4afe8157…`) + hub **v0.117.0** (`37ae31f`): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth `slow_crashloop_since` moves (at most once per 24 h) and the hub mints `controller_slow_crashloop` (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed `agent_update` to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`. **Proven live with the PRODUCTION window (no test-only interval):** five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 `SLOW CRASH-LOOP … restarts_24h=5`; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). **Reasoning kept:** the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. `audits/evidence-chaos-fixes-2026-09-17/partC-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 3c1882a:documentation/backlog/OPEN-ITEMS.md` |
| **R-637** | **Build the undo (`09` §3 decision 15) — the box puts a failed update back by itself (P2).** Closed in controller **v0.263.2** (`8fc2b4a` v0.263.0, `5d38573` v0.263.1, `2cd6666` v0.263.2). Copy method chosen by the bake-off (decision 19): a folder copy of every NAMED volume, `cp -a` into `<vol>.pre-update-<stamp>` after the pull, finished-marker last; bind folders never touched. Undo in `failAndHold`: copies validated first, volumes refilled, definition + pin + the pinned version's `.felhom.yml` (`applied-meta/`) put back from the job's own copies, old probe, `undone` or a HOLD whose sentence says the undo failed and the data state. **Proven live on 9202 (0.263.2):** docmost, romm, vikunja undone by the product with seeds before the backup, after it, and seconds before the press all read back, ledgers equal, page line hu/en; a cut-off copy → HOLD `untouched`; a power cut (`pct stop`) during `undoing` → resumed and undone; a manual press after an undo → `done`, note cleared; removal deleted kept copies. **Two defects the live proof found and the unit tests could not:** the undo's probe was gated on `running` while the current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` was already the new one because it flows in on every sync (v0.263.2). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. **Reasoning kept:** *a copy counts only with its finished-marker, judged by the helper container's own exit — killing `docker run` does not stop the copy; the undo reads its own copies, never the recovery unit.* | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
| **R-639** | **After a held update the previous definition survived only in the recovery unit (P3).** Closed in controller **v0.263.0/v0.263.2**: the pre-update compose, applied definition, pin and the pinned version's `.felhom.yml` are kept until the undo is over; the undo never reads the unit (R-645). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — SHIPPED** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
| **R-641** | **An app with no database server had no last-second copy (P2).** Closed in controller **v0.263.0**: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
| **R-642** | **Start answered 200 "start completed" over a crash loop (P3).** Closed in controller **v0.263.0**: start/restart answer `requested — state now: <state>`, never "completed" (`startAnswer`, pinned by `TestR642_*`, red-proofed); live on 9202: `Stack romm start requested — state now: running`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
File diff suppressed because one or more lines are too long