The undo, built and proven live: controller v0.263.2 (09 decision 15)
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two live-only defects, §6.4 part 1 SHIPPED. - Capability map: a failed update is undone by the box - PROVEN-LIVE. - Live evidence on 9202: three apps undone by the product with seeds before the backup, after it and seconds before the press read back; cut-off copy held honestly; power cut during the undo resumed; manual press after undo. - Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643 ruled; R-646 opened. STATUS asks the floor question. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
File diff suppressed because one or more lines are too long
@@ -744,9 +744,12 @@ closed by construction: nothing reports an update complete on the compose exit c
|
||||
| 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves |
|
||||
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
|
||||
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
|
||||
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
||||
| 8 | `done` | installed images recorded, journal cleared | — |
|
||||
| 5a | `copying` **(v0.263.0)** | `compose stop`, then each NAMED volume `cp -a` into `<volume>.pre-update-<stamp>` by a helper that writes a finished-marker last (bind folders never) | copies removed, **pin put back, the old version started again** — nothing new ran |
|
||||
| 6 | `starting` | `compose up -d --remove-orphans` | **UNDO** (below), HOLD only if the undo fails |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **UNDO** (below), HOLD only if the undo fails. Before v0.263.0: stop + HOLD, the pin stays (Scenario F) |
|
||||
| 7a | `undoing` **(v0.263.0)** | every copy validated (marker) BEFORE anything is poured back; volumes emptied and refilled; definition, pin and the pinned version's `.felhom.yml` record put back from the job's own copies; `up`; health with the OLD version's probe | HOLD, the sentence prefixed *„A frissítés nem sikerült, és az automatikus visszaállítás sem."* + the data state (`untouched` / `half` / `not_started`); copies kept |
|
||||
| 7b | `undone` **(v0.263.0)** | installed images recorded, copies removed, `app.yaml` `last_update_undone`, journal cleared | — |
|
||||
| 8 | `done` | installed images recorded, copies removed, `last_update_undone` cleared, journal cleared | — |
|
||||
|
||||
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
|
||||
`update.health_timeout` (default `5m`).
|
||||
@@ -792,7 +795,21 @@ the harness has proven it, is slice 6's.
|
||||
|
||||
**Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6.
|
||||
|
||||
### 6.1a The undo (decision 15) — SPIKED BY HAND 2026-09-23, not built
|
||||
### 6.1a The undo (decision 15) — SPIKED BY HAND, then BUILT: controller v0.263.2 (2026-09-23)
|
||||
|
||||
**SHIPPED AND PROVEN LIVE on 9202** (`audits/undo-live-2026-09-23/README.md`): docmost, romm and
|
||||
vikunja each made a real migrating update fail its (deliberately wrong) probe; the product undid all
|
||||
three in 30–52 s, with data written before the backup, after it, and seconds before the press all read
|
||||
back through each app's front door and the ledgers equal; the page carries one line in the request's
|
||||
language. A cut-off copy → HOLD saying *the data is as the new version left it*; a power cut during
|
||||
the undo → resumed after boot and completed; a person's press after an undo → `done`. **Two defects
|
||||
only the live box could show, fixed the same day:** the undo's probe was never asked while the
|
||||
current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` taken at update time
|
||||
was already the new one, because `.felhom.yml` flows in on every catalog sync (v0.263.2: the pinned
|
||||
version's file is now recorded in `applied-meta/` whenever a version is pinned). **Residual (R-646):**
|
||||
an app pinned before v0.263.2 has no such record until its next pin.
|
||||
|
||||
The spike, as it was run by hand before any build:
|
||||
|
||||
Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made
|
||||
to fail a deliberately wrong probe, each held by today's product, each then undone by hand.
|
||||
@@ -1025,7 +1042,7 @@ what the part can do to a household's data if it is wrong, not how likely that i
|
||||
|
||||
| # | part | rulings / rows | cost | depends on | risk to customer data |
|
||||
|---|---|---|---|---|---|
|
||||
| **1** | **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
|
||||
| **1** | **SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23** (`audits/undo-live-2026-09-23/`). **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. |
|
||||
| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none |
|
||||
| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none |
|
||||
| **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.263.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.263.0 Up 25 seconds (healthy)
|
||||
@@ -0,0 +1,23 @@
|
||||
11:15:35 === docmost: live undo by the product
|
||||
11:15:36 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cd8c-8fb0-741d-88b3-097dc97c4ec6","name":"drillBf2da7a","description":"","slug":"drillbf2da7a","logo":null,"visibility":"private","defaultRol
|
||||
11:15:36 seed C written right before the Update: True
|
||||
11:15:38 db before: 42 48 ledger, newest 20260620T010047-personal-spaces
|
||||
11:15:39 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
11:15:39 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
11:15:41 + 2.1s phase=pulling label=Új verzió letöltése…
|
||||
11:15:42 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
11:15:44 + 5.8s phase=starting label=Indítás az új verzióval…
|
||||
11:15:55 + 16.8s phase=verifying label=Működés ellenőrzése…
|
||||
11:17:26 + 107.1s phase=undoing label=Visszaállítás az előző változatra…
|
||||
11:19:13 + 214.3s phase=failed label=A frissítés nem sikerült
|
||||
11:19:13 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
|
||||
11:19:15 observables: {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": []}
|
||||
11:21:17 app never answered on docs/ (last rc=0 code=404)
|
||||
11:21:17 docmost: login as the seeded user http=404 ok=False
|
||||
11:21:17 docmost: login body 404 page not found
|
||||
|
||||
11:21:17 docmost B: cannot log in (http=404)
|
||||
11:21:17 docmost B: cannot log in (http=404)
|
||||
11:21:19 READBACK A=False B=False C=False db after: Error response from daemon: No such container: docmost-postgres Error response from daemon: No such container: docmost-postgres (before: 42 48 ledger, newest 20260620T010047-personal-spaces)
|
||||
11:21:19 PAGE: {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}}
|
||||
11:21:21 leftover copies: 3
|
||||
@@ -0,0 +1 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.263.1 Up 25 seconds (healthy)
|
||||
@@ -0,0 +1 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 25 seconds (healthy)
|
||||
@@ -0,0 +1,22 @@
|
||||
11:37:09 === romm: live undo by the product
|
||||
11:37:10 romm seed B: POST /api/users (as the admin) http=201 {"id":3,"username":"drillb5fbc5e","email":"drillb5fbc5e@gate.invalid","enabled":true,"role":"user","permission_group_id"
|
||||
11:37:10 seed C written right before the Update: True
|
||||
11:37:13 db before: tables=24 alembic=0095_virtual_collections_source
|
||||
11:37:13 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
11:37:13 + 0.0s phase=checking label=Ellenőrzés…
|
||||
11:37:14 + 0.6s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
11:37:15 + 1.6s phase=pulling label=Új verzió letöltése…
|
||||
11:37:17 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
11:37:21 + 7.9s phase=starting label=Indítás az új verzióval…
|
||||
11:37:32 + 19.0s phase=verifying label=Működés ellenőrzése…
|
||||
11:39:03 + 109.5s phase=undoing label=Visszaállítás az előző változatra…
|
||||
11:40:58 + 225.0s phase=failed label=A frissítés nem sikerült
|
||||
11:40:58 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.'
|
||||
11:41:01 observables: {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": []}
|
||||
11:43:02 app never answered on arcade/api/heartbeat (last rc=0 code=404)
|
||||
11:53:05 app never answered on arcade/api/heartbeat (last rc=0 code=404)
|
||||
11:53:05 romm B: GET /api/users http=404 seeded-user-listed=False (negative control listed=False)
|
||||
11:53:05 romm B: GET /api/users http=404 seeded-user-listed=False (negative control listed=False)
|
||||
11:53:08 READBACK A=False B=False C=False db after: tables=Error response from daemon: No such container: romm-db alembic=Error response from daemon: No such container: romm-db (before: tables=24 alembic=0095_virtual_collections_source)
|
||||
11:53:08 PAGE: {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."}}
|
||||
11:53:10 leftover copies: 3
|
||||
@@ -0,0 +1,17 @@
|
||||
11:55:43 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
|
||||
11:56:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
|
||||
11:56:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
|
||||
11:56:22 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'}
|
||||
11:56:27 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
|
||||
11:56:27 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
|
||||
11:56:54 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat
|
||||
11:57:01 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm'
|
||||
11:57:02 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
|
||||
11:57:33 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
|
||||
11:57:41 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
|
||||
0
|
||||
romm-drive-folder-removed
|
||||
|
||||
sync 200
|
||||
1a7a491 DRILL: FROM states again (correct probes) for the live undo proof
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
11:57:46 === prep docmost: deploy at the drill FROM pin
|
||||
11:57:46 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
11:58:16 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
11:58:16 deployed: True
|
||||
11:58:16 docmost: /api/auth/setup http=200 rc=0
|
||||
11:58:17 docmost: login as the seeded user http=200 ok=True
|
||||
11:58:17 C1 A: True
|
||||
11:58:17 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
11:58:57 [4] backup idle; last=None
|
||||
11:58:58 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cdb4-4376-7727-a390-a5b9e7f6c581","name":"drillB059474","description":"","slug":"drillb059474","logo":null,"visibility":"private","defaultRol
|
||||
11:58:58 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
11:58:58 B reads back: True
|
||||
11:59:01 named volumes: ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage']
|
||||
port: 3000
|
||||
@@ -0,0 +1,16 @@
|
||||
11:59:08 === prep romm: deploy at the drill FROM pin
|
||||
11:59:11 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm']
|
||||
11:59:11 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH']
|
||||
11:59:11 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
12:00:16 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}
|
||||
12:00:16 deployed: True
|
||||
12:00:18 romm: POST /api/users http=201
|
||||
12:00:19 romm: login as the seeded user http=200 ok=True
|
||||
12:00:19 C1 A: True
|
||||
12:00:19 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
12:01:19 [4] backup idle; last=None
|
||||
12:01:51 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillbd5863e","email":"drillbd5863e@gate.invalid","enabled":true,"role":"user","permission_group_id"
|
||||
12:01:51 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
12:01:51 B reads back: True
|
||||
12:01:54 named volumes: ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data']
|
||||
port: 8080
|
||||
@@ -0,0 +1,15 @@
|
||||
12:02:01 === prep vikunja: deploy at the drill FROM pin
|
||||
12:02:01 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
12:02:06 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
|
||||
12:02:06 deployed: True
|
||||
12:02:07 vikunja: register http=200
|
||||
12:02:07 vikunja: create project http=201
|
||||
12:02:07 vikunja: readback of the seeded project http=200 ok=True
|
||||
12:02:07 C1 A: True
|
||||
12:02:07 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
12:03:08 [4] backup idle; last=None
|
||||
12:03:08 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drilld918b0","created":"2026-09
|
||||
12:03:08 vikunja B: project readback=True attachment content readback=True
|
||||
12:03:08 B reads back: True
|
||||
12:03:11 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
|
||||
port: 3456
|
||||
@@ -0,0 +1,2 @@
|
||||
11:59:04 [5] drill commit 58a76b596330: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0)
|
||||
11:59:08 badge caught up after 4.5 s
|
||||
@@ -0,0 +1,2 @@
|
||||
12:01:57 [5] drill commit 9f9283d26b22: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0)
|
||||
12:02:01 badge caught up after 4.4 s
|
||||
@@ -0,0 +1,2 @@
|
||||
12:03:14 [5] drill commit 6075526e6b27: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
|
||||
12:03:19 badge caught up after 4.5 s
|
||||
@@ -0,0 +1,20 @@
|
||||
12:03:41 === docmost: live undo by the product
|
||||
12:03:42 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cdb8-9866-7a10-bde0-42e824b55295","name":"drillBbdf31d","description":"","slug":"drillbbdf31d","logo":null,"visibility":"private","defaultRol
|
||||
12:03:42 seed C written right before the Update: True
|
||||
12:03:44 db before: 42 48 ledger, newest 20260620T010047-personal-spaces
|
||||
12:03:44 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
12:03:44 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
12:03:46 + 1.6s phase=pulling label=Új verzió letöltése…
|
||||
12:03:48 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
12:03:50 + 5.3s phase=starting label=Indítás az új verzióval…
|
||||
12:04:01 + 16.3s phase=verifying label=Működés ellenőrzése…
|
||||
12:05:32 + 107.2s phase=undoing label=Visszaállítás az előző változatra…
|
||||
12:06:02 + 137.6s phase=undone label=Visszaállítva az előző változatra
|
||||
12:06:02 END phase=undone err=None hold=None
|
||||
12:06:05 observables: {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": ["docmost docmost/docmost:0.95.0 running=true restarts=0", "docmost-postgres postgres:16-alpine running=true restarts=0", "docmost-redis redis:7-alpine running=true restarts=0"]}
|
||||
12:06:05 docmost: login as the seeded user http=200 ok=True
|
||||
12:06:06 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
12:06:06 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
12:06:08 READBACK A=True B=True C=True db after: 42 48 ledger, newest 20260620T010047-personal-spaces (before: 42 48 ledger, newest 20260620T010047-personal-spaces)
|
||||
12:06:08 PAGE: {"hu": {"undone_line": "A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
|
||||
12:06:10 leftover copies: 0
|
||||
@@ -0,0 +1,20 @@
|
||||
12:06:10 === romm: live undo by the product
|
||||
12:06:12 romm seed B: POST /api/users (as the admin) http=201 {"id":3,"username":"drillb2d189e","email":"drillb2d189e@gate.invalid","enabled":true,"role":"user","permission_group_id"
|
||||
12:06:12 seed C written right before the Update: True
|
||||
12:06:14 db before: tables=24 alembic=0095_virtual_collections_source
|
||||
12:06:14 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
12:06:14 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
12:06:16 + 1.6s phase=pulling label=Új verzió letöltése…
|
||||
12:06:18 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
12:06:24 + 10.0s phase=starting label=Indítás az új verzióval…
|
||||
12:06:36 + 21.0s phase=verifying label=Működés ellenőrzése…
|
||||
12:08:06 + 111.9s phase=undoing label=Visszaállítás az előző változatra…
|
||||
12:08:58 + 163.9s phase=undone label=Visszaállítva az előző változatra
|
||||
12:08:58 END phase=undone err=None hold=None
|
||||
12:09:01 observables: {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": ["romm rommapp/romm:5.0.0 running=true restarts=0", "romm-db mariadb:11.4 running=true restarts=0", "romm-redis redis:7-alpine running=true restarts=0"]}
|
||||
12:09:04 romm: login as the seeded user http=200 ok=True
|
||||
12:09:04 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
12:09:05 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
12:09:07 READBACK A=True B=True C=True db after: tables=24 alembic=0095_virtual_collections_source (before: tables=24 alembic=0095_virtual_collections_source)
|
||||
12:09:07 PAGE: {"hu": {"undone_line": "A(z) romm frissítése 2026-09-23 12:08-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of romm at 2026-09-23 12:08 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
|
||||
12:09:09 leftover copies: 0
|
||||
@@ -0,0 +1,19 @@
|
||||
12:09:09 === vikunja: live undo by the product
|
||||
12:09:09 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drilld918b0","created":"2026-09
|
||||
12:09:09 seed C written right before the Update: True
|
||||
12:09:11 db before: tables=36 ledger=117 newest=SCHEMA_INIT
|
||||
12:09:12 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
12:09:12 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
12:09:13 + 1.1s phase=pulling label=Új verzió letöltése…
|
||||
12:09:14 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
12:09:15 + 3.7s phase=verifying label=Működés ellenőrzése…
|
||||
12:10:46 + 94.5s phase=undoing label=Visszaállítás az előző változatra…
|
||||
12:10:49 + 97.1s phase=undone label=Visszaállítva az előző változatra
|
||||
12:10:49 END phase=undone err=None hold=None
|
||||
12:10:51 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]}
|
||||
12:10:52 vikunja: readback of the seeded project http=200 ok=True
|
||||
12:10:52 vikunja B: project readback=True attachment content readback=True
|
||||
12:10:52 vikunja B: project readback=True attachment content readback=True
|
||||
12:10:54 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT)
|
||||
12:10:54 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 12:10-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 12:10 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
|
||||
12:10:56 leftover copies: 0
|
||||
@@ -0,0 +1,20 @@
|
||||
12:11:18 === romm: power cut DURING the undo
|
||||
12:11:18 romm seed B: POST /api/users (as the admin) http=201 {"id":4,"username":"drillb48b886","email":"drillb48b886@gate.invalid","enabled":true,"role":"user","permission_group_id"
|
||||
12:11:18 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
12:11:18 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
12:11:20 + 1.6s phase=pulling label=Új verzió letöltése…
|
||||
12:11:21 + 2.9s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
12:11:28 + 9.5s phase=starting label=Indítás az új verzióval…
|
||||
12:11:39 + 20.5s phase=verifying label=Működés ellenőrzése…
|
||||
12:13:09 + 111.0s phase=undoing label=Visszaállítás az előző változatra…
|
||||
12:13:15 >>> POWER CUT (pct stop 9202) in phase undoing: rc=0 in 5.2s
|
||||
12:13:18 guest started again: rc=0
|
||||
12:13:31 2026/09/23 10:13:25 update.go:1171: [WARN] [stacks] update recovery: romm was interrupted while UNDOING (started 2026-09-23T10:11:18Z) — marking it Updating and RESUMING the undo
|
||||
2026/09/23 10:13:25 update.go:1224: [INFO] [stacks] update romm: resuming the UNDO after a controller restart
|
||||
|
||||
12:14:22 after the restart: phase=undone hold=None
|
||||
12:14:27 romm: login as the seeded user http=200 ok=True
|
||||
12:14:28 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
12:14:28 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
12:14:31 READBACK A=True B=True C=True db: tables=24 alembic=0095_virtual_collections_source
|
||||
12:14:33 leftover copies: 0
|
||||
@@ -0,0 +1,26 @@
|
||||
12:14:33 === vikunja: a cut-off undo copy
|
||||
12:14:33 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
12:14:33 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
12:14:34 + 1.1s phase=pulling label=Új verzió letöltése…
|
||||
12:14:35 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
12:14:37 + 3.7s phase=starting label=Indítás az új verzióval…
|
||||
12:14:38 + 4.2s phase=verifying label=Működés ellenőrzése…
|
||||
12:14:40 >>> finished-marker removed from one copy:
|
||||
copy: vikunja_vikunja_data.pre-update-20260923T101436Z
|
||||
total 12
|
||||
drwxr-xr-x 3 root root 4096 Sep 23 10:14 .
|
||||
drwxr-xr-x 1 root root 4096 Sep 23 10:14 ..
|
||||
drwxr-xr-x 2 root root 4096 Sep 23 10:13 data
|
||||
-rw-r--r-- 1 root root 0 Sep 23 10:14 felhom-undo-complete
|
||||
data
|
||||
|
||||
12:16:08 + 94.8s phase=undoing label=Visszaállítás az előző változatra…
|
||||
12:16:09 + 95.4s phase=failed label=A frissítés nem sikerült
|
||||
12:16:09 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.'
|
||||
12:16:11 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}}
|
||||
12:16:11 box language -> en: http 302
|
||||
12:16:11 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}}
|
||||
12:16:11 box language -> hu: http 302
|
||||
12:16:13 copies kept: vikunja_vikunja_data.pre-update-20260923T101436Z
|
||||
vikunja_vikunja_db.pre-update-20260923T101436Z
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
12:16:42 === docmost: a person presses Update again after the undo (probe fixed in the catalog)
|
||||
12:16:42 before, the page: {"hu": {"undone_line": "A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}}
|
||||
12:16:42 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
12:16:42 + 0.0s phase=safety-dump label=Adatbázis pillanatkép…
|
||||
12:16:44 + 1.6s phase=pulling label=Új verzió letöltése…
|
||||
12:16:45 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt…
|
||||
12:16:48 + 5.3s phase=starting label=Indítás az új verzióval…
|
||||
12:16:59 + 16.3s phase=verifying label=Működés ellenőrzése…
|
||||
12:17:19 + 36.8s phase=done label=Frissítve
|
||||
12:17:19 END phase=done err=None hold=None
|
||||
12:17:20 docmost: login as the seeded user http=200 ok=True
|
||||
12:17:20 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
12:17:20 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
12:17:23 READBACK A=True B=True C=True installed={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
12:17:23 after, the page: {"hu": {"undone_line": null, "hold": null}, "en": {"undone_line": null, "hold": null}}
|
||||
12:17:25 app.yaml last_update_undone: 0
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
12:17:57 POST /api/stacks/romm/start -> 200 {'ok': True, 'data': {'state': 'running'}, 'message': 'Stack romm start requested — state now: running'}
|
||||
@@ -0,0 +1,50 @@
|
||||
2026/09/23 10:13:25 update.go:1171: [WARN] [stacks] update recovery: romm was interrupted while UNDOING (started 2026-09-23T10:11:18Z) — marking it Updating and RESUMING the undo
|
||||
2026/09/23 10:13:25 update.go:1224: [INFO] [stacks] update romm: resuming the UNDO after a controller restart
|
||||
2026/09/23 10:13:25 update.go:826: [ERROR] [stacks] update romm FAILED after the new version was started: resumed after a restart during the undo
|
||||
2026/09/23 10:13:25 update.go:818: [INFO] [stacks] update romm: kept 24813 bytes of the app's own log at /opt/docker/stacks/romm/hold-logs/20260923T101325Z/compose-logs.txt before stopping it (R-621)
|
||||
2026/09/23 10:13:25 update.go:1106: [INFO] [stacks] update romm: phase undoing
|
||||
2026/09/23 10:13:25 undo.go:357: [WARN] [stacks] update romm: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: resumed after a restart during the undo)
|
||||
2026/09/23 10:14:20 undo.go:404: [INFO] [stacks] update romm: UNDONE in 55s — the previous version is running on the data from before the update (the app's health check passed)
|
||||
2026/09/23 10:14:33 update.go:487: [INFO] [stacks] update vikunja: accepted — guarded update started
|
||||
2026/09/23 10:14:33 update.go:1106: [INFO] [stacks] update vikunja: phase checking
|
||||
2026/09/23 10:14:33 update.go:632: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T10:04:52Z (10m0s old, limit 24h0m0s)
|
||||
2026/09/23 10:14:33 update.go:1106: [INFO] [stacks] update vikunja: phase safety-dump
|
||||
2026/09/23 10:14:33 update.go:664: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) []
|
||||
2026/09/23 10:14:34 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB
|
||||
2026/09/23 10:14:34 update.go:1106: [INFO] [stacks] update vikunja: phase pinning
|
||||
2026/09/23 10:14:34 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0)
|
||||
2026/09/23 10:14:34 update.go:1106: [INFO] [stacks] update vikunja: phase pulling
|
||||
2026/09/23 10:14:35 update.go:1106: [INFO] [stacks] update vikunja: phase copying
|
||||
2026/09/23 10:14:36 update.go:1106: [INFO] [stacks] update vikunja: phase copying
|
||||
2026/09/23 10:14:36 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T101436Z in 469ms
|
||||
2026/09/23 10:14:36 update.go:1106: [INFO] [stacks] update vikunja: phase copying
|
||||
2026/09/23 10:14:37 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T101436Z in 444ms
|
||||
2026/09/23 10:14:37 update.go:1106: [INFO] [stacks] update vikunja: phase starting
|
||||
2026/09/23 10:14:37 update.go:1106: [INFO] [stacks] update vikunja: phase verifying
|
||||
2026/09/23 10:16:07 update.go:826: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy)
|
||||
2026/09/23 10:16:08 update.go:818: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T101607Z/compose-logs.txt before stopping it (R-621)
|
||||
2026/09/23 10:16:08 update.go:1106: [INFO] [stacks] update vikunja: phase undoing
|
||||
2026/09/23 10:16:08 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy))
|
||||
2026/09/23 10:16:08 undo.go:363: [ERROR] [stacks] update vikunja: the undo copy vikunja_vikunja_data.pre-update-20260923T101436Z has no finished-marker — it is cut off or missing; NOTHING is put back
|
||||
2026/09/23 10:16:08 update.go:838: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T101436Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T101436Z}]
|
||||
2026/09/23 10:16:08 update_guard.go:536: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T10:04:52Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched")
|
||||
2026/09/23 10:16:42 update.go:487: [INFO] [stacks] update docmost: accepted — guarded update started
|
||||
2026/09/23 10:16:42 update.go:1106: [INFO] [stacks] update docmost: phase checking
|
||||
2026/09/23 10:16:42 update.go:632: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T10:03:45Z (13m0s old, limit 24h0m0s)
|
||||
2026/09/23 10:16:42 update.go:1106: [INFO] [stacks] update docmost: phase safety-dump
|
||||
2026/09/23 10:16:43 update.go:664: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260923T101642Z-docmost-postgres.sql]
|
||||
2026/09/23 10:16:44 undo.go:252: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 64.8 MiB
|
||||
2026/09/23 10:16:44 update.go:1106: [INFO] [stacks] update docmost: phase pinning
|
||||
2026/09/23 10:16:44 pin.go:365: [INFO] [stacks] update docmost: pin advanced to the catalog's current definition (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:16-alpine, docmost-redis=redis:7-alpine)
|
||||
2026/09/23 10:16:44 update.go:1106: [INFO] [stacks] update docmost: phase pulling
|
||||
2026/09/23 10:16:45 update.go:1106: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/23 10:16:46 update.go:1106: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/23 10:16:46 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260923T101646Z in 673ms
|
||||
2026/09/23 10:16:46 update.go:1106: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/23 10:16:47 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260923T101646Z in 409ms
|
||||
2026/09/23 10:16:47 update.go:1106: [INFO] [stacks] update docmost: phase copying
|
||||
2026/09/23 10:16:47 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260923T101646Z in 415ms
|
||||
2026/09/23 10:16:47 update.go:1106: [INFO] [stacks] update docmost: phase starting
|
||||
2026/09/23 10:16:58 update.go:1106: [INFO] [stacks] update docmost: phase verifying
|
||||
2026/09/23 10:17:19 update.go:785: [INFO] [stacks] update docmost: healthy after 20s (the app's health check passed)
|
||||
2026/09/23 10:17:19 update.go:793: [INFO] [stacks] update docmost: DONE in 37s
|
||||
@@ -0,0 +1,59 @@
|
||||
12:18:27 undo copies before removal: vikunja_vikunja_data.pre-update-20260923T101436Z vikunja_vikunja_db.pre-update-20260923T101436Z
|
||||
12:18:28 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
|
||||
12:19:00 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
|
||||
12:19:07 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
|
||||
12:19:10 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'}
|
||||
12:19:15 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
|
||||
12:19:15 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
|
||||
12:19:42 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat
|
||||
12:19:49 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm'
|
||||
12:19:49 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'}
|
||||
12:20:21 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [
|
||||
12:20:28 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja'
|
||||
12:20:30 undo copies after removal (must be none): 0
|
||||
|
||||
12:20:45 docmost: containers=0 volumes=0
|
||||
romm: containers=0 volumes=0
|
||||
vikunja: containers=0 volumes=0
|
||||
romm drive folder removed
|
||||
removed docmost/docmost:0.95.0
|
||||
removed docmost/docmost:0.96.0
|
||||
removed rommapp/romm:5.0.0
|
||||
removed rommapp/romm:5.3.0
|
||||
removed vikunja/vikunja:2.3.0
|
||||
removed vikunja/vikunja:2.6.0
|
||||
removed nextcloud:34.0.1-apache
|
||||
removed gitea.dooplex.hu/admin/felhom-controller:0.263.0
|
||||
removed gitea.dooplex.hu/admin/felhom-controller:0.263.1
|
||||
controller.yaml.pre-lock-probe
|
||||
dload.out
|
||||
pre-docmost-compose.yml
|
||||
pre-romm-compose.yml
|
||||
pre-vikunja-compose.yml
|
||||
|
||||
git:
|
||||
branch: main
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: ""
|
||||
hub:
|
||||
0
|
||||
|
||||
sync 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'me
|
||||
cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
|
||||
https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
|
||||
controller.yaml == saved pre-bakeoff copy
|
||||
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 48 seconds (healthy)
|
||||
filebrowser gtstef/filebrowser:1.3.3-stable Up 8 minutes (healthy)
|
||||
gokapi f0rc3/gokapi:v1.9.6 Restarting (1) 25 seconds ago
|
||||
paperless-postgres postgres:16-alpine Up 8 minutes (healthy)
|
||||
paperless-redis redis:7-alpine Up 8 minutes (healthy)
|
||||
paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 8 minutes (healthy)
|
||||
privatebin privatebin/pdo:2.0.6 Up 8 minutes (healthy)
|
||||
traefik traefik:v3.6.7 Up 8 minutes
|
||||
|
||||
cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main
|
||||
|
||||
controller.yaml.pre-lock-probe
|
||||
image lines identical
|
||||
@@ -0,0 +1,52 @@
|
||||
# The undo, live on 9202 — controller v0.263.0 → v0.263.2, 2026-09-23
|
||||
|
||||
Venue: scratch guest **9202** (demo-hp), drill catalog (`admin/app-catalog-drill`, reset to live
|
||||
`cfcfe5278428` before and after), `update.health_timeout: 90s`. **Method: endpoint-level** — every
|
||||
act is the endpoint the UI invokes (deploy, backup, sync, rescan, update, start, remove, the language
|
||||
switch) and every page is fetched as HTML (`/apps/<app>?lang=hu|en`); no browser. The product performs
|
||||
the undo; nothing here does. Each edge is a real migrating image move with, in the DRILL template only,
|
||||
a health probe on a port the app does not answer. **Seed A** before the backup, **seed B** after it,
|
||||
**seed C** seconds before the press — C exists only in the undo's own last-second copy.
|
||||
|
||||
## Not done, or changed
|
||||
|
||||
- **Three releases, not one.** v0.263.0 was built, deployed to 9202 only, and its first two live
|
||||
proofs FAILED honestly (HELD, data put back, old version judged "did not start"). Two defects the unit
|
||||
tests could not see; each fixed, red-proofed and released the same session: **v0.263.1** (the undo's
|
||||
probe was gated on `running` while the current probe held the app `unhealthy`) and **v0.263.2** (the
|
||||
"old" `.felhom.yml` saved at update time was already the new one — it flows in on every catalog sync).
|
||||
Only 9202 ever ran 0.263.0/0.263.1; both images were removed from it by name. **No floor was raised.**
|
||||
- **The mail and the operator event** of decision 15 are not built (`09` §6.4 part 2).
|
||||
- **The hold sentence body stays Hungarian on an English box** (R-606); only the new prefix follows
|
||||
the box language.
|
||||
|
||||
## Results (all on v0.263.2 unless stated)
|
||||
|
||||
| proof | result | evidence |
|
||||
|---|---|---|
|
||||
| docmost (PostgreSQL) undone by the product | failed probe at +107 s → `undoing` → **`undone` at +137.6 s**; A, B, C read back; `42 tables, 48 ledger` before and after; 0 copies left | `40-undo-docmost.txt` |
|
||||
| romm (MariaDB) | `undone` at +163.9 s (undo 52 s); A, B, C; `24 tables, alembic 0095` before and after | `40-undo-romm.txt` |
|
||||
| vikunja (SQLite in a volume) | `undone` at +97.1 s (undo 3 s); A, B, C; `36 tables, ledger 117` before and after | `40-undo-vikunja.txt` |
|
||||
| the page line, both languages | hu: *„A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en: *"The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost."* (same for romm, vikunja) | `40-undo-*.txt` |
|
||||
| a cut-off copy (vikunja: the finished-marker taken from one copy during `verifying`) | `undoing` → **`failed` in 0.6 s, nothing poured back**; hold (hu box): *„A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése …"*; en box: *"The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja …"*; both copies kept | `60-cutoff-vikunja.txt` |
|
||||
| a power cut during the undo (romm: `pct stop 9202` the moment the phase read `undoing`) | after boot: *update recovery: romm was interrupted while UNDOING … RESUMING the undo* → **`undone`**; A, B, C; ledger equal; 0 copies left | `50-powercut-romm.txt` |
|
||||
| a person presses Update after an undo (docmost; the drill catalog then fixed the probe) | `done` at +36.8 s on 0.96.0; A, B, C; the undone line gone from both pages; `last_update_undone` gone from `app.yaml` | `70-manual-press-docmost.txt` |
|
||||
| a failed undo holds honestly (v0.263.0, docmost and romm — the defect above) | HOLD *„… Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el."* — true: the data had been put back; the old version was never probed | `10-docmost-undo.txt`, `20-romm-undo.txt` |
|
||||
| removal deletes kept copies (vikunja, held with 2 copies) | 2 → **0** after `POST /api/stacks/vikunja/remove` | `95-teardown.txt` |
|
||||
| R-642, the Start answer | `Stack romm start requested — state now: running` | `80-r642-start-answer.txt` |
|
||||
|
||||
Extra downtime of the copy (the app is stopped there anyway to be recreated): the `copying` phase
|
||||
lasted 2.1 s (docmost), 6.8 s (romm, incl. a 3.8 s stop), 1.6 s (vikunja).
|
||||
|
||||
## Teardown — three layers
|
||||
|
||||
- **machine (9202):** docmost, romm, vikunja removed through the product (no containers, no volumes,
|
||||
no undo copies); romm's drive folder removed by name (the product kept it, R-442 as in every drill);
|
||||
test images removed by name (docmost 0.95.0/0.96.0, romm 5.0.0/5.3.0, vikunja 2.3.0/2.6.0, nextcloud
|
||||
34.0.1-apache, controller 0.263.0/0.263.1); `controller.yaml` restored from the saved copy and read
|
||||
back identical; catalog cache re-cloned from the live repo (`cfcfe52`). **9202 stays on controller
|
||||
0.263.2** (self-update off; the fleet floor is untouched at 0.262.1). Apps afterwards: the same three
|
||||
as at the start of the day (gokapi still crash-looping, R-644).
|
||||
- **host (demo-hp):** nothing provisioned; the guest was stopped and started once for the power cut.
|
||||
- **hub:** nothing touched.
|
||||
- **drill repo:** reset to live `main`, image lines identical, `has_actions: false`.
|
||||
@@ -0,0 +1,266 @@
|
||||
#!/usr/bin/env python3
|
||||
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
|
||||
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
|
||||
drill catalog.
|
||||
|
||||
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
|
||||
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
|
||||
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
|
||||
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
|
||||
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
|
||||
Start supplies the app's secrets (they are never decrypted here).
|
||||
|
||||
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
|
||||
"""
|
||||
import json, os, sys, time
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
|
||||
romm_verify_b, vik_seed_b, vik_verify_b
|
||||
|
||||
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
|
||||
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
|
||||
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
|
||||
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
|
||||
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
|
||||
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
|
||||
SUFFIX = ".pre-undo"
|
||||
|
||||
|
||||
def seed_b(app, sub, A):
|
||||
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
|
||||
|
||||
|
||||
def verify_b(app, sub, A, B):
|
||||
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
|
||||
return all(r.values()) if isinstance(r, dict) else r
|
||||
|
||||
|
||||
def db_state(app):
|
||||
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
|
||||
if app == "docmost":
|
||||
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
|
||||
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
|
||||
if app == "romm":
|
||||
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
|
||||
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
|
||||
alembic = q("select version_num from alembic_version")
|
||||
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
|
||||
return w.guest("""python3 - <<'PY'
|
||||
import sqlite3
|
||||
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
|
||||
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
|
||||
m=c.execute("select count(*), max(id) from migration").fetchone()
|
||||
print(f"tables={t} ledger={m[0]} newest={m[1]}")
|
||||
PY""").strip()
|
||||
|
||||
|
||||
def volumes(app):
|
||||
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
|
||||
if not v.endswith(SUFFIX)]
|
||||
|
||||
|
||||
def containers(app):
|
||||
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
|
||||
|
||||
|
||||
def stage_prep(app):
|
||||
s = {"app": app, "sub": SUB[app]}
|
||||
w.say(f"=== prep {app}: deploy at the drill FROM pin")
|
||||
w.say("deployed:", w.deploy(app, s["sub"]))
|
||||
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
|
||||
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
|
||||
w.backup_now(app)
|
||||
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
|
||||
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
|
||||
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
|
||||
s["volumes"] = volumes(app)
|
||||
w.say("named volumes:", s["volumes"])
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_fcopy(app):
|
||||
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
|
||||
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
|
||||
s = load(app)
|
||||
s["state_before_update"] = db_state(app)
|
||||
w.say("db state before the update:", s["state_before_update"])
|
||||
vols = s["volumes"]
|
||||
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
|
||||
script = f"""
|
||||
set -u
|
||||
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
|
||||
t0=$(date +%s.%N)
|
||||
docker stop $cs >/dev/null
|
||||
t1=$(date +%s.%N)
|
||||
for v in {' '.join(vols)}; do
|
||||
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
|
||||
docker volume create "$v{SUFFIX}" >/dev/null
|
||||
a=$(date +%s.%N)
|
||||
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
|
||||
b=$(date +%s.%N)
|
||||
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
|
||||
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
|
||||
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
|
||||
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
|
||||
mkdir -p /root/bk; c=$(date +%s.%N)
|
||||
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
|
||||
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
|
||||
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
|
||||
done
|
||||
t2=$(date +%s.%N)
|
||||
docker start $cs >/dev/null
|
||||
t3=$(date +%s.%N)
|
||||
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
|
||||
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
|
||||
"""
|
||||
out = w.guest(script, timeout=1800)
|
||||
w.say(out)
|
||||
t0 = time.time()
|
||||
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
|
||||
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
|
||||
s["fcopy"] = out
|
||||
s["fcopy_at"] = ts()
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_break(app):
|
||||
s = load(app)
|
||||
frm, to, p0, p1 = EDGE[app]
|
||||
s["break_commit"] = break_edge(app, frm, to, p0, p1)
|
||||
w.say("badge caught up after", w.sync_rescan(app, to), "s")
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_update(app):
|
||||
s = load(app)
|
||||
r = w.press_update(app, poll=1.0)
|
||||
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
|
||||
s.setdefault("updates", []).append(r)
|
||||
save(app, s)
|
||||
|
||||
|
||||
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
|
||||
docker restart felhom-controller >/dev/null; sleep 15"""
|
||||
|
||||
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
|
||||
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
|
||||
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
|
||||
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
|
||||
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
|
||||
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
|
||||
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
|
||||
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
|
||||
|
||||
|
||||
def lift_and_pinback(app):
|
||||
frm, to, _, _ = EDGE[app]
|
||||
t = time.time()
|
||||
w.say(w.guest(LIFT.format(app=app)))
|
||||
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
|
||||
w.login()
|
||||
return round(time.time() - t, 1)
|
||||
|
||||
|
||||
def stage_undoF(app):
|
||||
s = load(app)
|
||||
w.say(f"=== undo by F: {app}")
|
||||
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
|
||||
vols = s["volumes"]
|
||||
script = f"""
|
||||
t0=$(date +%s.%N)
|
||||
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
|
||||
for v in {' '.join(vols)}; do
|
||||
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
|
||||
done
|
||||
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
|
||||
"""
|
||||
w.say(w.guest(script, timeout=1800))
|
||||
t0 = time.time()
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
|
||||
w.say("product start ->", c)
|
||||
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
|
||||
th = round(time.time() - t0, 1)
|
||||
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
|
||||
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
|
||||
st = db_state(app)
|
||||
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
|
||||
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_undoD(app):
|
||||
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
|
||||
s = load(app)
|
||||
w.say(f"=== undo by D: {app}")
|
||||
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
|
||||
if app == "vikunja":
|
||||
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
|
||||
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
|
||||
return
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
|
||||
time.sleep(8)
|
||||
if app == "docmost":
|
||||
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
|
||||
marker = "-- PostgreSQL database dump complete"
|
||||
else:
|
||||
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
|
||||
marker = "-- Dump completed"
|
||||
script = f"""
|
||||
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
|
||||
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
|
||||
t0=$(date +%s.%N)
|
||||
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
|
||||
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
|
||||
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
|
||||
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
|
||||
"""
|
||||
w.say(w.guest(script, timeout=900))
|
||||
t0 = time.time()
|
||||
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
|
||||
th = round(time.time() - t0, 1)
|
||||
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
|
||||
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
|
||||
st = db_state(app)
|
||||
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
|
||||
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_cutoff(app):
|
||||
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
|
||||
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
|
||||
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
|
||||
s = load(app)
|
||||
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
|
||||
script = f"""
|
||||
v={big}
|
||||
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
|
||||
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
|
||||
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
|
||||
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
|
||||
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
|
||||
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
|
||||
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
|
||||
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
|
||||
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
|
||||
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
|
||||
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
|
||||
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
|
||||
rm -f /root/cut.sql
|
||||
else echo "D: no safety dump exists for this app (R-641)"; fi
|
||||
"""
|
||||
w.say(f"=== cut-off copy: {app} (largest volume {big})")
|
||||
out = w.guest(script, timeout=600)
|
||||
w.say(out)
|
||||
s["cutoff"] = out
|
||||
save(app, s)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
|
||||
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
|
||||
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,194 @@
|
||||
#!/usr/bin/env python3
|
||||
"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes.
|
||||
|
||||
Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/<n>,
|
||||
fetches the app page in both languages, and reads the seeds back through each app's own front door.
|
||||
Seed C is written immediately before each Update, after every backup — so only the undo's own
|
||||
last-second copy can bring it back.
|
||||
"""
|
||||
import json, re, sys, time, html as H
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
from spike import load, save, ts
|
||||
import bakeoff as b
|
||||
|
||||
HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}
|
||||
|
||||
|
||||
def seed_c(app, s):
|
||||
return b.seed_b(app, s["sub"], s["seedA"])
|
||||
|
||||
|
||||
def page_lines(app):
|
||||
out = {}
|
||||
for lang in ("hu", "en"):
|
||||
h = H.unescape(w.page(f"/apps/{app}?lang={lang}"))
|
||||
m = re.search(r'data-update-undone="true">([^<]*)<', h)
|
||||
hold = re.search(r'data-held="true">([^<]*)<', h)
|
||||
out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None}
|
||||
return out
|
||||
|
||||
|
||||
def press(app, poll=0.5, on_phase=None):
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/update")
|
||||
w.say(f" Update -> {code} {str(d)[:140]}")
|
||||
phases, seen, t0 = [], None, time.time()
|
||||
while time.time() - t0 < 1500:
|
||||
try:
|
||||
st = w.stack(app)
|
||||
except Exception as e:
|
||||
st = {}
|
||||
ph = st.get("update_phase")
|
||||
if ph != seen and ph is not None:
|
||||
seen = ph
|
||||
phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label")))
|
||||
w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}")
|
||||
if on_phase and on_phase(ph):
|
||||
return phases, "interrupted"
|
||||
if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2:
|
||||
break
|
||||
time.sleep(poll)
|
||||
st = w.stack(app)
|
||||
w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}")
|
||||
return phases, st
|
||||
|
||||
|
||||
def readback(app, s, with_c=True):
|
||||
sub = s["sub"]
|
||||
w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2)
|
||||
A = b.FX[app].verify(w, sub, s["seedA"], w.say)
|
||||
B = b.verify_b(app, sub, s["seedA"], s["seedB"])
|
||||
C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None
|
||||
return {"A": A, "B": B, "C": C}
|
||||
|
||||
|
||||
def stage_undo(app):
|
||||
s = load(app)
|
||||
w.say(f"=== {app}: live undo by the product")
|
||||
s["seedC"] = seed_c(app, s); s["seedC_at"] = ts()
|
||||
w.say(" seed C written right before the Update:", bool(s["seedC"]))
|
||||
before = b.db_state(app); w.say(" db before:", before)
|
||||
phases, st = press(app)
|
||||
obs = w.observables(app)
|
||||
w.say(" observables:", json.dumps(obs))
|
||||
rb = readback(app, s)
|
||||
after = b.db_state(app)
|
||||
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})")
|
||||
pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False))
|
||||
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
|
||||
s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")},
|
||||
"readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs}
|
||||
save(app, s)
|
||||
|
||||
|
||||
|
||||
|
||||
def set_box_language(lang):
|
||||
"""POST /settings/language — the household's language switch (form + the dashboard's CSRF)."""
|
||||
sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip()
|
||||
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}",
|
||||
"-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}",
|
||||
f"{w.BASE}/settings/language"])
|
||||
w.say(f" box language -> {lang}: http {r.stdout.strip()}")
|
||||
|
||||
|
||||
def stage_powercut(app):
|
||||
"""Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard
|
||||
stop); boot it again and let the controller resume the undo."""
|
||||
import subprocess
|
||||
s = load(app)
|
||||
w.say(f"=== {app}: power cut DURING the undo")
|
||||
s["seedC"] = seed_c(app, s)
|
||||
cut = {}
|
||||
|
||||
def on_phase(ph):
|
||||
if ph == "undoing":
|
||||
t = time.time()
|
||||
r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120)
|
||||
cut["at"] = ts(); cut["rc"] = r.returncode
|
||||
w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s")
|
||||
return True
|
||||
return False
|
||||
phases, _ = press(app, poll=0.3, on_phase=on_phase)
|
||||
if not cut:
|
||||
w.say(" the cut never landed in `undoing` — recorded as a MISS"); return
|
||||
r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180)
|
||||
w.say(f" guest started again: rc={r.returncode}")
|
||||
for i in range(60):
|
||||
time.sleep(5)
|
||||
try:
|
||||
w.login(); st = w.stack(app)
|
||||
if st:
|
||||
break
|
||||
except SystemExit:
|
||||
continue
|
||||
w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6"))
|
||||
t0 = time.time()
|
||||
while time.time() - t0 < 600:
|
||||
st = w.stack(app)
|
||||
if not st.get("updating") and st.get("update_phase") in ("undone", "failed"):
|
||||
break
|
||||
time.sleep(3)
|
||||
w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}")
|
||||
rb = readback(app, s)
|
||||
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}")
|
||||
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
|
||||
s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_cutoff(app):
|
||||
"""Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of
|
||||
the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it
|
||||
back and HOLD with the new prefix."""
|
||||
s = load(app)
|
||||
w.say(f"=== {app}: a cut-off undo copy")
|
||||
done = {}
|
||||
|
||||
def on_phase(ph):
|
||||
if ph == "verifying" and not done:
|
||||
out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c"
|
||||
docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""")
|
||||
done["out"] = out
|
||||
w.say(" >>> finished-marker removed from one copy:\n" + out)
|
||||
return False
|
||||
phases, st = press(app, poll=0.5, on_phase=on_phase)
|
||||
before = b.db_state(app)
|
||||
pl = page_lines(app)
|
||||
w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False))
|
||||
set_box_language("en")
|
||||
pl_en = page_lines(app)
|
||||
w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False))
|
||||
set_box_language("hu")
|
||||
w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}"))
|
||||
s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_fixprobe_and_press(app):
|
||||
"""After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work
|
||||
and end the undone note."""
|
||||
s = load(app)
|
||||
frm, to, p0, p1 = b.EDGE[app]
|
||||
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
|
||||
f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f)
|
||||
w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"])
|
||||
w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120)
|
||||
w.sync_rescan(app, to)
|
||||
time.sleep(20)
|
||||
w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan")
|
||||
w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)")
|
||||
w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False))
|
||||
phases, st = press(app)
|
||||
rb = readback(app, s)
|
||||
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}")
|
||||
pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False))
|
||||
w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml"))
|
||||
s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl}
|
||||
save(app, s)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
{"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff,
|
||||
"fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2])
|
||||
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
|
||||
|
||||
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
|
||||
`controller.yaml.pre-bakeoff` (NOT the older `.pre-28`, which a restore must never pick up).
|
||||
"""
|
||||
import re, sys, io
|
||||
sys.path.insert(0, '.')
|
||||
import walk as w
|
||||
|
||||
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
|
||||
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
|
||||
|
||||
|
||||
def creds():
|
||||
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
|
||||
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
|
||||
if m:
|
||||
return m.group(1), m.group(2)
|
||||
raise SystemExit("no admin credential")
|
||||
|
||||
|
||||
def to_drill():
|
||||
u, t = creds()
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
test -f {VOL}/controller.yaml.pre-bakeoff || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-bakeoff
|
||||
python3 - <<'PY'
|
||||
import re
|
||||
p = "{VOL}/controller.yaml"
|
||||
s = open(p).read()
|
||||
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
|
||||
if not re.search(r'^update:', s, re.M):
|
||||
s += "update:\\n health_timeout: 90s\\n"
|
||||
open(p, "w").write(s)
|
||||
PY
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -A2 '^update:' {VOL}/controller.yaml
|
||||
"""))
|
||||
|
||||
|
||||
def restore():
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
cp -p {VOL}/controller.yaml.pre-bakeoff {VOL}/controller.yaml
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -c '^update:' {VOL}/controller.yaml || true
|
||||
"""))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
to_drill() if sys.argv[1] == "drill" else restore()
|
||||
@@ -0,0 +1,20 @@
|
||||
set -u
|
||||
python3 - <<'PY' > 30-reset-apps.txt 2>&1
|
||||
import walk as w, time
|
||||
w.login()
|
||||
for app in ("docmost","romm","vikunja"):
|
||||
w.remove(app)
|
||||
print(w.guest("docker volume ls -q --filter label=felhom.undo-copy-of | wc -l; for a in docmost romm vikunja; do docker volume ls -q --filter label=felhom.undo-copy-of=$a; done; rm -rf /mnt/felhom-drives/scratch_hdd/userdata/romm && echo romm-drive-folder-removed"))
|
||||
for i in range(10):
|
||||
c,d=w.ctl("POST","/api/sync")
|
||||
if c=="200": print("sync",c); break
|
||||
time.sleep(15)
|
||||
w.ctl("POST","/api/stacks/rescan")
|
||||
print(w.guest("C=/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache; git -C $C log --oneline -1"))
|
||||
PY
|
||||
for app in docmost romm vikunja; do
|
||||
timeout 1200 python3 bakeoff.py prep $app > 31-prep-$app.txt 2>&1; echo "prep $app rc=$?"
|
||||
echo "cat /opt/docker/stacks/$app/applied-meta/.felhom.yml 2>&1 | grep -A3 healthcheck | grep port" | /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad/g.sh >> 31-prep-$app.txt 2>&1
|
||||
timeout 600 python3 bakeoff.py break $app > 32-break-$app.txt 2>&1; echo "break $app rc=$?"
|
||||
done
|
||||
echo SETUP-DONE
|
||||
@@ -0,0 +1,160 @@
|
||||
#!/usr/bin/env python3
|
||||
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
|
||||
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
|
||||
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
|
||||
steps the product would take, each one timed.
|
||||
|
||||
State between stages lives in state-<app>.json so each stage can be run, read, and only then
|
||||
followed by the next (the hand undo needs a person looking at what the previous step left).
|
||||
"""
|
||||
import json, os, sys, time, re
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
import fixtures as fx
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
|
||||
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
|
||||
|
||||
|
||||
def st_path(app):
|
||||
return os.path.join(HERE, f"state-{app}.json")
|
||||
|
||||
|
||||
def load(app):
|
||||
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
|
||||
|
||||
|
||||
def save(app, s):
|
||||
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
|
||||
|
||||
|
||||
def ts():
|
||||
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||
|
||||
|
||||
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
|
||||
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
|
||||
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
|
||||
def docmost_seed_b(sub, A):
|
||||
jar = "/tmp/dm.jar"
|
||||
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
|
||||
name = "drillB" + os.urandom(3).hex()
|
||||
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
|
||||
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
|
||||
return {"space": name} if code in ("200", "201") else None
|
||||
|
||||
|
||||
def docmost_verify_b(sub, A, B):
|
||||
jar = "/tmp/dm.jar"
|
||||
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
|
||||
if code not in ("200", "201"):
|
||||
w.say(f" docmost B: cannot log in (http={code})"); return False
|
||||
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
|
||||
data="{}", method="POST")
|
||||
ok = code in ("200", "201") and B["space"] in out
|
||||
neg = "drillBnever" in out
|
||||
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
|
||||
return ok and not neg
|
||||
|
||||
|
||||
def break_edge(app, frm, to, port_from, port_to):
|
||||
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
|
||||
health probe pointed at a port the app does not answer. Both in one drill commit."""
|
||||
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
|
||||
f = open(fy).read()
|
||||
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
|
||||
assert m, "probe port not found"
|
||||
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
|
||||
open(fy, "w").write(f)
|
||||
h = w.drill_bump(app, frm, to)
|
||||
return h
|
||||
|
||||
|
||||
def pg_state(container, db, user):
|
||||
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
|
||||
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
|
||||
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
|
||||
|
||||
|
||||
def romm_seed_b(sub, A):
|
||||
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
|
||||
jar, tok = FX["romm"]._csrf(w, sub)
|
||||
user = "drillb" + os.urandom(3).hex()
|
||||
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
|
||||
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
|
||||
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
|
||||
method="POST")
|
||||
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
|
||||
return {"user": user} if code in ("200", "201") else None
|
||||
|
||||
|
||||
def romm_verify_b(sub, A, B):
|
||||
jar, tok = FX["romm"]._csrf(w, sub)
|
||||
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
|
||||
"-u", f"{A['user']}:{A['pw']}")
|
||||
ok = code == "200" and B["user"] in body
|
||||
neg = "drillbnever" in body
|
||||
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
|
||||
return ok and not neg
|
||||
|
||||
|
||||
def tree(paths):
|
||||
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
|
||||
return w.guest(cmd)
|
||||
|
||||
|
||||
def my_state(container="romm-db", db="romm"):
|
||||
# The root password is used INSIDE the container from its own env — it never leaves it.
|
||||
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
|
||||
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
|
||||
echo -n "alembic=$({q("select version_num from alembic_version")}) "
|
||||
echo -n "users=$({q("select count(*) from users")})"
|
||||
""")
|
||||
|
||||
|
||||
def vik_seed_b(sub, A):
|
||||
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
|
||||
app writes into its files volume), uploaded through the app's own attachment API."""
|
||||
tok, why = FX["vikunja"]._token(w, sub, A)
|
||||
title = "drillB-" + os.urandom(4).hex()
|
||||
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
|
||||
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
|
||||
if code not in ("200", "201"):
|
||||
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
|
||||
pid = json.loads(body)["id"]
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
|
||||
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
|
||||
tid = json.loads(body)["id"] if code in ("200", "201") else None
|
||||
content = "drill attachment " + os.urandom(8).hex()
|
||||
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
|
||||
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
|
||||
"-F", f"files=@{fn}", method="PUT")
|
||||
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
|
||||
return {"title": title, "pid": pid, "tid": tid, "att": content}
|
||||
|
||||
|
||||
def vik_verify_b(sub, A, B):
|
||||
tok, why = FX["vikunja"]._token(w, sub, A)
|
||||
if not tok:
|
||||
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
|
||||
proj = code == "200" and B["title"] in body
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
|
||||
att_ok = False
|
||||
try:
|
||||
atts = json.loads(body)
|
||||
if atts:
|
||||
aid = atts[0]["id"]
|
||||
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
|
||||
att_ok = c3 == "200" and B["att"] in b3
|
||||
except Exception as e:
|
||||
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
|
||||
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
|
||||
return {"project": proj, "attachment": att_ok}
|
||||
@@ -0,0 +1,493 @@
|
||||
#!/usr/bin/env python3
|
||||
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
|
||||
POST /api/stacks/<n>/deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan
|
||||
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
|
||||
and reads GET /api/stacks/<n>. No controller code exists for it.
|
||||
|
||||
The walk, per `09` §6.4 and the update-night brief §4:
|
||||
1 deploy from the DRILL catalog at the LIVE pin
|
||||
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
|
||||
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
|
||||
4 „Mentés most"
|
||||
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
|
||||
6 press the guarded Update, record every phase with timestamps
|
||||
7 read the seed back through the front door
|
||||
8 the four version observables side by side
|
||||
9 write the verdict record in `09`'s JSON shape
|
||||
|
||||
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
|
||||
"""
|
||||
import argparse, json, os, re, subprocess, sys, time
|
||||
from datetime import datetime, timezone
|
||||
|
||||
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
|
||||
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-live-2026-09-23"
|
||||
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
|
||||
BASE = "https://192.168.0.114"
|
||||
HOSTHDR = "Host: felhom.enkisfelhom.hu"
|
||||
DOMAIN = "enkisfelhom.hu"
|
||||
HP = "demo-hp"
|
||||
|
||||
LOG = []
|
||||
|
||||
|
||||
def say(*a):
|
||||
line = " ".join(str(x) for x in a)
|
||||
ts = datetime.now().strftime("%H:%M:%S")
|
||||
print(f"{ts} {line}", flush=True)
|
||||
LOG.append(f"{ts} {line}")
|
||||
|
||||
|
||||
def sh(args, timeout=300, inp=None):
|
||||
try:
|
||||
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
|
||||
except (subprocess.TimeoutExpired, OSError) as e:
|
||||
return subprocess.CompletedProcess(args, 124, "", f"{e}")
|
||||
|
||||
|
||||
def guest(script, timeout=600):
|
||||
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
|
||||
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
|
||||
"cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; "
|
||||
"pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"],
|
||||
timeout=timeout, inp=script)
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def login():
|
||||
pw = open(f"{SC}/.ctlpw").read().strip()
|
||||
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
|
||||
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
|
||||
h = open(f"{SC}/hdr.txt").read()
|
||||
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
|
||||
if not m:
|
||||
sys.exit("login failed: no session cookie")
|
||||
open(f"{SC}/sess.txt", "w").write(m.group(0))
|
||||
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
|
||||
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
|
||||
if not c:
|
||||
sys.exit("login failed: no csrf token")
|
||||
open(f"{SC}/csrf.txt", "w").write(c.group(1))
|
||||
|
||||
|
||||
def ctl(method, path, data=None, raw=False, tries=2):
|
||||
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
|
||||
in memory, so any controller restart during the night invalidates it silently."""
|
||||
for attempt in range(tries):
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf.txt").read().strip()
|
||||
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
|
||||
if method != "GET":
|
||||
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
|
||||
"-X", method]
|
||||
if data is not None:
|
||||
args += ["--data", json.dumps(data)]
|
||||
args.append(f"{BASE}{path}")
|
||||
r = sh(args)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
if code.strip() in ("302", "401") and attempt + 1 < tries:
|
||||
login()
|
||||
continue
|
||||
if raw:
|
||||
return code.strip(), body
|
||||
try:
|
||||
return code.strip(), json.loads(body)
|
||||
except Exception:
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
|
||||
|
||||
def page(path):
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
|
||||
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
|
||||
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
|
||||
"-w", "\n%{http_code}"]
|
||||
if method:
|
||||
args += ["-X", method]
|
||||
if data is not None:
|
||||
args += ["--data-binary", "@-"]
|
||||
args += list(extra) + [f"{BASE}{path}"]
|
||||
r = sh(args, timeout=timeout + 30, inp=data)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
return r.returncode, code.strip(), body
|
||||
|
||||
|
||||
def stack(name):
|
||||
_, d = ctl("GET", f"/api/stacks/{name}")
|
||||
return (d.get("data") or {}) if isinstance(d, dict) else {}
|
||||
|
||||
|
||||
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
|
||||
"""Settling says the container runs; this says the APP answers. Not the same thing."""
|
||||
last = None
|
||||
for _ in range(tries):
|
||||
rc, code, _ = app_curl(sub, path, timeout=15)
|
||||
last = (rc, code)
|
||||
if rc == 0 and code in want:
|
||||
return True
|
||||
time.sleep(delay)
|
||||
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
|
||||
return False
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the walk
|
||||
|
||||
|
||||
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
|
||||
|
||||
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
|
||||
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
|
||||
# password back off the box — the household sees it once. So the value the harness itself
|
||||
# generated is kept here for the life of the run, and nowhere else.
|
||||
GENERATED = {}
|
||||
|
||||
|
||||
def deploy_values(name, sub):
|
||||
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
|
||||
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
|
||||
|
||||
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
|
||||
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
|
||||
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
|
||||
reason this was safe to discover by running it (live-probes rule).
|
||||
|
||||
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
|
||||
the scratch drive first — the same act the drive browser performs for a household.
|
||||
"""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
|
||||
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
|
||||
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
|
||||
made = []
|
||||
for f in fields:
|
||||
ev, ty = f.get("env_var"), f.get("type")
|
||||
if ev in values:
|
||||
continue
|
||||
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
|
||||
# when the caller sends none, deliberately ("the user needs to know their password"),
|
||||
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
|
||||
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
|
||||
if not f.get("required") and ty != "password":
|
||||
continue # the controller generates the optional secrets itself
|
||||
if ty == "path":
|
||||
p = f"{DRIVE}/{name}"
|
||||
values[ev] = p
|
||||
made.append(p)
|
||||
elif ty in ("secret", "password"):
|
||||
import secrets as _s
|
||||
values[ev] = "Drill-" + _s.token_hex(12)
|
||||
GENERATED.setdefault(name, {})[ev] = values[ev]
|
||||
elif f.get("default"):
|
||||
values[ev] = f["default"]
|
||||
else:
|
||||
values[ev] = f"drill-{name}"
|
||||
if made:
|
||||
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
|
||||
say(f" [1] made the drive paths this app requires: {made}")
|
||||
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
|
||||
if extra:
|
||||
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
|
||||
return values
|
||||
|
||||
|
||||
def deploy(name, sub, extra_values=None):
|
||||
st = stack(name)
|
||||
if st.get("deployed"):
|
||||
say(f" [1] {name} already deployed — reusing")
|
||||
return True
|
||||
values = deploy_values(name, sub)
|
||||
if extra_values:
|
||||
values.update(extra_values)
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
|
||||
say(f" [1] deploy -> {code} {str(d)[:120]}")
|
||||
if code != "202":
|
||||
return False
|
||||
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
|
||||
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
|
||||
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
|
||||
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
|
||||
seen = None
|
||||
for _ in range(90):
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
seen = st.get("state")
|
||||
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
|
||||
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
|
||||
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
|
||||
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
|
||||
# (`runComposeDeploy` writes it), so that is what to wait for.
|
||||
pins = (st.get("app_config") or {}).get("pinned_images")
|
||||
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
|
||||
say(f" [1] deployed, controller state={seen}, "
|
||||
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
|
||||
if seen != "running":
|
||||
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
|
||||
f"not treated as a failure; the fixture's front-door wait is the real gate")
|
||||
return True
|
||||
say(f" [1] never became deployed (last controller state={seen!r})")
|
||||
return False
|
||||
|
||||
|
||||
def backup_now(name):
|
||||
code, d = ctl("POST", "/api/backup/run")
|
||||
say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}")
|
||||
for _ in range(90):
|
||||
time.sleep(5)
|
||||
c2, s = ctl("GET", "/api/backup/status")
|
||||
dd = s.get("data") or {}
|
||||
if not dd.get("running", False):
|
||||
say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}")
|
||||
return True
|
||||
say(" [4] backup still running after 7.5 min — carrying on")
|
||||
return False
|
||||
|
||||
|
||||
def drill_bump(app, frm, to, service_hint=None):
|
||||
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
|
||||
|
||||
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
|
||||
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
|
||||
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
|
||||
every service that carries the app's own version and no others.
|
||||
"""
|
||||
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
|
||||
fy = f"{DRILL}/templates/{app}/.felhom.yml"
|
||||
s = open(comp).read()
|
||||
froms = [x.strip() for x in frm.split(",") if x.strip()]
|
||||
tos = [x.strip() for x in to.split(",") if x.strip()]
|
||||
if len(froms) != len(tos):
|
||||
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
|
||||
return None
|
||||
for f1, t1 in zip(froms, tos):
|
||||
if f"image: {f1}" not in s:
|
||||
say(f" [5] FROM ref not found in compose: {f1}")
|
||||
return None
|
||||
s = s.replace(f"image: {f1}", f"image: {t1}")
|
||||
open(comp, "w").write(s)
|
||||
f = open(fy).read()
|
||||
today = datetime.now().strftime("%Y-%m-%d")
|
||||
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
|
||||
open(fy, "w").write(f)
|
||||
sh(["git", "-C", DRILL, "add", "-A"])
|
||||
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
|
||||
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
|
||||
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
|
||||
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
|
||||
return h
|
||||
|
||||
|
||||
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
|
||||
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
|
||||
|
||||
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
|
||||
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
|
||||
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
|
||||
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
|
||||
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
|
||||
NUMBER R-607 asks for and has never had.
|
||||
"""
|
||||
t0 = time.time()
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(2)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
time.sleep(2)
|
||||
if not expect_app or not expect_ref:
|
||||
return None
|
||||
for i in range(tries):
|
||||
cat = stack(expect_app).get("catalog_images") or {}
|
||||
if expect_ref in cat.values():
|
||||
waited = round(time.time() - t0, 1)
|
||||
if i:
|
||||
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
|
||||
f"to {expect_ref} — R-607's window, measured")
|
||||
return waited
|
||||
time.sleep(delay)
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(1)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
|
||||
f"catalog_images = {stack(expect_app).get('catalog_images')}")
|
||||
return None
|
||||
|
||||
|
||||
def badges(name):
|
||||
out = {}
|
||||
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
|
||||
h = page(f"/apps/{name}{suffix}")
|
||||
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
|
||||
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
|
||||
return out
|
||||
|
||||
|
||||
def press_update(name, poll=1.0, cap_s=1800):
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/update")
|
||||
say(f" [6] Update -> {code} {str(d)[:220]}")
|
||||
if code not in ("202", "200"):
|
||||
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
|
||||
phases, seen, t0 = [], None, time.time()
|
||||
while time.time() - t0 < cap_s:
|
||||
st = stack(name)
|
||||
ph = st.get("update_phase")
|
||||
if ph != seen:
|
||||
seen = ph
|
||||
rec = {"t": round(time.time() - t0, 1), "phase": ph,
|
||||
"label": st.get("update_phase_label"), "updating": st.get("updating"),
|
||||
"error": st.get("update_error"), "hold": st.get("hold_reason")}
|
||||
phases.append(rec)
|
||||
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
|
||||
f"err={rec['error']} hold={rec['hold']}")
|
||||
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
|
||||
break
|
||||
time.sleep(poll)
|
||||
st = stack(name)
|
||||
return {"accepted": True, "http": code, "phases": phases,
|
||||
"duration_s": round(time.time() - t0, 1),
|
||||
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
|
||||
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
|
||||
|
||||
|
||||
def observables(name):
|
||||
st = stack(name)
|
||||
ac = st.get("app_config") or {}
|
||||
live = guest(f"""
|
||||
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
|
||||
echo '---inspect---'
|
||||
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
|
||||
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
|
||||
done
|
||||
""")
|
||||
a, _, b = live.partition("---inspect---")
|
||||
return {
|
||||
"pinned_images": ac.get("pinned_images"),
|
||||
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
|
||||
for k, v in (ac.get("installed_images") or {}).items()},
|
||||
"catalog_images": st.get("catalog_images"),
|
||||
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
|
||||
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
|
||||
}
|
||||
|
||||
|
||||
def app_logs(name, lines=400):
|
||||
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
|
||||
one string with escaped newlines — a scan over the envelope sees a single enormous line and
|
||||
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
|
||||
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
|
||||
if isinstance(d, dict):
|
||||
data = d.get("data")
|
||||
if isinstance(data, dict) and isinstance(data.get("logs"), str):
|
||||
return data["logs"]
|
||||
if isinstance(d.get("_raw"), str):
|
||||
return d["_raw"]
|
||||
return str(d)
|
||||
|
||||
|
||||
def write_verdict(rec, appdir):
|
||||
os.makedirs(appdir, exist_ok=True)
|
||||
p = os.path.join(appdir, "verdict.json")
|
||||
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
|
||||
say(f" [9] verdict {rec['verdict']} -> {p}")
|
||||
|
||||
|
||||
def remove(name):
|
||||
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
|
||||
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
|
||||
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
|
||||
say(f" [X] stop -> {c1} {str(d1)[:100]}")
|
||||
for _ in range(24):
|
||||
time.sleep(5)
|
||||
if stack(name).get("state") != "running":
|
||||
break
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": True, "remove_backups": True})
|
||||
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
|
||||
if code == "409":
|
||||
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
|
||||
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
|
||||
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
|
||||
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
|
||||
# harness takes it, then tidies its own directory by name at teardown.
|
||||
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
|
||||
" — removing the app and KEEPING the drive data instead")
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": False, "remove_backups": True})
|
||||
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
|
||||
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
|
||||
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
|
||||
return code
|
||||
|
||||
|
||||
def app_env(name, key):
|
||||
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
|
||||
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
|
||||
shows them the same value. Data still goes in through the app's own front door."""
|
||||
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
|
||||
if ":" in out:
|
||||
return out.split(":", 1)[1].strip().strip('"').strip("'")
|
||||
return ""
|
||||
|
||||
|
||||
def snapshots(name):
|
||||
"""The restorable copies the backups page offers for this app."""
|
||||
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
|
||||
data = d.get("data") if isinstance(d, dict) else None
|
||||
if isinstance(data, dict):
|
||||
for k in ("snapshots", "items", "restore_points"):
|
||||
if isinstance(data.get(k), list):
|
||||
return data[k]
|
||||
return data if isinstance(data, list) else []
|
||||
|
||||
|
||||
def restore(name, snapshot_id=None, wait_s=1200):
|
||||
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
|
||||
|
||||
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
|
||||
`snapshot_id` — because that is the button the sentence tells them to press.
|
||||
"""
|
||||
snaps = snapshots(name)
|
||||
if snapshot_id is None:
|
||||
if not snaps:
|
||||
say(f" [R] no restorable copy offered for {name}")
|
||||
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
|
||||
first = snaps[0]
|
||||
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
|
||||
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
|
||||
"-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}",
|
||||
"--data-urlencode", f"stack_name={name}",
|
||||
"--data-urlencode", f"snapshot_id={snapshot_id}",
|
||||
f"{BASE}/backup/restore"], timeout=180)
|
||||
head = (r.stdout or "").split("\n")[0].strip()
|
||||
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
|
||||
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
|
||||
t0 = time.time()
|
||||
last = None
|
||||
while time.time() - t0 < wait_s:
|
||||
code, d = ctl("GET", "/api/backup/restore-status")
|
||||
dd = d.get("data") or {}
|
||||
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
|
||||
if cur != last:
|
||||
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
|
||||
last = cur
|
||||
if not dd.get("running", False) and time.time() - t0 > 5:
|
||||
break
|
||||
time.sleep(2)
|
||||
st = stack(name)
|
||||
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
|
||||
f"phase={st.get('update_phase')}")
|
||||
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
|
||||
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
|
||||
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
|
||||
"observables_after": observables(name)}
|
||||
@@ -335,3 +335,7 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-539** | **The restart brake caught a FAST crash loop and was blind to a SLOW one (P3).** Closed by operator ruling 3 of 2026-09-16 in agent **v0.132.0** (`18d03bd`, tag `v0.132.0`, sha256 `4afe8157…`) + hub **v0.117.0** (`37ae31f`): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth `slow_crashloop_since` moves (at most once per 24 h) and the hub mints `controller_slow_crashloop` (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed `agent_update` to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`. **Proven live with the PRODUCTION window (no test-only interval):** five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 `SLOW CRASH-LOOP … restarts_24h=5`; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). **Reasoning kept:** the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. `audits/evidence-chaos-fixes-2026-09-17/partC-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 3c1882a:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-637** | **Build the undo (`09` §3 decision 15) — the box puts a failed update back by itself (P2).** Closed in controller **v0.263.2** (`8fc2b4a` v0.263.0, `5d38573` v0.263.1, `2cd6666` v0.263.2). Copy method chosen by the bake-off (decision 19): a folder copy of every NAMED volume, `cp -a` into `<vol>.pre-update-<stamp>` after the pull, finished-marker last; bind folders never touched. Undo in `failAndHold`: copies validated first, volumes refilled, definition + pin + the pinned version's `.felhom.yml` (`applied-meta/`) put back from the job's own copies, old probe, `undone` or a HOLD whose sentence says the undo failed and the data state. **Proven live on 9202 (0.263.2):** docmost, romm, vikunja undone by the product with seeds before the backup, after it, and seconds before the press all read back, ledgers equal, page line hu/en; a cut-off copy → HOLD `untouched`; a power cut (`pct stop`) during `undoing` → resumed and undone; a manual press after an undo → `done`, note cleared; removal deleted kept copies. **Two defects the live proof found and the unit tests could not:** the undo's probe was gated on `running` while the current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` was already the new one because it flows in on every sync (v0.263.2). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. **Reasoning kept:** *a copy counts only with its finished-marker, judged by the helper container's own exit — killing `docker run` does not stop the copy; the undo reads its own copies, never the recovery unit.* | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-639** | **After a held update the previous definition survived only in the recovery unit (P3).** Closed in controller **v0.263.0/v0.263.2**: the pre-update compose, applied definition, pin and the pinned version's `.felhom.yml` are kept until the undo is over; the undo never reads the unit (R-645). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — SHIPPED** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-641** | **An app with no database server had no last-second copy (P2).** Closed in controller **v0.263.0**: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-642** | **Start answered 200 "start completed" over a crash loop (P3).** Closed in controller **v0.263.0**: start/restart answer `requested — state now: <state>`, never "completed" (`startAnswer`, pinned by `TestR642_*`, red-proofed); live on 9202: `Stack romm start requested — state now: running`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user