From 21f17ed32bdc88dfd2a87a96ebf806e304ab8982 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 23 Sep 2026 14:39:31 +0200 Subject: [PATCH] The household is told: controller v0.264.0 + hub v0.120.0 proven live, floor 0.264.0 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 09 §6.4 parts 2-3 SHIPPED. R-606, R-620, R-646 closed; R-647 (three leftovers) and R-648 (whole-box backup press in the harness) opened. Open rows 335 -> 334. Evidence: audits/undo-fleet-2026-09-23/. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 22 + REPORT.md | 159 ++- STATUS.md | 31 +- .../architecture/00-capability-map.md | 2 +- .../architecture/09-update-architecture.md | 4 +- .../00-floor-0.263.2.txt | 16 + .../10-enabled-events-before.txt | 5 + .../11-hub-0.120.0-deploy-and-migration.txt | 17 + .../undo-fleet-2026-09-23/20-9202-setup.txt | 17 + .../21-9202-prep-vikunja.txt | 16 + .../undo-fleet-2026-09-23/22-9202-undone.txt | 34 + .../undo-fleet-2026-09-23/23-9202-held.txt | 34 + .../24-9202-r620-warn-then-debug.txt | 10 + .../28-9202-controller-log.txt | 57 + .../29-9202-teardown.txt | 27 + .../30-9201-before-and-upgrade.txt | 25 + .../30-9201-standing-before.txt | 11 + .../31-9201-repoint-and-controls.txt | 28 + .../undo-fleet-2026-09-23/32-9201-prep.txt | 24 + .../33-9201-hu-vikunja.txt | 54 + .../34-9201-language-en.txt | 4 + .../35-9201-en-glance.txt | 35 + .../40-notification-log.txt | 11 + .../41-mails-as-received.md | 55 + .../48-9201-controller-log.txt | 12 + .../49-9201-teardown.txt | 43 + .../50-floor-0.264.0.txt | 15 + .../audits/undo-fleet-2026-09-23/README.md | 50 + .../audits/undo-fleet-2026-09-23/bakeoff.py | 266 +++++ .../audits/undo-fleet-2026-09-23/fixtures.py | 1008 +++++++++++++++++ .../audits/undo-fleet-2026-09-23/live.py | 194 ++++ .../audits/undo-fleet-2026-09-23/repoint.py | 60 + .../audits/undo-fleet-2026-09-23/spike.py | 160 +++ .../state-9201-vikunja.json | 172 +++ .../state-9202-vikunja.json | 167 +++ .../audits/undo-fleet-2026-09-23/walk.py | 495 ++++++++ documentation/backlog/CLOSED-ITEMS.md | 9 + documentation/backlog/OPEN-ITEMS.md | 5 +- 38 files changed, 3269 insertions(+), 85 deletions(-) create mode 100644 documentation/audits/undo-fleet-2026-09-23/00-floor-0.263.2.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/10-enabled-events-before.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/11-hub-0.120.0-deploy-and-migration.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/20-9202-setup.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/21-9202-prep-vikunja.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/22-9202-undone.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/23-9202-held.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/24-9202-r620-warn-then-debug.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/28-9202-controller-log.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/29-9202-teardown.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/30-9201-before-and-upgrade.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/30-9201-standing-before.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/31-9201-repoint-and-controls.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/32-9201-prep.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/33-9201-hu-vikunja.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/34-9201-language-en.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/35-9201-en-glance.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/40-notification-log.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/41-mails-as-received.md create mode 100644 documentation/audits/undo-fleet-2026-09-23/48-9201-controller-log.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/49-9201-teardown.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/50-floor-0.264.0.txt create mode 100644 documentation/audits/undo-fleet-2026-09-23/README.md create mode 100644 documentation/audits/undo-fleet-2026-09-23/bakeoff.py create mode 100644 documentation/audits/undo-fleet-2026-09-23/fixtures.py create mode 100644 documentation/audits/undo-fleet-2026-09-23/live.py create mode 100644 documentation/audits/undo-fleet-2026-09-23/repoint.py create mode 100644 documentation/audits/undo-fleet-2026-09-23/spike.py create mode 100644 documentation/audits/undo-fleet-2026-09-23/state-9201-vikunja.json create mode 100644 documentation/audits/undo-fleet-2026-09-23/state-9202-vikunja.json create mode 100644 documentation/audits/undo-fleet-2026-09-23/walk.py diff --git a/CONTEXT.md b/CONTEXT.md index 254c8a70..a6dbd571 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,28 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-23 (evening) — the household is told: controller v0.264.0 + hub v0.120.0 (`09` §6.4 parts 2–3); floor 0.264.0 + +**Events.** `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, sent by the +controller through `Manager.SetUpdateEventSink` (wired in main.go). Details `{app, stack_name, from, to, at, +copy_tier, copy_date, copy_holds}`. The held mail's message line is the hold sentence rendered in the +household's language (`pushEventBoth`); the operator gets the Hungarian wire text. +**Hub.** Allow-listed, NOT operator-only, in `defaultSeedEvents`; `mail.event.*` hu+en name the app via the +new register `appNamedMailEvents` (reads `details.stack_name`). **Existing households:** a one-time add-only +migration `Store.SeedEventTypesOnce` guarded by `hub_settings.seed_app_update_events_v1` — it changed all +three rows (`demo-felhom`, `demo-hp`, `_resend-rotation-test`; before/after in +`audits/undo-fleet-2026-09-23/10-*`, `11-*`). **Cooldown:** per app on both legs — operator via R-389's +`perAppCooldownEvents` (widened on purpose), household via a NEW sibling `perAppCustomerCooldownEvents` +(the customer key was `customer:type`). `app_start_failed`'s household grain is unchanged (pinned). +**Why the box also seeds:** the controller pushes its own `EnabledEvents` to the hub on every notifications +page save, so a hub-only migration would be undone by the next save; the box seeds once +(`app_update_events_seeded`) and the page has the two toggles. +**R-606:** update sentences are keys + args, rendered per reader; leftovers R-647. **R-620:** a disabled +notifier WARNs once per type. **R-646:** startup `BackfillAppliedMeta` for apps current with the catalog. +**Voice:** the undone line ends „nincs teendőd" (the product's „te"), not the brief's „teendője nincs". +**Harness lesson (R-648):** the drill's „Mentés most" is whole-box and restarts standing apps. +Evidence: `audits/undo-fleet-2026-09-23/README.md`. + ## 2026-09-23 (afternoon) — the undo is built: controller v0.263.2 (decisions 15, 19, 20) **Operator rulings 19 and 20** recorded in `09` §3: the copy method is chosen by a bake-off; on update diff --git a/REPORT.md b/REPORT.md index 656d5c54..f7b187e1 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,11 +1,11 @@ -# REPORT — the undo: bake-off, build (controller v0.263.0 → v0.263.2), live proof +# REPORT — the undo reaches the fleet, and the household is TOLD (controller v0.264.0, hub v0.120.0) -2026-09-23 (afternoon). Repos touched: **felhom-controller** (v0.263.0, v0.263.1, v0.263.2), -**felhom.eu** (docs, register, capability map, STATUS, CONTEXT, evidence), **admin/app-catalog-drill** -(drill commits, reset to live `main` at the end). **app-catalog-felhom.eu, felhom-agent, hub: untouched.** -Architecture read first and named: `documentation/architecture/09-update-architecture.md` (§3 decisions -11–20, §4, §6.1, §6.1a, §6.4). Baselines verified live: controller `b9deec19077b`, agent -`d9864a94bf62`, felhom.eu `4c92beab8fdb`, catalog `cfcfe5278428` — all as the brief said. +2026-09-23 (evening). Repos touched: **felhom-controller** (v0.264.0, `bc27894`), **felhom.eu** (hub +v0.120.0 `d3b2848`, manifest `7caa24f`, docs, register, evidence), **admin/app-catalog-drill** (drill +commits, reset to live `main` at the end). **felhom-agent, app-catalog-felhom.eu: untouched.** +Architecture read and named: `documentation/architecture/09-update-architecture.md` (§3 decision 15, +§6.4). Evidence: `documentation/audits/undo-fleet-2026-09-23/README.md`. **Method: endpoint-level** +(the endpoints the UI invokes; mails read from the catch-all mailbox; no browser). --- @@ -13,69 +13,104 @@ Architecture read first and named: `documentation/architecture/09-update-archite | item | state | |---|---| -| rulings 19–20 | **recorded** (`09` §3), commit `5a349d9` | -| Part 1 — bake-off | **done in ~40 min** of the 2 h cap; folder copy chosen | -| Part 2 — build | **done — but in THREE releases, not one.** v0.263.0 failed its first two live proofs honestly (HELD, data put back). Two defects only the live box could show; each fixed + red-proofed + released: **v0.263.1** (the undo's probe never ran on an app the current probe held `unhealthy`), **v0.263.2** (the "old" `.felhom.yml` was already the new one — it flows in on every catalog sync; now recorded at pin time). 0.263.0/0.263.1 ran only on 9202 and were removed from it. | -| the household mail + operator event of decision 15 | **not built** — `09` §6.4 part 2 (R-606). The page line and the hold sentence are built. | -| the floor | **not raised** — the operator's question (STATUS) | -| immich/nextcloud rate test | **nextcloud deployed and was used** (185 MB MariaDB volume, 300 files through WebDAV); immich not tried | -| cut-off copy on vikunja in the bake-off | **could not be cut** (2.9 MB finishes before a kill lands); proven on docmost and romm instead; in the LIVE proof the marker was removed from a vikunja copy mid-update | +| Part 0 — floor 0.263.2 | **done**, read back (`00-floor-0.263.2.txt`) | +| Part 1 — R-646 startup pass | **done**; live: 3 recorded on 9202, 10 on 9201, 0 skipped | +| Part 2 — two events, hub v0.120.0 first | **done**; hub deployed 11:49Z, before any box ran v0.264.0 | +| Part 3 — R-606 | **done, with leftovers → R-647** (a reader in the other language than the box sees a HELD sentence in the box's language; the English mail's raw `Note:` line keeps the Hungarian `copy_holds` phrase) | +| Part 4 — R-620 | **done** | +| Part 5 — live proof | **done**; all proofs passed → **floor raised to 0.264.0** | +| "standing apps untouched" | **NOT fully met.** The harness's „Mentés most" is a whole-box backup: on 9201 it stopped and restarted 9 of the 10 standing apps for their volume dumps, twice (a few seconds each). All healthy afterwards, files byte-identical. **R-648** filed. | +| the brief's Hungarian undone line | **changed:** it ends „semmi nem veszett el, nincs teendőd." — the brief's „teendője nincs" is the formal „ön" voice; the product speaks „te" everywhere and the voice gate ratchets the formal form (18). | +| the customer cooldown | **changed shape:** a NEW customer-side register (`perAppCustomerCooldownEvents`) instead of putting the customer leg of R-389's register to work — that would also have changed `app_start_failed`'s household grain, which nobody asked for (pinned by a test). | +| a box-side seed + two page toggles | **added, not in the brief:** the box pushes its own `EnabledEvents` to the hub on every notifications-page save, so the hub migration alone would be undone by the next save. | -**Claims in the brief that turned out wrong or unmeasured:** -1. *"All three apps keep their data only in named volumes"* — true from compose AND on disk for these - three; but romm also has two BIND folders (roms, resources), empty throughout — never copied, by rule. -2. *"A volume copy with the containers stopped is consistent for both engines"* — **measured true**: - docmost (PostgreSQL) and romm (MariaDB) came back with ledgers equal, three times each. -3. *"Immich or Nextcloud deploys on 9202"* — Nextcloud did; Immich was not tried. -4. **Not in the brief, and the most important:** *"the old `.felhom.yml`"* does not exist at update - time. It is replaced on every catalog sync (`09` §5.4). The build now records it when a version is - pinned. R-646 for apps pinned before that. +**Claims in the brief, checked:** +1. *"No app-update event exists today"* — **true for apps.** The only near-misses: `controller_updated` is + the controller's OWN self-update, and `update_available` exists only as a debug-page test option — never + emitted, and not in the hub's allow-list (it would 400). +2. *"`enabled_events` is the household's only mail gate"* — **wrong.** Four more gates stand in the way: + the `operatorOnlyEvents` register, a blocked customer, a missing e-mail address, and the cooldown — + keyed `customer:type`, which would have merged two apps' mails into one. And the box OVERWRITES + the hub's list on each page save (above). +3. *"9201 can follow the drill catalog without disturbing its standing apps"* — **true for the catalog, + false for the harness.** Measured: the 10 standing apps' compose + `.felhom.yml` were byte-identical + before, during and after (the drill differed from live only in the throwaway apps). The disturbance + came from the harness's whole-box backup press (R-648), not from the catalog. -## 2. The bake-off (`audits/undo-bakeoff-2026-09-23/README.md`) +## 2. Floor raises, read back from the hub -Both methods passed every row on docmost, romm, vikunja (seeds A and B back, ledgers equal, cut-off -detected before anything moves, ≤ 5.3 s extra downtime). **Folder copy chosen (decision 19):** an app -with no database server gets no dump, so dump-and-load would need the folder copy anyway. Rate: 185 MB -in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s (warm cache); 5 GB ≈ 12 s here, 25–50 s on a cold or spinning -disk. Found on the way: the undo must never read the recovery unit (R-645), and killing `docker run` -does not stop the copy container. - -## 3. Red-proofs (each seen failing, then restored) - -| # | mutation | test that failed | +| raise | read back | fleet | |---|---|---| -| 1 | the undo call removed (v0.262.1's shape) | `TestUndo_FailedUpdateIsPutBackWithItsData` — `held=true phase="failed"` | -| 2 | copy validation removed | `TestUndo_CutOffCopyIsRefusedBeforeAnythingMoves` — state `half` | -| 3 | old-probe rule removed | `TestUndo_UsesTheOldProbe` — `not_started` | -| 3b | the old probe read from the stack dir (v0.263.1's shape) | `TestUndo_UsesTheOldProbe` | -| 4 | the `undoing` recovery arm removed | `TestUndo_PowerCutDuringTheUndoResumesIt` — `resumed=[]` | -| 5a | `last_update_undone` not recorded | `TestUndo_FailedUpdateIsPutBackWithItsData` — `got ` | -| 5b | the page line removed from the handler | `TestUndo_PageSaysTheBoxPutTheAppBack` (hu and en) | -| 6 | the hold prefix ignored | `TestUndo_HoldSentenceSaysTheUndoWasTriedAndTheDataState` | -| 7 | R-642: "completed" back | `TestR642_StartIsNeverReportedCompleted` | -| 8 | a successful update keeps the undone note | `TestUndo_ASuccessfulUpdateEndsTheUndoneNote` | -| 9 | the undo probe switched off on an `unhealthy` app (v0.263.0's shape) | `TestUndo_OldProbeRunsOnAnAppTheCurrentProbeMarkedUnhealthy` — `not healthy within 1s (last: state unhealthy)`, the live message verbatim | -| 10 | the pin advance records no probe | `TestUndo_PinningRecordsThatVersionsProbe` | +| 0.263.2 / MinAgent 0.131.0 | `min_controller_version=0.263.2`, `min_agent=0.131.0` | `00-floor-0.263.2.txt` | +| **0.264.0 / MinAgent 0.131.0** (14:35:08) | `min_controller_version=0.264.0`, `min_agent=0.131.0`; hub log `managed floor SERVED for demo-felhom: floor 0.264.0 … from declared` | demo-felhom 0.263.2 → **0.264.0 in 20 s**; demo-hp 9201 already 0.264.0 (`50-floor-0.264.0.txt`) | -`go build ./... && go vet ./... && go test ./...` rc=0 before each of the three commits; controller -gates rc=0 (the `docker -v` gate needed four named-volume mounts allowlisted with their why). +## 3. The hub migration — rows changed (one-time, add-only) -## 4. Live proof (`audits/undo-live-2026-09-23/README.md`) — endpoint-level, both languages +| customer | before | after | +|---|---|---| +| `demo-felhom` | 10 types | the same 10 + `app_update_undone`, `app_update_held` | +| `demo-hp` | `null` (nothing on) | `["app_update_undone","app_update_held"]` | +| `_resend-rotation-test` | `[]` | `["app_update_undone","app_update_held"]` | -Three apps undone by the product with seeds A, B, C back and ledgers equal (30–52 s of undo); page -line hu/en; cut-off copy → HOLD `untouched` (prefix in the box's language, hu and en quoted); power cut -during `undoing` → resumed and undone; a person's press after an undo → `done`, note cleared; removal -deleted kept copies; R-642 Start answer. +Guard row `seed_app_update_events_v1 = 2026-09-23T11:49:03Z`. E-mail column never selected. -## 5. Rows +## 4. The mails (9201, customer demo-hp) and their `notification_log` rows -**Closed (4, moved to `CLOSED-ITEMS.md`):** R-637, R-639, R-641, R-642. **Narrowed:** R-638, R-640 (to -the restore paths). **Ruled:** R-643 (decision 20). **Noted:** R-645. **Opened:** R-645 (earlier this -session, filed with the bake-off) and R-646 (apps pinned before v0.263.2). **Open rows 337 → 335.** +| row | time (UTC) | type | leg | subject as received | +|---|---|---|---|---| +| 937 / 938 | 12:15 | undone | op / household | op `⚠️ demo-hp: app_update_undone` · hh **`Figyelmeztetés: vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut`** | +| 939 / 940 | 12:16 | held | op / household | op `🔴 demo-hp: app_update_held` · hh **`Hiba: vikunja: az alkalmazás leállítva, visszaállítás szükséges`** | +| 942 / 943 | 12:29 | undone | op / household | hh **`Warning: glance: the update did not work; the app runs on its previous version`** | +| 944 / 946 | 12:31 | held | op / household | hh **`Error: glance: the app is stopped and needs a restore`** | -## 6. Teardown — three layers +All 8 `sent`. The box language was switched to English once, between the two rounds (the hub had two +reports in English before the second round). Household bodies carry the box's own sentence (the undone +line / the hold sentence) in the household's language; the operator's carry the Hungarian wire text. +Verbatim bodies: `41-mails-as-received.md`. **Per-app cooldown, live:** row 943 went out 14 minutes after +row 938, inside the 6-hour customer cooldown. -Machine (9202): the three apps and every copy removed through the product; test images removed by -name; `controller.yaml` restored and read back; catalog cache on live `cfcfe52`; **9202 stays on -controller 0.263.2** (self-update off, fleet floor untouched). Host: nothing provisioned; 9202 stopped -and started once for the power cut. Hub: untouched. Drill repo: reset to live `main`. +## 5. The pages in both languages (9202, then 9201) + +- undone: hu *„A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan + visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en *"The update of vikunja + at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically — + nothing was lost."* +- held, box hu: `?lang=hu` → the Hungarian hold sentence; `?lang=en` → *"The update did not succeed, and the + automatic undo did not either. … own drive, 2026-09-23 13:58 — this copy holds the settings, the database + and the data volumes."* — the whole sentence English now (before v0.264.0 only the prefix was). +- `GET /api/stacks/vikunja?lang=en`: label *"The update did not succeed"*, error and hold in English. +- **Leftover (R-647):** box en + `?lang=hu` showed the English hold. + +## 6. Red-proofs (each seen failing, then restored) + +**Controller (9):** backfill stores nothing → `TestR646_…` *"the current app must have its record"*; +`dropped()` not called → `TestR620_…` *"want exactly TWO WARN lines, got 0"*; undone event not emitted → +*"exactly ONE app_update_undone event, got []"*; held event not emitted → *"…got []"*; page +`updateErrorText` not localised → `TestR606_UpdateSentencesFollowTheReader`; hold sentence ignores the +language → `undo_hold_test.go:62`; `UpdateErrorIn` returns the Hungarian → *"English = „A frissítés nem +indult el…""*; box seed not run → *"got [backup_failed offbox_enlarge_blocked]"*; sink not wired in main → +*"SetUpdateEventSink is never called"* (first mutation did not compile; retried with a compiling one). +**Hub (5):** seed call does nothing → *"c1: enabled_events = [backup_failed], want […]"*; customer per-app +suffix removed → *"want 3 household mails …, got 2"*; allow-list entry removed → *"must be in +allowedEventTypes"*; app-named headline removed → subject *"%s: a frissítés…"*; seed default removed → +*"defaultSeedEvents must include app_update_undone"* (first attempt matched no test; retried with a new +test that names both types). Every `-run` was checked for `=== RUN`. + +Gates: controller `go test ./...` rc=0, `controller_gates.py` rc=0; hub `go build/vet/test` rc=0, +`repo_gates.py` rc=0. The R-389 fence test was widened on purpose, with its reason; the settings-page +parity fixture was regenerated (measured diff: exactly the two new toggles). + +## 7. Rows + +**Closed (3, moved to `CLOSED-ITEMS.md`):** R-606, R-620, R-646. **Opened (2):** R-647 (the three +leftovers), R-648 (the whole-box backup press). **Open rows 335 → 334.** `09` §6.4 parts 2 and 3 marked +SHIPPED; the capability map's undo row now carries the mail; STATUS rewritten. + +## 8. Teardown — three layers + +- **machine:** 9202 and 9201 — throwaway apps removed through the product (0 volumes, 0 undo copies), + test images removed, `controller.yaml` identical to the saved copy, catalog on live `cfcfe52`, 9201's + language back to `hu`, standing apps deployed and healthy. Both on 0.264.0 (= the floor). +- **host:** nothing provisioned on demo-hp. +- **hub:** v0.120.0 (planned); floor 0.264.0. Scratchpad password files deleted; hub DB copies deleted. +- **drill repo:** reset to live `main`, read back from the remote. diff --git a/STATUS.md b/STATUS.md index b300d958..7a5930a1 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,26 +1,23 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-23 (afternoon) — a failed update now puts the app back by itself. Built, and proven on the scratch machine. One question for you: whether the fleet gets it.** +**Updated 2026-09-23 (evening) — the undo is on every demo machine now, and the household gets a mail when an update fails. The floor is raised.** -**Decisions I took on my own: none.** Your two afternoon rulings are written down: the bake-off picks the copy method, and on update nights the full-system backup waits for the updates. +**Decisions I took on my own: none.** One small deviation from your brief: the Hungarian mail line ends "nincs teendőd", not "teendője nincs", because the product speaks to the household as "te" everywhere. -**The bake-off.** Both ways of keeping the last-second copy passed every test on three apps. Copying the app's data folders won, because one of the apps has no database server and so gets no database copy at all. The extra downtime was 1 to 5 seconds. +**What the household gets now.** When an update fails and the machine puts the app back, the household gets one mail: "docmost: the update did not work; the app runs on its previous version". When the undo fails too, they get one mail that says the app is stopped and needs a restore, and which backup to use. The mail is in the household's own language. The operator gets the same two events. Two apps on the same night give two mails, not one. -**What the machine does now.** When an update fails its health check, the machine puts back the previous version and the data exactly as it was seconds before the update. It stops the app only if that undo fails too, and then it says so. It never touches the household's own folders (photos, documents). +**Proven on real machines, through the same buttons the page uses:** +- On the demo-hp machine: 4 mails to the household and 4 to the operator, first in Hungarian, then in English after one language switch. All arrived. The hub log shows all 8 as sent. +- On the scratch machine: the update messages on both app pages show in Hungarian and English. A machine with no hub now writes one warning line per kind of lost message. +- At start, the new version stores the health-check file of each up-to-date app, so its first undo asks the right question. It did this for 3 apps on one machine and 10 on the other. +- The floor is 0.264.0. The N100 demo machine updated itself within 20 seconds. -**Proven on the scratch machine, through the same buttons the page uses:** -- Three apps undone by the machine. Data written before the backup, after it, and seconds before the update all came back. It took 30 to 52 seconds. -- The app page shows one line in Hungarian or English: the update failed, the machine put the app back, nothing was lost. -- A damaged copy is caught before anything is poured back, and the app is held with an honest sentence. -- A power cut in the middle of the undo: after restart, the machine finished the undo. -- A person pressed Update again after the catalogue was fixed, and it worked. +**What was not clean.** +- My test "back up now" button backs up the whole machine. On demo-hp it stopped and restarted 9 of the 10 standing apps for a few seconds, twice. All came back healthy, with unchanged files. So "standing apps untouched" is not fully true. I filed a row to fix the test method. +- Three small leftovers, filed as one row: a reader who forces the other language sees a stopped app's sentence in the machine's language; the raw detail line of an English mail has one Hungarian phrase; and two of my log lines use unclear words. -**What went wrong on my side.** My first build failed its own live test twice. Both times the app stayed stopped with an honest sentence, and the data was safe. The first fault: the machine never asked the old version the right health question. The second: it kept the wrong copy of that question. My unit tests had passed both times. I fixed both the same afternoon and proved the fixes. The released version is the third build. +**Rows.** Three closed, two new. The list went from 335 to 334. -**Rows.** Four closed, two narrowed, one new. The list went from 337 to 335. +**What needs you: nothing.** If you do nothing, the fleet keeps the new version. The leftovers are small and wait for a free evening. -**What needs you — one question.** Should the demo machines and the fleet get this version now? -- **Yes (my pick):** I raise the floor, and both demo machines update themselves within a minute. A failed update then puts the app back instead of stopping it. -- **Not yet:** nothing changes. A failed update keeps stopping the app until someone restores it. - -**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. The scratch machine runs the new version and is back on the real catalogue.** +**Nothing on your own machine, Peti's machine or the off-site box was touched, except the hub update. The test apps are gone from both demo machines. Both are back on the real catalogue.** diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 5b26c150..8da88294 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -103,7 +103,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | | **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. **THE NARROWING ABOVE WAS CLOSED THE NEXT DAY, 2026-09-22 — and re-widened the row.** `audits/PROBE-FIX-2026-09-22.md`. All three wrong probes were corrected in the catalog (`app-catalog-felhom.eu@793c4fb`: tandoor `8080→80`, wger `80→8000`, zipline `/api/health→/api/healthcheck`) and **red-proofed live on 9202 through the product in both directions**: at the live pin all three read `Nem egészséges` / `Not healthy` on their own app page while docker reported every container healthy and the front door served a real page; after the real sync all three read `Fut` / `Running` with no redeploy. **tandoor's edge was then re-walked with nothing else changed and ended `done` at +41.1 s**, seed read back, where the identical edge had ended `failed` at +361.9 s with the app stopped — so the tally is now **15 proven, 2 failed, 4 inconclusive**, and all fifteen are on the live catalog. A `--fast` catalog gate (`check-probe-matches-compose.py`) now refuses a probe that does not match the same service's own compose healthcheck, with four red-proofs and ten decoys including the no-PyYAML mode CI actually runs. **WHAT THIS ROW STILL CANNOT CLAIM, and the reason is exactly R-96 rule 3:** the guard is now shown correct for **47 of 53** templates. `paperless-ngx`'s probe has **never run on any box** — no container name matches its stack name, so it is silently skipped and its badge can never go red (**R-630**); and five more cannot be judged statically, one of which (`home-assistant`) is right only because its check type cannot fail (**R-631**). An absent alarm is equally consistent with healthy and with never checked. **And the sweep's ceiling, counted: 28 of the 53 templates have never been deployed by any drill (R-632).** **THAT CEILING WAS REMOVED THE SAME NIGHT, 2026-09-22 — all 28 walked (`audits/DRILL-the-28-2026-09-22.md`), so every template in the catalog has now been attempted at least once.** 26 of 28 deployed, **6 proven**, 5 inconclusive, 14 with no within-a-major edge upstream, 1 failed honestly and 2 undeployable — one of those (`plant-it`) **by design**, refused by the product's lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a **restore from its own copy, with the seed read back again** — 21 restored, and **2 were correctly REFUSED** with the sentence `07` §6.2 predicts for a class-A app whose local copy holds no file leg. **AND THE NIGHT NARROWED THIS ROW AGAIN, in the place the probe work could not reach.** `paperless-ngx` has no container matching its stack name, so no probe is ever built for it — and `verifying` does not skip: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three containers all read `healthy`. The controller's own words: *`not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app`*, at **+313.0 s**. **R-630, raised to P1.** So "the update is guarded" is now shown correct for 47 of 53 templates, wrong for none, and **actively harmful for the one template that has no probe at all**. **Two further limits on what this row may claim, both about STATE rather than health:** a `remove` sent while a restore is still running reports success and leaves a container restarting with a live public route (**R-633**) — while the product already refuses exactly that clash for `update` and for `restore`, naming the blocking operation; and an app can be **running, healthy and serving while recorded as `deployed: false`**, in which state the product refuses to remove it at all (**R-634**). In both, a person needed a shell to clear what the product could not. **ALL THREE ARE FIXED IN CONTROLLER v0.262.0 (2026-09-22), and the first is PROVEN LIVE.** *The stopped app:* `verifying` no longer loops on a probe that resolves to nothing — it settles on container state, the way an app declaring no check is judged, and says which it did. **Measured on paperless-ngx: the identical Update that ended `failed` at +313.0 s with the app stopped now ends `done` at +53.4 s**, with no `no probe container` warning in the log because the explicit `healthcheck.container` resolved the target. *The ghost:* `RemoveStack` consults the backup side's `Busy` guard — which the product already applied to `update` and to `restore` — and then WATCHES the compose project for 25 s after `down`, removing anything that carries its label and answering `verified: true/false`, because `down` returning 0 is a request rather than a result. *The unremovable app:* the refusal now asks whether anything EXISTS (containers, a compose file, an `app.yaml`) instead of reading a flag. **WHAT THIS ROW STILL MAY NOT CLAIM:** R-634's MECHANISM — why `deployed` goes false while containers run — **is not diagnosed**; only the consequence is fixed. And a probe can be right about the port and still wrong about what a 200 means: `romm` answered 200 from nginx for six hours while its workers were OOM-killed behind it (**R-635**). **"The update is guarded" has never meant "the new version runs".** | -| **A failed update is UNDONE by the box itself — the previous version back with its data from seconds before the update; the app is held only if that undo fails too** | controller **v0.263.2** | **PROVEN-LIVE (2026-09-23)** | `audits/undo-live-2026-09-23/README.md` (+ `audits/undo-bakeoff-2026-09-23/` for the method). On scratch guest 9202, through the endpoints the UI invokes: docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite in a volume), each a real migrating edge failing a deliberately wrong probe, **undone in 30–52 s** with seeds written before the backup, after it and seconds before the press ALL read back through each app's front door, ledgers equal to before; page line in hu and en. Cut-off copy → HOLD (`untouched`), prefix in the box's language; power cut (`pct stop`) during `undoing` → resumed after boot and undone; a person's press after an undo → `done`; removal deletes kept copies. | Folder copy of NAMED volumes only (`09` §3 decision 19); bind folders never touched. **Not built:** the mail to the household and the operator event (`09` §6.4 part 2), and the automatic caller (part 7) — this protects the manual button today. **R-646:** an app pinned before v0.263.2 undoes with the current `.felhom.yml` until its next pin. | +| **A failed update is UNDONE by the box itself — the previous version back with its data from seconds before the update; the app is held only if that undo fails too** | controller **v0.263.2** (the undo), **v0.264.0** + hub **v0.120.0** (the mail) | **PROVEN-LIVE (2026-09-23)** | `audits/undo-live-2026-09-23/README.md` (+ `audits/undo-bakeoff-2026-09-23/` for the method). On scratch guest 9202, through the endpoints the UI invokes: docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite in a volume), each a real migrating edge failing a deliberately wrong probe, **undone in 30–52 s** with seeds written before the backup, after it and seconds before the press ALL read back through each app's front door, ledgers equal to before; page line in hu and en. Cut-off copy → HOLD (`untouched`), prefix in the box's language; power cut (`pct stop`) during `undoing` → resumed after boot and undone; a person's press after an undo → `done`; removal deletes kept copies. | Folder copy of NAMED volumes only (`09` §3 decision 19); bind folders never touched. **The household is told (2026-09-23, `audits/undo-fleet-2026-09-23/`):** on guest 9201 the household and the operator each received ONE mail per app per outcome — undone and held, in Hungarian (vikunja) and, after one language switch, in English (glance) — with the app named in the subject; `notification_log` rows 937–946 all `sent`. **Not built:** the automatic caller (part 7) — this protects the manual button today. Leftovers: R-647. | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index ee99d8f9..a65b71b6 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -1043,8 +1043,8 @@ what the part can do to a household's data if it is wrong, not how likely that i | # | part | rulings / rows | cost | depends on | risk to customer data | |---|---|---|---|---|---| | **1** | **SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23** (`audits/undo-live-2026-09-23/`). **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. | -| **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none | -| **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none | +| **2** | **SHIPPED — controller v0.264.0 + hub v0.120.0, proven live on 9202 and 9201 2026-09-23** (`audits/undo-fleet-2026-09-23/`). **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held: events `app_update_undone` (warning) and `app_update_held` (error), one per app per outcome, on by default, mailed in the household's language with the app named in the subject; per-app cooldown on both legs. Leftovers: R-647. | R-606, 15 | **1** | — | none | +| **3** | **SHIPPED — controller v0.264.0, proven live on 9202 2026-09-23.** **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none | | **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only | | **5** | **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape | | **6** | **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low | diff --git a/documentation/audits/undo-fleet-2026-09-23/00-floor-0.263.2.txt b/documentation/audits/undo-fleet-2026-09-23/00-floor-0.263.2.txt new file mode 100644 index 00000000..e9ab7167 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/00-floor-0.263.2.txt @@ -0,0 +1,16 @@ +=== FLOOR -> 0.263.2, MinAgent 0.131.0 declared (2026-09-23 13:07:31) +POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set +--- read back from the hub: +"min_controller_version" value="0.263.2" +"min_agent" value="0.131.0" +"min_agent" value="0.131.0" +--- watching the two demo boxes (up to 3 min) + +10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.262.1 Up 15 hours (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.262.1 Up 15 hours (healthy) + +20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 5 seconds (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 8 seconds (healthy) +--- hub log: +2026/09/23 13:07:34 [INFO] managed floor SERVED for demo-felhom: floor 0.263.2, agent requirement "0.131.0" from declared (golden 0.258.0) +2026/09/23 13:07:34 [INFO] managed floor SERVED for demo-hp: floor 0.263.2, agent requirement "0.131.0" from declared (golden 0.258.0) +--- hosts page: +demo-felhom-8363b5 +demo-hp-bb76ea +drill-r50-0a4f9a diff --git a/documentation/audits/undo-fleet-2026-09-23/10-enabled-events-before.txt b/documentation/audits/undo-fleet-2026-09-23/10-enabled-events-before.txt new file mode 100644 index 00000000..fb1ec983 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/10-enabled-events-before.txt @@ -0,0 +1,5 @@ +# customer_notifications BEFORE hub v0.120.0 (read-only, 2026-09-23T11:47:29Z); e-mail column deliberately not selected +_resend-rotation-test|[] +demo-felhom|["backup_failed","db_dump_failed","offbox_enlarge_blocked","storage_disconnected","node_down","health_critical","disk_warning","disk_critical","expected_backup_missed","expected_dbdump_missed"] +demo-hp|null +# guard row: diff --git a/documentation/audits/undo-fleet-2026-09-23/11-hub-0.120.0-deploy-and-migration.txt b/documentation/audits/undo-fleet-2026-09-23/11-hub-0.120.0-deploy-and-migration.txt new file mode 100644 index 00000000..e81c2e54 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/11-hub-0.120.0-deploy-and-migration.txt @@ -0,0 +1,17 @@ +# hub v0.120.0 rollout 2026-09-23T11:49:50Z pod hub-b6ff59c47-n9vp8 +gitea.dooplex.hu/admin/felhom-hub:0.120.0 +# startup log lines (migration): +2026/09/23 13:49:03 [INFO] [store] app-update event types added to 3 household(s)' notification prefs (one-time, add-only): [demo-felhom _resend-rotation-test demo-hp] +2026/09/23 13:49:03 [INFO] Default controller-version floor: 0.120.0 +2026/09/23 13:49:04 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000 +2026/09/23 13:49:04 [INFO] Registry version checker started (every 6h) +2026/09/23 13:49:04 [DEBUG] Registry version check: latest = 0.263.2 +2026/09/23 13:49:35 [INFO] Listening on :8080 +# customer_notifications AFTER (read-only; e-mail column not selected) +# guard row: +# customer_notifications AFTER (copy of hub.db+wal+shm taken 2026-09-23T11:50:35Z; e-mail column not selected) +_resend-rotation-test|["app_update_undone","app_update_held"] +demo-felhom|["backup_failed","db_dump_failed","offbox_enlarge_blocked","storage_disconnected","node_down","health_critical","disk_warning","disk_critical","expected_backup_missed","expected_dbdump_missed","app_update_undone","app_update_held"] +demo-hp|["app_update_undone","app_update_held"] +# guard row: +seed_app_update_events_v1|2026-09-23T11:49:03Z diff --git a/documentation/audits/undo-fleet-2026-09-23/20-9202-setup.txt b/documentation/audits/undo-fleet-2026-09-23/20-9202-setup.txt new file mode 100644 index 00000000..4f402376 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/20-9202-setup.txt @@ -0,0 +1,17 @@ +13:54:56 === 9202 setup: drill vikunja at the FROM pin, 9202 onto the drill catalog +13:54:57 [5] drill commit 847d544b8064: vikunja vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.3.0 (push rc=0) +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + +13:57:23 [sync] the badge NEVER caught up to vikunja/vikunja:2.3.0 in 127.6s — catalog_images = None +13:57:23 badge caught up after None s +13:57:25 catalog cache: 847d544 DRILL vikunja: vikunja/vikunja:2.6.0 -> vikunja/vikunja:2.3.0 +https://@gitea.dooplex.hu/admin/app-catalog-drill.git + diff --git a/documentation/audits/undo-fleet-2026-09-23/21-9202-prep-vikunja.txt b/documentation/audits/undo-fleet-2026-09-23/21-9202-prep-vikunja.txt new file mode 100644 index 00000000..e81a7978 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/21-9202-prep-vikunja.txt @@ -0,0 +1,16 @@ +13:57:31 === prep vikunja: deploy at the drill FROM pin +13:57:31 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +13:57:36 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'} +13:57:36 deployed: True +13:57:41 vikunja: register http=200 +13:57:41 vikunja: create project http=201 +13:57:42 vikunja: readback of the seeded project http=200 ok=True +13:57:42 C1 A: True +13:57:42 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +13:58:07 [4] backup idle; last=None +13:58:07 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill73bafb","created":"2026-09 +13:58:07 vikunja B: project readback=True attachment content readback=True +13:58:07 B reads back: True +13:58:10 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db'] + port: 3456 + diff --git a/documentation/audits/undo-fleet-2026-09-23/22-9202-undone.txt b/documentation/audits/undo-fleet-2026-09-23/22-9202-undone.txt new file mode 100644 index 00000000..c94c6999 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/22-9202-undone.txt @@ -0,0 +1,34 @@ +13:58:21 [5] drill commit 073bc953e182: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0) +13:58:25 badge caught up after 4.5 s +13:58:25 === vikunja: live undo by the product +13:58:25 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drill73bafb","created":"2026-09 +13:58:25 seed C written right before the Update: True +13:58:29 db before: tables=36 ledger=117 newest=SCHEMA_INIT +13:58:29 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +13:58:29 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +13:58:30 + 1.1s phase=pulling label=Új verzió letöltése… +13:58:33 + 4.2s phase=copying label=Az adatok másolása a frissítés előtt… +13:58:35 + 5.8s phase=starting label=Indítás az új verzióval… +13:58:35 + 6.3s phase=verifying label=Működés ellenőrzése… +14:00:06 + 96.7s phase=undoing label=Visszaállítás az előző változatra… +14:00:08 + 99.3s phase=undone label=Visszaállítva az előző változatra +14:00:08 END phase=undone err=None hold=None +14:00:11 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]} +14:00:11 vikunja: readback of the seeded project http=200 ok=True +14:00:12 vikunja B: project readback=True attachment content readback=True +14:00:12 vikunja B: project readback=True attachment content readback=True +14:00:14 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT) +14:00:14 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +14:00:16 leftover copies: 0 +14:00:16 GET /api/stacks/vikunja: phase='undone' label='Visszaállítva az előző változatra' err=None +14:00:16 GET /api/stacks/vikunja?lang=en: label='Put back to the previous version' err=None +14:00:16 --- controller log (positive observables): +14:00:18 2026/09/23 11:52:40 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only +2026/09/23 11:55:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only +2026/09/23 11:57:31 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deploy_started (severity info) — further app_deploy_started events are logged at DEBUG only +2026/09/23 11:57:34 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deployed (severity info) — further app_deployed events are logged at DEBUG only +2026/09/23 12:00:00 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event health_change (severity warn) — further health_change events are logged at DEBUG only +2026/09/23 12:00:08 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true) +2026/09/23 12:00:08 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only +2026/09/23 12:00:08 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 3s — the previous version is running on the data from before the update (the app's health check passed) + diff --git a/documentation/audits/undo-fleet-2026-09-23/23-9202-held.txt b/documentation/audits/undo-fleet-2026-09-23/23-9202-held.txt new file mode 100644 index 00000000..d4ab496d --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/23-9202-held.txt @@ -0,0 +1,34 @@ +14:00:30 === vikunja: a cut-off undo copy +14:00:31 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +14:00:31 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +14:00:32 + 1.1s phase=pulling label=Új verzió letöltése… +14:00:33 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt… +14:00:34 + 3.7s phase=verifying label=Működés ellenőrzése… +14:00:37 >>> finished-marker removed from one copy: +copy: vikunja_vikunja_data.pre-update-20260923T120033Z +total 12 +drwxr-xr-x 3 root root 4096 Sep 23 12:00 . +drwxr-xr-x 1 root root 4096 Sep 23 12:00 .. +drwxr-xr-x 2 root root 4096 Sep 23 12:00 data +-rw-r--r-- 1 root root 0 Sep 23 12:00 felhom-undo-complete +data + +14:02:05 + 94.7s phase=undoing label=Visszaállítás az előző változatra… +14:02:06 + 95.2s phase=failed label=A frissítés nem sikerült +14:02:06 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' +14:02:08 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."}} +14:02:08 box language -> en: http 302 +14:02:08 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes."}} +14:02:08 box language -> hu: http 302 +14:02:10 copies kept: vikunja_vikunja_data.pre-update-20260923T120033Z +vikunja_vikunja_db.pre-update-20260923T120033Z + +14:02:10 GET /api/stacks/vikunja: phase='failed' label='A frissítés nem sikerült' err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' +14:02:10 GET /api/stacks/vikunja?lang=en: label='The update did not succeed' err='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes.' hold='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes.' +14:02:10 --- controller log: +14:02:13 2026/09/23 12:00:08 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only +2026/09/23 12:02:05 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T120033Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T120033Z}] +2026/09/23 12:02:05 update_guard.go:553: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T11:58:06Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched") +2026/09/23 12:02:05 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true) +2026/09/23 12:02:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only + diff --git a/documentation/audits/undo-fleet-2026-09-23/24-9202-r620-warn-then-debug.txt b/documentation/audits/undo-fleet-2026-09-23/24-9202-r620-warn-then-debug.txt new file mode 100644 index 00000000..d62a728a --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/24-9202-r620-warn-then-debug.txt @@ -0,0 +1,10 @@ +http 200 type list n 2000 +--- every DROPPED line in the debug ring (WARN once per type, then DEBUG): +{"timestamp": "2026-09-23T11:57:31Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_deploy_started (severity info) — further app_deploy_started events are logged at DEBUG only", "source": "notifier.go:1163"} +{"timestamp": "2026-09-23T11:57:34Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_deployed (severity info) — further app_deployed events are logged at DEBUG only", "source": "notifier.go:1163"} +{"timestamp": "2026-09-23T12:00:00Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event health_change (severity warn) — further health_change events are logged at DEBUG only", "source": "notifier.go:1163"} +{"timestamp": "2026-09-23T12:00:08Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only", "source": "notifier.go:1163"} +{"timestamp": "2026-09-23T12:02:05Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only", "source": "notifier.go:1163"} +{"timestamp": "2026-09-23T12:02:30Z", "level": "WARN", "message": "notifier disabled (no hub configured): DROPPED event app_start_failed (severity warning) — further app_start_failed events are logged at DEBUG only", "source": "notifier.go:1163"} +{"timestamp": "2026-09-23T12:03:00Z", "level": "DEBUG", "message": "notifier disabled: dropped event app_start_failed (severity warning)", "source": "notifier.go:1167"} +--- counts per (type, level): {('app_deploy_started', 'WARN'): 1, ('app_deployed', 'WARN'): 1, ('health_change', 'WARN'): 1, ('app_update_undone', 'WARN'): 1, ('app_update_held', 'WARN'): 1} diff --git a/documentation/audits/undo-fleet-2026-09-23/28-9202-controller-log.txt b/documentation/audits/undo-fleet-2026-09-23/28-9202-controller-log.txt new file mode 100644 index 00000000..5d00d8be --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/28-9202-controller-log.txt @@ -0,0 +1,57 @@ +2026/09/23 11:52:35 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 3 [gokapi paperless-ngx privatebin]; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box +2026/09/23 11:52:40 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only +2026/09/23 11:55:00 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box +2026/09/23 11:55:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event controller_started (severity info) — further controller_started events are logged at DEBUG only +2026/09/23 11:57:31 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deploy_started (severity info) — further app_deploy_started events are logged at DEBUG only +2026/09/23 11:57:34 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_deployed (severity info) — further app_deployed events are logged at DEBUG only +2026/09/23 11:58:29 update.go:488: [INFO] [stacks] update vikunja: accepted — guarded update started +2026/09/23 11:58:29 update.go:1145: [INFO] [stacks] update vikunja: phase checking +2026/09/23 11:58:29 update.go:665: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T11:58:06Z (0s old, limit 24h0m0s) +2026/09/23 11:58:29 update.go:1145: [INFO] [stacks] update vikunja: phase safety-dump +2026/09/23 11:58:29 update.go:697: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) [] +2026/09/23 11:58:30 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.5 MiB +2026/09/23 11:58:30 update.go:1145: [INFO] [stacks] update vikunja: phase pinning +2026/09/23 11:58:30 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0) +2026/09/23 11:58:30 update.go:1145: [INFO] [stacks] update vikunja: phase pulling +2026/09/23 11:58:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying +2026/09/23 11:58:34 update.go:1145: [INFO] [stacks] update vikunja: phase copying +2026/09/23 11:58:34 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T115834Z in 471ms +2026/09/23 11:58:34 update.go:1145: [INFO] [stacks] update vikunja: phase copying +2026/09/23 11:58:35 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T115834Z in 431ms +2026/09/23 11:58:35 update.go:1145: [INFO] [stacks] update vikunja: phase starting +2026/09/23 11:58:35 update.go:1145: [INFO] [stacks] update vikunja: phase verifying +2026/09/23 12:00:00 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event health_change (severity warn) — further health_change events are logged at DEBUG only +2026/09/23 12:00:06 update.go:859: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy) +2026/09/23 12:00:06 update.go:851: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T120006Z/compose-logs.txt before stopping it (R-621) +2026/09/23 12:00:06 update.go:1145: [INFO] [stacks] update vikunja: phase undoing +2026/09/23 12:00:06 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/23 12:00:08 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true) +2026/09/23 12:00:08 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_undone (severity warning) — further app_update_undone events are logged at DEBUG only +2026/09/23 12:00:08 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 3s — the previous version is running on the data from before the update (the app's health check passed) +2026/09/23 12:00:31 update.go:488: [INFO] [stacks] update vikunja: accepted — guarded update started +2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase checking +2026/09/23 12:00:31 update.go:665: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T11:58:06Z (2m0s old, limit 24h0m0s) +2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase safety-dump +2026/09/23 12:00:31 update.go:697: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) [] +2026/09/23 12:00:31 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB +2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase pinning +2026/09/23 12:00:31 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0) +2026/09/23 12:00:31 update.go:1145: [INFO] [stacks] update vikunja: phase pulling +2026/09/23 12:00:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying +2026/09/23 12:00:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying +2026/09/23 12:00:33 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T120033Z in 476ms +2026/09/23 12:00:33 update.go:1145: [INFO] [stacks] update vikunja: phase copying +2026/09/23 12:00:34 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T120033Z in 413ms +2026/09/23 12:00:34 update.go:1145: [INFO] [stacks] update vikunja: phase starting +2026/09/23 12:00:34 update.go:1145: [INFO] [stacks] update vikunja: phase verifying +2026/09/23 12:02:05 update.go:859: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy) +2026/09/23 12:02:05 update.go:851: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T120205Z/compose-logs.txt before stopping it (R-621) +2026/09/23 12:02:05 update.go:1145: [INFO] [stacks] update vikunja: phase undoing +2026/09/23 12:02:05 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/23 12:02:05 undo.go:363: [ERROR] [stacks] update vikunja: the undo copy vikunja_vikunja_data.pre-update-20260923T120033Z has no finished-marker — it is cut off or missing; NOTHING is put back +2026/09/23 12:02:05 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T120033Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T120033Z}] +2026/09/23 12:02:05 update_guard.go:553: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T11:58:06Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched") +2026/09/23 12:02:05 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true) +2026/09/23 12:02:05 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_update_held (severity error) — further app_update_held events are logged at DEBUG only +2026/09/23 12:02:30 notifier.go:1163: [WARN] notifier disabled (no hub configured): DROPPED event app_start_failed (severity warning) — further app_start_failed events are logged at DEBUG only +2026/09/23 12:03:00 notifier.go:1167: [DEBUG] notifier disabled: dropped event app_start_failed (severity warning) diff --git a/documentation/audits/undo-fleet-2026-09-23/29-9202-teardown.txt b/documentation/audits/undo-fleet-2026-09-23/29-9202-teardown.txt new file mode 100644 index 00000000..06985d7b --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/29-9202-teardown.txt @@ -0,0 +1,27 @@ +14:03:26 --- evidence off the machine FIRST (R-320): the controller log of this phase +14:03:28 saved 57 lines to 28-9202-controller-log.txt +14:03:29 === teardown 9202: vikunja removed through the product +14:03:29 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +14:04:00 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +14:04:08 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +14:04:10 undo copies left: 0 +14:04:12 vikunja volumes left: 0 +14:04:14 images: Deleted: sha256:8711aa5bb344152bb7043c3c8a1b1013379492383d68414b3150fe3ae192ec51 +Deleted: sha256:3c889b99e3ccd594794c0c3738db5c8150535d0894b764667dfa041bb51c8dfe +Deleted: sha256:cf151044fdc4de3c2d630d3aea2fe854f70c0be91239a085c55e4fa8f16bef03 + +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +14:05:04 catalog cache: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462) +https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + +14:05:06 controller.yaml vs saved copy: IDENTICAL + +14:05:06 deployed apps now: ['gokapi', 'paperless-ngx', 'privatebin'] diff --git a/documentation/audits/undo-fleet-2026-09-23/30-9201-before-and-upgrade.txt b/documentation/audits/undo-fleet-2026-09-23/30-9201-before-and-upgrade.txt new file mode 100644 index 00000000..18d04b3c --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/30-9201-before-and-upgrade.txt @@ -0,0 +1,25 @@ +14:06:24 === 9201 BEFORE: standing apps, their compose + .felhom.yml checksums, image, controller +14:06:27 adventurelog 947bcad8f388 54cbd75330e3 Up 7 hours (healthy) +bentopdf 39679e28cdd6 3f63855ee03f Up 7 hours (healthy) +bookstack 190e20714347 7b57e5c06fb7 Up 7 hours (healthy) +calibre-web 295cee174968 8c3b9bec9c3a Up 7 hours (healthy) +docmost 9d01fd6fac49 da820f01b95c Up 7 hours (healthy) +kimai c17a7027190c 1fae6f4d5478 Up 7 hours (healthy) +opengist a44845f6db19 17d8ac371c07 Up 7 hours (healthy) +paperless-ngx ac1dd1354afb 19d1960ff453 Up 7 hours (healthy) +privatebin 79e7877dc3e8 eb4a2c7cc42a Up 7 hours (healthy) +romm cab4100bf0b0 2671d4102d5f Up 7 hours (healthy) +gitea.dooplex.hu/admin/felhom-controller:0.263.2 + +14:06:27 === upgrade 9201 to 0.264.0 (hub v0.120.0 is already live) +14:06:56 gitea.dooplex.hu/admin/felhom-controller:0.264.0 +gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up Less than a second (health: starting) +2026/09/23 12:06:31 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 10 [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin romm]; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box +"app_update_events_seeded": + +seeded: True | language: hu +enabled_events: None +top keys with notif: ['notifications', 'app_update_events_seeded'] + +{'cooldown_hours': 6} + diff --git a/documentation/audits/undo-fleet-2026-09-23/30-9201-standing-before.txt b/documentation/audits/undo-fleet-2026-09-23/30-9201-standing-before.txt new file mode 100644 index 00000000..ed77a19f --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/30-9201-standing-before.txt @@ -0,0 +1,11 @@ +adventurelog 947bcad8f388 54cbd75330e3 Up 7 hours (healthy) +bentopdf 39679e28cdd6 3f63855ee03f Up 7 hours (healthy) +bookstack 190e20714347 7b57e5c06fb7 Up 7 hours (healthy) +calibre-web 295cee174968 8c3b9bec9c3a Up 7 hours (healthy) +docmost 9d01fd6fac49 da820f01b95c Up 7 hours (healthy) +kimai c17a7027190c 1fae6f4d5478 Up 7 hours (healthy) +opengist a44845f6db19 17d8ac371c07 Up 7 hours (healthy) +paperless-ngx ac1dd1354afb 19d1960ff453 Up 7 hours (healthy) +privatebin 79e7877dc3e8 eb4a2c7cc42a Up 7 hours (healthy) +romm cab4100bf0b0 2671d4102d5f Up 7 hours (healthy) +gitea.dooplex.hu/admin/felhom-controller:0.263.2 diff --git a/documentation/audits/undo-fleet-2026-09-23/31-9201-repoint-and-controls.txt b/documentation/audits/undo-fleet-2026-09-23/31-9201-repoint-and-controls.txt new file mode 100644 index 00000000..ac274e24 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/31-9201-repoint-and-controls.txt @@ -0,0 +1,28 @@ +14:07:37 === 9201 onto the drill catalog (saved: controller.yaml.pre-undofleet) +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + +14:08:23 catalog cache: a8d8da9 DRILL glance: glanceapp/glance:v0.8.5 -> glanceapp/glance:v0.8.4 +https://@gitea.dooplex.hu/admin/app-catalog-drill.git + +14:08:25 control 1 — standing apps' compose + .felhom.yml unchanged by the drill sync: True (10 apps) +14:08:25 adventurelog 947bcad8f388 54cbd75330e3 Up 7 hours (healthy) +bentopdf 39679e28cdd6 3f63855ee03f Up 7 hours (healthy) +bookstack 190e20714347 7b57e5c06fb7 Up 7 hours (healthy) +calibre-web 295cee174968 8c3b9bec9c3a Up 7 hours (healthy) +docmost 9d01fd6fac49 da820f01b95c Up 7 hours (healthy) +kimai c17a7027190c 1fae6f4d5478 Up 7 hours (healthy) +opengist a44845f6db19 17d8ac371c07 Up 7 hours (healthy) +paperless-ngx ac1dd1354afb 19d1960ff453 Up 7 hours (healthy) +privatebin 79e7877dc3e8 eb4a2c7cc42a Up 7 hours (healthy) +romm cab4100bf0b0 2671d4102d5f Up 7 hours (healthy) + +14:08:25 control 2 — standing apps' badges: {'adventurelog': (None, None), 'bentopdf': (None, None), 'bookstack': (None, None), 'calibre-web': (None, None), 'docmost': (None, None), 'kimai': (None, None), 'opengist': (None, None), 'paperless-ngx': (None, None), 'privatebin': (None, None), 'romm': (None, None)} +14:08:25 control 3 — live catalog main: cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main diff --git a/documentation/audits/undo-fleet-2026-09-23/32-9201-prep.txt b/documentation/audits/undo-fleet-2026-09-23/32-9201-prep.txt new file mode 100644 index 00000000..3f5b381c --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/32-9201-prep.txt @@ -0,0 +1,24 @@ +14:08:40 === prep vikunja: deploy at the drill FROM pin +14:08:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +14:08:45 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'} +14:08:45 deployed: True +14:08:46 vikunja: register http=200 +14:08:46 vikunja: create project http=201 +14:08:46 vikunja: readback of the seeded project http=200 ok=True +14:08:46 C1 A: True +14:08:46 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +14:10:42 [4] backup idle; last=None +14:10:42 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill229ad3","created":"2026-09 +14:10:42 vikunja B: project readback=True attachment content readback=True +14:10:42 B reads back: True +14:10:46 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db'] +14:10:46 === glance: deploy at the drill FROM pin (v0.8.4) +14:10:46 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +14:11:06 [1] deployed, controller state=running, pinned={'glance': 'glanceapp/glance:v0.8.4'} +14:11:06 deployed: True +14:11:06 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +14:13:07 [4] backup idle; last=None +14:13:07 glance answers: True +14:13:09 vikunja probe: port: 3456 +glance probe: port: 8080 + diff --git a/documentation/audits/undo-fleet-2026-09-23/33-9201-hu-vikunja.txt b/documentation/audits/undo-fleet-2026-09-23/33-9201-hu-vikunja.txt new file mode 100644 index 00000000..caef2ac1 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/33-9201-hu-vikunja.txt @@ -0,0 +1,54 @@ +14:13:18 [5] drill commit 88b959602f91: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0) +14:13:23 badge caught up after 4.5 s +14:13:23 === vikunja: live undo by the product +14:13:23 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drill229ad3","created":"2026-09 +14:13:23 seed C written right before the Update: True +14:13:25 db before: tables=36 ledger=117 newest=SCHEMA_INIT +14:13:26 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +14:13:26 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +14:13:27 + 1.1s phase=pulling label=Új verzió letöltése… +14:13:31 + 4.8s phase=copying label=Az adatok másolása a frissítés előtt… +14:13:33 + 6.9s phase=starting label=Indítás az új verzióval… +14:13:33 + 7.4s phase=verifying label=Működés ellenőrzése… +14:15:03 + 97.7s phase=undoing label=Visszaállítás az előző változatra… +14:15:06 + 100.4s phase=undone label=Visszaállítva az előző változatra +14:15:06 END phase=undone err=None hold=None +14:15:09 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]} +14:15:09 vikunja: readback of the seeded project http=200 ok=True +14:15:09 vikunja B: project readback=True attachment content readback=True +14:15:09 vikunja B: project readback=True attachment content readback=True +14:15:11 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT) +14:15:11 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 14:15 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +14:15:13 leftover copies: 0 +14:15:13 === held: the same bad step pressed again, with one undo copy cut off +14:15:13 === vikunja: a cut-off undo copy +14:15:14 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +14:15:14 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +14:15:15 + 1.1s phase=pulling label=Új verzió letöltése… +14:15:16 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt… +14:15:17 + 3.2s phase=starting label=Indítás az új verzióval… +14:15:17 + 3.7s phase=verifying label=Működés ellenőrzése… +14:15:20 >>> finished-marker removed from one copy: +copy: vikunja_vikunja_data.pre-update-20260923T121516Z +total 12 +drwxr-xr-x 3 root root 4096 Sep 23 12:15 . +drwxr-xr-x 1 root root 4096 Sep 23 12:15 .. +drwxr-xr-x 2 root root 4096 Sep 23 12:15 data +-rw-r--r-- 1 root root 0 Sep 23 12:15 felhom-undo-complete +data + +14:16:48 + 94.3s phase=undoing label=Visszaállítás az előző változatra… +14:16:49 + 95.3s phase=failed label=A frissítés nem sikerült +14:16:49 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' +14:16:51 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."}} +14:16:51 box language -> en: http 302 +14:16:51 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes."}} +14:16:51 box language -> hu: http 302 +14:16:53 copies kept: vikunja_vikunja_data.pre-update-20260923T121516Z +vikunja_vikunja_db.pre-update-20260923T121516Z + +14:16:56 2026/09/23 12:15:06 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true) +2026/09/23 12:15:06 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 2s — the previous version is running on the data from before the update (the app's health check passed) +2026/09/23 12:16:48 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T121516Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T121516Z}] +2026/09/23 12:16:48 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true) + diff --git a/documentation/audits/undo-fleet-2026-09-23/34-9201-language-en.txt b/documentation/audits/undo-fleet-2026-09-23/34-9201-language-en.txt new file mode 100644 index 00000000..68e3add0 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/34-9201-language-en.txt @@ -0,0 +1,4 @@ +14:17:15 box language -> en: http 302 + +language = en + diff --git a/documentation/audits/undo-fleet-2026-09-23/35-9201-en-glance.txt b/documentation/audits/undo-fleet-2026-09-23/35-9201-en-glance.txt new file mode 100644 index 00000000..b1e3d8b2 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/35-9201-en-glance.txt @@ -0,0 +1,35 @@ +14:27:49 === glance (box language en): the failing edge v0.8.4 -> v0.8.5, probe 8080 -> 8999 (drill only) +14:27:49 [5] drill commit d36a5fc5294b: glance glanceapp/glance:v0.8.4 -> glanceapp/glance:v0.8.5 (push rc=0) +14:27:54 badge caught up after 4.4 s +14:27:54 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'The update has started – the card shows how it is going'} +14:27:54 + 0.0s phase=safety-dump label=Database snapshot… +14:27:55 + 1.1s phase=pulling label=Downloading the new version… +14:27:56 + 2.1s phase=copying label=Copying the data before the update… +14:27:57 + 2.7s phase=starting label=Starting the new version… +14:27:57 + 3.2s phase=verifying label=Checking that it works… +14:29:28 + 93.5s phase=undoing label=Putting the previous version back… +14:29:35 + 100.4s phase=undone label=Put back to the previous version +14:29:35 END phase=undone err=None hold=None +14:29:35 PAGE: {"hu": {"undone_line": "A(z) glance frissítése 2026-09-23 14:29-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of glance at 2026-09-23 14:29 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +14:29:37 leftover copies: 0 +14:29:37 === glance held: pressed again, one undo copy cut off in `verifying` +14:29:37 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'The update has started – the card shows how it is going'} +14:29:37 + 0.0s phase=safety-dump label=Database snapshot… +14:29:38 + 0.5s phase=pulling label=Downloading the new version… +14:29:39 + 2.1s phase=copying label=Copying the data before the update… +14:29:40 + 2.6s phase=starting label=Starting the new version… +14:29:41 + 3.2s phase=verifying label=Checking that it works… +14:29:43 >>> finished-marker removed from one copy: +copy: glance_glance_config.pre-update-20260923T122939Z +data + +14:31:11 + 93.8s phase=undoing label=Putting the previous version back… +14:31:12 + 94.3s phase=failed label=The update did not succeed +14:31:12 END phase=failed err='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.' hold='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.' +14:31:12 PAGE: {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes."}} +14:31:12 GET /api/stacks/glance (box en): label='The update did not succeed' err='The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes.' +14:31:14 2026/09/23 12:15:06 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true) +2026/09/23 12:16:48 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true) +2026/09/23 12:29:34 undo.go:474: [INFO] [stacks] update glance: event app_update_undone (hold recorded: true) +2026/09/23 12:31:11 undo.go:474: [INFO] [stacks] update glance: event app_update_held (hold recorded: true) + diff --git a/documentation/audits/undo-fleet-2026-09-23/40-notification-log.txt b/documentation/audits/undo-fleet-2026-09-23/40-notification-log.txt new file mode 100644 index 00000000..69d71e76 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/40-notification-log.txt @@ -0,0 +1,11 @@ +# notification_log, customer demo-hp, app_update_* (copy of hub.db+wal+shm 2026-09-23T12:31:46Z); message cut to 90 chars +name id customer_id event_type severity message status error_message created_at channel +id|created_at|event_type|severity|channel|status|error_message|message +937|2026-09-23 12:15:06|app_update_undone|warning|operator|sent||A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaáll +938|2026-09-23 12:15:07|app_update_undone|warning|customer|sent||A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaáll +939|2026-09-23 12:16:49|app_update_held|error|operator|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál +940|2026-09-23 12:16:49|app_update_held|error|customer|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál +942|2026-09-23 12:29:35|app_update_undone|warning|operator|sent||A(z) glance frissítése 2026-09-23 14:29-kor nem sikerült. A doboz automatikusan visszaállí +943|2026-09-23 12:29:35|app_update_undone|warning|customer|sent||A(z) glance frissítése 2026-09-23 14:29-kor nem sikerült. A doboz automatikusan visszaállí +944|2026-09-23 12:31:12|app_update_held|error|operator|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál +946|2026-09-23 12:31:12|app_update_held|error|customer|sent||A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat ál diff --git a/documentation/audits/undo-fleet-2026-09-23/41-mails-as-received.md b/documentation/audits/undo-fleet-2026-09-23/41-mails-as-received.md new file mode 100644 index 00000000..f7201908 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/41-mails-as-received.md @@ -0,0 +1,55 @@ +# The mails as received — 2026-09-23, guest 9201 (customer demo-hp), read with the Gmail connector + +Household address = the demo household's catch-all address (`drill@`); operator = `admin@`. Both land in the +@felhom.eu catch-all mailbox. Bodies quoted verbatim (plain-text part). The `notification_log` rows are in +`40-notification-log.txt` (ids 937–946, all `sent`). + +## Round 1 — box language Hungarian, app vikunja + +**Household, undone** (12:15:08Z, log row 938) +Subject: `[Felhom] Figyelmeztetés: vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut` +``` +vikunja: a frissítés nem sikerült, az alkalmazás a korábbi változattal fut +- Szint: Figyelmeztetés +- Típus: app_update_undone +- Üzenet: A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el, nincs teendőd. +- Megjegyzés: {"app":"vikunja","stack_name":"vikunja","from":{"vikunja":"vikunja/vikunja:2.3.0"},"to":{"vikunja":"vikunja/vikunja:2.6.0"},"at":"2026-09-23T12:15:06Z"} +``` +**Operator, undone** (12:15:07Z, row 937) Subject: `[Felhom] ⚠️ demo-hp: app_update_undone` — Message: the same Hungarian sentence. + +**Household, held** (12:16:50Z, row 940) +Subject: `[Felhom] Hiba: vikunja: az alkalmazás leállítva, visszaállítás szükséges` +``` +- Üzenet: A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. +- Megjegyzés: {...,"copy_tier":1,"copy_date":"2026-09-23T12:13:01Z","copy_holds":"a beállításokat, az adatbázist és az adatköteteket tartalmazza"} +``` +**Operator, held** (12:16:49Z, row 939) Subject: `[Felhom] 🔴 demo-hp: app_update_held` — Message: the same Hungarian hold sentence. + +## Box language switched to English once (14:17:15 local); hub reports received 14:17:18 and 14:22:42 + +## Round 2 — box language English, app glance + +**Household, undone** (12:29:35Z, row 943) +Subject: `[Felhom] Warning: glance: the update did not work; the app runs on its previous version` +``` +glance: the update did not work; the app runs on its previous version +- Level: Warning +- Type: app_update_undone +- Message: The update of glance at 2026-09-23 14:29 did not work. The box put back the previous version and its data automatically — nothing was lost, and there is nothing you need to do. +- Note: {"app":"glance","stack_name":"glance","from":{"glance":"glanceapp/glance:v0.8.4"},"to":{"glance":"glanceapp/glance:v0.8.5"},"at":"2026-09-23T12:29:34Z"} +``` +**Operator, undone** (12:29:35Z, row 942) — Message in Hungarian (the operator's wire text), as designed. + +**Household, held** (12:31:12Z, row 946) +Subject: `[Felhom] Error: glance: the app is stopped and needs a restore` +``` +- Message: The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of glance at 2026-09-23 14:31 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:12 — this copy holds the settings, the database and the data volumes. +- Note: {...,"copy_holds":"a beállításokat, az adatbázist és az adatköteteket tartalmazza"} <- Hungarian in the raw Note line (row filed) +``` +**Operator, held** (12:31:12Z, row 944) Subject: `[Felhom] 🔴 demo-hp: app_update_held` — Hungarian hold sentence. + +## What else arrived +- The operator also received `app_start_failed` for each held app (12:17:12Z vikunja, 12:31:12Z glance) — pre-existing + behaviour: a held app is a stopped app. The household did not (the type is not in its list). +- **Per-app cooldown, live:** glance's undone mail (row 943) went out 14 minutes after vikunja's (row 938) inside + the 6-hour customer cooldown. Before hub v0.120.0 the key was `customer:type`, and it would have been swallowed. diff --git a/documentation/audits/undo-fleet-2026-09-23/48-9201-controller-log.txt b/documentation/audits/undo-fleet-2026-09-23/48-9201-controller-log.txt new file mode 100644 index 00000000..dce62a4b --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/48-9201-controller-log.txt @@ -0,0 +1,12 @@ +2026/09/23 12:06:31 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 10 [adventurelog bentopdf bookstack calibre-web docmost kimai opengist paperless-ngx privatebin romm]; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box +2026/09/23 12:07:41 undo.go:508: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box +2026/09/23 12:15:06 undo.go:474: [INFO] [stacks] update vikunja: event app_update_undone (hold recorded: true) +2026/09/23 12:15:06 undo.go:405: [INFO] [stacks] update vikunja: UNDONE in 2s — the previous version is running on the data from before the update (the app's health check passed) +2026/09/23 12:16:48 update.go:871: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T121516Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T121516Z}] +2026/09/23 12:16:48 update_guard.go:553: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T12:13:01Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched") +2026/09/23 12:16:48 undo.go:474: [INFO] [stacks] update vikunja: event app_update_held (hold recorded: true) +2026/09/23 12:29:34 undo.go:474: [INFO] [stacks] update glance: event app_update_undone (hold recorded: true) +2026/09/23 12:29:34 undo.go:405: [INFO] [stacks] update glance: UNDONE in 7s — the previous version is running on the data from before the update (the app's health check passed) +2026/09/23 12:31:11 update.go:871: [ERROR] [stacks] update glance: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{glance_glance_config glance_glance_config.pre-update-20260923T122939Z}] +2026/09/23 12:31:11 update_guard.go:553: [WARN] [backup] glance is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T12:12:42Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched") +2026/09/23 12:31:11 undo.go:474: [INFO] [stacks] update glance: event app_update_held (hold recorded: true) diff --git a/documentation/audits/undo-fleet-2026-09-23/49-9201-teardown.txt b/documentation/audits/undo-fleet-2026-09-23/49-9201-teardown.txt new file mode 100644 index 00000000..bc609ee5 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/49-9201-teardown.txt @@ -0,0 +1,43 @@ +14:32:19 evidence first: 12 log lines saved +14:32:19 === remove vikunja through the product +14:32:19 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +14:32:51 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +14:32:59 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +14:32:59 === remove glance through the product +14:32:59 [X] stop -> 200 {'ok': True, 'message': 'Stack glance stop completed'} +14:33:30 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'glance', 'volumes_removed': ['glance_glance_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alka +14:33:38 [X] after remove: deployed=False leftovers='/opt/docker/stacks/glance' +14:33:40 undo copies left: 0 +14:33:42 volumes left: 0 +14:33:44 images: 18 + +14:33:44 box language -> hu: http 302 +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +14:34:35 catalog cache: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462) +https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + +14:34:37 controller.yaml vs saved: IDENTICAL + +14:34:39 language: hu + +14:34:41 standing apps' files identical to before: True +14:34:41 adventurelog 947bcad8f388 54cbd75330e3 Up 23 minutes (healt +bentopdf 39679e28cdd6 3f63855ee03f Up 8 hours (healthy) +bookstack 190e20714347 7b57e5c06fb7 Up 23 minutes (healt +calibre-web 295cee174968 8c3b9bec9c3a Up 22 minutes (healt +docmost 9d01fd6fac49 da820f01b95c Up 22 minutes (healt +kimai c17a7027190c 1fae6f4d5478 Up 22 minutes (healt +opengist a44845f6db19 17d8ac371c07 Up 22 minutes (healt +paperless-ngx ac1dd1354afb 19d1960ff453 Up 22 minutes (healt +privatebin 79e7877dc3e8 eb4a2c7cc42a Up 22 minutes (healt +romm cab4100bf0b0 2671d4102d5f Up 21 minutes (healt + +14:34:41 deployed: ['adventurelog', 'bentopdf', 'bookstack', 'calibre-web', 'docmost', 'kimai', 'opengist', 'paperless-ngx', 'privatebin', 'romm'] diff --git a/documentation/audits/undo-fleet-2026-09-23/50-floor-0.264.0.txt b/documentation/audits/undo-fleet-2026-09-23/50-floor-0.264.0.txt new file mode 100644 index 00000000..5e63656d --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/50-floor-0.264.0.txt @@ -0,0 +1,15 @@ +=== FLOOR -> 0.264.0, MinAgent 0.131.0 declared (2026-09-23 14:35:08) +POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set +--- read back from the hub: +"min_controller_version" value="0.264.0" +"min_agent" value="0.131.0" +"min_agent" value="0.131.0" +--- watching the two demo boxes (up to 3 min) + +10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up About a minute (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up About an hour (healthy) + +20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up About a minute (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.264.0 Up 10 seconds (healthy) +--- hub log: +2026/09/23 14:35:11 [INFO] managed floor SERVED for demo-felhom: floor 0.264.0, agent requirement "0.131.0" from declared (golden 0.258.0) +--- hosts page: +demo-felhom-8363b5 +demo-hp-bb76ea +drill-r50-0a4f9a diff --git a/documentation/audits/undo-fleet-2026-09-23/README.md b/documentation/audits/undo-fleet-2026-09-23/README.md new file mode 100644 index 00000000..e8cf3a18 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/README.md @@ -0,0 +1,50 @@ +# The undo reaches the fleet, and the household is told — controller v0.264.0 + hub v0.120.0, 2026-09-23 + +**Method: endpoint-level.** Every act is the endpoint the UI invokes (deploy, backup, sync, rescan, update, +remove, the language switch); pages are fetched as HTML (`/apps/?lang=hu|en`); the mails are read from +the @felhom.eu catch-all mailbox with the Gmail connector; the hub's rows from a copy of `hub.db`+`-wal`+`-shm`. +No browser. The product performs every undo and every hold; the harness (`walk.py`, `live.py`, `bakeoff.py`, +copied from `undo-live-2026-09-23/`, now selectable by `GUEST=9201|9202`) only presses and reads. + +## Order of events + +| # | what | evidence | +|---|---|---| +| 0 | floor 0.263.2 (Part 0) | `00-floor-0.263.2.txt` | +| 1 | hub v0.120.0 deployed by manifest bump + ArgoCD, **before any box ran v0.264.0**; one-time add-only migration changed all 3 `enabled_events` rows | `10-enabled-events-before.txt`, `11-hub-0.120.0-deploy-and-migration.txt` | +| 2 | 9202 (no hub): v0.264.0; R-646 backfill recorded 3; vikunja undone then held; pages hu/en; WARN once per type, then DEBUG | `20-*` … `29-*` | +| 3 | 9201 (hub on): v0.264.0; backfill recorded 10; drill repoint with three controls; vikunja undone+held (box hu); language → en; glance undone+held (box en) | `30-*` … `35-*` | +| 4 | 8 mails as received + 8 `notification_log` rows | `40-notification-log.txt`, `41-mails-as-received.md` | +| 5 | teardown of 9201, floor 0.264.0 read back | `48-*`, `49-*`, `50-floor-0.264.0.txt` | + +## Results + +| proof | result | +|---|---| +| migration (hub) | `demo-felhom` kept its 10 types and gained 2; `demo-hp` and `_resend-rotation-test` (`null` / `[]` = nothing on) gained the 2; guard row set 11:49:03Z; log names all three | +| R-646 | `applied-meta backfill (R-646): recorded 3 [gokapi paperless-ngx privatebin]` (9202); `recorded 10 […]` (9201); skipped 0 on both | +| undone (9202, vikunja) | `undone` at +99 s; seeds A, B, C back; ledger equal; page hu *„A(z) vikunja frissítése … — semmi nem veszett el."* / en *"The update of vikunja … — nothing was lost."*; API `?lang=en` label *"Put back to the previous version"* | +| held (9202, cut-off copy) | `failed` 1 s into `undoing`; hold sentence hu on the hu page and **en on the en page (R-606)**; both copies kept | +| R-620 | WARN once each for `controller_started`, `app_deploy_started`, `app_deployed`, `health_change`, `app_update_undone`, `app_update_held`, `app_start_failed`; the second `app_start_failed` at 12:03:00Z is DEBUG | +| mails (9201) | 4 household + 4 operator, rows 937–946 all `sent`; subjects name the app; household text in the box language (hu for vikunja, en for glance); operator text Hungarian. Quoted in `41-mails-as-received.md` | +| per-app cooldown | glance's undone mail went 14 min after vikunja's, inside the 6 h customer cooldown | +| standing apps (9201) | compose + `.felhom.yml` byte-identical before/after (10 apps), catalog back on live `cfcfe52`, `controller.yaml` identical to the saved copy. **But not untouched:** the harness's two „Mentés most" presses (whole-box backups) stopped and restarted 9 of 10 standing apps for their volume dumps — R-648 | + +## Found + +- **R-647** — (1) a reader in the other language than the box reads a HELD sentence in the box's language + (seen: 9202 box en, `?lang=hu` page showed English); (2) the English mail's raw `Note:` line carries the + Hungarian `copy_holds` phrase; (3) two log wordings (`hold recorded: true` on an undone event; `severity warn` + printing a health status). +- **R-648** — the drill's backup press is whole-box. + +## Teardown — three layers + +- **machine:** 9202 — vikunja removed through the product (0 volumes, 0 undo copies), test images removed, + `controller.yaml` identical to `controller.yaml.pre-undofleet`, catalog on live `cfcfe52`, apps as at the + start. 9201 — vikunja and glance removed through the product (0 volumes, 0 copies), test images removed, + language back to `hu`, config identical to the saved copy, catalog on live `cfcfe52`, the 10 standing apps + deployed and healthy. Both stay on controller 0.264.0 (= the floor). +- **host:** nothing provisioned on demo-hp. +- **hub:** v0.120.0 deployed (the one planned change); floor 0.264.0 / MinAgent 0.131.0; no customer reset. +- **drill repo:** reset to live `main` (`cfcfe52`), read back from the remote. diff --git a/documentation/audits/undo-fleet-2026-09-23/bakeoff.py b/documentation/audits/undo-fleet-2026-09-23/bakeoff.py new file mode 100644 index 00000000..7bd759f9 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/bakeoff.py @@ -0,0 +1,266 @@ +#!/usr/bin/env python3 +"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on +the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202, +drill catalog. + + F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure. + D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and + recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads. + +EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync, +rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is +lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own +Start supplies the app's secrets (they are never decrypted here). + +Usage: python3 bakeoff.py stages: prep fcopy break update undoF undoD cutoff state +""" +import json, os, sys, time +sys.path.insert(0, ".") +import walk as w +from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \ + romm_verify_b, vik_seed_b, vik_verify_b + +EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999), + "romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999), + "vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)} +UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost", + "vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja", + "romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"} +SUFFIX = ".pre-undo" + + +def seed_b(app, sub, A): + return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A) + + +def verify_b(app, sub, A, B): + r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B) + return all(r.values()) if isinstance(r, dict) else r + + +def db_state(app): + """The table count and the migration ledger, asked of the engine (or the SQLite file, read-only).""" + if app == "docmost": + return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; " + "docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip() + if app == "romm": + q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '" + base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45") + alembic = q("select version_num from alembic_version") + return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip() + return w.guest("""python3 - <<'PY' +import sqlite3 +c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True) +t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0] +m=c.execute("select count(*), max(id) from migration").fetchone() +print(f"tables={t} ledger={m[0]} newest={m[1]}") +PY""").strip() + + +def volumes(app): + return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split() + if not v.endswith(SUFFIX)] + + +def containers(app): + return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split() + + +def stage_prep(app): + s = {"app": app, "sub": SUB[app]} + w.say(f"=== prep {app}: deploy at the drill FROM pin") + w.say("deployed:", w.deploy(app, s["sub"])) + s["seedA"] = FX[app].seed(w, s["sub"], w.say) + w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say)) + w.backup_now(app) + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60) + s["seedB"] = seed_b(app, s["sub"], s["seedA"]) + w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"])) + s["volumes"] = volumes(app) + w.say("named volumes:", s["volumes"]) + save(app, s) + + +def stage_fcopy(app): + """F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and + separately a tar, for the rate), start. Measures the EXTRA downtime and the disk.""" + s = load(app) + s["state_before_update"] = db_state(app) + w.say("db state before the update:", s["state_before_update"]) + vols = s["volumes"] + w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml")) + script = f""" +set -u +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}') +t0=$(date +%s.%N) +docker stop $cs >/dev/null +t1=$(date +%s.%N) +for v in {' '.join(vols)}; do + docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1 + docker volume create "$v{SUFFIX}" >/dev/null + a=$(date +%s.%N) + docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v" + b=$(date +%s.%N) + bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1) + files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l') + cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1) + cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l') + mkdir -p /root/bk; c=$(date +%s.%N) + docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N) + tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar + python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))" +done +t2=$(date +%s.%N) +docker start $cs >/dev/null +t3=$(date +%s.%N) +python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))" +df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}' +""" + out = w.guest(script, timeout=1800) + w.say(out) + t0 = time.time() + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + w.say(f"front door answering again {round(time.time()-t0,1)}s after the start") + s["fcopy"] = out + s["fcopy_at"] = ts() + save(app, s) + + +def stage_break(app): + s = load(app) + frm, to, p0, p1 = EDGE[app] + s["break_commit"] = break_edge(app, frm, to, p0, p1) + w.say("badge caught up after", w.sync_rescan(app, to), "s") + save(app, s) + + +def stage_update(app): + s = load(app) + r = w.press_update(app, poll=1.0) + w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False)) + s.setdefault("updates", []).append(r) + save(app, s) + + +LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold' +docker restart felhom-controller >/dev/null; sleep 15""" + +# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre--compose.yml, taken +# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s +# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new +# version again. That is R-639 seen live, and why the product undo keeps its own copies. +PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml +grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }} +cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml +sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4""" + + +def lift_and_pinback(app): + frm, to, _, _ = EDGE[app] + t = time.time() + w.say(w.guest(LIFT.format(app=app))) + w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to))) + w.login() + return round(time.time() - t, 1) + + +def stage_undoF(app): + s = load(app) + w.say(f"=== undo by F: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + vols = s["volumes"] + script = f""" +t0=$(date +%s.%N) +cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null +for v in {' '.join(vols)}; do + docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v" +done +python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))" +""" + w.say(w.guest(script, timeout=1800)) + t0 = time.time() + c, d = w.ctl("POST", f"/api/stacks/{app}/start") + w.say("product start ->", c) + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_undoD(app): + """D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy.""" + s = load(app) + w.say(f"=== undo by D: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + if app == "vikunja": + w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it " + "is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.") + return + c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse + time.sleep(8) + if app == "docmost": + load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1""" + marker = "-- PostgreSQL database dump complete" + else: + load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1""" + marker = "-- Dump completed" + script = f""" +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B" +docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null +t0=$(date +%s.%N) +if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi +{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out +python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))" +docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null +""" + w.say(w.guest(script, timeout=900)) + t0 = time.time() + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_cutoff(app): + """The cut-off copy, both methods, DETECTED before anything is loaded or swapped. + F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker + the build would write only after cp exits 0. D: the safety dump cut in half — the marker check.""" + s = load(app) + big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0)) + script = f""" +v={big} +docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null +# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured +# on docmost 2026-09-23, rc 137 with a complete copy and a marker). +docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null +sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1 +echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)" +echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap" +docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1) +if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql + echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load" + rm -f /root/cut.sql +else echo "D: no safety dump exists for this app (R-641)"; fi +""" + w.say(f"=== cut-off copy: {app} (largest volume {big})") + out = w.guest(script, timeout=600) + w.say(out) + s["cutoff"] = out + save(app, s) + + +if __name__ == "__main__": + w.login() + {"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update, + "undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff, + "state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/undo-fleet-2026-09-23/fixtures.py b/documentation/audits/undo-fleet-2026-09-23/fixtures.py new file mode 100644 index 00000000..3a681e96 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/fixtures.py @@ -0,0 +1,1008 @@ +#!/usr/bin/env python3 +"""Box-side seed/verify fixtures for walk.py, guest 9202. + +THE ONE RULE (R-156), carried verbatim from `app-catalog-felhom.eu/scripts/upgrade_fixtures.py`: +*nothing is ever seeded into a volume by hand.* Every seed here goes in through the app's OWN +interface — its HTTP API through the household's real front door (traefik, `Host: .`), +or its own CLI running inside its own container. A raw SQL INSERT or a planted file is never used. + +If an app has no non-browser route, its fixture returns None and the edge is recorded +`inconclusive — no non-browser seed route`, WITH WHAT WAS TRIED. That is a result, not a gap. + +Each fixture: + seed(w, sub, say) -> an opaque token, or None + verify(w, sub, tok, say) -> True / False +verify() must ask the APP, never the filesystem: a migration is supposed to rewrite files. +Where a fixture can prove itself (a negative control that must read as absent) it does so on EVERY +call, so a readback that has broken into always saying "found" fails instead of passing everything. +""" +import base64, json, re, secrets, time + + +def _gx(w, container, *cmd, timeout=240): + """Run a command inside the app's OWN container on 9202 (its own CLI, not our SQL).""" + import shlex + line = " ".join(shlex.quote(c) for c in cmd) + return w.guest(f"docker exec {container} {line} 2>&1", timeout=timeout) + + +# ============================================================================================= +class PrivateBin: + """PrivateBin's own JSON API. A paste is a POST and reading it back is a GET — an + application-level round trip. File-backed, no database: this single seed IS the file half.""" + sub = "paste" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200",)): + return None + marker = "upg-" + secrets.token_hex(8) + ct = base64.b64encode(marker.encode()).decode() + body = json.dumps({ + "v": 2, + "adata": [[base64.b64encode(secrets.token_bytes(16)).decode(), + base64.b64encode(secrets.token_bytes(8)).decode(), + 100000, 256, 128, "aes", "gcm", "none"], "plaintext", 0, 0], + "ct": ct, "meta": {"expire": "never"}}) + rc, code, out = w.app_curl(sub, "/", "-H", "X-Requested-With: JSONHttpRequest", + "-H", "Content-Type: application/json", + data=body, method="POST") + try: + j = json.loads(out) + except Exception: + say(f" privatebin: POST returned non-JSON (http {code}): {out[:200]}") + return None + if j.get("status") != 0 or not j.get("id"): + say(f" privatebin: POST refused: {out[:250]}") + return None + say(f" privatebin: seeded paste id={j['id']}") + return {"id": j["id"], "marker": ct} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200",), tries=36): + return False + # negative control, every call: a paste id that cannot exist must NOT read back + rc, code, out = w.app_curl(sub, "/?pasteid=" + secrets.token_hex(8), + "-H", "X-Requested-With: JSONHttpRequest") + if t["marker"] in out: + say(" privatebin: READBACK UNUSABLE — a paste id that cannot exist returned the marker") + return False + rc, code, out = w.app_curl(sub, "/?pasteid=" + t["id"], + "-H", "X-Requested-With: JSONHttpRequest") + got = code == "200" and t["marker"] in out + say(f" privatebin: readback http={code} marker_present={got}") + return got + + +# ============================================================================================= +class Docmost: + """Docmost's own REST API: create the first workspace+user, then prove the account survives by + asking the app to AUTHENTICATE it. Login is version-stable across the API churn.""" + sub = "docs" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302", "404")): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + body = json.dumps({"workspaceName": "drill", "name": "drill", "email": email, "password": pw}) + rc, code, out = w.app_curl(sub, "/api/auth/setup", "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" docmost: /api/auth/setup http={code} rc={rc}") + if code not in ("200", "201"): + say(f" docmost: setup refused: {out[:250]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "404"), tries=36): + return False + # negative control: a password that was never set must NOT authenticate + bad = json.dumps({"email": t["email"], "password": "definitely-" + secrets.token_hex(8)}) + rc, code, _ = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" docmost: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"email": t["email"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" docmost: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" docmost: login body {out[:200]}") + return ok + + +# ============================================================================================= +class BookStack: + """BookStack mints no API token without a browser, so BOTH halves go through `php artisan` — + BookStack's OWN CLI, inside its own container, against its own User model. + + The exit code carries no information here (`bookstack:reset-mfa` exits 1 for a user it FOUND + and for one it did not), so the discriminator is the OUTPUT: the positive sentence required and + the not-found sentence required absent. The negative control runs on every verify. + + LIMITATION (R-460): this seeds the DATABASE half only. The FILE half needs the API token the + app cannot mint headlessly — so a bookstack edge is at best HALF-proven here. + """ + sub = "wiki" + + def _artisan(self, w, *args): + for path in ("/app/www/artisan", "/var/www/html/artisan"): + out = _gx(w, "bookstack", "php", path, *args) + if "Could not open input file" not in out: + return " ".join(out.split()) + return " ".join(out.split()) + + def _lookup(self, w, email): + out = self._artisan(w, "bookstack:reset-mfa", f"--email={email}") + found = f"Email: {email}" in out + missing = "could not be found" in out + if found == missing: + return None, out + return found, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + out = self._artisan(w, "bookstack:create-admin", f"--email={email}", + f"--name=drill-{secrets.token_hex(3)}", f"--password={pw}") + say(f" bookstack: artisan create-admin :: {out[:140]}") + if "successfully created" not in out: + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" bookstack: the app never served /login") + return False + absent, _ = self._lookup(w, f"nobody-{secrets.token_hex(6)}@gate.invalid") + if absent is not False: + say(f" bookstack: READBACK UNUSABLE — an email that cannot exist did not read absent ({absent})") + return False + found, out = self._lookup(w, t["email"]) + say(f" bookstack: readback of the seeded account found={found} :: {out[:140]}") + return found is True + + +# ============================================================================================= +class Gitea: + """Gitea's own admin CLI creates the first user; its own REST API (basic auth) then creates a + repository and reads it back. Both are the app's own interfaces.""" + sub = "git" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = _gx(w, "gitea", "su", "git", "-c", + f"gitea admin user create --username {user} --password {pw} " + f"--email {user}@gate.invalid --admin --must-change-password=false") + say(f" gitea: admin user create :: {' '.join(out.split())[:140]}") + if "has been successfully created" not in out and "successfully created" not in out: + return None + repo = "drillrepo" + secrets.token_hex(3) + rc, code, body = w.app_curl(sub, "/api/v1/user/repos", "-u", f"{user}:{pw}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": repo, "private": True}), method="POST") + say(f" gitea: create repo http={code}") + if code not in ("201", "200"): + say(f" gitea: repo refused {body[:200]}") + return None + return {"user": user, "pw": pw, "repo": repo} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + rc, code, _ = w.app_curl(sub, f"/api/v1/repos/{t['user']}/nope{secrets.token_hex(4)}", + "-u", f"{t['user']}:{t['pw']}") + if code == "200": + say(" gitea: READBACK UNUSABLE — a repo that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/repos/{t['user']}/{t['repo']}", + "-u", f"{t['user']}:{t['pw']}") + ok = code == "200" and t["repo"] in body + say(f" gitea: readback of the seeded repo http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Navidrome: + """Navidrome's own REST API: create the first admin through /auth/createAdmin, then prove the + account survives by logging in through the same door.""" + sub = "music" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, out = w.app_curl(sub, "/auth/createAdmin", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), method="POST") + say(f" navidrome: createAdmin http={code}") + if code not in ("200", "201"): + say(f" navidrome: refused {out[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" navidrome: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"username": t["user"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" navidrome: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vaultwarden: + """Vaultwarden's own account API: register an account, then prove it survives by asking the app + to issue a token for it (its own login endpoint, the household's own route).""" + sub = "vault" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/alive", want=("200",)): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + # Vaultwarden stores an already-hashed master key; the value is opaque to the server. + key = base64.b64encode(secrets.token_bytes(32)).decode() + body = json.dumps({"email": email, "name": "drill", "masterPasswordHash": key, + "key": "0." + base64.b64encode(secrets.token_bytes(48)).decode(), + "kdf": 0, "kdfIterations": 600000}) + rc, code, out = w.app_curl(sub, "/api/accounts/register", + "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" vaultwarden: register http={code}") + if code not in ("200", "204"): + say(f" vaultwarden: refused {out[:250]}") + return None + return {"email": email, "key": key} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/alive", want=("200",), tries=36): + return False + def login(pwhash): + return w.app_curl(sub, "/identity/connect/token", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=("grant_type=password&scope=api%20offline_access" + f"&client_id=web&deviceType=9&deviceIdentifier=drill" + f"&deviceName=drill&username={t['email']}&password={pwhash}"), + method="POST") + rc, code, _ = login(base64.b64encode(secrets.token_bytes(32)).decode()) + if code == "200": + say(" vaultwarden: READBACK UNUSABLE — a wrong master key authenticated") + return False + rc, code, out = login(t["key"].replace("+", "%2B").replace("=", "%3D").replace("/", "%2F")) + ok = code == "200" and "access_token" in out + say(f" vaultwarden: token for the seeded account http={code} ok={ok}") + if not ok: + say(f" vaultwarden: body {out[:200]}") + return ok + + +# ============================================================================================= +class Django: + """A Django app's OWN management CLI, inside its own container, against its own User model. + + Same category as BookStack's `php artisan`: the app's own code and its own ORM, never a raw SQL + INSERT and never a planted file (R-156). `createsuperuser --noinput` is Django's own documented + non-interactive route, and the readback asks the SAME ORM whether the account exists. + + THE FIXTURE PROVES ITSELF ON EVERY CALL: each verify() also asks for a username that cannot + exist and requires the answer False. A readback that has broken into always saying True + therefore fails instead of passing everything. + + LIMITATION, recorded rather than papered over: this seeds the DATABASE half only. An app whose + data is also FILES (adventurelog's images) has a file half this fixture does not touch. + """ + + def __init__(self, container, sub, ready_path="/", ready=("200", "302", "301", "404"), + python="python", workdir=None): + # `python` and `workdir` are per-app because the image decides them: adventurelog's + # interpreter is on PATH, tandoor ships a VENV and the bare `python` cannot import Django + # at all ("Couldn't import Django. Are you sure it's installed…"). Measured, not guessed. + self.container = container + self.sub = sub + self.ready_path = ready_path + self.ready = ready + self.python = python + self.workdir = workdir + + def _wd(self): + return f"-w {self.workdir} " if self.workdir else "" + + def _manage(self, w, code): + # -c is passed to `manage.py shell`; the app's own shell, its own ORM. + return w.guest( + f"docker exec {self._wd()}{self.container} {self.python} manage.py shell " + f"-c {json.dumps(code)} 2>&1", timeout=300) + + def _exists(self, w, username): + # ONE LINE, semicolon-separated. A `\n` inside a double-quoted shell argument reaches + # python as a literal backslash-n and is a SyntaxError — which is exactly how the first + # adventurelog run read as `inconclusive`. The fixture refused to guess, which is right, + # but the instrument was the thing that was broken. + out = self._manage(w, ( + "from django.contrib.auth import get_user_model; " + f"print('DRILL_ANSWER=' + str(get_user_model().objects.filter(username={username!r}).exists()))" + )) + m = re.search(r"DRILL_ANSWER=(True|False)", out) + return (m.group(1) == "True") if m else None, " ".join(out.split())[-300:] + + def seed(self, w, sub, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -e DJANGO_SUPERUSER_PASSWORD={pw} {self._wd()}{self.container} " + f"{self.python} manage.py createsuperuser --noinput " + f"--username {user} --email {user}@gate.invalid 2>&1", timeout=300) + say(f" {self.container}: createsuperuser :: {' '.join(out.split())[:160]}") + got, detail = self._exists(w, user) + if got is not True: + say(f" {self.container}: the account did not appear in the app's own ORM :: {detail[:200]}") + return None + say(f" {self.container}: seeded superuser {user}") + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + say(f" {self.container}: the app never served {self.ready_path}") + return False + absent, detail = self._exists(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" {self.container}: READBACK UNUSABLE — a username that cannot exist did not " + f"read as absent ({absent}) :: {detail[:200]}") + return False + found, detail = self._exists(w, t["user"]) + say(f" {self.container}: readback of the seeded account found={found}") + if found is not True: + say(f" {self.container}: :: {detail[:250]}") + return found is True + + +# ============================================================================================= +class Nextcloud: + """Nextcloud's OWN admin CLI, `occ`, inside its own container: its own code, its own user + backend. Not a SQL INSERT and not a planted file (R-156). + + `occ user:info` is the readback, and it PROVES ITSELF on every call: a uid that cannot exist + must answer "user not found". A readback that has broken into always succeeding therefore + fails instead of passing everything. + + This is the app chosen for the MariaDB engine-major edge (`09` §3 decision 5, R-469 lifted): + the app image does NOT move, only the `mariadb:` sidecar, so the edge carries exactly one + migration and a failure is readable. + """ + sub = "cloud" + + def _occ(self, w, *args, timeout=420): + import shlex + line = " ".join(shlex.quote(a) for a in args) + return w.guest(f"docker exec -u www-data nextcloud php occ {line} 2>&1", timeout=timeout) + + def _info(self, w, uid): + out = self._occ(w, "user:info", uid) + flat = " ".join(out.split()) + if "user not found" in flat.lower() or "could not be found" in flat.lower(): + return False, flat + if f"user_id: {uid}" in flat or f"- user_id: {uid}" in flat or f"user_id: {uid}" in out: + return True, flat + return None, flat + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + return None + uid = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -u www-data -e OC_PASS={pw} nextcloud php occ user:add " + f"--password-from-env --display-name={uid} {uid} 2>&1", timeout=420) + say(f" nextcloud: occ user:add :: {' '.join(out.split())[:160]}") + got, flat = self._info(w, uid) + if got is not True: + say(f" nextcloud: the account did not appear via occ user:info :: {flat[:220]}") + return None + say(f" nextcloud: seeded user {uid}") + return {"uid": uid, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + say(" nextcloud: the app never served /status.php") + return False + absent, flat = self._info(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" nextcloud: READBACK UNUSABLE — a uid that cannot exist did not read absent " + f"({absent}) :: {flat[:200]}") + return False + found, flat = self._info(w, t["uid"]) + say(f" nextcloud: readback of the seeded user found={found}") + if found is not True: + say(f" nextcloud: :: {flat[:250]}") + return found is True + + +# ============================================================================================= +class Grafana: + """Grafana's own HTTP API as the admin the DEPLOY created. The password is the one the + controller showed the household — read from the app's own `app.yaml`, not invented — and the + data (a folder) goes in and comes back through the app's own REST API.""" + sub = "grafana" + + def _auth(self, w, name="grafana"): + # app.yaml stores this ENCRYPTED (`ENC:…`), so it cannot be read back off the box — which + # is correct, and is why the harness uses the value IT generated for the deploy. + pw = (w.GENERATED.get(name) or {}).get("GF_SECURITY_ADMIN_PASSWORD") or "admin" + return f"admin:{pw}" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return None + au = self._auth(w) + title = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/folders", "-u", au, + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="POST") + say(f" grafana: create folder http={code}") + if code not in ("200", "201"): + say(f" grafana: refused {body[:220]}") + return None + try: + uid = json.loads(body)["uid"] + except Exception: + say(f" grafana: no uid in {body[:200]}") + return None + return {"uid": uid, "title": title} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return False + au = self._auth(w) + rc, code, _ = w.app_curl(sub, "/api/folders/nope" + secrets.token_hex(5), "-u", au) + if code == "200": + say(" grafana: READBACK UNUSABLE — a folder uid that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/folders/{t['uid']}", "-u", au) + ok = code == "200" and t["title"] in body + say(f" grafana: readback of the seeded folder http={code} ok={ok}") + return ok + + +# ============================================================================================= +class AudiobookShelf: + """audiobookshelf's own /init endpoint creates the first root account; its own /login proves + the account survived. Both are the app's own API.""" + sub = "audiobooks" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/init", "-H", "Content-Type: application/json", + data=json.dumps({"newRoot": {"username": user, "password": pw}}), + method="POST") + say(f" audiobookshelf: /init http={code}") + if code not in ("200", "204"): + say(f" audiobookshelf: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code == "200": + say(" audiobookshelf: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": t["pw"]}), + method="POST") + ok = code == "200" and t["user"] in body + say(f" audiobookshelf: login as the seeded root http={code} ok={ok}") + return ok + + +# ============================================================================================= +class ActualBudget: + """Actual's own bootstrap API sets the server password; its own login proves it survived.""" + sub = "budget" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/account/bootstrap", + "-H", "Content-Type: application/json", + data=json.dumps({"password": pw}), method="POST") + say(f" actualbudget: /account/bootstrap http={code} :: {body[:140]}") + if code not in ("200", "201") or '"status":"ok"' not in body: + return None + return {"pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def login(p): + return w.app_curl(sub, "/account/login", "-H", "Content-Type: application/json", + data=json.dumps({"loginMethod": "password", "password": p}), + method="POST") + rc, code, body = login("wrong-" + secrets.token_hex(6)) + if '"status":"ok"' in body: + say(" actualbudget: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = '"status":"ok"' in body + say(f" actualbudget: login with the seeded password http={code} ok={ok}") + if not ok: + say(f" actualbudget: body {body[:200]}") + return ok + + +# ============================================================================================= +class Mealie: + """Mealie ships a documented first-run admin. We log in as it through the app's own OAuth-style + token endpoint, create a recipe through the app's own API, and read the recipe back.""" + sub = "recipes" + + def _token(self, w, sub, pw="MyPassword"): + rc, code, body = w.app_curl( + sub, "/api/auth/token", "-H", "Content-Type: application/x-www-form-urlencoded", + data=f"username=changeme%40example.com&password={pw}", method="POST") + if code != "200": + return None, f"http={code} {body[:200]}" + try: + return json.loads(body)["access_token"], "" + except Exception: + return None, body[:200] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return None + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate as the first-run admin :: {why}") + return None + name = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/recipes", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": name}), method="POST") + say(f" mealie: create recipe http={code}") + if code not in ("200", "201"): + say(f" mealie: refused {body[:220]}") + return None + slug = body.strip().strip('"') + return {"slug": slug, "name": name} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return False + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate after the update :: {why}") + return False + rc, code, _ = w.app_curl(sub, "/api/recipes/nope" + secrets.token_hex(5), + "-H", f"Authorization: Bearer {tok}") + if code == "200": + say(" mealie: READBACK UNUSABLE — a slug that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/recipes/{t['slug']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["name"] in body + say(f" mealie: readback of the seeded recipe http={code} ok={ok}") + return ok + + +# ============================================================================================= +class N8n: + """n8n's own owner-setup API creates the first account; its own login proves it survived.""" + sub = "auto" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill" + secrets.token_hex(8) + "1" + rc, code, body = w.app_curl(sub, "/rest/owner/setup", "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "firstName": "drill", + "lastName": "drill", "password": pw}), + method="POST") + say(f" n8n: /rest/owner/setup http={code}") + if code not in ("200", "201"): + say(f" n8n: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return False + def login(p): + return w.app_curl(sub, "/rest/login", "-H", "Content-Type: application/json", + data=json.dumps({"emailOrLdapLoginId": t["email"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" n8n: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" and t["email"] in body + say(f" n8n: login as the seeded owner http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Zipline: + """Zipline's own setup/login API. Zipline 4 creates the first user through its own endpoint.""" + sub = "img" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/healthcheck", want=("200",), tries=90): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=30): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + for path in ("/api/auth/register", "/api/auth/setup"): + rc, code, body = w.app_curl(sub, path, "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), + method="POST") + say(f" zipline: {path} http={code} :: {body[:160]}") + if code in ("200", "201"): + return {"user": user, "pw": pw} + say(" zipline: neither register nor setup accepted a first user") + return None + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=60): + return False + def login(p): + return w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" zipline: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" + say(f" zipline: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vikunja: + """Vikunja's own REST API: register a user, log in, create a project, read the project back. + Four calls, all the app's own front door.""" + sub = "tasks" + + def _token(self, w, sub, t, pw=None): + rc, code, body = w.app_curl(sub, "/api/v1/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], + "password": pw or t["pw"]}), method="POST") + if code != "200": + return None, f"http={code} {body[:160]}" + try: + return json.loads(body)["token"], "" + except Exception: + return None, body[:160] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/v1/register", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw, + "email": f"{user}@gate.invalid"}), + method="POST") + say(f" vikunja: register http={code}") + if code not in ("200", "201"): + say(f" vikunja: refused {body[:220]}") + return None + t = {"user": user, "pw": pw} + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: could not log in after registering :: {why}") + return None + title = "drill-" + secrets.token_hex(5) + # Vikunja CREATES with PUT, not POST — a POST answers `405 Method Not Allowed`, which + # reads like a broken fixture and is really the wrong verb. Measured 2026-09-21. + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="PUT") + say(f" vikunja: create project http={code}") + if code not in ("200", "201"): + say(f" vikunja: project refused {body[:220]}") + return None + t["title"] = title + t["pid"] = json.loads(body).get("id") + return t + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return False + bad, why = self._token(w, sub, t, pw="wrong-" + secrets.token_hex(6)) + if bad: + say(" vikunja: READBACK UNUSABLE — a wrong password authenticated") + return False + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: the seeded account no longer authenticates :: {why}") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{t['pid']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["title"] in body + say(f" vikunja: readback of the seeded project http={code} ok={ok}") + return ok + + +# ============================================================================================= +class OpenGist: + """Opengist's own sign-up and sign-in FORMS. + + Two things had to be measured. Its sign-up is CSRF-protected: a bare POST answers 500 with an + HTML page, which reads like a broken app and is really a missing token — fetch the form, keep + its cookie, send its `_csrf` back. And its REST API refuses the account's own password + (`401 {"message":"Bad crendentials"}`) because it wants a token the app will not mint without a + browser. So the SEEDED DATA is the account itself and the READBACK is a real sign-in, which is + the same shape the docmost and navidrome fixtures use. + + LIMITATION, recorded rather than papered over: this is the DATABASE half. A gist's CONTENT is + not seeded, because that needs the API token above. + """ + sub = "gist" + + def _form(self, w, sub, path, jar, fields): + rc, code, html = w.app_curl(sub, path, "-b", jar, "-c", jar) + m = re.search(r'name="_csrf"[^>]*value="([^"]+)"', html or "") + if not m: + return None, f"no _csrf on {path} (http={code})" + body = "&".join([f"_csrf={m.group(1)}"] + [f"{k}={v}" for k, v in fields.items()]) + rc, code, out = w.app_curl(sub, path, "-b", jar, "-c", jar, + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + return code, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, out = self._form(w, sub, "/register", jar, {"username": user, "password": pw}) + say(f" opengist: /register (with its own _csrf) http={code}") + if code not in ("200", "302", "303"): + say(f" opengist: refused {str(out)[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + # Wait for the LOGIN FORM, not for the root page. Measured 2026-09-21: immediately after a + # successful update the root answers while /login does not yet carry its `_csrf`, so the + # sign-in silently fails and the app looks like it lost the account. It had not. + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" opengist: /login never came back after the update") + return False + for _ in range(24): + rc, code, html = w.app_curl(sub, "/login") + if code == "200" and '_csrf' in (html or ""): + break + time.sleep(5) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar, + {"username": t["user"], "password": "wrong-" + secrets.token_hex(5)}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar) + if t["user"] in (home or ""): + say(" opengist: READBACK UNUSABLE — a wrong password signed in") + return False + jar2 = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar2, {"username": t["user"], "password": t["pw"]}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar2) + ok = t["user"] in (home or "") + say(f" opengist: sign-in as the seeded account http={code} name_on_page={ok}") + return ok + + +# ============================================================================================= +class Papra: + """Papra's own e-mail sign-up and sign-in endpoints.""" + sub = "papra" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/auth/sign-up/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "password": pw, + "name": "drill"}), method="POST") + say(f" papra: sign-up http={code}") + if code not in ("200", "201"): + say(f" papra: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def signin(p): + return w.app_curl(sub, "/api/auth/sign-in/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": t["email"], "password": p}), method="POST") + rc, code, _ = signin("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" papra: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = signin(t["pw"]) + ok = code == "200" + say(f" papra: sign-in as the seeded account http={code} ok={ok}") + return ok + + +# ============================================================================================= +class HomeAssistant: + """Home Assistant's own onboarding API creates the owner account and hands back a code the + same API exchanges for a token. Both are the app's own documented non-browser route.""" + sub = "ha" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/onboarding/users", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "name": "drill", "username": user, + "password": pw, "language": "en"}), + method="POST") + say(f" home-assistant: /api/onboarding/users http={code}") + if code not in ("200", "201"): + say(f" home-assistant: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def _login(self, w, sub, user, pw): + """The app's own login flow: start it, then answer it. A 200 with a step_id of + `mfa`/`init` means the credentials were REFUSED; only `create_entry` is a pass.""" + rc, code, body = w.app_curl(sub, "/auth/login_flow", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "handler": ["homeassistant", None], + "redirect_uri": f"https://{sub}.felhom.invalid/", + "type": "authorize"}), method="POST") + if code not in ("200", "201"): + return None, f"flow start http={code} {body[:160]}" + try: + fid = json.loads(body)["flow_id"] + except Exception: + return None, body[:160] + rc, code, body = w.app_curl(sub, f"/auth/login_flow/{fid}", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "username": user, "password": pw}), + method="POST") + try: + j = json.loads(body) + except Exception: + return None, body[:160] + return (j.get("result") if j.get("type") == "create_entry" else None), body[:200] + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return False + bad, why = self._login(w, sub, t["user"], "wrong-" + secrets.token_hex(6)) + if bad: + say(" home-assistant: READBACK UNUSABLE — a wrong password authenticated") + return False + good, why = self._login(w, sub, t["user"], t["pw"]) + ok = bool(good) + say(f" home-assistant: login as the seeded owner ok={ok}") + if not ok: + say(f" home-assistant: {why}") + return ok + + +# ============================================================================================= +class Romm: + """RomM's own user API, driven the way RomM's own front end drives it. + + Three things had to be measured rather than guessed, and each one answered a 403 or a 422 that + looked like a different fault: RomM sets a **`romm_csrftoken` cookie** on any GET and requires + it back in an **`x-csrftoken` header** (a bare POST is `403 CSRF token verification failed`, + which reads like an auth problem); the fields go in the **JSON body**, not the query string (a + query-string POST is `422 Field required` for every field it was just given); and `email` is + required alongside username, password and role. + + On a fresh install with no admin the first `POST /api/users` is accepted unauthenticated; + afterwards it is not — which is what makes the readback (`POST /api/login` as that user) a real + authentication rather than a repeat of the seed. + + LIMITATION: this is the DATABASE half. RomM's other half is the ROM library on the drive, which + this does not populate. + """ + sub = "arcade" + + def _csrf(self, w, sub): + jar = f"/tmp/romm-{secrets.token_hex(4)}.jar" + w.app_curl(sub, "/api/heartbeat", "-c", jar) + out = w.sh(["bash", "-lc", f"grep -i csrf {jar} | awk '{{print $7}}'"]).stdout or "" + return jar, out.strip() + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + jar, tok = self._csrf(w, sub) + if not tok: + say(" romm: no romm_csrftoken cookie was set on /api/heartbeat") + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl( + sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": pw, "role": "admin"}), method="POST") + say(f" romm: POST /api/users http={code}") + if code not in ("200", "201"): + say(f" romm: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + return False + jar, tok = self._csrf(w, sub) + rc, code, _ = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:wrong-{secrets.token_hex(5)}", method="POST") + if code == "200": + say(" romm: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:{t['pw']}", method="POST") + ok = code == "200" + say(f" romm: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" romm: body {body[:200]}") + return ok + + +FIXTURES = { + "home-assistant": HomeAssistant(), + "romm": Romm(), + "vikunja": Vikunja(), + "opengist": OpenGist(), + "papra": Papra(), + "mealie": Mealie(), + "n8n": N8n(), + "zipline": Zipline(), + "grafana": Grafana(), + "audiobookshelf": AudiobookShelf(), + "actualbudget": ActualBudget(), + "nextcloud": Nextcloud(), + "adventurelog": Django("adventurelog", "travel", "/admin/login/"), + "tandoor": Django("tandoor", "recipes", "/accounts/login/", + python="/opt/recipes/venv/bin/python", workdir="/opt/recipes"), + "privatebin": PrivateBin(), + "docmost": Docmost(), + "bookstack": BookStack(), + "gitea": Gitea(), + "navidrome": Navidrome(), + "vaultwarden": Vaultwarden(), +} diff --git a/documentation/audits/undo-fleet-2026-09-23/live.py b/documentation/audits/undo-fleet-2026-09-23/live.py new file mode 100644 index 00000000..8a024100 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/live.py @@ -0,0 +1,194 @@ +#!/usr/bin/env python3 +"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes. + +Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/, +fetches the app page in both languages, and reads the seeds back through each app's own front door. +Seed C is written immediately before each Update, after every backup — so only the undo's own +last-second copy can bring it back. +""" +import json, re, sys, time, html as H +sys.path.insert(0, ".") +import walk as w +from spike import load, save, ts +import bakeoff as b + +HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"} + + +def seed_c(app, s): + return b.seed_b(app, s["sub"], s["seedA"]) + + +def page_lines(app): + out = {} + for lang in ("hu", "en"): + h = H.unescape(w.page(f"/apps/{app}?lang={lang}")) + m = re.search(r'data-update-undone="true">([^<]*)<', h) + hold = re.search(r'data-held="true">([^<]*)<', h) + out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None} + return out + + +def press(app, poll=0.5, on_phase=None): + code, d = w.ctl("POST", f"/api/stacks/{app}/update") + w.say(f" Update -> {code} {str(d)[:140]}") + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < 1500: + try: + st = w.stack(app) + except Exception as e: + st = {} + ph = st.get("update_phase") + if ph != seen and ph is not None: + seen = ph + phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label"))) + w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}") + if on_phase and on_phase(ph): + return phases, "interrupted" + if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2: + break + time.sleep(poll) + st = w.stack(app) + w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}") + return phases, st + + +def readback(app, s, with_c=True): + sub = s["sub"] + w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2) + A = b.FX[app].verify(w, sub, s["seedA"], w.say) + B = b.verify_b(app, sub, s["seedA"], s["seedB"]) + C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None + return {"A": A, "B": B, "C": C} + + +def stage_undo(app): + s = load(app) + w.say(f"=== {app}: live undo by the product") + s["seedC"] = seed_c(app, s); s["seedC_at"] = ts() + w.say(" seed C written right before the Update:", bool(s["seedC"])) + before = b.db_state(app); w.say(" db before:", before) + phases, st = press(app) + obs = w.observables(app) + w.say(" observables:", json.dumps(obs)) + rb = readback(app, s) + after = b.db_state(app) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})") + pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False)) + w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip()) + s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")}, + "readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs} + save(app, s) + + + + +def set_box_language(lang): + """POST /settings/language — the household's language switch (form + the dashboard's CSRF).""" + sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip() + r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", + "-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}", + f"{w.BASE}/settings/language"]) + w.say(f" box language -> {lang}: http {r.stdout.strip()}") + + +def stage_powercut(app): + """Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard + stop); boot it again and let the controller resume the undo.""" + import subprocess + s = load(app) + w.say(f"=== {app}: power cut DURING the undo") + s["seedC"] = seed_c(app, s) + cut = {} + + def on_phase(ph): + if ph == "undoing": + t = time.time() + r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120) + cut["at"] = ts(); cut["rc"] = r.returncode + w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s") + return True + return False + phases, _ = press(app, poll=0.3, on_phase=on_phase) + if not cut: + w.say(" the cut never landed in `undoing` — recorded as a MISS"); return + r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180) + w.say(f" guest started again: rc={r.returncode}") + for i in range(60): + time.sleep(5) + try: + w.login(); st = w.stack(app) + if st: + break + except SystemExit: + continue + w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6")) + t0 = time.time() + while time.time() - t0 < 600: + st = w.stack(app) + if not st.get("updating") and st.get("update_phase") in ("undone", "failed"): + break + time.sleep(3) + w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}") + rb = readback(app, s) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}") + w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip()) + s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb} + save(app, s) + + +def stage_cutoff(app): + """Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of + the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it + back and HOLD with the new prefix.""" + s = load(app) + w.say(f"=== {app}: a cut-off undo copy") + done = {} + + def on_phase(ph): + if ph == "verifying" and not done: + out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c" +docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""") + done["out"] = out + w.say(" >>> finished-marker removed from one copy:\n" + out) + return False + phases, st = press(app, poll=0.5, on_phase=on_phase) + before = b.db_state(app) + pl = page_lines(app) + w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False)) + set_box_language("en") + pl_en = page_lines(app) + w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False)) + set_box_language("hu") + w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}")) + s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done} + save(app, s) + + +def stage_fixprobe_and_press(app): + """After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work + and end the undone note.""" + s = load(app) + frm, to, p0, p1 = b.EDGE[app] + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f) + w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"]) + w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + w.sync_rescan(app, to) + time.sleep(20) + w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan") + w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)") + w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False)) + phases, st = press(app) + rb = readback(app, s) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}") + pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False)) + w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml")) + s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl} + save(app, s) + + +if __name__ == "__main__": + w.login() + {"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff, + "fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/undo-fleet-2026-09-23/repoint.py b/documentation/audits/undo-fleet-2026-09-23/repoint.py new file mode 100644 index 00000000..cb28637e --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/repoint.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config. + +`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is +`controller.yaml.pre-undofleet` (NOT the older `.pre-28`, which a restore must never pick up). +""" +import re, sys, io +sys.path.insert(0, '.') +import walk as w + +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" + + +def creds(): + for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"): + m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l) + if m: + return m.group(1), m.group(2) + raise SystemExit("no admin credential") + + +def to_drill(): + u, t = creds() + print(w.guest(f""" +set -e +test -f {VOL}/controller.yaml.pre-undofleet || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-undofleet +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml" +s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +if not re.search(r'^update:', s, re.M): + s += "update:\\n health_timeout: 90s\\n" +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -A2 '^update:' {VOL}/controller.yaml +""")) + + +def restore(): + print(w.guest(f""" +set -e +cp -p {VOL}/controller.yaml.pre-undofleet {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -c '^update:' {VOL}/controller.yaml || true +""")) + + +if __name__ == "__main__": + to_drill() if sys.argv[1] == "drill" else restore() diff --git a/documentation/audits/undo-fleet-2026-09-23/spike.py b/documentation/audits/undo-fleet-2026-09-23/spike.py new file mode 100644 index 00000000..369d1f64 --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/spike.py @@ -0,0 +1,160 @@ +#!/usr/bin/env python3 +"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202. + +EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup, +sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being +spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the +steps the product would take, each one timed. + +State between stages lives in state-.json so each stage can be run, read, and only then +followed by the next (the hand undo needs a person looking at what the previous step left). +""" +import json, os, sys, time, re +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +HERE = os.path.dirname(os.path.abspath(__file__)) +FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()} +SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"} + + +def st_path(app): + return os.path.join(HERE, f"state-{os.environ.get('GUEST', '9202')}-{app}.json") + + +def load(app): + return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {} + + +def save(app, s): + json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False) + + +def ts(): + return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + + +# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only +# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it +# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier. +def docmost_seed_b(sub, A): + jar = "/tmp/dm.jar" + w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + name = "drillB" + os.urandom(3).hex() + rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json", + data=json.dumps({"name": name, "slug": name.lower()}), method="POST") + w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}") + return {"space": name} if code in ("200", "201") else None + + +def docmost_verify_b(sub, A, B): + jar = "/tmp/dm.jar" + rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + if code not in ("200", "201"): + w.say(f" docmost B: cannot log in (http={code})"); return False + rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json", + data="{}", method="POST") + ok = code in ("200", "201") and B["space"] in out + neg = "drillBnever" in out + w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def break_edge(app, frm, to, port_from, port_to): + """The failing edge: a REAL migrating image move, plus — in the DRILL template only — the + health probe pointed at a port the app does not answer. Both in one drill commit.""" + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read() + m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f) + assert m, "probe port not found" + f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():] + open(fy, "w").write(f) + h = w.drill_bump(app, frm, to) + return h + + +def pg_state(container, db, user): + return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1") + + +def romm_seed_b(sub, A): + """A SECOND RomM user, created by the first (admin) one — written after the backup.""" + jar, tok = FX["romm"]._csrf(w, sub) + user = "drillb" + os.urandom(3).hex() + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}), + method="POST") + w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}") + return {"user": user} if code in ("200", "201") else None + + +def romm_verify_b(sub, A, B): + jar, tok = FX["romm"]._csrf(w, sub) + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}") + ok = code == "200" and B["user"] in body + neg = "drillbnever" in body + w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def tree(paths): + cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths) + return w.guest(cmd) + + +def my_state(container="romm-db", db="romm"): + # The root password is used INSIDE the container from its own env — it never leaves it. + q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '" + return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) " +echo -n "alembic=$({q("select version_num from alembic_version")}) " +echo -n "users=$({q("select count(*) from users")})" +""") + + +def vik_seed_b(sub, A): + """A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the + app writes into its files volume), uploaded through the app's own attachment API.""" + tok, why = FX["vikunja"]._token(w, sub, A) + title = "drillB-" + os.urandom(4).hex() + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT") + if code not in ("200", "201"): + w.say(f" vikunja B: project refused {code} {body[:120]}"); return None + pid = json.loads(body)["id"] + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT") + tid = json.loads(body)["id"] if code in ("200", "201") else None + content = "drill attachment " + os.urandom(8).hex() + fn = "/tmp/vik-att.txt"; open(fn, "w").write(content) + rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}", + "-F", f"files=@{fn}", method="PUT") + w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}") + return {"title": title, "pid": pid, "tid": tid, "att": content} + + +def vik_verify_b(sub, A, B): + tok, why = FX["vikunja"]._token(w, sub, A) + if not tok: + w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False} + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}") + proj = code == "200" and B["title"] in body + rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}") + att_ok = False + try: + atts = json.loads(body) + if atts: + aid = atts[0]["id"] + rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}") + att_ok = c3 == "200" and B["att"] in b3 + except Exception as e: + w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}") + w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}") + return {"project": proj, "attachment": att_ok} diff --git a/documentation/audits/undo-fleet-2026-09-23/state-9201-vikunja.json b/documentation/audits/undo-fleet-2026-09-23/state-9201-vikunja.json new file mode 100644 index 00000000..2888ab2d --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/state-9201-vikunja.json @@ -0,0 +1,172 @@ +{ + "app": "vikunja", + "sub": "tasks", + "seedA": { + "user": "drill229ad3", + "pw": "", + "title": "drill-251d05a1ca", + "pid": 2 + }, + "seedB": { + "title": "drillB-f2e2cc24", + "pid": 3, + "tid": 1, + "att": "drill attachment c2abe62af7869525" + }, + "volumes": [ + "vikunja_vikunja_data", + "vikunja_vikunja_db" + ], + "break_commit": "88b959602f91", + "seedC": { + "title": "drillB-935bb593", + "pid": 4, + "tid": 2, + "att": "drill attachment 2b1d7877c1929b80" + }, + "seedC_at": "2026-09-23T12:13:23Z", + "live_undo": { + "phases": [ + [ + 0.0, + "safety-dump", + "Adatbázis pillanatkép…" + ], + [ + 1.1, + "pulling", + "Új verzió letöltése…" + ], + [ + 4.8, + "copying", + "Az adatok másolása a frissítés előtt…" + ], + [ + 6.9, + "starting", + "Indítás az új verzióval…" + ], + [ + 7.4, + "verifying", + "Működés ellenőrzése…" + ], + [ + 97.7, + "undoing", + "Visszaállítás az előző változatra…" + ], + [ + 100.4, + "undone", + "Visszaállítva az előző változatra" + ] + ], + "end": { + "update_phase": "undone", + "update_error": null, + "hold_reason": null + }, + "readback": { + "A": true, + "B": true, + "C": true + }, + "db_before": "tables=36 ledger=117 newest=SCHEMA_INIT", + "db_after": "tables=36 ledger=117 newest=SCHEMA_INIT", + "page": { + "hu": { + "undone_line": "A(z) vikunja frissítése 2026-09-23 14:15-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", + "hold": null + }, + "en": { + "undone_line": "The update of vikunja at 2026-09-23 14:15 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", + "hold": null + } + }, + "obs": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.3.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.3.0" + }, + "catalog_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.3.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.3.0 running=true restarts=0" + ] + } + }, + "cutoff_live": { + "phases": [ + [ + 0.0, + "safety-dump", + "Adatbázis pillanatkép…" + ], + [ + 1.1, + "pulling", + "Új verzió letöltése…" + ], + [ + 2.1, + "copying", + "Az adatok másolása a frissítés előtt…" + ], + [ + 3.2, + "starting", + "Indítás az új verzióval…" + ], + [ + 3.7, + "verifying", + "Működés ellenőrzése…" + ], + [ + 94.3, + "undoing", + "Visszaállítás az előző változatra…" + ], + [ + 95.3, + "failed", + "A frissítés nem sikerült" + ] + ], + "end": { + "update_phase": "failed", + "hold_reason": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." + }, + "page_hu_box": { + "hu": { + "undone_line": null, + "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 14:13 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." + }, + "en": { + "undone_line": null, + "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes." + } + }, + "page_en_box": { + "hu": { + "undone_line": null, + "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes." + }, + "en": { + "undone_line": null, + "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:16 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 14:13 — this copy holds the settings, the database and the data volumes." + } + }, + "marker": { + "out": "copy: vikunja_vikunja_data.pre-update-20260923T121516Z\ntotal 12\ndrwxr-xr-x 3 root root 4096 Sep 23 12:15 .\ndrwxr-xr-x 1 root root 4096 Sep 23 12:15 ..\ndrwxr-xr-x 2 root root 4096 Sep 23 12:15 data\n-rw-r--r-- 1 root root 0 Sep 23 12:15 felhom-undo-complete\ndata\n" + } + } +} \ No newline at end of file diff --git a/documentation/audits/undo-fleet-2026-09-23/state-9202-vikunja.json b/documentation/audits/undo-fleet-2026-09-23/state-9202-vikunja.json new file mode 100644 index 00000000..6e25aaec --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/state-9202-vikunja.json @@ -0,0 +1,167 @@ +{ + "app": "vikunja", + "sub": "tasks", + "seedA": { + "user": "drill73bafb", + "pw": "", + "title": "drill-94f9145a94", + "pid": 2 + }, + "seedB": { + "title": "drillB-b6e6bb28", + "pid": 3, + "tid": 1, + "att": "drill attachment 12d38cc6ade81d67" + }, + "volumes": [ + "vikunja_vikunja_data", + "vikunja_vikunja_db" + ], + "break_commit": "073bc953e182", + "seedC": { + "title": "drillB-6d4934fb", + "pid": 4, + "tid": 2, + "att": "drill attachment d4f2ee5a03683645" + }, + "seedC_at": "2026-09-23T11:58:25Z", + "live_undo": { + "phases": [ + [ + 0.0, + "safety-dump", + "Adatbázis pillanatkép…" + ], + [ + 1.1, + "pulling", + "Új verzió letöltése…" + ], + [ + 4.2, + "copying", + "Az adatok másolása a frissítés előtt…" + ], + [ + 5.8, + "starting", + "Indítás az új verzióval…" + ], + [ + 6.3, + "verifying", + "Működés ellenőrzése…" + ], + [ + 96.7, + "undoing", + "Visszaállítás az előző változatra…" + ], + [ + 99.3, + "undone", + "Visszaállítva az előző változatra" + ] + ], + "end": { + "update_phase": "undone", + "update_error": null, + "hold_reason": null + }, + "readback": { + "A": true, + "B": true, + "C": true + }, + "db_before": "tables=36 ledger=117 newest=SCHEMA_INIT", + "db_after": "tables=36 ledger=117 newest=SCHEMA_INIT", + "page": { + "hu": { + "undone_line": "A(z) vikunja frissítése 2026-09-23 14:00-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", + "hold": null + }, + "en": { + "undone_line": "The update of vikunja at 2026-09-23 14:00 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", + "hold": null + } + }, + "obs": { + "pinned_images": { + "vikunja": "vikunja/vikunja:2.3.0" + }, + "installed_images": { + "vikunja": "vikunja/vikunja:2.3.0" + }, + "catalog_images": { + "vikunja": "vikunja/vikunja:2.6.0" + }, + "live_compose_image_lines": [ + "image: vikunja/vikunja:2.3.0" + ], + "docker_inspect": [ + "vikunja vikunja/vikunja:2.3.0 running=true restarts=0" + ] + } + }, + "cutoff_live": { + "phases": [ + [ + 0.0, + "safety-dump", + "Adatbázis pillanatkép…" + ], + [ + 1.1, + "pulling", + "Új verzió letöltése…" + ], + [ + 2.1, + "copying", + "Az adatok másolása a frissítés előtt…" + ], + [ + 3.7, + "verifying", + "Működés ellenőrzése…" + ], + [ + 94.7, + "undoing", + "Visszaállítás az előző változatra…" + ], + [ + 95.2, + "failed", + "A frissítés nem sikerült" + ] + ], + "end": { + "update_phase": "failed", + "hold_reason": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." + }, + "page_hu_box": { + "hu": { + "undone_line": null, + "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 14:02-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 13:58 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." + }, + "en": { + "undone_line": null, + "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes." + } + }, + "page_en_box": { + "hu": { + "undone_line": null, + "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes." + }, + "en": { + "undone_line": null, + "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. The update of vikunja at 2026-09-23 14:02 did not succeed, and the app did not start on the new version. The app stays stopped for safety, so that its data is not damaged. It can be restored on the Backups page from this backup: own drive, 2026-09-23 13:58 — this copy holds the settings, the database and the data volumes." + } + }, + "marker": { + "out": "copy: vikunja_vikunja_data.pre-update-20260923T120033Z\ntotal 12\ndrwxr-xr-x 3 root root 4096 Sep 23 12:00 .\ndrwxr-xr-x 1 root root 4096 Sep 23 12:00 ..\ndrwxr-xr-x 2 root root 4096 Sep 23 12:00 data\n-rw-r--r-- 1 root root 0 Sep 23 12:00 felhom-undo-complete\ndata\n" + } + } +} \ No newline at end of file diff --git a/documentation/audits/undo-fleet-2026-09-23/walk.py b/documentation/audits/undo-fleet-2026-09-23/walk.py new file mode 100644 index 00000000..fa89ae2b --- /dev/null +++ b/documentation/audits/undo-fleet-2026-09-23/walk.py @@ -0,0 +1,495 @@ +#!/usr/bin/env python3 +"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints. + +EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses: + POST /api/stacks//deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan + POST /api/stacks//update · POST /api/stacks//remove +and reads GET /api/stacks/. No controller code exists for it. + +The walk, per `09` §6.4 and the update-night brief §4: + 1 deploy from the DRILL catalog at the LIVE pin + 2 seed through the app's OWN front door (R-156: never a volume, never SQL) + 3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing + 4 „Mentés most" + 5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages + 6 press the guarded Update, record every phase with timestamps + 7 read the seed back through the front door + 8 the four version observables side by side + 9 write the verdict record in `09`'s JSON shape + +`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`. +""" +import argparse, json, os, re, subprocess, sys, time +from datetime import datetime, timezone + +SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad" +EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-fleet-2026-09-23" +DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest. +GUEST = os.environ.get("GUEST", "9202") +BASE = {"9202": "https://192.168.0.114", "9201": "https://192.168.0.138"}[GUEST] +DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu") +HOSTHDR = f"Host: felhom.{DOMAIN}" +HP = "demo-hp" + +LOG = [] + + +def say(*a): + line = " ".join(str(x) for x in a) + ts = datetime.now().strftime("%H:%M:%S") + print(f"{ts} {line}", flush=True) + LOG.append(f"{ts} {line}") + + +def sh(args, timeout=300, inp=None): + try: + return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp) + except (subprocess.TimeoutExpired, OSError) as e: + return subprocess.CompletedProcess(args, 124, "", f"{e}") + + +def guest(script, timeout=600): + """Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting).""" + r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP, + f"cat > /tmp/w{GUEST}.sh; pct push {GUEST} /tmp/w{GUEST}.sh /tmp/w.sh >/dev/null 2>&1; " + f"pct exec {GUEST} -- bash /tmp/w.sh; rm -f /tmp/w{GUEST}.sh"], + timeout=timeout, inp=script) + return r.stdout or "" + + +def login(): + pw = open(f"{SC}/.ctlpw").read().strip() + sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR, + "-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"]) + h = open(f"{SC}/hdr.txt").read() + m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I) + if not m: + sys.exit("login failed: no session cookie") + open(f"{SC}/sess.txt", "w").write(m.group(0)) + r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"]) + c = re.search(r'/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN. + + Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because + a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password + (grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only + reason this was safe to discover by running it (live-probes rule). + + A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on + the scratch drive first — the same act the drive browser performs for a household. + """ + code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields") + fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or [] + values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub} + made = [] + for f in fields: + ev, ty = f.get("env_var"), f.get("type") + if ev in values: + continue + # `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses + # when the caller sends none, deliberately ("the user needs to know their password"), + # while `.felhom.yml` declares `required: false` and the API serves that verbatim. A + # caller that trusts the contract gets a 400. Measured tonight on grafana; filed. + if not f.get("required") and ty != "password": + continue # the controller generates the optional secrets itself + if ty == "path": + p = f"{DRIVE}/{name}" + values[ev] = p + made.append(p) + elif ty in ("secret", "password"): + import secrets as _s + values[ev] = "Drill-" + _s.token_hex(12) + GENERATED.setdefault(name, {})[ev] = values[ev] + elif f.get("default"): + values[ev] = f["default"] + else: + values[ev] = f"drill-{name}" + if made: + guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made)) + say(f" [1] made the drive paths this app requires: {made}") + extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")] + if extra: + say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}") + return values + + +def deploy(name, sub, extra_values=None): + st = stack(name) + if st.get("deployed"): + say(f" [1] {name} already deployed — reusing") + return True + values = deploy_values(name, sub) + if extra_values: + values.update(extra_values) + code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values}) + say(f" [1] deploy -> {code} {str(d)[:120]}") + if code != "202": + return False + # WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the + # container `healthy` while the controller's own state read `unhealthy` — a gate on `running` + # alone therefore times out on an app that is up. The state is RECORDED rather than required; + # the real gate is the fixture's own `wait_app`, which asks whether the APP answers. + seen = None + for _ in range(90): + time.sleep(5) + st = stack(name) + seen = st.get("state") + # `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21: + # tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm + # read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the + # deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark + # (`runComposeDeploy` writes it), so that is what to wait for. + pins = (st.get("app_config") or {}).get("pinned_images") + if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"): + say(f" [1] deployed, controller state={seen}, " + f"pinned={(st.get('app_config') or {}).get('pinned_images')}") + if seen != "running": + say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, " + f"not treated as a failure; the fixture's front-door wait is the real gate") + return True + say(f" [1] never became deployed (last controller state={seen!r})") + return False + + +def backup_now(name): + code, d = ctl("POST", "/api/backup/run") + say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}") + for _ in range(90): + time.sleep(5) + c2, s = ctl("GET", "/api/backup/status") + dd = s.get("data") or {} + if not dd.get("running", False): + say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}") + return True + say(" [4] backup still running after 7.5 min — carrying on") + return False + + +def drill_bump(app, frm, to, service_hint=None): + """Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates). + + `frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in + two images (adventurelog's backend and frontend) moves both in one edge, while its engine + sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move + every service that carries the app's own version and no others. + """ + comp = f"{DRILL}/templates/{app}/docker-compose.yml" + fy = f"{DRILL}/templates/{app}/.felhom.yml" + s = open(comp).read() + froms = [x.strip() for x in frm.split(",") if x.strip()] + tos = [x.strip() for x in to.split(",") if x.strip()] + if len(froms) != len(tos): + say(f" [5] from/to lists differ in length: {froms} vs {tos}") + return None + for f1, t1 in zip(froms, tos): + if f"image: {f1}" not in s: + say(f" [5] FROM ref not found in compose: {f1}") + return None + s = s.replace(f"image: {f1}", f"image: {t1}") + open(comp, "w").write(s) + f = open(fy).read() + today = datetime.now().strftime("%Y-%m-%d") + f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M) + open(fy, "w").write(f) + sh(["git", "-C", DRILL, "add", "-A"]) + sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"]) + r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})") + return h + + +def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5): + """Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP. + + R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and + `catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not + enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the + Update that followed moved nothing and still reported "Frissitve". So when the caller knows + which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the + NUMBER R-607 asks for and has never had. + """ + t0 = time.time() + ctl("POST", "/api/sync") + time.sleep(2) + ctl("POST", "/api/stacks/rescan") + time.sleep(2) + if not expect_app or not expect_ref: + return None + for i in range(tries): + cat = stack(expect_app).get("catalog_images") or {} + if expect_ref in cat.values(): + waited = round(time.time() - t0, 1) + if i: + say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up " + f"to {expect_ref} — R-607's window, measured") + return waited + time.sleep(delay) + ctl("POST", "/api/sync") + time.sleep(1) + ctl("POST", "/api/stacks/rescan") + say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — " + f"catalog_images = {stack(expect_app).get('catalog_images')}") + return None + + +def badges(name): + out = {} + for lang, suffix in (("hu", ""), ("en", "?lang=en")): + h = page(f"/apps/{name}{suffix}") + m = re.findall(r']*title="([^"]*)"[^>]*>([^<]*)<', h) + out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3] + return out + + +def press_update(name, poll=1.0, cap_s=1800): + code, d = ctl("POST", f"/api/stacks/{name}/update") + say(f" [6] Update -> {code} {str(d)[:220]}") + if code not in ("202", "200"): + return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0} + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < cap_s: + st = stack(name) + ph = st.get("update_phase") + if ph != seen: + seen = ph + rec = {"t": round(time.time() - t0, 1), "phase": ph, + "label": st.get("update_phase_label"), "updating": st.get("updating"), + "error": st.get("update_error"), "hold": st.get("hold_reason")} + phases.append(rec) + say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} " + f"err={rec['error']} hold={rec['hold']}") + if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3: + break + time.sleep(poll) + st = stack(name) + return {"accepted": True, "http": code, "phases": phases, + "duration_s": round(time.time() - t0, 1), + "final_phase": st.get("update_phase"), "update_error": st.get("update_error"), + "hold_reason": st.get("hold_reason"), "state": st.get("state")} + + +def observables(name): + st = stack(name) + ac = st.get("app_config") or {} + live = guest(f""" +grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//' +echo '---inspect---' +for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do + echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}' +done +""") + a, _, b = live.partition("---inspect---") + return { + "pinned_images": ac.get("pinned_images"), + "installed_images": {k: (v.get("ref") if isinstance(v, dict) else v) + for k, v in (ac.get("installed_images") or {}).items()}, + "catalog_images": st.get("catalog_images"), + "live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()], + "docker_inspect": [x for x in b.strip().splitlines() if x.strip()], + } + + +def app_logs(name, lines=400): + """The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is + one string with escaped newlines — a scan over the envelope sees a single enormous line and + finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96 + rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines.""" + code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}") + if isinstance(d, dict): + data = d.get("data") + if isinstance(data, dict) and isinstance(data.get("logs"), str): + return data["logs"] + if isinstance(d.get("_raw"), str): + return d["_raw"] + return str(d) + + +def write_verdict(rec, appdir): + os.makedirs(appdir, exist_ok=True) + p = os.path.join(appdir, "verdict.json") + json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False) + say(f" [9] verdict {rec['verdict']} -> {p}") + + +def remove(name): + """Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint + refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up.""" + c1, d1 = ctl("POST", f"/api/stacks/{name}/stop") + say(f" [X] stop -> {c1} {str(d1)[:100]}") + for _ in range(24): + time.sleep(5) + if stack(name).get("state") != "running": + break + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": True, "remove_backups": True}) + say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}") + if code == "409": + # R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive + # path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202 + # `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits + # this. The household's other choice — remove the app, KEEP the data — is accepted, and the + # harness takes it, then tidies its own directory by name at teardown. + say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)" + " — removing the app and KEEPING the drive data instead") + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": False, "remove_backups": True}) + say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}") + time.sleep(5) + st = stack(name) + left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; " + f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'") + say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}") + return code + + +def app_env(name, key): + """Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the + app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller + shows them the same value. Data still goes in through the app's own front door.""" + out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1") + if ":" in out: + return out.split(":", 1)[1].strip().strip('"').strip("'") + return "" + + +def snapshots(name): + """The restorable copies the backups page offers for this app.""" + code, d = ctl("GET", f"/api/backup/snapshots?stack={name}") + data = d.get("data") if isinstance(d, dict) else None + if isinstance(data, dict): + for k in ("snapshots", "items", "restore_points"): + if isinstance(data.get(k), list): + return data[k] + return data if isinstance(data, list) else [] + + +def restore(name, snapshot_id=None, wait_s=1200): + """The household's own way out: the „Visszaállítás a mentésből" button on the backups page. + + A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`, + `snapshot_id` — because that is the button the sentence tells them to press. + """ + snaps = snapshots(name) + if snapshot_id is None: + if not snaps: + say(f" [R] no restorable copy offered for {name}") + return {"ok": False, "why": "no snapshot offered", "snapshots": snaps} + first = snaps[0] + snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id") + say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)") + sess = open(f"{SC}/sess.txt").read().strip() + csrf = open(f"{SC}/csrf.txt").read().strip() + r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}", + "-X", "POST", + "--data-urlencode", f"_csrf={csrf}", + "--data-urlencode", f"stack_name={name}", + "--data-urlencode", f"snapshot_id={snapshot_id}", + f"{BASE}/backup/restore"], timeout=180) + head = (r.stdout or "").split("\n")[0].strip() + loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")] + say(f" [R] POST /backup/restore -> {head} {loc[:1]}") + t0 = time.time() + last = None + while time.time() - t0 < wait_s: + code, d = ctl("GET", "/api/backup/restore-status") + dd = d.get("data") or {} + cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message")) + if cur != last: + say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}") + last = cur + if not dd.get("running", False) and time.time() - t0 > 5: + break + time.sleep(2) + st = stack(name) + say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} " + f"phase={st.get('update_phase')}") + return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps, + "http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1), + "state_after": st.get("state"), "hold_after": st.get("hold_reason"), + "observables_after": observables(name)} diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 1c18276a..258f3f5c 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -339,3 +339,12 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-639** | **After a held update the previous definition survived only in the recovery unit (P3).** Closed in controller **v0.263.0/v0.263.2**: the pre-update compose, applied definition, pin and the pinned version's `.felhom.yml` are kept until the undo is over; the undo never reads the unit (R-645). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — SHIPPED** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | | **R-641** | **An app with no database server had no last-second copy (P2).** Closed in controller **v0.263.0**: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | | **R-642** | **Start answered 200 "start completed" over a crash loop (P3).** Closed in controller **v0.263.0**: start/restart answer `requested — state now: `, never "completed" (`startAnswer`, pinned by `TestR642_*`, red-proofed); live on 9202: `Stack romm start requested — state now: running`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | + +## 2026-09-23 — the household is told when an update is undone or held (controller v0.264.0, hub v0.120.0) + +| Row | What | Closed | Full text | +|---|---|---|---| +| **R-606** | **Every sentence the update path shows a household was Hungarian-only (P2).** Closed in controller **v0.264.0** (`bc27894`): `UpdateError` is stored as a key + args; phase labels, refusals, failure lines, the undone line, the hold sentence and its prefix and `UpdateCopyHolds` render in the reader's language on both pages and in `GET /api/stacks/`; the stored Hungarian is byte-identical (parity gate green). Live: 9202 and 9201, both pages, both languages. Three leftovers → **R-647**. | **CLOSED 2026-09-23 — PROVEN-LIVE (endpoint-level; `audits/undo-fleet-2026-09-23/22-*`, `23-*`, `35-*`)** | `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` | +| **R-620** | **A disabled notifier dropped every event with no local trace (P3).** Closed in controller **v0.264.0**: one WARN per event type per process naming the dropped event, then DEBUG. Live on 9202: six types each WARNed once; `app_start_failed` WARN at 12:02:30Z then DEBUG at 12:03:00Z. | **CLOSED 2026-09-23 — PROVEN-LIVE (`24-9202-r620-warn-then-debug.txt`)** | as above | +| **R-646** | **An app pinned before v0.263.2 had no record of its own `.felhom.yml` (P3).** Closed in controller **v0.264.0**: a startup pass records `applied-meta/` for every deployed, pinned app CURRENT with the catalog; a behind app is skipped by name (its file is gone — not backfillable by design). Live: 9202 recorded 3, 9201 recorded 10, skipped 0. | **CLOSED 2026-09-23 — PROVEN-LIVE (`28-9202-controller-log.txt`, `30-9201-before-and-upgrade.txt`)** | as above | + diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 2f516632..b4e35c17 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -764,7 +764,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** | | **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | | **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** | -| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** **— UPDATE NIGHT 2026-09-21:** **CONFIRMED 2026-09-21 (update night) on the HOLD sentence, which this row's own text ranks highest** — *„a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK"*. Read off `/apps/adventurelog?lang=en` while the app was genuinely held after a real failed upstream edge. **Everything around it is correctly English** — the nav, „An installed app is not running: AdventureLog (stopped)", „Update available — today", „Move to another storage" — **and the two sentences that matter are Hungarian**: „A(z) adventurelog frissítése … nem sikerült, és az alkalmazás nem indult el az új verzióval." and „Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." So an English household is told in English that their app is stopped, and in Hungarian **which copy brings their data back, when it was taken and what is inside it.** Positive and negative controls both quoted. **Also settled by the same reading, and it belongs to `09` Q4:** the held app is surfaced on EVERY authenticated page, not only its own — the banner carried BOTH held apps at once. Evidence: `audits/update-night-2026-09-21/15-r606-hold-sentence-on-the-english-page.txt` and `apps/adventurelog/held-page-en.html`. **— UPDATE NIGHT 2026-09-21:** **the row records v0.260.0 as having made the pre-flight REFUSALS reach an English household in English. Measured 2026-09-21: it did not.** Three refusals were requested with `?lang=en`, with the Hungarian request as the control, and all three came back **identical Hungarian**: `held` (the one that names which copy holds what), `not_deployed`, and `disk` („Nincs elég szabad hely a frissítéshez: 1.4 GB szabad…"). **The mechanism is not a regression — it is that the pipe was built and the sentences never entered it.** `Router.langFor` DOES honour `?lang=` (`api/i18n_api.go:30-40`) and `errText` calls it; but `MsgUpdateDiskFmt`, `MsgUpdateNotDeployed`, `MsgUpdateBusy` and their siblings (`stacks/update.go` L81-96) are **finished Hungarian string constants** raised with `fmt.Sprintf`, and `errText` correctly renders "its own text" for an error carrying no bundle message. Only the sentences BORN as keys — v0.260.0's `downgrade`, v0.261.0's `self_updating` — actually translate. **A row that records something as fixed when it is not is worse than an open row**, which is why this correction is here rather than in prose. Evidence: `audits/update-night-2026-09-21/20-refusals-in-english.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-607** | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. **— UPDATE NIGHT 2026-09-21:** **Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one.** On `mealie` the bump was pushed, `POST /api/sync` AND `POST /api/stacks/rescan` were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached `catalog_images` yet. The guarded Update was then pressed and **reported „Frissitve" after 2.1 seconds having moved nothing at all**: pinned, installed, the live compose line and `docker inspect` all still read `v3.20.1`. That is honest given a stale cache — the pin is written from *the catalog's current definition*, which was still the old one — but what the household sees is a button that says it updated them and did not. **A NUMBER, at last, which is what this row asks for:** the night's harness was changed to poll `catalog_images` until the pushed reference appears and to report how long that took; those figures are each edge's `badge_catchup_seconds`, and here they are: **4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s.** The four fast ones are one sync+rescan round; the 29-second one (`nextcloud`, an engine-sidecar bump) needed **additional** sync+rescan rounds before `catalog_images` carried the pushed reference. **So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour**, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-608** | **[P2-MEDIUM] The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the window proposed for automatic app updates.** FOUND 2026-09-21 by reading the clock, not by a failure. The controller self-updates daily at `self_update.auto_update_time` — **default 04:30** (`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365) — and again from `MaybeAutoUpdate` after ANY hub report once a floor sits above the box, so at any hour. The swap restarts the controller container. `09` §3b **Q1** proposes **02:30–05:00** for automatic app updates. **It contains 04:30.** **MEASURED, and the gap was NARROWER than it first looked — which is why the fix is where it is:** the updater's only busy gate was `backupRunning` (`updater.go` L61, read at L487 dry-run, L512 `TriggerUpdate`, L660 `maybeAutoUpdate`), wired in `main.go` L659 to `backupMgr.IsRunning()`. The guarded update's **`backing-up` phase DOES take the backup single-flight** (`RunAppBackupNow` → `acquireRunning`, `backup/update_guard.go:333`), so that ONE phase was already covered. `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` were not — and the last two are exactly where the new version may already have touched the customer's data. The reverse was absent too: `UpdatePreflight` never asked whether a swap was running. **CLOSED 2026-09-21 — controller v0.261.0.** `stacks.Manager.AnyUpdating()` → `Updater.SetAppUpdatingCheck`, a deliberate sibling of `SetBackupRunningCheck` consulted in the SAME three places; `Updater.IsUpdateRunning` → `Manager.SetSelfUpdatingCheck`, and `UpdatePreflight` refuses `self_updating`. Both wired in `main.go`, the only place holding both objects — **`stacks` never imports `selfupdate`.** Two sentences born as bundle keys. **THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH:** `Stack.Updating` is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it. A latching gate would be a worse failure than the one prevented, and a silent one. Pinned by `TestR608_LockReleasesAfterHold`. Four red-proofs, each seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** | | **R-609** | **[P3-LOW] An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never".** `UpdateRefusal.Reason` has existed since v0.237.0 (`busy`, `deploying`, `updating`, `held`, `migrating`, `memory`, `disk`, `no_backup`, `downgrade`, and `self_updating` since v0.261.0) and never left the process: the 409 body carried only the translated sentence. **The distinction is not decorative** — `busy`/`updating`/`deploying`/`migrating`/`self_updating` are TRANSIENT and `held`/`downgrade` are TERMINAL until a person acts. A caller that cannot tell them apart either gives up on a passing backup window or presses a terminally-refused button on every pass for ever. `09` §6.2's unattended caller reads exactly this. **CLOSED 2026-09-21 — controller v0.261.0.** The body gains `data: {"reason": ""}`, ADDITIVELY; the sentence is unchanged and no page moves. Table-driven test over five reachable paths plus a control that a non-refusal carries none. **FOUND WHILE WRITING THE TEST, NOT BY READING — and it was the reason that matters most:** `actionStack` refuses a HELD app on **its own line, BEFORE `UpdatePreflight`** (`api/router.go` ~L601), so `held` would have been the one reason missing from the wire. That line now carries it too. Red-proof seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** | @@ -789,7 +788,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-617** | **[P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through `repos/migrate` — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it".** FOUND 2026-09-21 creating the drill catalog. Both `~/.git-credentials` tokens carry `write:misc,write:notification,write:package,write:issue,write:repository`; `POST /api/v1/user/repos` requires `write:user` and answers **403**, and `POST /api/v1/admin/users//repos` requires `write:admin` and answers 403 too. `POST /api/v1/repos/migrate` with the same token answered **201** and created the private repository. **Why this is a row and not a note:** a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. **What it needs:** one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: `audits/update-night-2026-09-21/03-drill-repo.txt`. | **READY — rank P3-LOW; owner: CC (docs)** | | **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. **CLOSED 2026-09-22.** All three fixed in one commit (`app-catalog-felhom.eu@793c4fb`): tandoor `8080 -> 80`, wger `80 -> 8000`, zipline `/api/health -> /api/healthcheck`. No `image:` line moved, so no `catalog_since` moved. **RED-PROOFED LIVE ON 9202 THROUGH THE PRODUCT, BOTH DIRECTIONS.** Before the fix, at the LIVE pin, all three read **`Nem egészséges` / `Not healthy`** on their own app page while docker reported every container healthy and each front door served a real page through the household's own route — tandoor 200 `Login / Sign In`, zipline 200 `Zipline`, wger 200 `wger Workout Manager`. The fix was applied through the REAL sync (`POST /api/sync` answered *frissítve: tandoor, wger, zipline*) and all three read **`Fut` / `Running`** at the next poll, with no redeploy and no restart. **AND THE EDGE THAT FAILED WAS RE-WALKED AND PASSED:** tandoor `2.6.13 -> 2.6.15` via the drill catalog ended **`done` at +41.1 s** with the seed read back through tandoor's own front door and both containers running with zero restarts — where the identical edge on 2026-09-21 entered `verifying` at +58.4 s and ended `failed` at **+361.9 s** with the app stopped. **Same app, same versions, same button; the only change is one port number.** tandoor's verdict moved `failed -> proven` and it is now on the live catalog. **THE GATE SHIPPED WITH IT:** `scripts/check-probe-matches-compose.py`, a `--fast` row in `catalog_gates.py`, comparing the probe against the SAME service's own compose healthcheck. The rule was read out of `healthprobe.go` rather than guessed: a wrong PORT refuses for every check type; a wrong PATH refuses only for `type: api` WITH an `expect` block and WARNS otherwise, which is why `home-assistant` is warned about and not convicted. Four red-proofs and five decoys, plus five more with PyYAML shadowed out (R-630's sibling problem — see below). **Residual, filed separately:** R-630 (paperless-ngx's probe can never run at all) and R-631 (five templates the static rule cannot judge). | **CLOSED 2026-09-22 — three probes fixed, red-proofed live both ways, gate shipped with decoys; tandoor re-walked `failed -> proven`** | | **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** | -| **R-620** | **[P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so.** FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. `hub.enabled: false` there, and `Notifier.Publish` returns at `notify/notifier.go:269` — **before** any log line — as do `NotifyHealthChange` (:359) and four more entry points. Startup says it once (`[INFO] Notifier disabled (hub not configured)`) and then every later event, of every severity up to `critical`, vanishes without a word. **The measurable consequence tonight:** the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. **The consequence on a real box is smaller but not zero:** the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. **Fix shape:** one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: `audits/update-night-2026-09-21/11-notifier-disabled.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. **FIXED in v0.262.0.** `failAndHold` now writes each service's log into `/hold-logs//compose-logs.txt` (`compose logs --no-color --tail 400`) **before** the `down` that destroys them, best-effort by design: a hold must never fail because its evidence could not be written. This is what `adventurelog` cost — nine migrations ran, the app never bound its port, and the only log that could have said why was gone before anyone looked. | **CLOSED 2026-09-22 — v0.262.0: the hold keeps the app's own log before stopping it** | | **R-622** | **[P2-MEDIUM] `adventurelog v0.13.0` migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe.** MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. `v0.12.1 → v0.13.0` (backend AND frontend together, PostGIS held constant). The backend applied **nine migrations, every one `... OK`** — `adventures.0072_trail_wanderer_author_fields` through `integrations.0009_alter_endurainintegration_auth_method`, plus `billing.0001_initial` — and then the container's own healthcheck failed with `URLError: [Errno 111] Connection refused` on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. **THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING:** the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." **This is exactly the case `09` §4 exists for:** the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. **What it needs:** adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/`. | **READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails** | | **R-623** | **[P3-LOW] The unattended-update caller turned every SUCCESS into a `timeout`, and then refused to press that app again — the instrument, not the box.** FOUND 2026-09-21 (update night) by reading `unattended-caller.py` before relying on it for the Q4 hold measurement. Its `call()` returns the API **envelope** — `{"ok": true, "data": {…}}` — and `follow()` read `update_phase` and `updating` **off the envelope**, where neither exists. Both were therefore always `None`; the end test `not updating and phase in ("done","failed")` could never fire; every followed update ran the full **900-second** timeout and was recorded `timeout`, which the caller treats as terminal and adds to `never_again`. `main()` unwraps `data` for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. **The 2026-09-21 run did not catch it because the only pass that reached `follow()` was Scenario F, whose log was lost to a buffering `tail`** — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. **This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement**, and worse, it is a measurement that says the box behaved badly when the box behaved well. **FIXED in the same file 2026-09-21** (unwrap `data`, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. **What it does NOT invalidate:** the G-b no-retry proof, which is entirely in the refusal path and never reached `follow()`. **What it DOES qualify:** any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | **CLOSED 2026-09-21 — fixed in `audits/update-arc-gaps-2026-09-21/unattended-caller.py`** | @@ -811,7 +809,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** | | **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** | | **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | -| **R-646** | **[P3-LOW] An app pinned before controller v0.263.2 has no record of its own `.felhom.yml`, so its FIRST undo probes with whatever file the catalog sync last put in place.** The undo judges the old version with the pinned version's probe, kept in `/applied-meta/` since v0.263.2 — written at deploy, pin adoption and each guarded update. Every app deployed and pinned before that release has no record until its next pin; its first update keeps the CURRENT `.felhom.yml` for the undo (logged: *no applied .felhom.yml for the pinned version*). When the new version changed its probe, that current file is the new one, and an undo that worked would be judged "did not start" and HELD — an honest hold (`not_started`), never a false success. MEASURED on 9202 2026-09-23 (romm, v0.263.1's shape). **A backfill is possible for the apps that are NOT behind** (their current file IS the pinned version's): a startup pass like `AdoptPins`, for apps whose catalog order is equal. Apps already behind cannot be backfilled — the pinned version's file is gone. `pin.go`, `undo.go` `savePreUpdateMeta`. | **OPEN — P3; owner: CC** | +| **R-647** | **[P3-LOW] Three small leftovers of the update mail and of R-606, found by the live proof of 2026-09-23.** (1) **A reader in the OTHER language than the box reads a held update's sentence in the box's language.** `fillHoldReason` and the stored `UpdateError` of a hold use the box language at hold time (the hold sentence is composed, not a key), so on an English box `?lang=hu` shows the English hold, and on a Hungarian box `UpdateErrorIn("en")` falls back to the stored Hungarian. The HOUSEHOLD always reads its own box language, which works — only the `?lang=` test switch and an operator reading a box in the other language see it. Fix shape: store a held update's error as a key the page can re-render (`RestoreHoldForLang`), and let the Hungarian `holdText` also call it. (2) **The `Note:` line of an English household mail prints the raw details, whose `copy_holds` is the Hungarian phrase** (`"copy_holds":"a beállításokat, …"`). Fix shape: send the phrase's key, or leave `copy_holds` out of the household copy. (3) **Two log wordings of my own:** the undone event logs `hold recorded: true` (meaningless for an undone update), and the disabled-notifier line for `health_change` prints the health STATUS where it says `severity` (`severity warn`) — the real events carry correct severities. Evidence: `audits/undo-fleet-2026-09-23/23-9202-held.txt`, `41-mails-as-received.md`, `24-9202-r620-warn-then-debug.txt`. | **OPEN — P3; owner: CC (controller)** | +| **R-648** | **[P3-LOW] The drill's „Mentés most" is a WHOLE-BOX backup, so every drill that seeds a throwaway app stops and restarts every standing app for its volume dump.** Found 2026-09-23 on 9201: two presses (one per throwaway app) restarted 9 of the 10 standing apps twice, a few seconds each (`backup.go` "Stopping for safe volume dump"); all came back healthy with byte-identical files, and the nightly does the same — but a brief that says "standing apps untouched" is not met. Fix shape: the harness backs up only the throwaway app (a per-app backup call, if the product has one; otherwise rely on the guarded update's own `backing-up` phase and skip the press). Evidence: `audits/undo-fleet-2026-09-23/49-9201-teardown.txt`. | **OPEN — P3; owner: CC (harness)** |