diff --git a/CONTEXT.md b/CONTEXT.md index 49237fec..254c8a70 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,21 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-23 (afternoon) — the undo is built: controller v0.263.2 (decisions 15, 19, 20) + +**Operator rulings 19 and 20** recorded in `09` §3: the copy method is chosen by a bake-off; on update +nights the full-system backup waits for the update leg inside its window (leg stops at W+5h; built with +§6.4 part 7). **Bake-off** (`audits/undo-bakeoff-2026-09-23/`): both methods passed; the folder copy +won because an app with no database server gets no dump. **Built** in controller v0.263.0–v0.263.2: +phase `copying` after the pull (named volumes `cp -a` into `.pre-update-`, marker last), +`undoing`/`undone` on a failed check, a HOLD only when the undo fails (prefix + data state), journal +recovery for both new phases, `last_update_undone` + a page line, R-642's honest Start answer. +**Proven live on 9202** (`audits/undo-live-2026-09-23/`). **Two live-only defects, fixed:** the +undo's probe must run on an app the current probe marks `unhealthy`; and the pinned version's +`.felhom.yml` must be recorded AT PIN TIME (`applied-meta/`), because `.felhom.yml` flows in on every +catalog sync. **Residual:** R-646 (apps pinned before v0.263.2), R-638/R-640 narrowed to the restore +paths, R-645. **The floor is NOT raised — that is the operator's question.** + ## 2026-09-23 — the operator rules on automatic updates (`09` §3 decisions 11–18); two spikes say what the build needs **Operator rulings, not CC decisions.** One window = a leg of the backup chain (11); automatic, per-box diff --git a/REPORT.md b/REPORT.md index 78f96570..656d5c54 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,119 +1,81 @@ -# REPORT — update arc: the operator's rulings recorded, the undo and ladder spiked, the memory watch, the build plan +# REPORT — the undo: bake-off, build (controller v0.263.0 → v0.263.2), live proof -2026-09-23. Repos touched: **felhom.eu** (docs, register, STATUS, CONTEXT, evidence), -**app-catalog-felhom.eu** (`scripts/` only — the memory watch), **admin/app-catalog-drill** (drill -commits, reset to live `main` at the end). **felhom-controller, felhom-agent, hub: read only.** -Architecture read first and named: `documentation/architecture/09-update-architecture.md` (all of it), -`07-backup-architecture.md` §6. -Baselines verified live before starting: controller `b9deec19077b`, agent `d9864a94bf62`, felhom.eu -`267dcad01bcf`, catalog `02844ae0a579` — all equal to the brief. +2026-09-23 (afternoon). Repos touched: **felhom-controller** (v0.263.0, v0.263.1, v0.263.2), +**felhom.eu** (docs, register, capability map, STATUS, CONTEXT, evidence), **admin/app-catalog-drill** +(drill commits, reset to live `main` at the end). **app-catalog-felhom.eu, felhom-agent, hub: untouched.** +Architecture read first and named: `documentation/architecture/09-update-architecture.md` (§3 decisions +11–20, §4, §6.1, §6.1a, §6.4). Baselines verified live: controller `b9deec19077b`, agent +`d9864a94bf62`, felhom.eu `4c92beab8fdb`, catalog `cfcfe5278428` — all as the brief said. --- -## 1. Not done, or changed from the brief — first +## 1. Not done, or changed | item | state | |---|---| -| Part 0 — rulings into `09` | **done**, commit `805ad1e` (documents only) | -| Part 1 — the undo, four cases + the wrong case | **done, with three changes.** (a) **The fourth case, "files on disk", was measured on vikunja's attachment (a file in a volume), not on a bind-mounted drive folder:** romm's two drive folders stayed EMPTY throughout — they measured nothing and are reported as unmeasured. (b) **Starting the app by hand needed its decrypted secrets; the session's safety guard refused that, and it was not worked around.** The product's own Start was used instead, after lifting the hold with the operator CLI + a controller restart. (c) The wrong case was run on BOTH engines, and on PostgreSQL with both the product's loader and the fixed one — the fixed one produced the session's most important finding (a truncated copy loads rc 0). | -| Part 2 — the ladder | **done** | -| Part 3 — the memory watch + red-proof | **done**; harness v2. `C3`, the harness's standing negative control, was **NOT run**: its template's `container_name: privatebin` collides with the privatebin the controller runs on 9202. The memory watch has its own pair instead — M1old must fail, M1 must pass. | -| Part 4 — the build plan | **done**, `09` §6.4 — with **one open point for the operator** the brief did not expect (§6) | +| rulings 19–20 | **recorded** (`09` §3), commit `5a349d9` | +| Part 1 — bake-off | **done in ~40 min** of the 2 h cap; folder copy chosen | +| Part 2 — build | **done — but in THREE releases, not one.** v0.263.0 failed its first two live proofs honestly (HELD, data put back). Two defects only the live box could show; each fixed + red-proofed + released: **v0.263.1** (the undo's probe never ran on an app the current probe held `unhealthy`), **v0.263.2** (the "old" `.felhom.yml` was already the new one — it flows in on every catalog sync; now recorded at pin time). 0.263.0/0.263.1 ran only on 9202 and were removed from it. | +| the household mail + operator event of decision 15 | **not built** — `09` §6.4 part 2 (R-606). The page line and the hold sentence are built. | +| the floor | **not raised** — the operator's question (STATUS) | +| immich/nextcloud rate test | **nextcloud deployed and was used** (185 MB MariaDB volume, 300 files through WebDAV); immich not tried | +| cut-off copy on vikunja in the bake-off | **could not be cut** (2.9 MB finishes before a kill lands); proven on docmost and romm instead; in the LIVE proof the marker was removed from a vikunja copy mid-update | -**Claims in the brief that turned out wrong:** -1. *"A safety dump exists for every app class"* — **false.** An app with no database server gets none - (`update safety dump for vikunja: the app has no database — nothing to copy (no-op)`), measured. -2. *"The box keeps a git clone of the catalog"* with history — **false.** Depth 1 on both demo - guests, `rev-list --count HEAD` = 1; `sync.go:283`/`:300` clone and fetch `--depth 1`. -3. *"`stacks.update_window` is unread"* — **true, and the grep was widened** from `config.go` + - `setup/handlers.go` to the whole controller repo: the only other hits are - `configs/controller.yaml.example` and the i18n base file. No Go code reads it. -4. The register held **329** row lines by `grep -c '^| \*\*R-'`, not 326; the highest id was R-636 as stated. +**Claims in the brief that turned out wrong or unmeasured:** +1. *"All three apps keep their data only in named volumes"* — true from compose AND on disk for these + three; but romm also has two BIND folders (roms, resources), empty throughout — never copied, by rule. +2. *"A volume copy with the containers stopped is consistent for both engines"* — **measured true**: + docmost (PostgreSQL) and romm (MariaDB) came back with ledgers equal, three times each. +3. *"Immich or Nextcloud deploys on 9202"* — Nextcloud did; Immich was not tried. +4. **Not in the brief, and the most important:** *"the old `.felhom.yml`"* does not exist at update + time. It is replaced on every catalog sync (`09` §5.4). The build now records it when a version is + pinned. R-646 for apps pinned before that. -## 2. Part 0 — the rulings +## 2. The bake-off (`audits/undo-bakeoff-2026-09-23/README.md`) -`09` §3 gains decisions **11–18** in the existing shape; decision 3's second half and §6.1's abort -paragraph are marked REPLACED with pointers; §4 says why the undo is not a rollback; §3b is kept, -headed ANSWERED, each question pointing at its decision; §6.2 rewritten to the ruled shape; the slices -table updated. Register: R-450, R-451, R-446, R-463 cite the decisions. +Both methods passed every row on docmost, romm, vikunja (seeds A and B back, ledgers equal, cut-off +detected before anything moves, ≤ 5.3 s extra downtime). **Folder copy chosen (decision 19):** an app +with no database server gets no dump, so dump-and-load would need the folder copy anyway. Rate: 185 MB +in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s (warm cache); 5 GB ≈ 12 s here, 25–50 s on a cold or spinning +disk. Found on the way: the undo must never read the recovery unit (R-645), and killing `docker run` +does not stop the copy container. -## 3. Part 1 — the undo, by hand +## 3. Red-proofs (each seen failing, then restored) -Full evidence and tables: `documentation/audits/update-rulings-2026-09-23/README.md`. +| # | mutation | test that failed | +|---|---|---| +| 1 | the undo call removed (v0.262.1's shape) | `TestUndo_FailedUpdateIsPutBackWithItsData` — `held=true phase="failed"` | +| 2 | copy validation removed | `TestUndo_CutOffCopyIsRefusedBeforeAnythingMoves` — state `half` | +| 3 | old-probe rule removed | `TestUndo_UsesTheOldProbe` — `not_started` | +| 3b | the old probe read from the stack dir (v0.263.1's shape) | `TestUndo_UsesTheOldProbe` | +| 4 | the `undoing` recovery arm removed | `TestUndo_PowerCutDuringTheUndoResumesIt` — `resumed=[]` | +| 5a | `last_update_undone` not recorded | `TestUndo_FailedUpdateIsPutBackWithItsData` — `got ` | +| 5b | the page line removed from the handler | `TestUndo_PageSaysTheBoxPutTheAppBack` (hu and en) | +| 6 | the hold prefix ignored | `TestUndo_HoldSentenceSaysTheUndoWasTriedAndTheDataState` | +| 7 | R-642: "completed" back | `TestR642_StartIsNeverReportedCompleted` | +| 8 | a successful update keeps the undone note | `TestUndo_ASuccessfulUpdateEndsTheUndoneNote` | +| 9 | the undo probe switched off on an `unhealthy` app (v0.263.0's shape) | `TestUndo_OldProbeRunsOnAnAppTheCurrentProbeMarkedUnhealthy` — `not healthy within 1s (last: state unhealthy)`, the live message verbatim | +| 10 | the pin advance records no probe | `TestUndo_PinningRecordsThatVersionsProbe` | -| | docmost (PostgreSQL) | romm (MariaDB) | vikunja (volume, no DB server) | -|---|---|---|---| -| held after | 95.3 s | 102.6 s | 93.3 s | -| old version on migrated data | **refuses** | **refuses** | starts | -| safety dump | 135 816 B, DB only | 62 943 B, DB only | **none** | -| product loader (`ImportDump` semantics) | **FAILS**, rc 3, foreign keys | rc 0, **12 tables left behind** | — | -| fixed load (empty schema + copy, one transaction) | rc 0, 1.38 s | (product loader sufficed) | — | -| undo → healthy on the OLD probe | ≈ 16 s | ≈ 38 s | ≈ 1 s | -| data before / after the backup | yes / **yes** | yes / **yes** | yes / **yes** (attachment too) | +`go build ./... && go vet ./... && go test ./...` rc=0 before each of the three commits; controller +gates rc=0 (the `docker -v` gate needed four named-volume mounts allowlisted with their why). -**Does a product path load a safety dump back?** Yes — `rollbackSafetyDump`, but only the off-site -restore calls it; the update never reads its own dump, and `failAndHold` deletes the pre-update -definition copies. +## 4. Live proof (`audits/undo-live-2026-09-23/README.md`) — endpoint-level, both languages -**The wrong case:** PostgreSQL + product loader → rc 3, nothing changed (honest hold). **PostgreSQL + -the fixed atomic loader + a half-length copy → rc 0 and an EMPTY database (0 users, 0 constraints, -0 indexes) that still shows 42 tables and 48 ledger rows** — no hold, a dishonest success. MariaDB + -half copy → rc 1, half the tables already replaced (not atomic). The whole copies carry an end marker -the truncated ones lack; `ValidateDump` does not check it. +Three apps undone by the product with seeds A, B, C back and ledgers equal (30–52 s of undo); page +line hu/en; cut-off copy → HOLD `untouched` (prefix in the box's language, hu and en quoted); power cut +during `undoing` → resumed and undone; a person's press after an undo → `done`, note cleared; removal +deleted kept copies; R-642 Start answer. -## 4. Part 2 — the ladder +## 5. Rows -One press on vikunja two steps behind: **2.3.0 → 2.5.0 in 9.5 s; 2.4.0 never ran.** The box cannot see -2.4.0 (depth-1 clone). Recommended format: `update_ladder:` in `.felhom.yml`, intermediate steps with -their own definition — **not** git history, because romm's image-moving commit is the definition that -OOM-looped on demo-hp. Full comparison in the audit. +**Closed (4, moved to `CLOSED-ITEMS.md`):** R-637, R-639, R-641, R-642. **Narrowed:** R-638, R-640 (to +the restore paths). **Ruled:** R-643 (decision 20). **Noted:** R-645. **Opened:** R-645 (earlier this +session, filed with the bake-off) and R-646 (apps pinned before v0.263.2). **Open rows 337 → 335.** -## 5. Part 3 — the memory watch +## 6. Teardown — three layers -`upgrade-test.py` v2: after a successful readback, `--soak` seconds (default 600) of light load -(4 callers), sampling every 15 s the kernel's `oom_kill` counter read host-side from the container's -cgroup, the peak, the limit, the restarts and Docker's OOMKilled flag. Kill or restart → `failed`; -peak > 80 % → mark `memory_tight`. New `Romm` fixture; edges `M1` (current template) and `M1old` -(the template as promoted, `15f9ebf`). - -- **M1old (red-proof): `failed` — first OOM kill at +76 s**, peak 512 MiB = 100 % of the limit, - restarts 0 (the container kept running — the shape that hid it on demo-hp), abort `refuses`. - **Ten minutes is ample for this failure under load.** -- **M1 (positive control):** **proven** + mark `memory_tight` — 608.5 s under 11 429 requests (5 712 × 200, 5 717 × 401): **0 kernel OOM kills, 0 restarts**, peak 621 MiB = **81 %** of 768 MiB — the watch passes the fix and still flags the thin headroom R-635 left open. - -## 6. Part 4 — the build plan - -`09` §6.4: ten parts, **≈ 22 evenings**, recommended order undo → sentences in the household's language -+ notifier honesty → test record + gate + memory check → ladder → digests → the automatic leg → R-636 → -R-625 → PostgreSQL conversion; fleet view deferred by ruling. **One open point needs the operator -(R-643):** as ruled, the update leg sits between the off-site leg (W+105m) and the full-system gate -(W+2h) — at most 15 minutes a night. Recommendation: the full-system backup waits for the leg inside -its own four-hour window. - -## 7. Rows - -**Opened (8):** R-637 (build the undo), R-638 (the loader cannot replay over a newer schema — and the -restore the hold names is UNMEASURED after a real schema migration), R-639 (pre-update copies deleted -on hold), R-640 (a truncated PostgreSQL copy loads rc 0 into an empty database), R-641 (no-DB apps have -no last-second copy), R-642 (Start returns 200 over a crash loop), R-643 (the ≤15-minute leg, -operator), R-644 (gokapi crash-looping on 9202 at session start, not caused here). -**Updated:** R-446, R-450, R-451, R-462, R-463. **Closed:** none. **329 → 337.** - -## 8. Teardown — three layers - -| layer | state | -|---|---| -| **machine — guest 9202** | `controller.yaml` restored from the saved copy and read back identical (live catalog, no `update:` block); catalog cache re-cloned from the live repo (`02844ae`); docmost, romm, vikunja removed through the product (no containers, no volumes); romm's drive folder (kept by the product, R-442) removed by name; `/opt/upg` and every temp file in `/root` removed; the eight test images removed **by name** — `vikunja:2.4.0` was **absent**, i.e. never pulled, which is the jump seen from a second side; **no `prune`**. Containers afterwards: the same three apps as at the start (gokapi still crash-looping — R-644, pre-existing). `82-teardown-guest.txt` | -| **host — demo-hp** | nothing provisioned; only transient `/tmp` files, removed | -| **hub** | nothing touched — no hub call was made | -| **drill repo** | reset to live `main` `02844ae0a579`; `has_actions: false`; image lines identical to live. The local clone's push URL to the LIVE catalog was disabled at the start. `81-teardown-drill-repo.txt` | - -**One instrumentation slip, recorded:** a background watcher and the first teardown both wrote the same -temporary script file on demo-hp at the same moment, so the first teardown never ran (its output file -held the watcher's lines). Caught by reading the file, re-run after the watcher ended; the second run -is the one recorded. - -**Fences:** DooPlex, Peti's box, ep0, the demo guests' apps, `drill-r50`, `tester-1` and the hub were -not touched. The live catalog's `main` was `02844ae0a579` before and after. - -**`unproven.py --summary`:** walked 20 / partial 17 / built 14 / missing 4 — **not walked 35 of 55, unchanged.** +Machine (9202): the three apps and every copy removed through the product; test images removed by +name; `controller.yaml` restored and read back; catalog cache on live `cfcfe52`; **9202 stays on +controller 0.263.2** (self-update off, fleet floor untouched). Host: nothing provisioned; 9202 stopped +and started once for the power cut. Hub: untouched. Drill repo: reset to live `main`. diff --git a/STATUS.md b/STATUS.md index 97da5e2c..b300d958 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,21 +1,26 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-23 — your seven answers on automatic updates are now written rules. I tested the two new pieces by hand on the scratch machine. The build plan is ready for you, part by part.** +**Updated 2026-09-23 (afternoon) — a failed update now puts the app back by itself. Built, and proven on the scratch machine. One question for you: whether the fleet gets it.** -**Decisions I took on my own: none.** +**Decisions I took on my own: none.** Your two afternoon rulings are written down: the bake-off picks the copy method, and on update nights the full-system backup waits for the updates. -**The automatic undo works — but not with the tool the machine has today.** I broke three real updates on purpose. For two of the apps, the old version refused to start on the data the new version had changed. After I loaded the copy taken seconds before the update, all three apps came back with all their data, including what was written after the nightly backup. It took 16 and 38 seconds. **But the machine's current way of loading a copy fails on one app and leaves junk behind on another.** A fix is known and tested: empty the database first, then load, as one step. +**The bake-off.** Both ways of keeping the last-second copy passed every test on three apps. Copying the app's data folders won, because one of the apps has no database server and so gets no database copy at all. The extra downtime was 1 to 5 seconds. -**One dangerous finding.** A cut-off database copy loads as a "success" — into an empty database. The machine's copy checker would accept it. The fix is simple: check that the copy has its end marker before loading. This also protects today's restores. +**What the machine does now.** When an update fails its health check, the machine puts back the previous version and the data exactly as it was seconds before the update. It stops the app only if that undo fails too, and then it says so. It never touches the household's own folders (photos, documents). -**The step-by-step climb does not exist yet.** Today a machine two versions behind jumps straight to the newest one; the middle version never runs. The machine also cannot see old versions: it keeps only the newest copy of the catalogue. I propose a list of tested steps in each app's catalogue file. +**Proven on the scratch machine, through the same buttons the page uses:** +- Three apps undone by the machine. Data written before the backup, after it, and seconds before the update all came back. It took 30 to 52 seconds. +- The app page shows one line in Hungarian or English: the update failed, the machine put the app back, nothing was lost. +- A damaged copy is caught before anything is poured back, and the app is held with an honest sentence. +- A power cut in the middle of the undo: after restart, the machine finished the undo. +- A person pressed Update again after the catalogue was fixed, and it worked. -**The memory lesson from RomM is now in the test bench.** After an update, the bench runs the app for 10 minutes and watches memory. RomM's old setup failed in 76 seconds, so the bench would have caught it. +**What went wrong on my side.** My first build failed its own live test twice. Both times the app stayed stopped with an honest sentence, and the data was safe. The first fault: the machine never asked the old version the right health question. The second: it kept the wrong copy of that question. My unit tests had passed both times. I fixed both the same afternoon and proved the fixes. The released version is the third build. -**What needs you — two things.** -1. **The build plan: about 22 evenings in 10 parts.** I recommend starting with the undo (4 evenings). It also makes the manual Update button safer on its own. If you do nothing, nothing is built and updates stay manual. -2. **One question about the night schedule.** As ruled, updates run after the off-site copy and before the full-system backup. That gap is at most 15 minutes a night. Option A (my pick): the full-system backup waits for updates, still inside its own 4-hour window. Option B: keep the 15 minutes; a machine far behind takes weeks to catch up. If you do nothing, part 7 of the plan waits. +**Rows.** Four closed, two narrowed, one new. The list went from 337 to 335. -**Rows opened and closed.** Eight opened, none closed. The list went from 329 to 337. +**What needs you — one question.** Should the demo machines and the fleet get this version now? +- **Yes (my pick):** I raise the floor, and both demo machines update themselves within a minute. A failed update then puts the app back instead of stopping it. +- **Not yet:** nothing changes. A failed update keeps stopping the app until someone restores it. -**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. Only the scratch machine was used, and it is back on the real catalogue.** +**Nothing on your own machine, Peti's machine or the off-site box was touched. The demo machines were not touched. The scratch machine runs the new version and is back on the real catalogue.** diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 4269c44b..5b26c150 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -103,6 +103,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | | **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. **THE NARROWING ABOVE WAS CLOSED THE NEXT DAY, 2026-09-22 — and re-widened the row.** `audits/PROBE-FIX-2026-09-22.md`. All three wrong probes were corrected in the catalog (`app-catalog-felhom.eu@793c4fb`: tandoor `8080→80`, wger `80→8000`, zipline `/api/health→/api/healthcheck`) and **red-proofed live on 9202 through the product in both directions**: at the live pin all three read `Nem egészséges` / `Not healthy` on their own app page while docker reported every container healthy and the front door served a real page; after the real sync all three read `Fut` / `Running` with no redeploy. **tandoor's edge was then re-walked with nothing else changed and ended `done` at +41.1 s**, seed read back, where the identical edge had ended `failed` at +361.9 s with the app stopped — so the tally is now **15 proven, 2 failed, 4 inconclusive**, and all fifteen are on the live catalog. A `--fast` catalog gate (`check-probe-matches-compose.py`) now refuses a probe that does not match the same service's own compose healthcheck, with four red-proofs and ten decoys including the no-PyYAML mode CI actually runs. **WHAT THIS ROW STILL CANNOT CLAIM, and the reason is exactly R-96 rule 3:** the guard is now shown correct for **47 of 53** templates. `paperless-ngx`'s probe has **never run on any box** — no container name matches its stack name, so it is silently skipped and its badge can never go red (**R-630**); and five more cannot be judged statically, one of which (`home-assistant`) is right only because its check type cannot fail (**R-631**). An absent alarm is equally consistent with healthy and with never checked. **And the sweep's ceiling, counted: 28 of the 53 templates have never been deployed by any drill (R-632).** **THAT CEILING WAS REMOVED THE SAME NIGHT, 2026-09-22 — all 28 walked (`audits/DRILL-the-28-2026-09-22.md`), so every template in the catalog has now been attempted at least once.** 26 of 28 deployed, **6 proven**, 5 inconclusive, 14 with no within-a-major edge upstream, 1 failed honestly and 2 undeployable — one of those (`plant-it`) **by design**, refused by the product's lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a **restore from its own copy, with the seed read back again** — 21 restored, and **2 were correctly REFUSED** with the sentence `07` §6.2 predicts for a class-A app whose local copy holds no file leg. **AND THE NIGHT NARROWED THIS ROW AGAIN, in the place the probe work could not reach.** `paperless-ngx` has no container matching its stack name, so no probe is ever built for it — and `verifying` does not skip: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three containers all read `healthy`. The controller's own words: *`not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app`*, at **+313.0 s**. **R-630, raised to P1.** So "the update is guarded" is now shown correct for 47 of 53 templates, wrong for none, and **actively harmful for the one template that has no probe at all**. **Two further limits on what this row may claim, both about STATE rather than health:** a `remove` sent while a restore is still running reports success and leaves a container restarting with a live public route (**R-633**) — while the product already refuses exactly that clash for `update` and for `restore`, naming the blocking operation; and an app can be **running, healthy and serving while recorded as `deployed: false`**, in which state the product refuses to remove it at all (**R-634**). In both, a person needed a shell to clear what the product could not. **ALL THREE ARE FIXED IN CONTROLLER v0.262.0 (2026-09-22), and the first is PROVEN LIVE.** *The stopped app:* `verifying` no longer loops on a probe that resolves to nothing — it settles on container state, the way an app declaring no check is judged, and says which it did. **Measured on paperless-ngx: the identical Update that ended `failed` at +313.0 s with the app stopped now ends `done` at +53.4 s**, with no `no probe container` warning in the log because the explicit `healthcheck.container` resolved the target. *The ghost:* `RemoveStack` consults the backup side's `Busy` guard — which the product already applied to `update` and to `restore` — and then WATCHES the compose project for 25 s after `down`, removing anything that carries its label and answering `verified: true/false`, because `down` returning 0 is a request rather than a result. *The unremovable app:* the refusal now asks whether anything EXISTS (containers, a compose file, an `app.yaml`) instead of reading a flag. **WHAT THIS ROW STILL MAY NOT CLAIM:** R-634's MECHANISM — why `deployed` goes false while containers run — **is not diagnosed**; only the consequence is fixed. And a probe can be right about the port and still wrong about what a 200 means: `romm` answered 200 from nginx for six hours while its workers were OOM-killed behind it (**R-635**). **"The update is guarded" has never meant "the new version runs".** | +| **A failed update is UNDONE by the box itself — the previous version back with its data from seconds before the update; the app is held only if that undo fails too** | controller **v0.263.2** | **PROVEN-LIVE (2026-09-23)** | `audits/undo-live-2026-09-23/README.md` (+ `audits/undo-bakeoff-2026-09-23/` for the method). On scratch guest 9202, through the endpoints the UI invokes: docmost (PostgreSQL), romm (MariaDB) and vikunja (SQLite in a volume), each a real migrating edge failing a deliberately wrong probe, **undone in 30–52 s** with seeds written before the backup, after it and seconds before the press ALL read back through each app's front door, ledgers equal to before; page line in hu and en. Cut-off copy → HOLD (`untouched`), prefix in the box's language; power cut (`pct stop`) during `undoing` → resumed after boot and undone; a person's press after an undo → `done`; removal deletes kept copies. | Folder copy of NAMED volumes only (`09` §3 decision 19); bind folders never touched. **Not built:** the mail to the household and the operator event (`09` §6.4 part 2), and the automatic caller (part 7) — this protects the manual button today. **R-646:** an app pinned before v0.263.2 undoes with the current `.felhom.yml` until its next pin. | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index ae8883ed..ee99d8f9 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -744,9 +744,12 @@ closed by construction: nothing reports an update complete on the compose exit c | 3 | `safety-dump` | `WriteUpdateSafetyDump` (R-361's undo copy) — **before the pin moves** | refused; nothing moves | | 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back | | 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) | -| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD | -| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) | -| 8 | `done` | installed images recorded, journal cleared | — | +| 5a | `copying` **(v0.263.0)** | `compose stop`, then each NAMED volume `cp -a` into `.pre-update-` by a helper that writes a finished-marker last (bind folders never) | copies removed, **pin put back, the old version started again** — nothing new ran | +| 6 | `starting` | `compose up -d --remove-orphans` | **UNDO** (below), HOLD only if the undo fails | +| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **UNDO** (below), HOLD only if the undo fails. Before v0.263.0: stop + HOLD, the pin stays (Scenario F) | +| 7a | `undoing` **(v0.263.0)** | every copy validated (marker) BEFORE anything is poured back; volumes emptied and refilled; definition, pin and the pinned version's `.felhom.yml` record put back from the job's own copies; `up`; health with the OLD version's probe | HOLD, the sentence prefixed *„A frissítés nem sikerült, és az automatikus visszaállítás sem."* + the data state (`untouched` / `half` / `not_started`); copies kept | +| 7b | `undone` **(v0.263.0)** | installed images recorded, copies removed, `app.yaml` `last_update_undone`, journal cleared | — | +| 8 | `done` | installed images recorded, copies removed, `last_update_undone` cleared, journal cleared | — | **The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and `update.health_timeout` (default `5m`). @@ -792,7 +795,21 @@ the harness has proven it, is slice 6's. **Not gated here:** a multi-major jump (R-40). It fails health and is held honestly; stepping is slice 6. -### 6.1a The undo (decision 15) — SPIKED BY HAND 2026-09-23, not built +### 6.1a The undo (decision 15) — SPIKED BY HAND, then BUILT: controller v0.263.2 (2026-09-23) + +**SHIPPED AND PROVEN LIVE on 9202** (`audits/undo-live-2026-09-23/README.md`): docmost, romm and +vikunja each made a real migrating update fail its (deliberately wrong) probe; the product undid all +three in 30–52 s, with data written before the backup, after it, and seconds before the press all read +back through each app's front door and the ledgers equal; the page carries one line in the request's +language. A cut-off copy → HOLD saying *the data is as the new version left it*; a power cut during +the undo → resumed after boot and completed; a person's press after an undo → `done`. **Two defects +only the live box could show, fixed the same day:** the undo's probe was never asked while the +current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` taken at update time +was already the new one, because `.felhom.yml` flows in on every catalog sync (v0.263.2: the pinned +version's file is now recorded in `applied-meta/` whenever a version is pinned). **Residual (R-646):** +an app pinned before v0.263.2 has no such record until its next pin. + +The spike, as it was run by hand before any build: Evidence: `audits/update-rulings-2026-09-23/README.md`. Three real migrating edges on 9202, each made to fail a deliberately wrong probe, each held by today's product, each then undone by hand. @@ -1025,7 +1042,7 @@ what the part can do to a household's data if it is wrong, not how likely that i | # | part | rulings / rows | cost | depends on | risk to customer data | |---|---|---|---|---|---| -| **1** | **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. | +| **1** | **SHIPPED — controller v0.263.2, proven live on 9202 2026-09-23** (`audits/undo-live-2026-09-23/`). **The undo.** Keep the pre-update copies (compose, applied, pin, **old `.felhom.yml`**) until the undo is over; in `failAndHold`: pin back → DB up alone → **validate the copy's completion marker** → **empty-then-load in one transaction** (PostgreSQL: the dump's schemas dropped and recreated inside the load's transaction; MariaDB: every table dropped first, and a failed load HOLDS with a sentence saying the database is in neither state) → full start → **health with the OLD probe** → `undone`, else HOLD. A volume tar at safety-dump time for apps with no database server. Household page + event; the mail rides part 2. | 15; the audit's 8-point list | **4** | — | **HIGH by nature** — it writes the customer's database. Bounded: it only ever loads the copy taken seconds before, validated first, atomically on PostgreSQL; every failure mode ends in today's hold. **It also makes the manual button safer on its own**, which is why it goes first. | | **2** | **The update sentences in the household's language** (R-606) and a mail when an automatic update is undone or held. | R-606, 15 | **1** | — | none | | **3** | **A disabled notifier says so** (R-620), so the mail of part 2 can be measured on a scratch box at all. | R-620 | **0.5** | — | none | | **4** | **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only | diff --git a/documentation/audits/undo-live-2026-09-23/00-deploy-0.263.0-on-9202.txt b/documentation/audits/undo-live-2026-09-23/00-deploy-0.263.0-on-9202.txt new file mode 100644 index 00000000..6e0e24ad --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/00-deploy-0.263.0-on-9202.txt @@ -0,0 +1,2 @@ +gitea.dooplex.hu/admin/felhom-controller:0.263.0 +gitea.dooplex.hu/admin/felhom-controller:0.263.0 Up 25 seconds (healthy) diff --git a/documentation/audits/undo-live-2026-09-23/10-docmost-undo.txt b/documentation/audits/undo-live-2026-09-23/10-docmost-undo.txt new file mode 100644 index 00000000..4a1ceb50 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/10-docmost-undo.txt @@ -0,0 +1,23 @@ +11:15:35 === docmost: live undo by the product +11:15:36 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cd8c-8fb0-741d-88b3-097dc97c4ec6","name":"drillBf2da7a","description":"","slug":"drillbf2da7a","logo":null,"visibility":"private","defaultRol +11:15:36 seed C written right before the Update: True +11:15:38 db before: 42 48 ledger, newest 20260620T010047-personal-spaces +11:15:39 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +11:15:39 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +11:15:41 + 2.1s phase=pulling label=Új verzió letöltése… +11:15:42 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt… +11:15:44 + 5.8s phase=starting label=Indítás az új verzióval… +11:15:55 + 16.8s phase=verifying label=Működés ellenőrzése… +11:17:26 + 107.1s phase=undoing label=Visszaállítás az előző változatra… +11:19:13 + 214.3s phase=failed label=A frissítés nem sikerült +11:19:13 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' +11:19:15 observables: {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": []} +11:21:17 app never answered on docs/ (last rc=0 code=404) +11:21:17 docmost: login as the seeded user http=404 ok=False +11:21:17 docmost: login body 404 page not found + +11:21:17 docmost B: cannot log in (http=404) +11:21:17 docmost B: cannot log in (http=404) +11:21:19 READBACK A=False B=False C=False db after: Error response from daemon: No such container: docmost-postgres Error response from daemon: No such container: docmost-postgres (before: 42 48 ledger, newest 20260620T010047-personal-spaces) +11:21:19 PAGE: {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) docmost frissítése 2026-09-23 11:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:14 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}} +11:21:21 leftover copies: 3 diff --git a/documentation/audits/undo-live-2026-09-23/11-deploy-0.263.1-on-9202.txt b/documentation/audits/undo-live-2026-09-23/11-deploy-0.263.1-on-9202.txt new file mode 100644 index 00000000..763c51ac --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/11-deploy-0.263.1-on-9202.txt @@ -0,0 +1 @@ +gitea.dooplex.hu/admin/felhom-controller:0.263.1 Up 25 seconds (healthy) diff --git a/documentation/audits/undo-live-2026-09-23/12-deploy-0.263.2-on-9202.txt b/documentation/audits/undo-live-2026-09-23/12-deploy-0.263.2-on-9202.txt new file mode 100644 index 00000000..b285f3da --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/12-deploy-0.263.2-on-9202.txt @@ -0,0 +1 @@ +gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 25 seconds (healthy) diff --git a/documentation/audits/undo-live-2026-09-23/20-romm-undo.txt b/documentation/audits/undo-live-2026-09-23/20-romm-undo.txt new file mode 100644 index 00000000..f2a116c9 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/20-romm-undo.txt @@ -0,0 +1,22 @@ +11:37:09 === romm: live undo by the product +11:37:10 romm seed B: POST /api/users (as the admin) http=201 {"id":3,"username":"drillb5fbc5e","email":"drillb5fbc5e@gate.invalid","enabled":true,"role":"user","permission_group_id" +11:37:10 seed C written right before the Update: True +11:37:13 db before: tables=24 alembic=0095_virtual_collections_source +11:37:13 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +11:37:13 + 0.0s phase=checking label=Ellenőrzés… +11:37:14 + 0.6s phase=safety-dump label=Adatbázis pillanatkép… +11:37:15 + 1.6s phase=pulling label=Új verzió letöltése… +11:37:17 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt… +11:37:21 + 7.9s phase=starting label=Indítás az új verzióval… +11:37:32 + 19.0s phase=verifying label=Működés ellenőrzése… +11:39:03 + 109.5s phase=undoing label=Visszaállítás az előző változatra… +11:40:58 + 225.0s phase=failed label=A frissítés nem sikerült +11:40:58 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.' +11:41:01 observables: {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": []} +11:43:02 app never answered on arcade/api/heartbeat (last rc=0 code=404) +11:53:05 app never answered on arcade/api/heartbeat (last rc=0 code=404) +11:53:05 romm B: GET /api/users http=404 seeded-user-listed=False (negative control listed=False) +11:53:05 romm B: GET /api/users http=404 seeded-user-listed=False (negative control listed=False) +11:53:08 READBACK A=False B=False C=False db after: tables=Error response from daemon: No such container: romm-db alembic=Error response from daemon: No such container: romm-db (before: tables=24 alembic=0095_virtual_collections_source) +11:53:08 PAGE: {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el. A(z) romm frissítése 2026-09-23 11:40-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 11:36 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem."}} +11:53:10 leftover copies: 3 diff --git a/documentation/audits/undo-live-2026-09-23/30-reset-apps.txt b/documentation/audits/undo-live-2026-09-23/30-reset-apps.txt new file mode 100644 index 00000000..35e15f23 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/30-reset-apps.txt @@ -0,0 +1,17 @@ +11:55:43 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'} +11:56:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ +11:56:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' +11:56:22 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'} +11:56:27 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +11:56:27 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +11:56:54 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat +11:57:01 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm' +11:57:02 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +11:57:33 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +11:57:41 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +0 +romm-drive-folder-removed + +sync 200 +1a7a491 DRILL: FROM states again (correct probes) for the live undo proof + diff --git a/documentation/audits/undo-live-2026-09-23/31-prep-docmost.txt b/documentation/audits/undo-live-2026-09-23/31-prep-docmost.txt new file mode 100644 index 00000000..6a07c280 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/31-prep-docmost.txt @@ -0,0 +1,14 @@ +11:57:46 === prep docmost: deploy at the drill FROM pin +11:57:46 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +11:58:16 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +11:58:16 deployed: True +11:58:16 docmost: /api/auth/setup http=200 rc=0 +11:58:17 docmost: login as the seeded user http=200 ok=True +11:58:17 C1 A: True +11:58:17 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +11:58:57 [4] backup idle; last=None +11:58:58 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cdb4-4376-7727-a390-a5b9e7f6c581","name":"drillB059474","description":"","slug":"drillb059474","logo":null,"visibility":"private","defaultRol +11:58:58 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +11:58:58 B reads back: True +11:59:01 named volumes: ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'] + port: 3000 diff --git a/documentation/audits/undo-live-2026-09-23/31-prep-romm.txt b/documentation/audits/undo-live-2026-09-23/31-prep-romm.txt new file mode 100644 index 00000000..d151782b --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/31-prep-romm.txt @@ -0,0 +1,16 @@ +11:59:08 === prep romm: deploy at the drill FROM pin +11:59:11 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm'] +11:59:11 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH'] +11:59:11 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +12:00:16 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'} +12:00:16 deployed: True +12:00:18 romm: POST /api/users http=201 +12:00:19 romm: login as the seeded user http=200 ok=True +12:00:19 C1 A: True +12:00:19 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +12:01:19 [4] backup idle; last=None +12:01:51 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillbd5863e","email":"drillbd5863e@gate.invalid","enabled":true,"role":"user","permission_group_id" +12:01:51 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +12:01:51 B reads back: True +12:01:54 named volumes: ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'] + port: 8080 diff --git a/documentation/audits/undo-live-2026-09-23/31-prep-vikunja.txt b/documentation/audits/undo-live-2026-09-23/31-prep-vikunja.txt new file mode 100644 index 00000000..6838553e --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/31-prep-vikunja.txt @@ -0,0 +1,15 @@ +12:02:01 === prep vikunja: deploy at the drill FROM pin +12:02:01 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +12:02:06 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'} +12:02:06 deployed: True +12:02:07 vikunja: register http=200 +12:02:07 vikunja: create project http=201 +12:02:07 vikunja: readback of the seeded project http=200 ok=True +12:02:07 C1 A: True +12:02:07 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +12:03:08 [4] backup idle; last=None +12:03:08 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drilld918b0","created":"2026-09 +12:03:08 vikunja B: project readback=True attachment content readback=True +12:03:08 B reads back: True +12:03:11 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db'] + port: 3456 diff --git a/documentation/audits/undo-live-2026-09-23/32-break-docmost.txt b/documentation/audits/undo-live-2026-09-23/32-break-docmost.txt new file mode 100644 index 00000000..ca78fe22 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/32-break-docmost.txt @@ -0,0 +1,2 @@ +11:59:04 [5] drill commit 58a76b596330: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0) +11:59:08 badge caught up after 4.5 s diff --git a/documentation/audits/undo-live-2026-09-23/32-break-romm.txt b/documentation/audits/undo-live-2026-09-23/32-break-romm.txt new file mode 100644 index 00000000..cf859b24 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/32-break-romm.txt @@ -0,0 +1,2 @@ +12:01:57 [5] drill commit 9f9283d26b22: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0) +12:02:01 badge caught up after 4.4 s diff --git a/documentation/audits/undo-live-2026-09-23/32-break-vikunja.txt b/documentation/audits/undo-live-2026-09-23/32-break-vikunja.txt new file mode 100644 index 00000000..963f3847 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/32-break-vikunja.txt @@ -0,0 +1,2 @@ +12:03:14 [5] drill commit 6075526e6b27: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0) +12:03:19 badge caught up after 4.5 s diff --git a/documentation/audits/undo-live-2026-09-23/40-undo-docmost.txt b/documentation/audits/undo-live-2026-09-23/40-undo-docmost.txt new file mode 100644 index 00000000..1d3584d8 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/40-undo-docmost.txt @@ -0,0 +1,20 @@ +12:03:41 === docmost: live undo by the product +12:03:42 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cdb8-9866-7a10-bde0-42e824b55295","name":"drillBbdf31d","description":"","slug":"drillbbdf31d","logo":null,"visibility":"private","defaultRol +12:03:42 seed C written right before the Update: True +12:03:44 db before: 42 48 ledger, newest 20260620T010047-personal-spaces +12:03:44 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +12:03:44 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +12:03:46 + 1.6s phase=pulling label=Új verzió letöltése… +12:03:48 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt… +12:03:50 + 5.3s phase=starting label=Indítás az új verzióval… +12:04:01 + 16.3s phase=verifying label=Működés ellenőrzése… +12:05:32 + 107.2s phase=undoing label=Visszaállítás az előző változatra… +12:06:02 + 137.6s phase=undone label=Visszaállítva az előző változatra +12:06:02 END phase=undone err=None hold=None +12:06:05 observables: {"pinned_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "installed_images": {"docmost": "docmost/docmost:0.95.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "catalog_images": {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: docmost/docmost:0.95.0", "image: postgres:16-alpine", "image: redis:7-alpine"], "docker_inspect": ["docmost docmost/docmost:0.95.0 running=true restarts=0", "docmost-postgres postgres:16-alpine running=true restarts=0", "docmost-redis redis:7-alpine running=true restarts=0"]} +12:06:05 docmost: login as the seeded user http=200 ok=True +12:06:06 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +12:06:06 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +12:06:08 READBACK A=True B=True C=True db after: 42 48 ledger, newest 20260620T010047-personal-spaces (before: 42 48 ledger, newest 20260620T010047-personal-spaces) +12:06:08 PAGE: {"hu": {"undone_line": "A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +12:06:10 leftover copies: 0 diff --git a/documentation/audits/undo-live-2026-09-23/40-undo-romm.txt b/documentation/audits/undo-live-2026-09-23/40-undo-romm.txt new file mode 100644 index 00000000..1c76210f --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/40-undo-romm.txt @@ -0,0 +1,20 @@ +12:06:10 === romm: live undo by the product +12:06:12 romm seed B: POST /api/users (as the admin) http=201 {"id":3,"username":"drillb2d189e","email":"drillb2d189e@gate.invalid","enabled":true,"role":"user","permission_group_id" +12:06:12 seed C written right before the Update: True +12:06:14 db before: tables=24 alembic=0095_virtual_collections_source +12:06:14 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +12:06:14 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +12:06:16 + 1.6s phase=pulling label=Új verzió letöltése… +12:06:18 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt… +12:06:24 + 10.0s phase=starting label=Indítás az új verzióval… +12:06:36 + 21.0s phase=verifying label=Működés ellenőrzése… +12:08:06 + 111.9s phase=undoing label=Visszaállítás az előző változatra… +12:08:58 + 163.9s phase=undone label=Visszaállítva az előző változatra +12:08:58 END phase=undone err=None hold=None +12:09:01 observables: {"pinned_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "installed_images": {"romm": "rommapp/romm:5.0.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "catalog_images": {"romm": "rommapp/romm:5.3.0", "romm-db": "mariadb:11.4", "romm-redis": "redis:7-alpine"}, "live_compose_image_lines": ["image: rommapp/romm:5.0.0", "image: mariadb:11.4", "image: redis:7-alpine"], "docker_inspect": ["romm rommapp/romm:5.0.0 running=true restarts=0", "romm-db mariadb:11.4 running=true restarts=0", "romm-redis redis:7-alpine running=true restarts=0"]} +12:09:04 romm: login as the seeded user http=200 ok=True +12:09:04 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +12:09:05 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +12:09:07 READBACK A=True B=True C=True db after: tables=24 alembic=0095_virtual_collections_source (before: tables=24 alembic=0095_virtual_collections_source) +12:09:07 PAGE: {"hu": {"undone_line": "A(z) romm frissítése 2026-09-23 12:08-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of romm at 2026-09-23 12:08 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +12:09:09 leftover copies: 0 diff --git a/documentation/audits/undo-live-2026-09-23/40-undo-vikunja.txt b/documentation/audits/undo-live-2026-09-23/40-undo-vikunja.txt new file mode 100644 index 00000000..a4cd5ed6 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/40-undo-vikunja.txt @@ -0,0 +1,19 @@ +12:09:09 === vikunja: live undo by the product +12:09:09 vikunja seed B: project http=200 task=2 attachment upload http=200 {"errors":null,"success":[{"id":2,"task_id":2,"created_by":{"id":1,"name":"","username":"drilld918b0","created":"2026-09 +12:09:09 seed C written right before the Update: True +12:09:11 db before: tables=36 ledger=117 newest=SCHEMA_INIT +12:09:12 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +12:09:12 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +12:09:13 + 1.1s phase=pulling label=Új verzió letöltése… +12:09:14 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt… +12:09:15 + 3.7s phase=verifying label=Működés ellenőrzése… +12:10:46 + 94.5s phase=undoing label=Visszaállítás az előző változatra… +12:10:49 + 97.1s phase=undone label=Visszaállítva az előző változatra +12:10:49 END phase=undone err=None hold=None +12:10:51 observables: {"pinned_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "installed_images": {"vikunja": "vikunja/vikunja:2.3.0"}, "catalog_images": {"vikunja": "vikunja/vikunja:2.6.0"}, "live_compose_image_lines": ["image: vikunja/vikunja:2.3.0"], "docker_inspect": ["vikunja vikunja/vikunja:2.3.0 running=true restarts=0"]} +12:10:52 vikunja: readback of the seeded project http=200 ok=True +12:10:52 vikunja B: project readback=True attachment content readback=True +12:10:52 vikunja B: project readback=True attachment content readback=True +12:10:54 READBACK A=True B=True C=True db after: tables=36 ledger=117 newest=SCHEMA_INIT (before: tables=36 ledger=117 newest=SCHEMA_INIT) +12:10:54 PAGE: {"hu": {"undone_line": "A(z) vikunja frissítése 2026-09-23 12:10-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of vikunja at 2026-09-23 12:10 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +12:10:56 leftover copies: 0 diff --git a/documentation/audits/undo-live-2026-09-23/50-powercut-romm.txt b/documentation/audits/undo-live-2026-09-23/50-powercut-romm.txt new file mode 100644 index 00000000..492bbcb9 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/50-powercut-romm.txt @@ -0,0 +1,20 @@ +12:11:18 === romm: power cut DURING the undo +12:11:18 romm seed B: POST /api/users (as the admin) http=201 {"id":4,"username":"drillb48b886","email":"drillb48b886@gate.invalid","enabled":true,"role":"user","permission_group_id" +12:11:18 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +12:11:18 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +12:11:20 + 1.6s phase=pulling label=Új verzió letöltése… +12:11:21 + 2.9s phase=copying label=Az adatok másolása a frissítés előtt… +12:11:28 + 9.5s phase=starting label=Indítás az új verzióval… +12:11:39 + 20.5s phase=verifying label=Működés ellenőrzése… +12:13:09 + 111.0s phase=undoing label=Visszaállítás az előző változatra… +12:13:15 >>> POWER CUT (pct stop 9202) in phase undoing: rc=0 in 5.2s +12:13:18 guest started again: rc=0 +12:13:31 2026/09/23 10:13:25 update.go:1171: [WARN] [stacks] update recovery: romm was interrupted while UNDOING (started 2026-09-23T10:11:18Z) — marking it Updating and RESUMING the undo +2026/09/23 10:13:25 update.go:1224: [INFO] [stacks] update romm: resuming the UNDO after a controller restart + +12:14:22 after the restart: phase=undone hold=None +12:14:27 romm: login as the seeded user http=200 ok=True +12:14:28 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +12:14:28 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False) +12:14:31 READBACK A=True B=True C=True db: tables=24 alembic=0095_virtual_collections_source +12:14:33 leftover copies: 0 diff --git a/documentation/audits/undo-live-2026-09-23/60-cutoff-vikunja.txt b/documentation/audits/undo-live-2026-09-23/60-cutoff-vikunja.txt new file mode 100644 index 00000000..20076460 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/60-cutoff-vikunja.txt @@ -0,0 +1,26 @@ +12:14:33 === vikunja: a cut-off undo copy +12:14:33 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +12:14:33 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +12:14:34 + 1.1s phase=pulling label=Új verzió letöltése… +12:14:35 + 2.1s phase=copying label=Az adatok másolása a frissítés előtt… +12:14:37 + 3.7s phase=starting label=Indítás az új verzióval… +12:14:38 + 4.2s phase=verifying label=Működés ellenőrzése… +12:14:40 >>> finished-marker removed from one copy: +copy: vikunja_vikunja_data.pre-update-20260923T101436Z +total 12 +drwxr-xr-x 3 root root 4096 Sep 23 10:14 . +drwxr-xr-x 1 root root 4096 Sep 23 10:14 .. +drwxr-xr-x 2 root root 4096 Sep 23 10:13 data +-rw-r--r-- 1 root root 0 Sep 23 10:14 felhom-undo-complete +data + +12:16:08 + 94.8s phase=undoing label=Visszaállítás az előző változatra… +12:16:09 + 95.4s phase=failed label=A frissítés nem sikerült +12:16:09 END phase=failed err='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' hold='A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.' +12:16:11 PAGE (box language hu): {"hu": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}} +12:16:11 box language -> en: http 302 +12:16:11 PAGE (box language en): {"hu": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}, "en": {"undone_line": null, "hold": "The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja frissítése 2026-09-23 12:16-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 12:04 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."}} +12:16:11 box language -> hu: http 302 +12:16:13 copies kept: vikunja_vikunja_data.pre-update-20260923T101436Z +vikunja_vikunja_db.pre-update-20260923T101436Z + diff --git a/documentation/audits/undo-live-2026-09-23/70-manual-press-docmost.txt b/documentation/audits/undo-live-2026-09-23/70-manual-press-docmost.txt new file mode 100644 index 00000000..811e8819 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/70-manual-press-docmost.txt @@ -0,0 +1,17 @@ +12:16:42 === docmost: a person presses Update again after the undo (probe fixed in the catalog) +12:16:42 before, the page: {"hu": {"undone_line": "A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el.", "hold": null}, "en": {"undone_line": "The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost.", "hold": null}} +12:16:42 Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +12:16:42 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… +12:16:44 + 1.6s phase=pulling label=Új verzió letöltése… +12:16:45 + 3.2s phase=copying label=Az adatok másolása a frissítés előtt… +12:16:48 + 5.3s phase=starting label=Indítás az új verzióval… +12:16:59 + 16.3s phase=verifying label=Működés ellenőrzése… +12:17:19 + 36.8s phase=done label=Frissítve +12:17:19 END phase=done err=None hold=None +12:17:20 docmost: login as the seeded user http=200 ok=True +12:17:20 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +12:17:20 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False) +12:17:23 READBACK A=True B=True C=True installed={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +12:17:23 after, the page: {"hu": {"undone_line": null, "hold": null}, "en": {"undone_line": null, "hold": null}} +12:17:25 app.yaml last_update_undone: 0 + diff --git a/documentation/audits/undo-live-2026-09-23/80-r642-start-answer.txt b/documentation/audits/undo-live-2026-09-23/80-r642-start-answer.txt new file mode 100644 index 00000000..27ca917b --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/80-r642-start-answer.txt @@ -0,0 +1 @@ +12:17:57 POST /api/stacks/romm/start -> 200 {'ok': True, 'data': {'state': 'running'}, 'message': 'Stack romm start requested — state now: running'} diff --git a/documentation/audits/undo-live-2026-09-23/90-controller-undo-log.txt b/documentation/audits/undo-live-2026-09-23/90-controller-undo-log.txt new file mode 100644 index 00000000..49149c60 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/90-controller-undo-log.txt @@ -0,0 +1,50 @@ +2026/09/23 10:13:25 update.go:1171: [WARN] [stacks] update recovery: romm was interrupted while UNDOING (started 2026-09-23T10:11:18Z) — marking it Updating and RESUMING the undo +2026/09/23 10:13:25 update.go:1224: [INFO] [stacks] update romm: resuming the UNDO after a controller restart +2026/09/23 10:13:25 update.go:826: [ERROR] [stacks] update romm FAILED after the new version was started: resumed after a restart during the undo +2026/09/23 10:13:25 update.go:818: [INFO] [stacks] update romm: kept 24813 bytes of the app's own log at /opt/docker/stacks/romm/hold-logs/20260923T101325Z/compose-logs.txt before stopping it (R-621) +2026/09/23 10:13:25 update.go:1106: [INFO] [stacks] update romm: phase undoing +2026/09/23 10:13:25 undo.go:357: [WARN] [stacks] update romm: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: resumed after a restart during the undo) +2026/09/23 10:14:20 undo.go:404: [INFO] [stacks] update romm: UNDONE in 55s — the previous version is running on the data from before the update (the app's health check passed) +2026/09/23 10:14:33 update.go:487: [INFO] [stacks] update vikunja: accepted — guarded update started +2026/09/23 10:14:33 update.go:1106: [INFO] [stacks] update vikunja: phase checking +2026/09/23 10:14:33 update.go:632: [INFO] [stacks] update vikunja: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T10:04:52Z (10m0s old, limit 24h0m0s) +2026/09/23 10:14:33 update.go:1106: [INFO] [stacks] update vikunja: phase safety-dump +2026/09/23 10:14:33 update.go:664: [INFO] [stacks] update vikunja: safety dump done (0 file(s)) [] +2026/09/23 10:14:34 undo.go:252: [INFO] [stacks] update vikunja: the undo copy will hold 2 named volume(s), 2.6 MiB +2026/09/23 10:14:34 update.go:1106: [INFO] [stacks] update vikunja: phase pinning +2026/09/23 10:14:34 pin.go:365: [INFO] [stacks] update vikunja: pin advanced to the catalog's current definition (vikunja=vikunja/vikunja:2.6.0) +2026/09/23 10:14:34 update.go:1106: [INFO] [stacks] update vikunja: phase pulling +2026/09/23 10:14:35 update.go:1106: [INFO] [stacks] update vikunja: phase copying +2026/09/23 10:14:36 update.go:1106: [INFO] [stacks] update vikunja: phase copying +2026/09/23 10:14:36 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_data → vikunja_vikunja_data.pre-update-20260923T101436Z in 469ms +2026/09/23 10:14:36 update.go:1106: [INFO] [stacks] update vikunja: phase copying +2026/09/23 10:14:37 undo.go:279: [INFO] [stacks] update vikunja: copied vikunja_vikunja_db → vikunja_vikunja_db.pre-update-20260923T101436Z in 444ms +2026/09/23 10:14:37 update.go:1106: [INFO] [stacks] update vikunja: phase starting +2026/09/23 10:14:37 update.go:1106: [INFO] [stacks] update vikunja: phase verifying +2026/09/23 10:16:07 update.go:826: [ERROR] [stacks] update vikunja FAILED after the new version was started: not healthy: not healthy within 1m30s (last: state unhealthy) +2026/09/23 10:16:08 update.go:818: [INFO] [stacks] update vikunja: kept 1299 bytes of the app's own log at /opt/docker/stacks/vikunja/hold-logs/20260923T101607Z/compose-logs.txt before stopping it (R-621) +2026/09/23 10:16:08 update.go:1106: [INFO] [stacks] update vikunja: phase undoing +2026/09/23 10:16:08 undo.go:357: [WARN] [stacks] update vikunja: UNDO — putting back the previous version and its 2 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: state unhealthy)) +2026/09/23 10:16:08 undo.go:363: [ERROR] [stacks] update vikunja: the undo copy vikunja_vikunja_data.pre-update-20260923T101436Z has no finished-marker — it is cut off or missing; NOTHING is put back +2026/09/23 10:16:08 update.go:838: [ERROR] [stacks] update vikunja: the UNDO failed too (untouched) — HOLDING the app; the undo copies are kept: [{vikunja_vikunja_data vikunja_vikunja_data.pre-update-20260923T101436Z} {vikunja_vikunja_db vikunja_vikunja_db.pre-update-20260923T101436Z}] +2026/09/23 10:16:08 update_guard.go:536: [WARN] [backup] vikunja is HELD STOPPED after a failed update (restore point: tier 1 "saját meghajtó", 2026-09-23T10:04:52Z; holds: "a beállításokat, az adatbázist és az adatköteteket tartalmazza"; undo: "untouched") +2026/09/23 10:16:42 update.go:487: [INFO] [stacks] update docmost: accepted — guarded update started +2026/09/23 10:16:42 update.go:1106: [INFO] [stacks] update docmost: phase checking +2026/09/23 10:16:42 update.go:632: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-23T10:03:45Z (13m0s old, limit 24h0m0s) +2026/09/23 10:16:42 update.go:1106: [INFO] [stacks] update docmost: phase safety-dump +2026/09/23 10:16:43 update.go:664: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260923T101642Z-docmost-postgres.sql] +2026/09/23 10:16:44 undo.go:252: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 64.8 MiB +2026/09/23 10:16:44 update.go:1106: [INFO] [stacks] update docmost: phase pinning +2026/09/23 10:16:44 pin.go:365: [INFO] [stacks] update docmost: pin advanced to the catalog's current definition (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:16-alpine, docmost-redis=redis:7-alpine) +2026/09/23 10:16:44 update.go:1106: [INFO] [stacks] update docmost: phase pulling +2026/09/23 10:16:45 update.go:1106: [INFO] [stacks] update docmost: phase copying +2026/09/23 10:16:46 update.go:1106: [INFO] [stacks] update docmost: phase copying +2026/09/23 10:16:46 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260923T101646Z in 673ms +2026/09/23 10:16:46 update.go:1106: [INFO] [stacks] update docmost: phase copying +2026/09/23 10:16:47 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260923T101646Z in 409ms +2026/09/23 10:16:47 update.go:1106: [INFO] [stacks] update docmost: phase copying +2026/09/23 10:16:47 undo.go:279: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260923T101646Z in 415ms +2026/09/23 10:16:47 update.go:1106: [INFO] [stacks] update docmost: phase starting +2026/09/23 10:16:58 update.go:1106: [INFO] [stacks] update docmost: phase verifying +2026/09/23 10:17:19 update.go:785: [INFO] [stacks] update docmost: healthy after 20s (the app's health check passed) +2026/09/23 10:17:19 update.go:793: [INFO] [stacks] update docmost: DONE in 37s diff --git a/documentation/audits/undo-live-2026-09-23/95-teardown.txt b/documentation/audits/undo-live-2026-09-23/95-teardown.txt new file mode 100644 index 00000000..84974938 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/95-teardown.txt @@ -0,0 +1,59 @@ +12:18:27 undo copies before removal: vikunja_vikunja_data.pre-update-20260923T101436Z vikunja_vikunja_db.pre-update-20260923T101436Z +12:18:28 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'} +12:19:00 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ +12:19:07 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' +12:19:10 [X] stop -> 200 {'ok': True, 'message': 'Stack romm stop completed'} +12:19:15 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/romm tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +12:19:15 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +12:19:42 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'romm', 'volumes_removed': ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data'], 'hdd_paths_removed': [], 'hdd_pat +12:19:49 [X] after remove: deployed=False leftovers='/opt/docker/stacks/romm' +12:19:49 [X] stop -> 200 {'ok': True, 'message': 'Stack vikunja stop completed'} +12:20:21 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vikunja', 'volumes_removed': ['vikunja_vikunja_data', 'vikunja_vikunja_db'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [ +12:20:28 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vikunja' +12:20:30 undo copies after removal (must be none): 0 + +12:20:45 docmost: containers=0 volumes=0 +romm: containers=0 volumes=0 +vikunja: containers=0 volumes=0 +romm drive folder removed +removed docmost/docmost:0.95.0 +removed docmost/docmost:0.96.0 +removed rommapp/romm:5.0.0 +removed rommapp/romm:5.3.0 +removed vikunja/vikunja:2.3.0 +removed vikunja/vikunja:2.6.0 +removed nextcloud:34.0.1-apache +removed gitea.dooplex.hu/admin/felhom-controller:0.263.0 +removed gitea.dooplex.hu/admin/felhom-controller:0.263.1 +controller.yaml.pre-lock-probe +dload.out +pre-docmost-compose.yml +pre-romm-compose.yml +pre-vikunja-compose.yml + +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: + username: "" +hub: +0 + +sync 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'me +cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462) +https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git +controller.yaml == saved pre-bakeoff copy +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.263.2 Up 48 seconds (healthy) +filebrowser gtstef/filebrowser:1.3.3-stable Up 8 minutes (healthy) +gokapi f0rc3/gokapi:v1.9.6 Restarting (1) 25 seconds ago +paperless-postgres postgres:16-alpine Up 8 minutes (healthy) +paperless-redis redis:7-alpine Up 8 minutes (healthy) +paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 8 minutes (healthy) +privatebin privatebin/pdo:2.0.6 Up 8 minutes (healthy) +traefik traefik:v3.6.7 Up 8 minutes + +cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main + +controller.yaml.pre-lock-probe +image lines identical diff --git a/documentation/audits/undo-live-2026-09-23/README.md b/documentation/audits/undo-live-2026-09-23/README.md new file mode 100644 index 00000000..213dd6a9 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/README.md @@ -0,0 +1,52 @@ +# The undo, live on 9202 — controller v0.263.0 → v0.263.2, 2026-09-23 + +Venue: scratch guest **9202** (demo-hp), drill catalog (`admin/app-catalog-drill`, reset to live +`cfcfe5278428` before and after), `update.health_timeout: 90s`. **Method: endpoint-level** — every +act is the endpoint the UI invokes (deploy, backup, sync, rescan, update, start, remove, the language +switch) and every page is fetched as HTML (`/apps/?lang=hu|en`); no browser. The product performs +the undo; nothing here does. Each edge is a real migrating image move with, in the DRILL template only, +a health probe on a port the app does not answer. **Seed A** before the backup, **seed B** after it, +**seed C** seconds before the press — C exists only in the undo's own last-second copy. + +## Not done, or changed + +- **Three releases, not one.** v0.263.0 was built, deployed to 9202 only, and its first two live + proofs FAILED honestly (HELD, data put back, old version judged "did not start"). Two defects the unit + tests could not see; each fixed, red-proofed and released the same session: **v0.263.1** (the undo's + probe was gated on `running` while the current probe held the app `unhealthy`) and **v0.263.2** (the + "old" `.felhom.yml` saved at update time was already the new one — it flows in on every catalog sync). + Only 9202 ever ran 0.263.0/0.263.1; both images were removed from it by name. **No floor was raised.** +- **The mail and the operator event** of decision 15 are not built (`09` §6.4 part 2). +- **The hold sentence body stays Hungarian on an English box** (R-606); only the new prefix follows + the box language. + +## Results (all on v0.263.2 unless stated) + +| proof | result | evidence | +|---|---|---| +| docmost (PostgreSQL) undone by the product | failed probe at +107 s → `undoing` → **`undone` at +137.6 s**; A, B, C read back; `42 tables, 48 ledger` before and after; 0 copies left | `40-undo-docmost.txt` | +| romm (MariaDB) | `undone` at +163.9 s (undo 52 s); A, B, C; `24 tables, alembic 0095` before and after | `40-undo-romm.txt` | +| vikunja (SQLite in a volume) | `undone` at +97.1 s (undo 3 s); A, B, C; `36 tables, ledger 117` before and after | `40-undo-vikunja.txt` | +| the page line, both languages | hu: *„A(z) docmost frissítése 2026-09-23 12:05-kor nem sikerült. A doboz automatikusan visszaállította az előző változatot és az adatokat — semmi nem veszett el."* · en: *"The update of docmost at 2026-09-23 12:05 did not succeed. The box put back the previous version and its data automatically — nothing was lost."* (same for romm, vikunja) | `40-undo-*.txt` | +| a cut-off copy (vikunja: the finished-marker taken from one copy during `verifying`) | `undoing` → **`failed` in 0.6 s, nothing poured back**; hold (hu box): *„A frissítés nem sikerült, és az automatikus visszaállítás sem. Az adatok az új változat által hagyott állapotban vannak. A(z) vikunja frissítése …"*; en box: *"The update did not succeed, and the automatic undo did not either. The data is as the new version left it. A(z) vikunja …"*; both copies kept | `60-cutoff-vikunja.txt` | +| a power cut during the undo (romm: `pct stop 9202` the moment the phase read `undoing`) | after boot: *update recovery: romm was interrupted while UNDOING … RESUMING the undo* → **`undone`**; A, B, C; ledger equal; 0 copies left | `50-powercut-romm.txt` | +| a person presses Update after an undo (docmost; the drill catalog then fixed the probe) | `done` at +36.8 s on 0.96.0; A, B, C; the undone line gone from both pages; `last_update_undone` gone from `app.yaml` | `70-manual-press-docmost.txt` | +| a failed undo holds honestly (v0.263.0, docmost and romm — the defect above) | HOLD *„… Az adatok a frissítés előtti állapotba kerültek vissza, de az előző változat nem indult el."* — true: the data had been put back; the old version was never probed | `10-docmost-undo.txt`, `20-romm-undo.txt` | +| removal deletes kept copies (vikunja, held with 2 copies) | 2 → **0** after `POST /api/stacks/vikunja/remove` | `95-teardown.txt` | +| R-642, the Start answer | `Stack romm start requested — state now: running` | `80-r642-start-answer.txt` | + +Extra downtime of the copy (the app is stopped there anyway to be recreated): the `copying` phase +lasted 2.1 s (docmost), 6.8 s (romm, incl. a 3.8 s stop), 1.6 s (vikunja). + +## Teardown — three layers + +- **machine (9202):** docmost, romm, vikunja removed through the product (no containers, no volumes, + no undo copies); romm's drive folder removed by name (the product kept it, R-442 as in every drill); + test images removed by name (docmost 0.95.0/0.96.0, romm 5.0.0/5.3.0, vikunja 2.3.0/2.6.0, nextcloud + 34.0.1-apache, controller 0.263.0/0.263.1); `controller.yaml` restored from the saved copy and read + back identical; catalog cache re-cloned from the live repo (`cfcfe52`). **9202 stays on controller + 0.263.2** (self-update off; the fleet floor is untouched at 0.262.1). Apps afterwards: the same three + as at the start of the day (gokapi still crash-looping, R-644). +- **host (demo-hp):** nothing provisioned; the guest was stopped and started once for the power cut. +- **hub:** nothing touched. +- **drill repo:** reset to live `main`, image lines identical, `has_actions: false`. diff --git a/documentation/audits/undo-live-2026-09-23/bakeoff.py b/documentation/audits/undo-live-2026-09-23/bakeoff.py new file mode 100644 index 00000000..7bd759f9 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/bakeoff.py @@ -0,0 +1,266 @@ +#!/usr/bin/env python3 +"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on +the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202, +drill catalog. + + F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure. + D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and + recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads. + +EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync, +rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is +lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own +Start supplies the app's secrets (they are never decrypted here). + +Usage: python3 bakeoff.py stages: prep fcopy break update undoF undoD cutoff state +""" +import json, os, sys, time +sys.path.insert(0, ".") +import walk as w +from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \ + romm_verify_b, vik_seed_b, vik_verify_b + +EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999), + "romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999), + "vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)} +UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost", + "vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja", + "romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"} +SUFFIX = ".pre-undo" + + +def seed_b(app, sub, A): + return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A) + + +def verify_b(app, sub, A, B): + r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B) + return all(r.values()) if isinstance(r, dict) else r + + +def db_state(app): + """The table count and the migration ledger, asked of the engine (or the SQLite file, read-only).""" + if app == "docmost": + return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; " + "docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip() + if app == "romm": + q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '" + base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45") + alembic = q("select version_num from alembic_version") + return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip() + return w.guest("""python3 - <<'PY' +import sqlite3 +c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True) +t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0] +m=c.execute("select count(*), max(id) from migration").fetchone() +print(f"tables={t} ledger={m[0]} newest={m[1]}") +PY""").strip() + + +def volumes(app): + return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split() + if not v.endswith(SUFFIX)] + + +def containers(app): + return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split() + + +def stage_prep(app): + s = {"app": app, "sub": SUB[app]} + w.say(f"=== prep {app}: deploy at the drill FROM pin") + w.say("deployed:", w.deploy(app, s["sub"])) + s["seedA"] = FX[app].seed(w, s["sub"], w.say) + w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say)) + w.backup_now(app) + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60) + s["seedB"] = seed_b(app, s["sub"], s["seedA"]) + w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"])) + s["volumes"] = volumes(app) + w.say("named volumes:", s["volumes"]) + save(app, s) + + +def stage_fcopy(app): + """F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and + separately a tar, for the rate), start. Measures the EXTRA downtime and the disk.""" + s = load(app) + s["state_before_update"] = db_state(app) + w.say("db state before the update:", s["state_before_update"]) + vols = s["volumes"] + w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml")) + script = f""" +set -u +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}') +t0=$(date +%s.%N) +docker stop $cs >/dev/null +t1=$(date +%s.%N) +for v in {' '.join(vols)}; do + docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1 + docker volume create "$v{SUFFIX}" >/dev/null + a=$(date +%s.%N) + docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v" + b=$(date +%s.%N) + bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1) + files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l') + cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1) + cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l') + mkdir -p /root/bk; c=$(date +%s.%N) + docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N) + tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar + python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))" +done +t2=$(date +%s.%N) +docker start $cs >/dev/null +t3=$(date +%s.%N) +python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))" +df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}' +""" + out = w.guest(script, timeout=1800) + w.say(out) + t0 = time.time() + w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + w.say(f"front door answering again {round(time.time()-t0,1)}s after the start") + s["fcopy"] = out + s["fcopy_at"] = ts() + save(app, s) + + +def stage_break(app): + s = load(app) + frm, to, p0, p1 = EDGE[app] + s["break_commit"] = break_edge(app, frm, to, p0, p1) + w.say("badge caught up after", w.sync_rescan(app, to), "s") + save(app, s) + + +def stage_update(app): + s = load(app) + r = w.press_update(app, poll=1.0) + w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False)) + s.setdefault("updates", []).append(r) + save(app, s) + + +LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold' +docker restart felhom-controller >/dev/null; sleep 15""" + +# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre--compose.yml, taken +# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s +# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new +# version again. That is R-639 seen live, and why the product undo keeps its own copies. +PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml +grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }} +cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml +sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4""" + + +def lift_and_pinback(app): + frm, to, _, _ = EDGE[app] + t = time.time() + w.say(w.guest(LIFT.format(app=app))) + w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to))) + w.login() + return round(time.time() - t, 1) + + +def stage_undoF(app): + s = load(app) + w.say(f"=== undo by F: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + vols = s["volumes"] + script = f""" +t0=$(date +%s.%N) +cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null +for v in {' '.join(vols)}; do + docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v" +done +python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))" +""" + w.say(w.guest(script, timeout=1800)) + t0 = time.time() + c, d = w.ctl("POST", f"/api/stacks/{app}/start") + w.say("product start ->", c) + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_undoD(app): + """D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy.""" + s = load(app) + w.say(f"=== undo by D: {app}") + w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)") + if app == "vikunja": + w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it " + "is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.") + return + c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse + time.sleep(8) + if app == "docmost": + load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1""" + marker = "-- PostgreSQL database dump complete" + else: + load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1""" + marker = "-- Dump completed" + script = f""" +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B" +docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null +t0=$(date +%s.%N) +if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi +{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out +python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))" +docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null +""" + w.say(w.guest(script, timeout=900)) + t0 = time.time() + up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2) + th = round(time.time() - t0, 1) + A = FX[app].verify(w, s["sub"], s["seedA"], w.say) + B = verify_b(app, s["sub"], s["seedA"], s["seedB"]) + st = db_state(app) + w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})") + s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st} + save(app, s) + + +def stage_cutoff(app): + """The cut-off copy, both methods, DETECTED before anything is loaded or swapped. + F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker + the build would write only after cp exits 0. D: the safety dump cut in half — the marker check.""" + s = load(app) + big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0)) + script = f""" +v={big} +docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null +cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null +# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured +# on docmost 2026-09-23, rc 137 with a complete copy and a marker). +docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null +sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1 +echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)" +echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap" +docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null +D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1) +if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql + echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load" + rm -f /root/cut.sql +else echo "D: no safety dump exists for this app (R-641)"; fi +""" + w.say(f"=== cut-off copy: {app} (largest volume {big})") + out = w.guest(script, timeout=600) + w.say(out) + s["cutoff"] = out + save(app, s) + + +if __name__ == "__main__": + w.login() + {"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update, + "undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff, + "state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/undo-live-2026-09-23/fixtures.py b/documentation/audits/undo-live-2026-09-23/fixtures.py new file mode 100644 index 00000000..3a681e96 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/fixtures.py @@ -0,0 +1,1008 @@ +#!/usr/bin/env python3 +"""Box-side seed/verify fixtures for walk.py, guest 9202. + +THE ONE RULE (R-156), carried verbatim from `app-catalog-felhom.eu/scripts/upgrade_fixtures.py`: +*nothing is ever seeded into a volume by hand.* Every seed here goes in through the app's OWN +interface — its HTTP API through the household's real front door (traefik, `Host: .`), +or its own CLI running inside its own container. A raw SQL INSERT or a planted file is never used. + +If an app has no non-browser route, its fixture returns None and the edge is recorded +`inconclusive — no non-browser seed route`, WITH WHAT WAS TRIED. That is a result, not a gap. + +Each fixture: + seed(w, sub, say) -> an opaque token, or None + verify(w, sub, tok, say) -> True / False +verify() must ask the APP, never the filesystem: a migration is supposed to rewrite files. +Where a fixture can prove itself (a negative control that must read as absent) it does so on EVERY +call, so a readback that has broken into always saying "found" fails instead of passing everything. +""" +import base64, json, re, secrets, time + + +def _gx(w, container, *cmd, timeout=240): + """Run a command inside the app's OWN container on 9202 (its own CLI, not our SQL).""" + import shlex + line = " ".join(shlex.quote(c) for c in cmd) + return w.guest(f"docker exec {container} {line} 2>&1", timeout=timeout) + + +# ============================================================================================= +class PrivateBin: + """PrivateBin's own JSON API. A paste is a POST and reading it back is a GET — an + application-level round trip. File-backed, no database: this single seed IS the file half.""" + sub = "paste" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200",)): + return None + marker = "upg-" + secrets.token_hex(8) + ct = base64.b64encode(marker.encode()).decode() + body = json.dumps({ + "v": 2, + "adata": [[base64.b64encode(secrets.token_bytes(16)).decode(), + base64.b64encode(secrets.token_bytes(8)).decode(), + 100000, 256, 128, "aes", "gcm", "none"], "plaintext", 0, 0], + "ct": ct, "meta": {"expire": "never"}}) + rc, code, out = w.app_curl(sub, "/", "-H", "X-Requested-With: JSONHttpRequest", + "-H", "Content-Type: application/json", + data=body, method="POST") + try: + j = json.loads(out) + except Exception: + say(f" privatebin: POST returned non-JSON (http {code}): {out[:200]}") + return None + if j.get("status") != 0 or not j.get("id"): + say(f" privatebin: POST refused: {out[:250]}") + return None + say(f" privatebin: seeded paste id={j['id']}") + return {"id": j["id"], "marker": ct} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200",), tries=36): + return False + # negative control, every call: a paste id that cannot exist must NOT read back + rc, code, out = w.app_curl(sub, "/?pasteid=" + secrets.token_hex(8), + "-H", "X-Requested-With: JSONHttpRequest") + if t["marker"] in out: + say(" privatebin: READBACK UNUSABLE — a paste id that cannot exist returned the marker") + return False + rc, code, out = w.app_curl(sub, "/?pasteid=" + t["id"], + "-H", "X-Requested-With: JSONHttpRequest") + got = code == "200" and t["marker"] in out + say(f" privatebin: readback http={code} marker_present={got}") + return got + + +# ============================================================================================= +class Docmost: + """Docmost's own REST API: create the first workspace+user, then prove the account survives by + asking the app to AUTHENTICATE it. Login is version-stable across the API churn.""" + sub = "docs" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302", "404")): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + body = json.dumps({"workspaceName": "drill", "name": "drill", "email": email, "password": pw}) + rc, code, out = w.app_curl(sub, "/api/auth/setup", "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" docmost: /api/auth/setup http={code} rc={rc}") + if code not in ("200", "201"): + say(f" docmost: setup refused: {out[:250]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "404"), tries=36): + return False + # negative control: a password that was never set must NOT authenticate + bad = json.dumps({"email": t["email"], "password": "definitely-" + secrets.token_hex(8)}) + rc, code, _ = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" docmost: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"email": t["email"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" docmost: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" docmost: login body {out[:200]}") + return ok + + +# ============================================================================================= +class BookStack: + """BookStack mints no API token without a browser, so BOTH halves go through `php artisan` — + BookStack's OWN CLI, inside its own container, against its own User model. + + The exit code carries no information here (`bookstack:reset-mfa` exits 1 for a user it FOUND + and for one it did not), so the discriminator is the OUTPUT: the positive sentence required and + the not-found sentence required absent. The negative control runs on every verify. + + LIMITATION (R-460): this seeds the DATABASE half only. The FILE half needs the API token the + app cannot mint headlessly — so a bookstack edge is at best HALF-proven here. + """ + sub = "wiki" + + def _artisan(self, w, *args): + for path in ("/app/www/artisan", "/var/www/html/artisan"): + out = _gx(w, "bookstack", "php", path, *args) + if "Could not open input file" not in out: + return " ".join(out.split()) + return " ".join(out.split()) + + def _lookup(self, w, email): + out = self._artisan(w, "bookstack:reset-mfa", f"--email={email}") + found = f"Email: {email}" in out + missing = "could not be found" in out + if found == missing: + return None, out + return found, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + out = self._artisan(w, "bookstack:create-admin", f"--email={email}", + f"--name=drill-{secrets.token_hex(3)}", f"--password={pw}") + say(f" bookstack: artisan create-admin :: {out[:140]}") + if "successfully created" not in out: + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" bookstack: the app never served /login") + return False + absent, _ = self._lookup(w, f"nobody-{secrets.token_hex(6)}@gate.invalid") + if absent is not False: + say(f" bookstack: READBACK UNUSABLE — an email that cannot exist did not read absent ({absent})") + return False + found, out = self._lookup(w, t["email"]) + say(f" bookstack: readback of the seeded account found={found} :: {out[:140]}") + return found is True + + +# ============================================================================================= +class Gitea: + """Gitea's own admin CLI creates the first user; its own REST API (basic auth) then creates a + repository and reads it back. Both are the app's own interfaces.""" + sub = "git" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = _gx(w, "gitea", "su", "git", "-c", + f"gitea admin user create --username {user} --password {pw} " + f"--email {user}@gate.invalid --admin --must-change-password=false") + say(f" gitea: admin user create :: {' '.join(out.split())[:140]}") + if "has been successfully created" not in out and "successfully created" not in out: + return None + repo = "drillrepo" + secrets.token_hex(3) + rc, code, body = w.app_curl(sub, "/api/v1/user/repos", "-u", f"{user}:{pw}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": repo, "private": True}), method="POST") + say(f" gitea: create repo http={code}") + if code not in ("201", "200"): + say(f" gitea: repo refused {body[:200]}") + return None + return {"user": user, "pw": pw, "repo": repo} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + rc, code, _ = w.app_curl(sub, f"/api/v1/repos/{t['user']}/nope{secrets.token_hex(4)}", + "-u", f"{t['user']}:{t['pw']}") + if code == "200": + say(" gitea: READBACK UNUSABLE — a repo that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/repos/{t['user']}/{t['repo']}", + "-u", f"{t['user']}:{t['pw']}") + ok = code == "200" and t["repo"] in body + say(f" gitea: readback of the seeded repo http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Navidrome: + """Navidrome's own REST API: create the first admin through /auth/createAdmin, then prove the + account survives by logging in through the same door.""" + sub = "music" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302")): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, out = w.app_curl(sub, "/auth/createAdmin", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), method="POST") + say(f" navidrome: createAdmin http={code}") + if code not in ("200", "201"): + say(f" navidrome: refused {out[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=36): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code in ("200", "201"): + say(" navidrome: READBACK UNUSABLE — a wrong password authenticated") + return False + body = json.dumps({"username": t["user"], "password": t["pw"]}) + rc, code, out = w.app_curl(sub, "/auth/login", "-H", "Content-Type: application/json", + data=body, method="POST") + ok = code in ("200", "201") + say(f" navidrome: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vaultwarden: + """Vaultwarden's own account API: register an account, then prove it survives by asking the app + to issue a token for it (its own login endpoint, the household's own route).""" + sub = "vault" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/alive", want=("200",)): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + # Vaultwarden stores an already-hashed master key; the value is opaque to the server. + key = base64.b64encode(secrets.token_bytes(32)).decode() + body = json.dumps({"email": email, "name": "drill", "masterPasswordHash": key, + "key": "0." + base64.b64encode(secrets.token_bytes(48)).decode(), + "kdf": 0, "kdfIterations": 600000}) + rc, code, out = w.app_curl(sub, "/api/accounts/register", + "-H", "Content-Type: application/json", + data=body, method="POST") + say(f" vaultwarden: register http={code}") + if code not in ("200", "204"): + say(f" vaultwarden: refused {out[:250]}") + return None + return {"email": email, "key": key} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/alive", want=("200",), tries=36): + return False + def login(pwhash): + return w.app_curl(sub, "/identity/connect/token", + "-H", "Content-Type: application/x-www-form-urlencoded", + data=("grant_type=password&scope=api%20offline_access" + f"&client_id=web&deviceType=9&deviceIdentifier=drill" + f"&deviceName=drill&username={t['email']}&password={pwhash}"), + method="POST") + rc, code, _ = login(base64.b64encode(secrets.token_bytes(32)).decode()) + if code == "200": + say(" vaultwarden: READBACK UNUSABLE — a wrong master key authenticated") + return False + rc, code, out = login(t["key"].replace("+", "%2B").replace("=", "%3D").replace("/", "%2F")) + ok = code == "200" and "access_token" in out + say(f" vaultwarden: token for the seeded account http={code} ok={ok}") + if not ok: + say(f" vaultwarden: body {out[:200]}") + return ok + + +# ============================================================================================= +class Django: + """A Django app's OWN management CLI, inside its own container, against its own User model. + + Same category as BookStack's `php artisan`: the app's own code and its own ORM, never a raw SQL + INSERT and never a planted file (R-156). `createsuperuser --noinput` is Django's own documented + non-interactive route, and the readback asks the SAME ORM whether the account exists. + + THE FIXTURE PROVES ITSELF ON EVERY CALL: each verify() also asks for a username that cannot + exist and requires the answer False. A readback that has broken into always saying True + therefore fails instead of passing everything. + + LIMITATION, recorded rather than papered over: this seeds the DATABASE half only. An app whose + data is also FILES (adventurelog's images) has a file half this fixture does not touch. + """ + + def __init__(self, container, sub, ready_path="/", ready=("200", "302", "301", "404"), + python="python", workdir=None): + # `python` and `workdir` are per-app because the image decides them: adventurelog's + # interpreter is on PATH, tandoor ships a VENV and the bare `python` cannot import Django + # at all ("Couldn't import Django. Are you sure it's installed…"). Measured, not guessed. + self.container = container + self.sub = sub + self.ready_path = ready_path + self.ready = ready + self.python = python + self.workdir = workdir + + def _wd(self): + return f"-w {self.workdir} " if self.workdir else "" + + def _manage(self, w, code): + # -c is passed to `manage.py shell`; the app's own shell, its own ORM. + return w.guest( + f"docker exec {self._wd()}{self.container} {self.python} manage.py shell " + f"-c {json.dumps(code)} 2>&1", timeout=300) + + def _exists(self, w, username): + # ONE LINE, semicolon-separated. A `\n` inside a double-quoted shell argument reaches + # python as a literal backslash-n and is a SyntaxError — which is exactly how the first + # adventurelog run read as `inconclusive`. The fixture refused to guess, which is right, + # but the instrument was the thing that was broken. + out = self._manage(w, ( + "from django.contrib.auth import get_user_model; " + f"print('DRILL_ANSWER=' + str(get_user_model().objects.filter(username={username!r}).exists()))" + )) + m = re.search(r"DRILL_ANSWER=(True|False)", out) + return (m.group(1) == "True") if m else None, " ".join(out.split())[-300:] + + def seed(self, w, sub, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -e DJANGO_SUPERUSER_PASSWORD={pw} {self._wd()}{self.container} " + f"{self.python} manage.py createsuperuser --noinput " + f"--username {user} --email {user}@gate.invalid 2>&1", timeout=300) + say(f" {self.container}: createsuperuser :: {' '.join(out.split())[:160]}") + got, detail = self._exists(w, user) + if got is not True: + say(f" {self.container}: the account did not appear in the app's own ORM :: {detail[:200]}") + return None + say(f" {self.container}: seeded superuser {user}") + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, self.ready_path, want=self.ready, tries=90): + say(f" {self.container}: the app never served {self.ready_path}") + return False + absent, detail = self._exists(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" {self.container}: READBACK UNUSABLE — a username that cannot exist did not " + f"read as absent ({absent}) :: {detail[:200]}") + return False + found, detail = self._exists(w, t["user"]) + say(f" {self.container}: readback of the seeded account found={found}") + if found is not True: + say(f" {self.container}: :: {detail[:250]}") + return found is True + + +# ============================================================================================= +class Nextcloud: + """Nextcloud's OWN admin CLI, `occ`, inside its own container: its own code, its own user + backend. Not a SQL INSERT and not a planted file (R-156). + + `occ user:info` is the readback, and it PROVES ITSELF on every call: a uid that cannot exist + must answer "user not found". A readback that has broken into always succeeding therefore + fails instead of passing everything. + + This is the app chosen for the MariaDB engine-major edge (`09` §3 decision 5, R-469 lifted): + the app image does NOT move, only the `mariadb:` sidecar, so the edge carries exactly one + migration and a failure is readable. + """ + sub = "cloud" + + def _occ(self, w, *args, timeout=420): + import shlex + line = " ".join(shlex.quote(a) for a in args) + return w.guest(f"docker exec -u www-data nextcloud php occ {line} 2>&1", timeout=timeout) + + def _info(self, w, uid): + out = self._occ(w, "user:info", uid) + flat = " ".join(out.split()) + if "user not found" in flat.lower() or "could not be found" in flat.lower(): + return False, flat + if f"user_id: {uid}" in flat or f"- user_id: {uid}" in flat or f"user_id: {uid}" in out: + return True, flat + return None, flat + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + return None + uid = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + out = w.guest( + f"docker exec -u www-data -e OC_PASS={pw} nextcloud php occ user:add " + f"--password-from-env --display-name={uid} {uid} 2>&1", timeout=420) + say(f" nextcloud: occ user:add :: {' '.join(out.split())[:160]}") + got, flat = self._info(w, uid) + if got is not True: + say(f" nextcloud: the account did not appear via occ user:info :: {flat[:220]}") + return None + say(f" nextcloud: seeded user {uid}") + return {"uid": uid, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status.php", want=("200",), tries=120): + say(" nextcloud: the app never served /status.php") + return False + absent, flat = self._info(w, "nobody" + secrets.token_hex(6)) + if absent is not False: + say(f" nextcloud: READBACK UNUSABLE — a uid that cannot exist did not read absent " + f"({absent}) :: {flat[:200]}") + return False + found, flat = self._info(w, t["uid"]) + say(f" nextcloud: readback of the seeded user found={found}") + if found is not True: + say(f" nextcloud: :: {flat[:250]}") + return found is True + + +# ============================================================================================= +class Grafana: + """Grafana's own HTTP API as the admin the DEPLOY created. The password is the one the + controller showed the household — read from the app's own `app.yaml`, not invented — and the + data (a folder) goes in and comes back through the app's own REST API.""" + sub = "grafana" + + def _auth(self, w, name="grafana"): + # app.yaml stores this ENCRYPTED (`ENC:…`), so it cannot be read back off the box — which + # is correct, and is why the harness uses the value IT generated for the deploy. + pw = (w.GENERATED.get(name) or {}).get("GF_SECURITY_ADMIN_PASSWORD") or "admin" + return f"admin:{pw}" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return None + au = self._auth(w) + title = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/folders", "-u", au, + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="POST") + say(f" grafana: create folder http={code}") + if code not in ("200", "201"): + say(f" grafana: refused {body[:220]}") + return None + try: + uid = json.loads(body)["uid"] + except Exception: + say(f" grafana: no uid in {body[:200]}") + return None + return {"uid": uid, "title": title} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + return False + au = self._auth(w) + rc, code, _ = w.app_curl(sub, "/api/folders/nope" + secrets.token_hex(5), "-u", au) + if code == "200": + say(" grafana: READBACK UNUSABLE — a folder uid that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/folders/{t['uid']}", "-u", au) + ok = code == "200" and t["title"] in body + say(f" grafana: readback of the seeded folder http={code} ok={ok}") + return ok + + +# ============================================================================================= +class AudiobookShelf: + """audiobookshelf's own /init endpoint creates the first root account; its own /login proves + the account survived. Both are the app's own API.""" + sub = "audiobooks" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/init", "-H", "Content-Type: application/json", + data=json.dumps({"newRoot": {"username": user, "password": pw}}), + method="POST") + say(f" audiobookshelf: /init http={code}") + if code not in ("200", "204"): + say(f" audiobookshelf: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/status", want=("200",), tries=72): + return False + bad = json.dumps({"username": t["user"], "password": "wrong-" + secrets.token_hex(6)}) + rc, code, _ = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=bad, method="POST") + if code == "200": + say(" audiobookshelf: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": t["pw"]}), + method="POST") + ok = code == "200" and t["user"] in body + say(f" audiobookshelf: login as the seeded root http={code} ok={ok}") + return ok + + +# ============================================================================================= +class ActualBudget: + """Actual's own bootstrap API sets the server password; its own login proves it survived.""" + sub = "budget" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/account/bootstrap", + "-H", "Content-Type: application/json", + data=json.dumps({"password": pw}), method="POST") + say(f" actualbudget: /account/bootstrap http={code} :: {body[:140]}") + if code not in ("200", "201") or '"status":"ok"' not in body: + return None + return {"pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def login(p): + return w.app_curl(sub, "/account/login", "-H", "Content-Type: application/json", + data=json.dumps({"loginMethod": "password", "password": p}), + method="POST") + rc, code, body = login("wrong-" + secrets.token_hex(6)) + if '"status":"ok"' in body: + say(" actualbudget: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = '"status":"ok"' in body + say(f" actualbudget: login with the seeded password http={code} ok={ok}") + if not ok: + say(f" actualbudget: body {body[:200]}") + return ok + + +# ============================================================================================= +class Mealie: + """Mealie ships a documented first-run admin. We log in as it through the app's own OAuth-style + token endpoint, create a recipe through the app's own API, and read the recipe back.""" + sub = "recipes" + + def _token(self, w, sub, pw="MyPassword"): + rc, code, body = w.app_curl( + sub, "/api/auth/token", "-H", "Content-Type: application/x-www-form-urlencoded", + data=f"username=changeme%40example.com&password={pw}", method="POST") + if code != "200": + return None, f"http={code} {body[:200]}" + try: + return json.loads(body)["access_token"], "" + except Exception: + return None, body[:200] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return None + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate as the first-run admin :: {why}") + return None + name = "drill-" + secrets.token_hex(5) + rc, code, body = w.app_curl(sub, "/api/recipes", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"name": name}), method="POST") + say(f" mealie: create recipe http={code}") + if code not in ("200", "201"): + say(f" mealie: refused {body[:220]}") + return None + slug = body.strip().strip('"') + return {"slug": slug, "name": name} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/app/about", want=("200",), tries=90): + return False + tok, why = self._token(w, sub) + if not tok: + say(f" mealie: could not authenticate after the update :: {why}") + return False + rc, code, _ = w.app_curl(sub, "/api/recipes/nope" + secrets.token_hex(5), + "-H", f"Authorization: Bearer {tok}") + if code == "200": + say(" mealie: READBACK UNUSABLE — a slug that cannot exist returned 200") + return False + rc, code, body = w.app_curl(sub, f"/api/recipes/{t['slug']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["name"] in body + say(f" mealie: readback of the seeded recipe http={code} ok={ok}") + return ok + + +# ============================================================================================= +class N8n: + """n8n's own owner-setup API creates the first account; its own login proves it survived.""" + sub = "auto" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill" + secrets.token_hex(8) + "1" + rc, code, body = w.app_curl(sub, "/rest/owner/setup", "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "firstName": "drill", + "lastName": "drill", "password": pw}), + method="POST") + say(f" n8n: /rest/owner/setup http={code}") + if code not in ("200", "201"): + say(f" n8n: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/healthz", want=("200",), tries=90): + return False + def login(p): + return w.app_curl(sub, "/rest/login", "-H", "Content-Type: application/json", + data=json.dumps({"emailOrLdapLoginId": t["email"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" n8n: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" and t["email"] in body + say(f" n8n: login as the seeded owner http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Zipline: + """Zipline's own setup/login API. Zipline 4 creates the first user through its own endpoint.""" + sub = "img" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/healthcheck", want=("200",), tries=90): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=30): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + for path in ("/api/auth/register", "/api/auth/setup"): + rc, code, body = w.app_curl(sub, path, "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw}), + method="POST") + say(f" zipline: {path} http={code} :: {body[:160]}") + if code in ("200", "201"): + return {"user": user, "pw": pw} + say(" zipline: neither register nor setup accepted a first user") + return None + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302", "307"), tries=60): + return False + def login(p): + return w.app_curl(sub, "/api/auth/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], "password": p}), + method="POST") + rc, code, _ = login("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" zipline: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = login(t["pw"]) + ok = code == "200" + say(f" zipline: login as the seeded user http={code} ok={ok}") + return ok + + +# ============================================================================================= +class Vikunja: + """Vikunja's own REST API: register a user, log in, create a project, read the project back. + Four calls, all the app's own front door.""" + sub = "tasks" + + def _token(self, w, sub, t, pw=None): + rc, code, body = w.app_curl(sub, "/api/v1/login", "-H", "Content-Type: application/json", + data=json.dumps({"username": t["user"], + "password": pw or t["pw"]}), method="POST") + if code != "200": + return None, f"http={code} {body[:160]}" + try: + return json.loads(body)["token"], "" + except Exception: + return None, body[:160] + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/v1/register", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "password": pw, + "email": f"{user}@gate.invalid"}), + method="POST") + say(f" vikunja: register http={code}") + if code not in ("200", "201"): + say(f" vikunja: refused {body[:220]}") + return None + t = {"user": user, "pw": pw} + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: could not log in after registering :: {why}") + return None + title = "drill-" + secrets.token_hex(5) + # Vikunja CREATES with PUT, not POST — a POST answers `405 Method Not Allowed`, which + # reads like a broken fixture and is really the wrong verb. Measured 2026-09-21. + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"title": title}), method="PUT") + say(f" vikunja: create project http={code}") + if code not in ("200", "201"): + say(f" vikunja: project refused {body[:220]}") + return None + t["title"] = title + t["pid"] = json.loads(body).get("id") + return t + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/v1/info", want=("200",), tries=72): + return False + bad, why = self._token(w, sub, t, pw="wrong-" + secrets.token_hex(6)) + if bad: + say(" vikunja: READBACK UNUSABLE — a wrong password authenticated") + return False + tok, why = self._token(w, sub, t) + if not tok: + say(f" vikunja: the seeded account no longer authenticates :: {why}") + return False + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{t['pid']}", + "-H", f"Authorization: Bearer {tok}") + ok = code == "200" and t["title"] in body + say(f" vikunja: readback of the seeded project http={code} ok={ok}") + return ok + + +# ============================================================================================= +class OpenGist: + """Opengist's own sign-up and sign-in FORMS. + + Two things had to be measured. Its sign-up is CSRF-protected: a bare POST answers 500 with an + HTML page, which reads like a broken app and is really a missing token — fetch the form, keep + its cookie, send its `_csrf` back. And its REST API refuses the account's own password + (`401 {"message":"Bad crendentials"}`) because it wants a token the app will not mint without a + browser. So the SEEDED DATA is the account itself and the READBACK is a real sign-in, which is + the same shape the docmost and navidrome fixtures use. + + LIMITATION, recorded rather than papered over: this is the DATABASE half. A gist's CONTENT is + not seeded, because that needs the API token above. + """ + sub = "gist" + + def _form(self, w, sub, path, jar, fields): + rc, code, html = w.app_curl(sub, path, "-b", jar, "-c", jar) + m = re.search(r'name="_csrf"[^>]*value="([^"]+)"', html or "") + if not m: + return None, f"no _csrf on {path} (http={code})" + body = "&".join([f"_csrf={m.group(1)}"] + [f"{k}={v}" for k, v in fields.items()]) + rc, code, out = w.app_curl(sub, path, "-b", jar, "-c", jar, + "-H", "Content-Type: application/x-www-form-urlencoded", + data=body, method="POST") + return code, out + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, out = self._form(w, sub, "/register", jar, {"username": user, "password": pw}) + say(f" opengist: /register (with its own _csrf) http={code}") + if code not in ("200", "302", "303"): + say(f" opengist: refused {str(out)[:200]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + # Wait for the LOGIN FORM, not for the root page. Measured 2026-09-21: immediately after a + # successful update the root answers while /login does not yet carry its `_csrf`, so the + # sign-in silently fails and the app looks like it lost the account. It had not. + if not w.wait_app(sub, "/login", want=("200",), tries=72): + say(" opengist: /login never came back after the update") + return False + for _ in range(24): + rc, code, html = w.app_curl(sub, "/login") + if code == "200" and '_csrf' in (html or ""): + break + time.sleep(5) + jar = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar, + {"username": t["user"], "password": "wrong-" + secrets.token_hex(5)}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar) + if t["user"] in (home or ""): + say(" opengist: READBACK UNUSABLE — a wrong password signed in") + return False + jar2 = f"/tmp/og-{secrets.token_hex(4)}.jar" + code, _ = self._form(w, sub, "/login", jar2, {"username": t["user"], "password": t["pw"]}) + rc, c2, home = w.app_curl(sub, "/", "-b", jar2) + ok = t["user"] in (home or "") + say(f" opengist: sign-in as the seeded account http={code} name_on_page={ok}") + return ok + + +# ============================================================================================= +class Papra: + """Papra's own e-mail sign-up and sign-in endpoints.""" + sub = "papra" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/health", want=("200",), tries=72): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + email = f"drill-{secrets.token_hex(4)}@gate.invalid" + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/auth/sign-up/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": email, "password": pw, + "name": "drill"}), method="POST") + say(f" papra: sign-up http={code}") + if code not in ("200", "201"): + say(f" papra: refused {body[:220]}") + return None + return {"email": email, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=72): + return False + def signin(p): + return w.app_curl(sub, "/api/auth/sign-in/email", + "-H", "Content-Type: application/json", + data=json.dumps({"email": t["email"], "password": p}), method="POST") + rc, code, _ = signin("wrong-" + secrets.token_hex(6)) + if code == "200": + say(" papra: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = signin(t["pw"]) + ok = code == "200" + say(f" papra: sign-in as the seeded account http={code} ok={ok}") + return ok + + +# ============================================================================================= +class HomeAssistant: + """Home Assistant's own onboarding API creates the owner account and hands back a code the + same API exchanges for a token. Both are the app's own documented non-browser route.""" + sub = "ha" + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl(sub, "/api/onboarding/users", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "name": "drill", "username": user, + "password": pw, "language": "en"}), + method="POST") + say(f" home-assistant: /api/onboarding/users http={code}") + if code not in ("200", "201"): + say(f" home-assistant: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def _login(self, w, sub, user, pw): + """The app's own login flow: start it, then answer it. A 200 with a step_id of + `mfa`/`init` means the credentials were REFUSED; only `create_entry` is a pass.""" + rc, code, body = w.app_curl(sub, "/auth/login_flow", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "handler": ["homeassistant", None], + "redirect_uri": f"https://{sub}.felhom.invalid/", + "type": "authorize"}), method="POST") + if code not in ("200", "201"): + return None, f"flow start http={code} {body[:160]}" + try: + fid = json.loads(body)["flow_id"] + except Exception: + return None, body[:160] + rc, code, body = w.app_curl(sub, f"/auth/login_flow/{fid}", + "-H", "Content-Type: application/json", + data=json.dumps({"client_id": f"https://{sub}.felhom.invalid/", + "username": user, "password": pw}), + method="POST") + try: + j = json.loads(body) + except Exception: + return None, body[:160] + return (j.get("result") if j.get("type") == "create_entry" else None), body[:200] + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/", want=("200", "302"), tries=120): + return False + bad, why = self._login(w, sub, t["user"], "wrong-" + secrets.token_hex(6)) + if bad: + say(" home-assistant: READBACK UNUSABLE — a wrong password authenticated") + return False + good, why = self._login(w, sub, t["user"], t["pw"]) + ok = bool(good) + say(f" home-assistant: login as the seeded owner ok={ok}") + if not ok: + say(f" home-assistant: {why}") + return ok + + +# ============================================================================================= +class Romm: + """RomM's own user API, driven the way RomM's own front end drives it. + + Three things had to be measured rather than guessed, and each one answered a 403 or a 422 that + looked like a different fault: RomM sets a **`romm_csrftoken` cookie** on any GET and requires + it back in an **`x-csrftoken` header** (a bare POST is `403 CSRF token verification failed`, + which reads like an auth problem); the fields go in the **JSON body**, not the query string (a + query-string POST is `422 Field required` for every field it was just given); and `email` is + required alongside username, password and role. + + On a fresh install with no admin the first `POST /api/users` is accepted unauthenticated; + afterwards it is not — which is what makes the readback (`POST /api/login` as that user) a real + authentication rather than a repeat of the seed. + + LIMITATION: this is the DATABASE half. RomM's other half is the ROM library on the drive, which + this does not populate. + """ + sub = "arcade" + + def _csrf(self, w, sub): + jar = f"/tmp/romm-{secrets.token_hex(4)}.jar" + w.app_curl(sub, "/api/heartbeat", "-c", jar) + out = w.sh(["bash", "-lc", f"grep -i csrf {jar} | awk '{{print $7}}'"]).stdout or "" + return jar, out.strip() + + def seed(self, w, sub, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + if not w.wait_app(sub, "/", want=("200", "302"), tries=30): + return None + jar, tok = self._csrf(w, sub) + if not tok: + say(" romm: no romm_csrftoken cookie was set on /api/heartbeat") + return None + user = "drill" + secrets.token_hex(3) + pw = "Drill-" + secrets.token_hex(10) + rc, code, body = w.app_curl( + sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": pw, "role": "admin"}), method="POST") + say(f" romm: POST /api/users http={code}") + if code not in ("200", "201"): + say(f" romm: refused {body[:220]}") + return None + return {"user": user, "pw": pw} + + def verify(self, w, sub, t, say): + if not w.wait_app(sub, "/api/heartbeat", want=("200",), tries=120): + return False + jar, tok = self._csrf(w, sub) + rc, code, _ = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:wrong-{secrets.token_hex(5)}", method="POST") + if code == "200": + say(" romm: READBACK UNUSABLE — a wrong password authenticated") + return False + rc, code, body = w.app_curl(sub, "/api/login", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{t['user']}:{t['pw']}", method="POST") + ok = code == "200" + say(f" romm: login as the seeded user http={code} ok={ok}") + if not ok: + say(f" romm: body {body[:200]}") + return ok + + +FIXTURES = { + "home-assistant": HomeAssistant(), + "romm": Romm(), + "vikunja": Vikunja(), + "opengist": OpenGist(), + "papra": Papra(), + "mealie": Mealie(), + "n8n": N8n(), + "zipline": Zipline(), + "grafana": Grafana(), + "audiobookshelf": AudiobookShelf(), + "actualbudget": ActualBudget(), + "nextcloud": Nextcloud(), + "adventurelog": Django("adventurelog", "travel", "/admin/login/"), + "tandoor": Django("tandoor", "recipes", "/accounts/login/", + python="/opt/recipes/venv/bin/python", workdir="/opt/recipes"), + "privatebin": PrivateBin(), + "docmost": Docmost(), + "bookstack": BookStack(), + "gitea": Gitea(), + "navidrome": Navidrome(), + "vaultwarden": Vaultwarden(), +} diff --git a/documentation/audits/undo-live-2026-09-23/live.py b/documentation/audits/undo-live-2026-09-23/live.py new file mode 100644 index 00000000..8a024100 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/live.py @@ -0,0 +1,194 @@ +#!/usr/bin/env python3 +"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes. + +Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/, +fetches the app page in both languages, and reads the seeds back through each app's own front door. +Seed C is written immediately before each Update, after every backup — so only the undo's own +last-second copy can bring it back. +""" +import json, re, sys, time, html as H +sys.path.insert(0, ".") +import walk as w +from spike import load, save, ts +import bakeoff as b + +HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"} + + +def seed_c(app, s): + return b.seed_b(app, s["sub"], s["seedA"]) + + +def page_lines(app): + out = {} + for lang in ("hu", "en"): + h = H.unescape(w.page(f"/apps/{app}?lang={lang}")) + m = re.search(r'data-update-undone="true">([^<]*)<', h) + hold = re.search(r'data-held="true">([^<]*)<', h) + out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None} + return out + + +def press(app, poll=0.5, on_phase=None): + code, d = w.ctl("POST", f"/api/stacks/{app}/update") + w.say(f" Update -> {code} {str(d)[:140]}") + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < 1500: + try: + st = w.stack(app) + except Exception as e: + st = {} + ph = st.get("update_phase") + if ph != seen and ph is not None: + seen = ph + phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label"))) + w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}") + if on_phase and on_phase(ph): + return phases, "interrupted" + if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2: + break + time.sleep(poll) + st = w.stack(app) + w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}") + return phases, st + + +def readback(app, s, with_c=True): + sub = s["sub"] + w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2) + A = b.FX[app].verify(w, sub, s["seedA"], w.say) + B = b.verify_b(app, sub, s["seedA"], s["seedB"]) + C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None + return {"A": A, "B": B, "C": C} + + +def stage_undo(app): + s = load(app) + w.say(f"=== {app}: live undo by the product") + s["seedC"] = seed_c(app, s); s["seedC_at"] = ts() + w.say(" seed C written right before the Update:", bool(s["seedC"])) + before = b.db_state(app); w.say(" db before:", before) + phases, st = press(app) + obs = w.observables(app) + w.say(" observables:", json.dumps(obs)) + rb = readback(app, s) + after = b.db_state(app) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})") + pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False)) + w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip()) + s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")}, + "readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs} + save(app, s) + + + + +def set_box_language(lang): + """POST /settings/language — the household's language switch (form + the dashboard's CSRF).""" + sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip() + r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", + "-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}", + f"{w.BASE}/settings/language"]) + w.say(f" box language -> {lang}: http {r.stdout.strip()}") + + +def stage_powercut(app): + """Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard + stop); boot it again and let the controller resume the undo.""" + import subprocess + s = load(app) + w.say(f"=== {app}: power cut DURING the undo") + s["seedC"] = seed_c(app, s) + cut = {} + + def on_phase(ph): + if ph == "undoing": + t = time.time() + r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120) + cut["at"] = ts(); cut["rc"] = r.returncode + w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s") + return True + return False + phases, _ = press(app, poll=0.3, on_phase=on_phase) + if not cut: + w.say(" the cut never landed in `undoing` — recorded as a MISS"); return + r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180) + w.say(f" guest started again: rc={r.returncode}") + for i in range(60): + time.sleep(5) + try: + w.login(); st = w.stack(app) + if st: + break + except SystemExit: + continue + w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6")) + t0 = time.time() + while time.time() - t0 < 600: + st = w.stack(app) + if not st.get("updating") and st.get("update_phase") in ("undone", "failed"): + break + time.sleep(3) + w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}") + rb = readback(app, s) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}") + w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip()) + s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb} + save(app, s) + + +def stage_cutoff(app): + """Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of + the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it + back and HOLD with the new prefix.""" + s = load(app) + w.say(f"=== {app}: a cut-off undo copy") + done = {} + + def on_phase(ph): + if ph == "verifying" and not done: + out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c" +docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""") + done["out"] = out + w.say(" >>> finished-marker removed from one copy:\n" + out) + return False + phases, st = press(app, poll=0.5, on_phase=on_phase) + before = b.db_state(app) + pl = page_lines(app) + w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False)) + set_box_language("en") + pl_en = page_lines(app) + w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False)) + set_box_language("hu") + w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}")) + s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done} + save(app, s) + + +def stage_fixprobe_and_press(app): + """After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work + and end the undone note.""" + s = load(app) + frm, to, p0, p1 = b.EDGE[app] + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f) + w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"]) + w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120) + w.sync_rescan(app, to) + time.sleep(20) + w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan") + w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)") + w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False)) + phases, st = press(app) + rb = readback(app, s) + w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}") + pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False)) + w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml")) + s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl} + save(app, s) + + +if __name__ == "__main__": + w.login() + {"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff, + "fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2]) diff --git a/documentation/audits/undo-live-2026-09-23/repoint.py b/documentation/audits/undo-live-2026-09-23/repoint.py new file mode 100644 index 00000000..ff32bbaa --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/repoint.py @@ -0,0 +1,60 @@ +#!/usr/bin/env python3 +"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config. + +`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is +`controller.yaml.pre-bakeoff` (NOT the older `.pre-28`, which a restore must never pick up). +""" +import re, sys, io +sys.path.insert(0, '.') +import walk as w + +VOL = "/var/lib/docker/volumes/felhom-controller-data/_data" +DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git" + + +def creds(): + for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"): + m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l) + if m: + return m.group(1), m.group(2) + raise SystemExit("no admin credential") + + +def to_drill(): + u, t = creds() + print(w.guest(f""" +set -e +test -f {VOL}/controller.yaml.pre-bakeoff || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-bakeoff +python3 - <<'PY' +import re +p = "{VOL}/controller.yaml" +s = open(p).read() +s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M) +s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M) +if not re.search(r'^update:', s, re.M): + s += "update:\\n health_timeout: 90s\\n" +open(p, "w").write(s) +PY +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -A2 '^update:' {VOL}/controller.yaml +""")) + + +def restore(): + print(w.guest(f""" +set -e +cp -p {VOL}/controller.yaml.pre-bakeoff {VOL}/controller.yaml +rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache +docker restart felhom-controller >/dev/null +sleep 15 +grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: /' +grep -c '^update:' {VOL}/controller.yaml || true +""")) + + +if __name__ == "__main__": + to_drill() if sys.argv[1] == "drill" else restore() diff --git a/documentation/audits/undo-live-2026-09-23/setup.sh b/documentation/audits/undo-live-2026-09-23/setup.sh new file mode 100644 index 00000000..b79c2b72 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/setup.sh @@ -0,0 +1,20 @@ +set -u +python3 - <<'PY' > 30-reset-apps.txt 2>&1 +import walk as w, time +w.login() +for app in ("docmost","romm","vikunja"): + w.remove(app) +print(w.guest("docker volume ls -q --filter label=felhom.undo-copy-of | wc -l; for a in docmost romm vikunja; do docker volume ls -q --filter label=felhom.undo-copy-of=$a; done; rm -rf /mnt/felhom-drives/scratch_hdd/userdata/romm && echo romm-drive-folder-removed")) +for i in range(10): + c,d=w.ctl("POST","/api/sync") + if c=="200": print("sync",c); break + time.sleep(15) +w.ctl("POST","/api/stacks/rescan") +print(w.guest("C=/var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache; git -C $C log --oneline -1")) +PY +for app in docmost romm vikunja; do + timeout 1200 python3 bakeoff.py prep $app > 31-prep-$app.txt 2>&1; echo "prep $app rc=$?" + echo "cat /opt/docker/stacks/$app/applied-meta/.felhom.yml 2>&1 | grep -A3 healthcheck | grep port" | /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad/g.sh >> 31-prep-$app.txt 2>&1 + timeout 600 python3 bakeoff.py break $app > 32-break-$app.txt 2>&1; echo "break $app rc=$?" +done +echo SETUP-DONE diff --git a/documentation/audits/undo-live-2026-09-23/spike.py b/documentation/audits/undo-live-2026-09-23/spike.py new file mode 100644 index 00000000..f0de8d13 --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/spike.py @@ -0,0 +1,160 @@ +#!/usr/bin/env python3 +"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202. + +EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup, +sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being +spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the +steps the product would take, each one timed. + +State between stages lives in state-.json so each stage can be run, read, and only then +followed by the next (the hand undo needs a person looking at what the previous step left). +""" +import json, os, sys, time, re +sys.path.insert(0, ".") +import walk as w +import fixtures as fx + +HERE = os.path.dirname(os.path.abspath(__file__)) +FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()} +SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"} + + +def st_path(app): + return os.path.join(HERE, f"state-{app}.json") + + +def load(app): + return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {} + + +def save(app, s): + json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False) + + +def ts(): + return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + + +# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only +# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it +# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier. +def docmost_seed_b(sub, A): + jar = "/tmp/dm.jar" + w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + name = "drillB" + os.urandom(3).hex() + rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json", + data=json.dumps({"name": name, "slug": name.lower()}), method="POST") + w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}") + return {"space": name} if code in ("200", "201") else None + + +def docmost_verify_b(sub, A, B): + jar = "/tmp/dm.jar" + rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json", + data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST") + if code not in ("200", "201"): + w.say(f" docmost B: cannot log in (http={code})"); return False + rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json", + data="{}", method="POST") + ok = code in ("200", "201") and B["space"] in out + neg = "drillBnever" in out + w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def break_edge(app, frm, to, port_from, port_to): + """The failing edge: a REAL migrating image move, plus — in the DRILL template only — the + health probe pointed at a port the app does not answer. Both in one drill commit.""" + fy = f"{w.DRILL}/templates/{app}/.felhom.yml" + f = open(fy).read() + m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f) + assert m, "probe port not found" + f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():] + open(fy, "w").write(f) + h = w.drill_bump(app, frm, to) + return h + + +def pg_state(container, db, user): + return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; " + f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1") + + +def romm_seed_b(sub, A): + """A SECOND RomM user, created by the first (admin) one — written after the backup.""" + jar, tok = FX["romm"]._csrf(w, sub) + user = "drillb" + os.urandom(3).hex() + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json", + data=json.dumps({"username": user, "email": f"{user}@gate.invalid", + "password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}), + method="POST") + w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}") + return {"user": user} if code in ("200", "201") else None + + +def romm_verify_b(sub, A, B): + jar, tok = FX["romm"]._csrf(w, sub) + rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}", + "-u", f"{A['user']}:{A['pw']}") + ok = code == "200" and B["user"] in body + neg = "drillbnever" in body + w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})") + return ok and not neg + + +def tree(paths): + cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths) + return w.guest(cmd) + + +def my_state(container="romm-db", db="romm"): + # The root password is used INSIDE the container from its own env — it never leaves it. + q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '" + return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) " +echo -n "alembic=$({q("select version_num from alembic_version")}) " +echo -n "users=$({q("select count(*) from users")})" +""") + + +def vik_seed_b(sub, A): + """A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the + app writes into its files volume), uploaded through the app's own attachment API.""" + tok, why = FX["vikunja"]._token(w, sub, A) + title = "drillB-" + os.urandom(4).hex() + rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT") + if code not in ("200", "201"): + w.say(f" vikunja B: project refused {code} {body[:120]}"); return None + pid = json.loads(body)["id"] + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}", + "-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT") + tid = json.loads(body)["id"] if code in ("200", "201") else None + content = "drill attachment " + os.urandom(8).hex() + fn = "/tmp/vik-att.txt"; open(fn, "w").write(content) + rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}", + "-F", f"files=@{fn}", method="PUT") + w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}") + return {"title": title, "pid": pid, "tid": tid, "att": content} + + +def vik_verify_b(sub, A, B): + tok, why = FX["vikunja"]._token(w, sub, A) + if not tok: + w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False} + rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}") + proj = code == "200" and B["title"] in body + rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}") + att_ok = False + try: + atts = json.loads(body) + if atts: + aid = atts[0]["id"] + rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}") + att_ok = c3 == "200" and B["att"] in b3 + except Exception as e: + w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}") + w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}") + return {"project": proj, "attachment": att_ok} diff --git a/documentation/audits/undo-live-2026-09-23/walk.py b/documentation/audits/undo-live-2026-09-23/walk.py new file mode 100644 index 00000000..a284840b --- /dev/null +++ b/documentation/audits/undo-live-2026-09-23/walk.py @@ -0,0 +1,493 @@ +#!/usr/bin/env python3 +"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints. + +EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses: + POST /api/stacks//deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan + POST /api/stacks//update · POST /api/stacks//remove +and reads GET /api/stacks/. No controller code exists for it. + +The walk, per `09` §6.4 and the update-night brief §4: + 1 deploy from the DRILL catalog at the LIVE pin + 2 seed through the app's OWN front door (R-156: never a volume, never SQL) + 3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing + 4 „Mentés most" + 5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages + 6 press the guarded Update, record every phase with timestamps + 7 read the seed back through the front door + 8 the four version observables side by side + 9 write the verdict record in `09`'s JSON shape + +`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`. +""" +import argparse, json, os, re, subprocess, sys, time +from datetime import datetime, timezone + +SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad" +EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-live-2026-09-23" +DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill" +BASE = "https://192.168.0.114" +HOSTHDR = "Host: felhom.enkisfelhom.hu" +DOMAIN = "enkisfelhom.hu" +HP = "demo-hp" + +LOG = [] + + +def say(*a): + line = " ".join(str(x) for x in a) + ts = datetime.now().strftime("%H:%M:%S") + print(f"{ts} {line}", flush=True) + LOG.append(f"{ts} {line}") + + +def sh(args, timeout=300, inp=None): + try: + return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp) + except (subprocess.TimeoutExpired, OSError) as e: + return subprocess.CompletedProcess(args, 124, "", f"{e}") + + +def guest(script, timeout=600): + """Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting).""" + r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP, + "cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; " + "pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"], + timeout=timeout, inp=script) + return r.stdout or "" + + +def login(): + pw = open(f"{SC}/.ctlpw").read().strip() + sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR, + "-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"]) + h = open(f"{SC}/hdr.txt").read() + m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I) + if not m: + sys.exit("login failed: no session cookie") + open(f"{SC}/sess.txt", "w").write(m.group(0)) + r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"]) + c = re.search(r'/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN. + + Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because + a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password + (grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only + reason this was safe to discover by running it (live-probes rule). + + A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on + the scratch drive first — the same act the drive browser performs for a household. + """ + code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields") + fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or [] + values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub} + made = [] + for f in fields: + ev, ty = f.get("env_var"), f.get("type") + if ev in values: + continue + # `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses + # when the caller sends none, deliberately ("the user needs to know their password"), + # while `.felhom.yml` declares `required: false` and the API serves that verbatim. A + # caller that trusts the contract gets a 400. Measured tonight on grafana; filed. + if not f.get("required") and ty != "password": + continue # the controller generates the optional secrets itself + if ty == "path": + p = f"{DRIVE}/{name}" + values[ev] = p + made.append(p) + elif ty in ("secret", "password"): + import secrets as _s + values[ev] = "Drill-" + _s.token_hex(12) + GENERATED.setdefault(name, {})[ev] = values[ev] + elif f.get("default"): + values[ev] = f["default"] + else: + values[ev] = f"drill-{name}" + if made: + guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made)) + say(f" [1] made the drive paths this app requires: {made}") + extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")] + if extra: + say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}") + return values + + +def deploy(name, sub, extra_values=None): + st = stack(name) + if st.get("deployed"): + say(f" [1] {name} already deployed — reusing") + return True + values = deploy_values(name, sub) + if extra_values: + values.update(extra_values) + code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values}) + say(f" [1] deploy -> {code} {str(d)[:120]}") + if code != "202": + return False + # WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the + # container `healthy` while the controller's own state read `unhealthy` — a gate on `running` + # alone therefore times out on an app that is up. The state is RECORDED rather than required; + # the real gate is the fixture's own `wait_app`, which asks whether the APP answers. + seen = None + for _ in range(90): + time.sleep(5) + st = stack(name) + seen = st.get("state") + # `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21: + # tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm + # read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the + # deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark + # (`runComposeDeploy` writes it), so that is what to wait for. + pins = (st.get("app_config") or {}).get("pinned_images") + if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"): + say(f" [1] deployed, controller state={seen}, " + f"pinned={(st.get('app_config') or {}).get('pinned_images')}") + if seen != "running": + say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, " + f"not treated as a failure; the fixture's front-door wait is the real gate") + return True + say(f" [1] never became deployed (last controller state={seen!r})") + return False + + +def backup_now(name): + code, d = ctl("POST", "/api/backup/run") + say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}") + for _ in range(90): + time.sleep(5) + c2, s = ctl("GET", "/api/backup/status") + dd = s.get("data") or {} + if not dd.get("running", False): + say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}") + return True + say(" [4] backup still running after 7.5 min — carrying on") + return False + + +def drill_bump(app, frm, to, service_hint=None): + """Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates). + + `frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in + two images (adventurelog's backend and frontend) moves both in one edge, while its engine + sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move + every service that carries the app's own version and no others. + """ + comp = f"{DRILL}/templates/{app}/docker-compose.yml" + fy = f"{DRILL}/templates/{app}/.felhom.yml" + s = open(comp).read() + froms = [x.strip() for x in frm.split(",") if x.strip()] + tos = [x.strip() for x in to.split(",") if x.strip()] + if len(froms) != len(tos): + say(f" [5] from/to lists differ in length: {froms} vs {tos}") + return None + for f1, t1 in zip(froms, tos): + if f"image: {f1}" not in s: + say(f" [5] FROM ref not found in compose: {f1}") + return None + s = s.replace(f"image: {f1}", f"image: {t1}") + open(comp, "w").write(s) + f = open(fy).read() + today = datetime.now().strftime("%Y-%m-%d") + f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M) + open(fy, "w").write(f) + sh(["git", "-C", DRILL, "add", "-A"]) + sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"]) + r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120) + h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip() + say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})") + return h + + +def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5): + """Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP. + + R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and + `catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not + enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the + Update that followed moved nothing and still reported "Frissitve". So when the caller knows + which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the + NUMBER R-607 asks for and has never had. + """ + t0 = time.time() + ctl("POST", "/api/sync") + time.sleep(2) + ctl("POST", "/api/stacks/rescan") + time.sleep(2) + if not expect_app or not expect_ref: + return None + for i in range(tries): + cat = stack(expect_app).get("catalog_images") or {} + if expect_ref in cat.values(): + waited = round(time.time() - t0, 1) + if i: + say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up " + f"to {expect_ref} — R-607's window, measured") + return waited + time.sleep(delay) + ctl("POST", "/api/sync") + time.sleep(1) + ctl("POST", "/api/stacks/rescan") + say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — " + f"catalog_images = {stack(expect_app).get('catalog_images')}") + return None + + +def badges(name): + out = {} + for lang, suffix in (("hu", ""), ("en", "?lang=en")): + h = page(f"/apps/{name}{suffix}") + m = re.findall(r']*title="([^"]*)"[^>]*>([^<]*)<', h) + out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3] + return out + + +def press_update(name, poll=1.0, cap_s=1800): + code, d = ctl("POST", f"/api/stacks/{name}/update") + say(f" [6] Update -> {code} {str(d)[:220]}") + if code not in ("202", "200"): + return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0} + phases, seen, t0 = [], None, time.time() + while time.time() - t0 < cap_s: + st = stack(name) + ph = st.get("update_phase") + if ph != seen: + seen = ph + rec = {"t": round(time.time() - t0, 1), "phase": ph, + "label": st.get("update_phase_label"), "updating": st.get("updating"), + "error": st.get("update_error"), "hold": st.get("hold_reason")} + phases.append(rec) + say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} " + f"err={rec['error']} hold={rec['hold']}") + if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3: + break + time.sleep(poll) + st = stack(name) + return {"accepted": True, "http": code, "phases": phases, + "duration_s": round(time.time() - t0, 1), + "final_phase": st.get("update_phase"), "update_error": st.get("update_error"), + "hold_reason": st.get("hold_reason"), "state": st.get("state")} + + +def observables(name): + st = stack(name) + ac = st.get("app_config") or {} + live = guest(f""" +grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//' +echo '---inspect---' +for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do + echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}' +done +""") + a, _, b = live.partition("---inspect---") + return { + "pinned_images": ac.get("pinned_images"), + "installed_images": {k: (v.get("ref") if isinstance(v, dict) else v) + for k, v in (ac.get("installed_images") or {}).items()}, + "catalog_images": st.get("catalog_images"), + "live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()], + "docker_inspect": [x for x in b.strip().splitlines() if x.strip()], + } + + +def app_logs(name, lines=400): + """The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is + one string with escaped newlines — a scan over the envelope sees a single enormous line and + finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96 + rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines.""" + code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}") + if isinstance(d, dict): + data = d.get("data") + if isinstance(data, dict) and isinstance(data.get("logs"), str): + return data["logs"] + if isinstance(d.get("_raw"), str): + return d["_raw"] + return str(d) + + +def write_verdict(rec, appdir): + os.makedirs(appdir, exist_ok=True) + p = os.path.join(appdir, "verdict.json") + json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False) + say(f" [9] verdict {rec['verdict']} -> {p}") + + +def remove(name): + """Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint + refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up.""" + c1, d1 = ctl("POST", f"/api/stacks/{name}/stop") + say(f" [X] stop -> {c1} {str(d1)[:100]}") + for _ in range(24): + time.sleep(5) + if stack(name).get("state") != "running": + break + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": True, "remove_backups": True}) + say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}") + if code == "409": + # R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive + # path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202 + # `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits + # this. The household's other choice — remove the app, KEEP the data — is accepted, and the + # harness takes it, then tidies its own directory by name at teardown. + say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)" + " — removing the app and KEEPING the drive data instead") + code, d = ctl("POST", f"/api/stacks/{name}/remove", + {"remove_hdd_data": False, "remove_backups": True}) + say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}") + time.sleep(5) + st = stack(name) + left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; " + f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'") + say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}") + return code + + +def app_env(name, key): + """Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the + app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller + shows them the same value. Data still goes in through the app's own front door.""" + out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1") + if ":" in out: + return out.split(":", 1)[1].strip().strip('"').strip("'") + return "" + + +def snapshots(name): + """The restorable copies the backups page offers for this app.""" + code, d = ctl("GET", f"/api/backup/snapshots?stack={name}") + data = d.get("data") if isinstance(d, dict) else None + if isinstance(data, dict): + for k in ("snapshots", "items", "restore_points"): + if isinstance(data.get(k), list): + return data[k] + return data if isinstance(data, list) else [] + + +def restore(name, snapshot_id=None, wait_s=1200): + """The household's own way out: the „Visszaállítás a mentésből" button on the backups page. + + A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`, + `snapshot_id` — because that is the button the sentence tells them to press. + """ + snaps = snapshots(name) + if snapshot_id is None: + if not snaps: + say(f" [R] no restorable copy offered for {name}") + return {"ok": False, "why": "no snapshot offered", "snapshots": snaps} + first = snaps[0] + snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id") + say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)") + sess = open(f"{SC}/sess.txt").read().strip() + csrf = open(f"{SC}/csrf.txt").read().strip() + r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}", + "-X", "POST", + "--data-urlencode", f"_csrf={csrf}", + "--data-urlencode", f"stack_name={name}", + "--data-urlencode", f"snapshot_id={snapshot_id}", + f"{BASE}/backup/restore"], timeout=180) + head = (r.stdout or "").split("\n")[0].strip() + loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")] + say(f" [R] POST /backup/restore -> {head} {loc[:1]}") + t0 = time.time() + last = None + while time.time() - t0 < wait_s: + code, d = ctl("GET", "/api/backup/restore-status") + dd = d.get("data") or {} + cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message")) + if cur != last: + say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}") + last = cur + if not dd.get("running", False) and time.time() - t0 > 5: + break + time.sleep(2) + st = stack(name) + say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} " + f"phase={st.get('update_phase')}") + return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps, + "http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1), + "state_after": st.get("state"), "hold_after": st.get("hold_reason"), + "observables_after": observables(name)} diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 46b34971..1c18276a 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -335,3 +335,7 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` | | **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` | | **R-539** | **The restart brake caught a FAST crash loop and was blind to a SLOW one (P3).** Closed by operator ruling 3 of 2026-09-16 in agent **v0.132.0** (`18d03bd`, tag `v0.132.0`, sha256 `4afe8157…`) + hub **v0.117.0** (`37ae31f`): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth `slow_crashloop_since` moves (at most once per 24 h) and the hub mints `controller_slow_crashloop` (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed `agent_update` to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`. **Proven live with the PRODUCTION window (no test-only interval):** five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 `SLOW CRASH-LOOP … restarts_24h=5`; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). **Reasoning kept:** the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. `audits/evidence-chaos-fixes-2026-09-17/partC-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 3c1882a:documentation/backlog/OPEN-ITEMS.md` | +| **R-637** | **Build the undo (`09` §3 decision 15) — the box puts a failed update back by itself (P2).** Closed in controller **v0.263.2** (`8fc2b4a` v0.263.0, `5d38573` v0.263.1, `2cd6666` v0.263.2). Copy method chosen by the bake-off (decision 19): a folder copy of every NAMED volume, `cp -a` into `.pre-update-` after the pull, finished-marker last; bind folders never touched. Undo in `failAndHold`: copies validated first, volumes refilled, definition + pin + the pinned version's `.felhom.yml` (`applied-meta/`) put back from the job's own copies, old probe, `undone` or a HOLD whose sentence says the undo failed and the data state. **Proven live on 9202 (0.263.2):** docmost, romm, vikunja undone by the product with seeds before the backup, after it, and seconds before the press all read back, ledgers equal, page line hu/en; a cut-off copy → HOLD `untouched`; a power cut (`pct stop`) during `undoing` → resumed and undone; a manual press after an undo → `done`, note cleared; removal deleted kept copies. **Two defects the live proof found and the unit tests could not:** the undo's probe was gated on `running` while the current probe held the app `unhealthy` (v0.263.1), and the "old" `.felhom.yml` was already the new one because it flows in on every sync (v0.263.2). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. **Reasoning kept:** *a copy counts only with its finished-marker, judged by the helper container's own exit — killing `docker run` does not stop the copy; the undo reads its own copies, never the recovery unit.* | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | +| **R-639** | **After a held update the previous definition survived only in the recovery unit (P3).** Closed in controller **v0.263.0/v0.263.2**: the pre-update compose, applied definition, pin and the pinned version's `.felhom.yml` are kept until the undo is over; the undo never reads the unit (R-645). Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — SHIPPED** (controller v0.263.2) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | +| **R-641** | **An app with no database server had no last-second copy (P2).** Closed in controller **v0.263.0**: the undo's folder copy covers every named volume, so an app with no database server gets its copy by construction — proven live on vikunja (SQLite in a volume), seeds before/after the backup and seconds before the press read back. Evidence: `audits/undo-bakeoff-2026-09-23/`, `audits/undo-live-2026-09-23/`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | +| **R-642** | **Start answered 200 "start completed" over a crash loop (P3).** Closed in controller **v0.263.0**: start/restart answer `requested — state now: `, never "completed" (`startAnswer`, pinned by `TestR642_*`, red-proofed); live on 9202: `Stack romm start requested — state now: running`. | **CLOSED 2026-09-23 — PROVEN-LIVE** (controller v0.263.0) | full text: `git show 4c92bea:documentation/backlog/OPEN-ITEMS.md` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 69129937..2f516632 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -806,15 +806,12 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. **THE BOUNDED HALF IS FIXED in v0.262.0; the MECHANISM IS STILL NOT DIAGNOSED, and this row stays open for it.** `RemoveStack` refused on `!stack.Deployed` — a FLAG — while the machine plainly had containers, a compose file and an `app.yaml`. It now asks whether anything EXISTS (`halfStateEvidence`): containers, a compose file, or an `app.yaml` are each enough, and it removes what exists with every other fence unchanged (R-442 stays fail-closed). **The household must always be able to remove what the box shows them** — that is true whatever the cause of the bad record. Red-proof seen failing. **What is NOT done:** why `deployed` stays false while containers run. `runComposeDeploy`'s pin-write path was not read against a fresh reproduction in this session. **THE UNREMOVABLE HALF IS PROVEN LIVE ON 9202:** `sparkyfitness` rebuilt in its exact measured shape — `app.yaml` and a compose file on disk, no containers, `deployed: false`, `state: not_deployed` — and the remove answered **200** with `leftovers: NONE`. Under v0.261.0 the identical call answered `stack "sparkyfitness" is not deployed`. | **OPEN — P1 for the MECHANISM only; the unremovable half is CLOSED in v0.262.0 and proven live** | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** FOUND 2026-09-22 17:37 **because the operator heard the fans**, which is the only reason it was found at all. `romm` was promoted `5.0.0 -> 5.3.0` on the live catalog that morning (`audits/PROBE-FIX-2026-09-22.md`) after the edge was PROVEN on scratch guest 9202, and the guarded Update was then pressed on demo-hp guest 9201 at **09:08:22Z**, reaching **`done` in 74.8 s** with the app `running`. **It ran clean for two hours.** The first worker kill is at **11:09:20Z**; by 15:38Z there had been **4,530** of them — `Worker (pid:…) was sent SIGKILL! Perhaps out of memory?` — with `docker inspect` reading **`OOMKilled: true`**, the container pinned at **457 MiB of its 512 MiB limit**, and `docker stats` showing **499.51% CPU**. The host's load average sat at **5.2 while otherwise idle**. `rq_cron` is killed and restarted every few seconds in a permanent storm. **THREE THINGS THIS ESTABLISHES, and the third is the one that changes how promotions are judged.** (1) The template's own comment says *`RAM: ~300MB (mem_limit: 1024M total — romm 512M + mariadb 384M + redis 128M)`*; **5.3.0 needs more than 512M and the template was not re-sized when the version moved.** A version move is not only an `image:` line. (2) **The update reported `done` and the app reads `running`**, because nginx answers `GET /` with 200 while the gunicorn workers behind it are being killed — a THIRD variant of the R-618/R-630 theme: the probe is right, the port is right, and the answer is still a false green. (3) **THE ALARM DID FIRE, AND MY FIRST WRITE-UP OF THIS ROW SAID IT DID NOT — CORRECTED 2026-09-22 BY THE OPERATOR, WHO PRODUCED THE MAILS.** The controller HAS an OOM detector (`main.go:821`, `[deadapp] romm: container romm was OOM-killed`, evaluated every 30 s), it emits **`app_oom`** (`notifier.go:726`), the hub allow-lists it (`dispatcher.go:636`) and dispatched it to the OPERATOR channel as **SENT** — `admin@felhom.eu` received **`[Felhom] ⚠️ demo-hp: app_oom`** at **11:09 CEST** and again at **17:48 CEST**, each carrying the app name, the Hungarian sentence and the dashboard link. The CUSTOMER channel is **SKIPPED**, which is correct. **I asserted an absence without looking at the instrument** — the hub's own Events and Notifications tabs show all of it — and I did it by reasoning from a memory note (`lxc-docker-oom-signals-unreliable`, R-528) instead of reading the hub. **That is R-628's shape exactly, four days on, and from the same hand.** **WHAT IS ACTUALLY WRONG, and it is narrower and real:** `notifier.go:715-726` keys the alarm on `container|startedAt` in an `oomSeen` map and emits **once per container lifetime**. So **six hours of continuous thrashing — 4,530 worker kills — produced exactly ONE mail**, severity `warning`, never escalating, while the app went on reading `running`. **A six-hour storm is indistinguishable from a single transient kill.** The hub's App Telemetry panel did carry the magnitude (RomM: **5,023 errors, 632 warnings**, peak 855 MB) but nothing turns that into a second, louder signal. So the fix worth having is not a detector — there is one — but an ESCALATION: a warning that repeats for hours should stop looking like a warning that happened once. **AND THE LESSON FOR R-462's METHOD, which is the real cost:** every `proven` verdict in the update night and in the twenty-eight measures the app for the **minutes of the walk**, not for a day of running. `romm` passed its walk, was seeded, read back and restored — and broke two hours later. **`proven` currently means "the update applied and the data survived", NOT "the new version runs".** Needs: decide between rolling the catalog back to 5.0.0 and raising romm's `mem_limit` (measured, not guessed); and a soak longer than a walk before any future promotion. **FIXED AND MEASURED 2026-09-22, in two steps, and the FIRST step was still a guess.** *Step 1 (operator's choice):* the limit was raised 512M → **768M** (`app-catalog-felhom.eu@886956d`). It slowed the kills from ~12/min to ~7/min and **stopped nothing** — 37 SIGKILLs in five minutes, `OOMKilled` still true, and the cgroup's own `memory.peak` read **exactly 768 MiB**: it hit the new ceiling and died there. *Step 2, from a MEASUREMENT instead:* the per-process RSS inside the container reads **~216 MiB per warm uvicorn worker**, so the image's default of four workers plus the master needs **~882 MiB** before nginx and the job runner — more than any sensible limit for this box. **The lever was in the image all along:** `/init:143` runs `--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a SERVER default on an appliance serving one household.** Setting **`WEB_SERVER_CONCURRENCY=2`** (`app-catalog-felhom.eu@f4eb94f`, limit left at 768M) and applying it through the product's own Update button fixed it. **PROVEN UNDER LOAD, not just at idle:** 6 concurrent callers driven at romm through the household's own route for 300 s — **26,645 requests** (9,687 × 200, 16,958 × 401 on the auth-gated endpoints), CPU a steady **~200%** (exactly two workers saturated, by design), memory oscillating **416–614 MiB against the 768 MiB limit and trending DOWN**, and **zero** SIGKILLs, `OOMKilled: false`, `RestartCount: 0` throughout. At idle afterwards: **1.64% CPU**, 610 MiB, host load falling from 5.2 to 2.4. **AND THE FIRST SOAK MEASURED NOTHING, which is worth more than the second one:** it was pointed at `arcade.enkisfelhom.hu` — the scratch-guest fixture's default subdomain — while this box deployed romm at `jatek`. Every request 404'd at traefik in 9 ms, romm sat idle at 0.64% CPU, and the counter cheerfully reported **14,026 successful requests**. It was caught only because 0.64% CPU under load is not believable. **A positive AND a negative control are now asserted before any load is driven** (`jatek` must not 404; a nonsense host must). R-96 rule 3, in a new surface. **WHAT STAYS OPEN, and it is the part that outlives romm:** 610 MiB of 768 MiB is **79%** — it works with ~158 MiB of headroom and the soak never exceeded 614 MiB, but it is not generous, and nothing watches it. **And the method lesson for R-462:** a version move is not only an `image:` line — the new version's SHAPE (worker counts, per-worker footprint) has to be measured too, and a walk lasting minutes cannot see a ceiling reached in two hours. Every `proven` verdict in the update night and in the twenty-eight means *"the update applied and the data survived"*, **not** *"the new version runs"*. Evidence: `audits/probe-fix-2026-09-22/romm-soak.json`, `romm-soak.out`. | **CLOSED 2026-09-22 — two workers, 768M, proven under 26,645 requests; the 79% headroom and the `proven`-means-minutes lesson are carried into R-462** | | **R-636** | **[P2-MEDIUM] An app that has been out of memory for six hours sends the same single warning an app that hiccuped once sends.** FOUND 2026-09-22 when the operator produced the alarm mails I had wrongly written off as absent (R-635). **The detector is correct and works.** `main.go:821` re-checks every 30 s and logs `[deadapp] romm: container romm was OOM-killed`; `notifier.go:726` emits **`app_oom`** at severity `warning`; the hub allow-lists it (`dispatcher.go:636`) and delivered it to the OPERATOR channel — two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER is SKIPPED, correctly. **The defect is the SHAPE of the signal, not its absence.** `notifier.go:715-724` keys on `container|startedAt` in an `oomSeen` map and returns early on a repeat, so the alarm fires **once per container lifetime**. romm's 09:08 container produced **4,530 worker kills over six hours and exactly one mail**; the 15:47 container produced one more. Severity never escalates, the app goes on reading `running`, and `08` §4 rightly does not put it in `IsDownState`. **So a six-hour storm that pins five cores is indistinguishable, in the operator's inbox, from one transient kill at 3 a.m.** — and the operator, who had two correct mails, still found the fault by hearing the fans. **The once-per-lifetime rule is right in itself** (it is what stops a crash loop from mailing 4,530 times, which is R-629's lesson) — what is missing is the second, louder signal when the same key keeps re-firing. **The magnitude IS already collected:** the hub's App Telemetry panel read **RomM: 5,023 errors, 632 warnings, peak 855 MB against a 1280M limit** while the same panel showed every other app at 0. Nothing turns that into an event. **Candidate shapes, none chosen here:** escalate `app_oom` to `error` when the same `container|startedAt` key re-fires past a threshold; or a periodic digest for a key still firing after N minutes; or let App Telemetry raise its own event when an app's error count crosses a bound. All are hub/controller code. Evidence: the operator's screenshots of the Events, Notifications and App Telemetry tabs plus the two mails; `felhom-controller/controller/internal/notify/notifier.go:715-726`. | **OPEN — P2; owner: CC; product code, so not fixed unattended** | -| **R-637** | **[P2-MEDIUM] BUILD THE UNDO (`09` §3 decision 15) — the eight things the 2026-09-23 spike says the product must add before a failed update can put itself back.** SPIKED BY HAND 2026-09-23 on 9202, three real migrating edges each made to fail a deliberately wrong probe: docmost (PostgreSQL) and romm (MariaDB) — whose OLD versions REFUSE the migrated data (*corrupted migrations …*, *Can't locate revision …*) — and vikunja (SQLite in a volume), whose old version starts. **The undo worked on all three**: data written before AND after the backup read back through each app's own front door, ≈16 s / ≈38 s / ≈1 s after the failing health wait. **The list:** (1) keep the pre-update copies — compose, applied, pin AND the old `.felhom.yml` — until the undo is over (`failAndHold` deletes the first three, the fourth was never kept, R-639); (2) the undo inside `failAndHold`: pin back → DB up alone → validated load → start → health with the OLD probe → `undone`, else HOLD; (3) empty-then-load atomically, because the loader the product has cannot undo a migration (R-638); (4) check the copy's completion marker first (R-640); (5) a volume copy at safety-dump time for apps with no database server (R-641); (6) files: nothing measured touched them — a step that does says so through decision 13's *files may change* mark; (7) the undo's success is the probe, never the Start's return (R-642); (8) remember the failed step so the caller never re-presses it. `09` §6.1a, §6.4 part 1 (4 evenings). Evidence: `audits/update-rulings-2026-09-23/README.md`. | **READY — owner: CC; `09` §6.4 part 1, returns to the operator for go/no-go** | -| **R-638** | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. | **OPEN — P2; owner: CC; measure the named restore after a real schema migration first** | -| **R-639** | **[P3-LOW] After a failed update is held, the previous definition survives only in the recovery unit — `failAndHold` deletes the journal's copies, and the old `.felhom.yml` was never kept.** READ FROM SOURCE and CONFIRMED 2026-09-23: `failAndHold` (`stacks/update.go:766`) ends with `clearJournal` + `removePreUpdateCopies`; in all three spike cases `ls /*pre-update*` counted 0 after the hold. The old definition was taken from the unit's `compose/` directory each time — present only because a backup preceded the update. The old `.felhom.yml` matters as much: its probe is the one the OLD version answers, and the catalog's copy flows to the app regardless (§5.4). Needed by R-637. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-41`, `romm-40`. | **OPEN — P3; owner: CC; lands with R-637** | -| **R-640** | **[P2-MEDIUM] A TRUNCATED PostgreSQL copy loads with exit 0 into an EMPTY database — and `ValidateDump` would accept it.** MEASURED 2026-09-23 on 9202 in a scratch database beside docmost's: the first half of the 135 816-byte undo copy, loaded with `psql -v ON_ERROR_STOP=1 --single-transaction`, returned **rc 0** and left **42 tables and 48 migration rows — and 0 users, 0 spaces, 0 constraints, 0 indexes**; psql treats end-of-file inside a `COPY` as end of data and commits. The whole copy ends with `-- PostgreSQL database dump complete` (plus a `\unrestrict` line); the truncated one does not. `ValidateDump` (`appbackup/dbdump.go:415`, read from source) checks the header and a `CREATE TABLE` only. **An undo or a restore fed a truncated copy would report success, start the app on an empty database, and pass a health check.** MariaDB's truncated load failed (rc 1) but NON-atomically — half the tables already replaced; its copy carries `-- Dump completed`. Fix: check the engine's completion marker before any load. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-46`, `docmost-47`, `romm-44`. | **OPEN — P2; owner: CC; lands with R-637, also guards the restore paths** | -| **R-641** | **[P2-MEDIUM] An app with no database server has NO last-second copy — the update's safety dump is a no-op — so if its old version refuses the migrated data, everything written since the last backup is lost by any route back.** MEASURED 2026-09-23 on 9202 with vikunja (SQLite in a volume): the controller logged *update safety dump for vikunja: the app has no database — nothing to copy (no-op)*. Vikunja 2.3.0 happens to start on 2.6.0's migrated SQLite, so nothing was lost; the fallback was measured anyway — the tier unit's volume tar put back in 0.86 s brought back the data written before the backup and **NOT the project and attachment written after it**. The fix is a tar of the data volumes at safety-dump time (2.2 MB here). Needed by R-637. Evidence: `audits/update-rulings-2026-09-23/README.md`, `vikunja-40`..`-44`. | **OPEN — P2; owner: CC; lands with R-637** | -| **R-642** | **[P3-LOW] `POST /api/stacks/{name}/start` answers 200 *Stack … start completed* while the app is crash-looping.** MEASURED 2026-09-23 on 9202 twice: docmost 0.95.0 and romm 5.0.0 started on data their newer versions had migrated — both refused and restarted in a loop (`Restarting (1)`, front door 404) behind a 200. The same false-green class as R-443 (closed for the Update) and R-635, on the Start. It matters now because an undo (R-637) must never read the start's return as success. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-43`, `romm-43`. | **OPEN — P3; owner: CC** | -| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. | **WAITING-ON-OPERATOR — `09` §6.4's one open point; owner: CC once answered** | +| **R-638** | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | +| **R-640** | **[P2-MEDIUM] A TRUNCATED PostgreSQL copy loads with exit 0 into an EMPTY database — and `ValidateDump` would accept it.** MEASURED 2026-09-23 on 9202 in a scratch database beside docmost's: the first half of the 135 816-byte undo copy, loaded with `psql -v ON_ERROR_STOP=1 --single-transaction`, returned **rc 0** and left **42 tables and 48 migration rows — and 0 users, 0 spaces, 0 constraints, 0 indexes**; psql treats end-of-file inside a `COPY` as end of data and commits. The whole copy ends with `-- PostgreSQL database dump complete` (plus a `\unrestrict` line); the truncated one does not. `ValidateDump` (`appbackup/dbdump.go:415`, read from source) checks the header and a `CREATE TABLE` only. **An undo or a restore fed a truncated copy would report success, start the app on an empty database, and pass a health check.** MariaDB's truncated load failed (rc 1) but NON-atomically — half the tables already replaced; its copy carries `-- Dump completed`. Fix: check the engine's completion marker before any load. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-46`, `docmost-47`, `romm-44`. **-- NARROWED 2026-09-23:** the undo is not exposed (a folder copy, validated by its own finished-marker). The restore paths still load dumps through `ValidateDump`, which does not check the engine's completion marker. | **OPEN — P2, narrowed to the restore paths; owner: CC** | +| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** | | **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** | -| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". | **OPEN — P3; owner: CC** | +| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | +| **R-646** | **[P3-LOW] An app pinned before controller v0.263.2 has no record of its own `.felhom.yml`, so its FIRST undo probes with whatever file the catalog sync last put in place.** The undo judges the old version with the pinned version's probe, kept in `/applied-meta/` since v0.263.2 — written at deploy, pin adoption and each guarded update. Every app deployed and pinned before that release has no record until its next pin; its first update keeps the CURRENT `.felhom.yml` for the undo (logged: *no applied .felhom.yml for the pinned version*). When the new version changed its probe, that current file is the new one, and an undo that worked would be judged "did not start" and HELD — an honest hold (`not_started`), never a false success. MEASURED on 9202 2026-09-23 (romm, v0.263.1's shape). **A backfill is possible for the apps that are NOT behind** (their current file IS the pinned version's): a startup pass like `AdoptPins`, for apps whose catalog order is equal. Apps already behind cannot be backfilled — the pinned version's file is gone. `pin.go`, `undo.go` `savePreUpdateMeta`. | **OPEN — P3; owner: CC** |